{"entries":[{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-121","CWE-119"],"fleet":{"pain_class":"firmware-flash"},"id":"CVE-2013-3607","cve":"CVE-2013-3607","aliases":["VU#648646"],"title":"Supermicro IPMI BMC web interface (login.cgi) - H8/X7/X8/X9 generation boards, firmware before SMT_X9_315","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro IPMI BMC web interface (login.cgi) - H8/X7/X8/X9 generation boards, firmware before SMT_X9_315","year":"2013","cvss_score":10,"severity":"critical","kev":false,"impact":"Unauthenticated root on the baseboard management controller, delivered through the login form. The BMC web server strcpy()s the submitted username and password into fixed 128-byte and 24-byte stack buffers, so an over-long login field executes code on the service processor before any credential is checked. Owning the BMC is strictly better than owning the host: it survives OS reinstall, it can reflash the BIOS, it can mount virtual media to boot arbitrary code, it can power-cycle the node, and it can read the host's console. In a GPU rental fleet, one compromised BMC is a persistent implant under every tenant that ever lands on that node. CERT counted roughly 135 affected board models across 30 firmware versions - this is not a niche SKU.","attack_vector":"Network, pre-auth. Anything that can reach the BMC's HTTP(S) port. Fatal when the management network is flat, reachable from tenant VLANs, or - as internet scans repeatedly found - exposed directly to the internet.","remediation":"Flash BMC firmware to SMT_X9_315 or later. Roughly 10-20 minutes per node including the BMC reset, and the node should be drained first because a BMC reset during a job risks losing console and out-of-band control mid-flight - and on some boards a failed flash bricks the management controller entirely, so this is a change window, not a rolling background task. The compensating control to apply immediately regardless of firmware state is network: BMCs belong on an isolated management VLAN reachable only from a bastion, never routable from tenant networks. Most operators with this exposure have a topology problem, not just a firmware problem.","references":["https://www.kb.cert.org/vuls/id/648646","https://www.rapid7.com/blog/post/2013/11/06/supermicro-ipmi-firmware-vulnerabilities/","https://www.usenix.org/system/files/conference/woot13/woot13-bonkoski_0.pdf"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2013-4782","cve":"CVE-2013-4782","aliases":[],"title":"Supermicro BMC (IPMI cipher suite 0): Authentication bypass","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC (IPMI cipher suite 0)","year":"2013","cvss_score":10,"severity":"critical","kev":false,"impact":"Authentication bypass — arbitrary IPMI commands with any password when cipher suite 0 is enabled. Still shipped enabled on some ODM builds a decade later","attack_vector":"Network / IPMI over LAN","remediation":"Disable cipher suite 0 in BMC config fleet-wide and assert it in a config scan; not fixable by firmware alone since it is a spec-permitted mode","references":["https://nvd.nist.gov/vuln/detail/CVE-2013-4782"],"status":"curated","published":"2013-07-08"},{"id":"CVE-2013-4783","cve":"CVE-2013-4783","aliases":[],"title":"Dell iDRAC (IPMI 1.5 cipher 0): Remote authentication bypass via cipher suite 0","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC (IPMI 1.5 cipher 0)","year":"2013","cvss_score":10,"severity":"critical","kev":false,"impact":"Remote authentication bypass via cipher suite 0","attack_vector":"Network / IPMI","remediation":"Disable cipher 0 in the iDRAC config baseline","references":["https://nvd.nist.gov/vuln/detail/CVE-2013-4783"],"status":"curated","published":"2013-07-08"},{"id":"CVE-2013-4784","cve":"CVE-2013-4784","aliases":[],"title":"HPE iLO (IPMI cipher 0): IPMI authentication bypass via cipher suite 0 on the iLO BMC","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HPE iLO (IPMI cipher 0)","year":"2013","cvss_score":10,"severity":"critical","kev":false,"impact":"IPMI authentication bypass via cipher suite 0 on the iLO BMC","attack_vector":"Network / IPMI","remediation":"Config baseline change disabling cipher 0","references":["https://nvd.nist.gov/vuln/detail/CVE-2013-4784"],"status":"curated","published":"2013-07-08"},{"id":"CVE-2014-2955","cve":"CVE-2014-2955","aliases":[],"title":"Raritan PX rack PDU (before firmware 1.5.11, DPXR20A-16 and related PX models): The PDU's IPMI interface accepts","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Raritan PX rack PDU (before firmware 1.5.11, DPXR20A-16 and related PX models)","year":"2014","cvss_score":10,"severity":"critical","kev":false,"impact":"The PDU's IPMI interface accepts 'cipher suite 0' (aka cipher zero), a known-broken IPMI auth mode that accepts any password. An attacker who reaches the IPMI port can issue arbitrary IPMI commands with no valid credential — including power-control commands to cut or cycle outlets feeding whatever racks that PDU serves, tenant-owned or not.","attack_vector":"Fully remote and unauthenticated over the network-reachable IPMI port — the attacker just needs to request cipher suite 0 and supply any password; the PDU accepts it.","remediation":"Firmware upgrade to 1.5.11 or later, which disables cipher-zero support. If a firmware upgrade isn't immediately possible, a network-segmentation change to firewall off the IPMI port from anything but a trusted management VLAN is the compensating control — this is the same 'IPMI cipher zero' class of bug that hit many BMC vendors around the same era, so audit for other devices on the same segment too. Flash each PDU one at a time; outlets keep powering their load during the update, but remote power-control briefly drops.","references":["https://nvd.nist.gov/vuln/detail/CVE-2014-2955"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2014-07-14"},{"id":"CVE-2015-5053","cve":"CVE-2015-5053","aliases":["GRID vGPU host memory mapping"],"title":"NVIDIA GPU driver / GRID vGPU and vSGA - host memory mapping path: The host memory mapping path did not restrict access","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU driver / GRID vGPU and vSGA - host memory mapping path","year":"2015","cvss_score":10,"severity":"critical","kev":false,"impact":"The host memory mapping path did not restrict access to third-party device I/O memory, giving a guest privilege escalation on the host. Scored 10.0. Included here despite its age because it is the archetype of the vGPU escape class and because GRID/vSGA stacks this old are still found running in long-lived VDI estates that were never in anyone's patch programme.","attack_vector":"A guest on a GRID vGPU or vSGA host. The mapping path is reachable from the guest driver.","remediation":"Upgrade the host driver to R346 346.87 / R352 352.41 (Linux) or R352 352.46 (GRID vGPU and vSGA) or, realistically, to a supported branch - anything on these versions is a decade out of support. Cost: full hypervisor host drain. If you find this in your estate the finding is not the CVE, it is that you have an unmanaged vGPU host.","references":["https://nvd.nist.gov/vuln/detail/CVE-2015-5053"],"status":"curated","fleet":{"pain_class":"node-drain"},"tags":["tenant-isolation"],"published":"2015-11-24"},{"id":"CVE-2017-12542","cve":"CVE-2017-12542","aliases":[],"title":"HPE iLO4: Authentication bypass and remote code execution — the \"29 A's\" `Connection` header bug","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HPE iLO4","year":"2017","cvss_score":10,"severity":"critical","kev":false,"impact":"Authentication bypass and remote code execution — the \"29 A's\" `Connection` header bug; trivially scriptable pre-auth root on the BMC","attack_vector":"Network, unauthenticated","remediation":"iLO4 firmware update to 2.53+; a node left unpatched here is fully owned by a single curl request","references":["https://nvd.nist.gov/vuln/detail/CVE-2017-12542"],"status":"curated","fleet":{"ubiquity":"common - iLO is HPE's BMC across ProLiant/Apollo, incl. GPU-dense SKUs","remediation_pain":"firmware-flash to iLO >= 2.54 on every node, out-of-band; the exploit is trivial (a long header) and public, so exposure windows are measured in hours","pain_class":"firmware-flash","why_fleet_wide":"Unauthenticated remote auth bypass into the BMC yields administrator on the management processor, virtual-media boot of attacker media, and firmware-level persistence under the OS - identical across every HPE node of that generation."},"published":"2018-02-15"},{"id":"CVE-2019-16649","cve":"CVE-2019-16649","aliases":[],"title":"Supermicro BMC virtual media (H11/H12/M11/X9/X10/X11): Virtual media service uses weak/absent encryption","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC virtual media (H11/H12/M11/X9/X10/X11)","year":"2019","cvss_score":10,"severity":"critical","kev":false,"impact":"Virtual media service uses weak/absent encryption and authentication — credential capture and attaching an arbitrary virtual USB device to the host, i.e. arbitrary boot media","attack_vector":"Network, unauthenticated","remediation":"BMC firmware update across every affected generation; interim control is blocking the virtual-media ports (623/5900/5901) at the management-network boundary","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-16649"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2019-09-21"},{"id":"CVE-2019-16650","cve":"CVE-2019-16650","aliases":["USBAnywhere"],"title":"Supermicro X10/X11 BMC (virtual media service): The BMC's virtual media service reuses socket file descriptors, so","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro X10/X11 BMC (virtual media service)","year":"2019","cvss_score":10,"severity":"critical","kev":false,"impact":"The BMC's virtual media service reuses socket file descriptors, so an unauthenticated attacker inherits an existing client's privileges and can attach a virtual USB device to the server. In practice that means booting your node from the attacker's image, or dropping files onto a running host, without ever having a BMC credential. On a bare-metal GPU fleet this is a direct tenant-to-tenant and outsider-to-host compromise, and the implanted image outlives any OS reinstall.","attack_vector":"Anything with a network route to the BMC's virtual media port. No credentials, no user interaction. Supermicro boards are the whitebox default under a large share of neocloud GPU capacity, and BMCs on these boards are frequently found directly on a routable network.","remediation":"BMC firmware flash per node, out-of-band, with the usual Supermicro caveat that the fixed version differs per board SKU - you need a per-model inventory before you can plan the rollout. Immediate config-only mitigation that actually works: block the virtual media ports (623, 5900, 623/udp and the 623x range Supermicro uses) at the network edge and put every BMC behind a jump host on a dedicated management VLAN. Do the network control first; the flash campaign will take weeks.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-16650","https://eclypsium.com/blog/usbanywhere/"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2019-09-21"},{"id":"CVE-2019-5684","cve":"CVE-2019-5684","aliases":[],"title":"NVIDIA Windows GPU Display Driver, DirectX driver (shader compiler/runtime): A crafted shader reads out of bounds on an","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver, DirectX driver (shader compiler/runtime)","year":"2019","cvss_score":10,"severity":"critical","kev":false,"impact":"A crafted shader reads out of bounds on an input texture array and gets code execution in the driver. VMware shipped its own advisory for this (VMSA-2019-0012) because in a virtualised graphics setup the shader comes from inside a guest VM - so this is a guest-to-host code execution path on a shared GPU host, which is why it carries a CVSS of 10.0. On a GPU cloud running vSGA/vGPU-adjacent graphics, one tenant's shader compromises the hypervisor host and therefore every other tenant on it.","attack_vector":"Anyone who can submit a shader to the host GPU: a tenant VM, a remote graphics session, or a local process. In the virtualised case, an unprivileged user inside any guest.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved. If you run VMware ESXi with virtualised graphics, also apply the fix VMware ships in VMSA-2019-0012 - the hypervisor-side package is separate from the in-guest driver, and both matter.","references":["http://www.vmware.com/security/advisories/VMSA-2019-0012.html","https://support.lenovo.com/us/en/product_security/LEN-28096","https://www.talosintelligence.com/vulnerability_reports/TALOS-2019-0779","https://nvd.nist.gov/vuln/detail/CVE-2019-5684"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2019-08-06"},{"id":"CVE-2021-1388","cve":"CVE-2021-1388","aliases":[],"title":"Cisco ACI Multi-Site Orchestrator (Application Services Engine): Complete unauthenticated authentication bypass on the","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Cisco ACI Multi-Site Orchestrator (Application Services Engine)","year":"2021","cvss_score":10,"severity":"critical","kev":false,"impact":"Complete unauthenticated authentication bypass on the controller that programs policy across every ACI site. Whoever gets this owns the tenant separation model for the entire multi-site fabric — they can write EPG and contract policy that stitches any tenant to any other. It is a 10.0 for a reason.","attack_vector":"Unauthenticated, remote — anything that can reach the MSO API endpoint. If the orchestrator's management interface is on a flat ops network, that is a very large set of machines.","remediation":"Upgrade the MSO application. Application-level upgrade rather than a switch reload, so the data plane stays up — but treat any pre-patch exposure as a policy compromise and re-audit every contract and EPG binding afterwards, which is the real cost.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1388"],"status":"curated","tags":["tenant-isolation"],"published":"2021-02-24"},{"id":"CVE-2021-22205","cve":"CVE-2021-22205","aliases":[],"title":"GitLab: Image files passed unvalidated to a file parser (ExifTool)","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"GitLab","year":"2021","cvss_score":10,"severity":"critical","kev":true,"impact":"Image files passed unvalidated to a file parser (ExifTool) -> unauthenticated remote command execution","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; assume compromise on any unpatched internet-facing instance","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-22205"],"status":"curated","published":"2021-04-23"},{"id":"CVE-2021-23281","cve":"CVE-2021-23281","aliases":[],"title":"Eaton Intelligent Power Manager (IPM) prior to 1.69: Unauthenticated remote code execution on Eaton's power-management","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Eaton Intelligent Power Manager (IPM) prior to 1.69","year":"2021","cvss_score":10,"severity":"critical","kev":false,"impact":"Unauthenticated remote code execution on Eaton's power-management platform, via unsanitised data reaching a Node.js code path. IPM is the software that monitors and shuts down infrastructure in response to power events across a site, so unauthenticated RCE here is total control of the power-response layer. A CVSS 10.0 on facility software is rare and this one earns it.","attack_vector":"Unauthenticated, remote, over the network to the IPM server.","remediation":"Upgrade IPM to 1.69 or later. Server-side upgrade, so cheap in maintenance terms - but assume compromise on any instance that was network-reachable and rebuild rather than patch. Rotate every device credential IPM held. IPM ships as part of several OEM power bundles, so check for it under other names before concluding you do not run it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-23281"],"status":"curated","tags":["physical-impact"],"published":"2021-04-13"},{"id":"CVE-2021-26728","cve":"CVE-2021-26728","aliases":[],"title":"Lanner IAC-AST2500A BMC standard firmware 1.10.0: Arbitrary code execution as root on the BMC, at the maximum severity","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lanner IAC-AST2500A BMC standard firmware 1.10.0","year":"2021","cvss_score":10,"severity":"critical","kev":false,"impact":"Arbitrary code execution as root on the BMC, at the maximum severity the scale allows. The operator outcome is complete out-of-band ownership of every node carrying this module: power control, boot-device selection, KVM into whatever the tenant has on screen, virtual-media boot of an attacker-supplied image, and an implant in BMC flash that survives host reimaging. Because the IAC-AST2500A is a module dropped into other people's board designs, an operator may be running it without the word 'Lanner' appearing anywhere in their asset inventory. Command injection plus stack buffer overflows in the KillDupUsr_func handler of spx_restservice, the REST service that fronts the BMC web interface. The IAC-AST2500A is a MegaRAC-derived BMC module resold into a wide range of whitebox and edge server designs.","attack_vector":"Network reachability to the BMC's REST service. Anything routable to the out-of-band management VLAN, which for whitebox and edge deployments is frequently less segmented than in a purpose-built datacenter.","remediation":"Firmware flash from Lanner, per module. This is the hardest remediation class in this database: Lanner's website returns a blanket Cloudflare 403 to automated clients, the advisories that exist are third-party (Nozomi Networks) rather than vendor-published, and firmware for a resold BMC module often has to be sourced through the board integrator rather than Lanner directly. For many operators the honest answer is that no fixed image is obtainable, and the only real mitigation is hard network isolation of the management VLAN with an explicit allowlist. Inventory first - identify which of your boards carry an IAC-AST2500A before assuming you are unaffected.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26728","https://www.nozominetworks.com/labs/vulnerability-advisories/cve-2021-26728/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-26729","cve":"CVE-2021-26729","aliases":[],"title":"Lanner IAC-AST2500A BMC firmware 1.10.0: Root on the BMC without any credential at all, because the vulnerable handler","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lanner IAC-AST2500A BMC firmware 1.10.0","year":"2021","cvss_score":10,"severity":"critical","kev":false,"impact":"Root on the BMC without any credential at all, because the vulnerable handler is the login handler. This is the single worst entry in the ODM section: an attacker with a network path and nothing else takes complete out-of-band control of the node - power, console, virtual media, and persistent firmware residency below every layer the operator can reimage. On bare-metal rental this destroys the tenant handoff guarantee, because a node compromised this way stays compromised through wipe and reprovision. Command injection and multiple stack buffer overflows in the Login_handler_func function of spx_restservice, i.e. in the code that runs before anyone has authenticated.","attack_vector":"Anything that can reach the BMC's REST service over the network, unauthenticated. No credential, no host access, no prior foothold - only routability to the management interface.","remediation":"Firmware flash from Lanner or your board integrator. As with the rest of this cluster, obtaining a fixed image is the real obstacle: the vendor's site is WAF-blocked to automated access and the public advisories are third-party. Given the unauthenticated nature of this one, treat network isolation as mandatory and immediate rather than as a stopgap - these BMCs must not be reachable from anything except a small, explicitly allowlisted set of management hosts, and never from a tenant or general corporate network. If you cannot obtain fixed firmware, document the residual risk and plan the hardware out.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26729","https://www.nozominetworks.com/labs/vulnerability-advisories/cve-2021-26729/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-32798","cve":"CVE-2021-32798","aliases":[],"title":"Jupyter Notebook (untrusted notebooks): Untrusted notebook executes JavaScript in the user's session on open","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Jupyter Notebook (untrusted notebooks)","year":"2021","cvss_score":10,"severity":"critical","kev":false,"impact":"Untrusted notebook executes JavaScript in the user's session on open","attack_vector":"Customer-supplied `.ipynb` opened by another user or an operator","remediation":"Upgrade. Notebook files shared between tenants are active content","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-32798"],"status":"curated","published":"2021-08-09"},{"id":"CVE-2021-39296","cve":"CVE-2021-39296","aliases":["GHSA-gg9x-v835-m48q"],"title":"OpenBMC phosphor-net-ipmid (IPMI 2.0 RMCP+ / IPMI over LAN): The headline OpenBMC bug","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"OpenBMC phosphor-net-ipmid (IPMI 2.0 RMCP+ / IPMI over LAN)","year":"2021","cvss_score":10,"severity":"critical","kev":false,"impact":"The headline OpenBMC bug. Crafted IPMI session-setup messages skip authentication entirely and hand the attacker full administrative control of the BMC - no credentials, no prior access, just UDP packets at the management interface. From there: power-cycle any node, mount virtual media, install BMC-resident firmware that survives host reimage, and pivot onto the host. On a GPU cluster where the management VLAN reaches every node, one packet source that can see that VLAN owns the fleet's out-of-band plane. CVSS 10.0 with scope change, which is rare and deserved. Google's security team reported it; Intel shipped it as SA-00737.","attack_vector":"Network access to the BMC's IPMI-over-LAN port (UDP 623). Unauthenticated. In practice: anyone who reaches the management VLAN - a misrouted tenant network, a jump host, a compromised switch, or a BMC accidentally exposed to the internet.","remediation":"Fixed in OpenBMC after 2.9. Getting the fix onto nodes is a BMC firmware flash: out-of-band, per node, ODM-rebase-lagged, brick risk. But the config-only mitigation here is strong and should be done first, today: disable IPMI over LAN entirely and use Redfish. OpenBMC's Redfish support has been production-ready for years and most fleets no longer need RMCP+. If you cannot disable it, ACL UDP 623 so only your management jump hosts can reach it. Check your fleet for BMCs reachable outside the management VLAN before doing anything else.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-39296","https://github.com/google/security-research/security/advisories/GHSA-gg9x-v835-m48q","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00737.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2021-09-09"},{"id":"CVE-2022-0543","cve":"CVE-2022-0543","aliases":[],"title":"Redis: Debian/Ubuntu packaging leaves a Lua sandbox escape","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Redis","year":"2022","cvss_score":10,"severity":"critical","kev":true,"impact":"Debian/Ubuntu packaging leaves a Lua sandbox escape -> RCE; mass-exploited by botnets","attack_vector":"Network (remote)","remediation":"Control-plane: distro package upgrade; never expose Redis to a tenant network","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-0543"],"status":"curated","published":"2022-02-18"},{"id":"CVE-2022-21941","cve":"CVE-2022-21941","aliases":["ICSA-22-242-11"],"title":"Software House iSTAR Ultra door controller (before 6.8.9.CU01): Unauthenticated command injection giving root","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Software House iSTAR Ultra door controller (before 6.8.9.CU01)","year":"2022","cvss_score":10,"severity":"critical","kev":false,"impact":"Unauthenticated command injection giving root on the door controller. The iSTAR Ultra is the panel that decides whether a door opens - including the doors into the datacenter hall and the cages inside it. Root on it means an attacker unlocks doors on demand, holds them unlocked, enrols their own credentials, and deletes or rewrites the access log so no record of the entry exists. What someone standing inside your cage can then do is the real impact: pull NVMe drives holding customer data and model weights, attach a laptop to a server's console or management port, plug into the out-of-band switch and reach every BMC in the row, install an interposer on a management link, or simply photograph the topology. For a bare-metal GPU provider that is also a tenant-handoff catastrophe - the boundary you sell your customers is precisely that nobody else can physically touch their machines, and this bug removes it silently. Root persistence on the controller means the compromise survives your incident response unless you re-image the panel.","attack_vector":"Unauthenticated, over the network, to the controller. iSTAR panels are typically on a dedicated physical-security VLAN, which sounds reassuring until you check who else is on it: the CCTV/VMS servers, the intercom system, the badge-office workstations, and the security integrator's remote-support path. Any of those is a stepping stone. Panels are also often mounted in unsecured back-of-house spaces, so physical access to the panel's Ethernet port is a parallel route.","remediation":"Firmware update to 6.8.9.CU01 or later, applied through the Software House/Johnson Controls integrator. Door controller firmware updates are disruptive in a specific way most operators underestimate: doors typically fall back to a fail-secure or fail-safe local mode during the flash, so you need security staff physically present at affected doors for the window. Plan it, do not skip it - a 10.0 on the panel guarding your GPUs is not something to defer. After patching, re-image rather than trust any panel you suspect was reachable while vulnerable, rotate the integrator's credentials, and put the physical-security VLAN behind a firewall with an explicit allow-list rather than treating it as inherently trusted.","references":["https://www.cisa.gov/news-events/ics-advisories/icsa-22-242-11","https://nvd.nist.gov/vuln/detail/CVE-2022-21941","https://www.johnsoncontrols.com/cyber-solutions/security-advisories"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-29165","cve":"CVE-2022-29165","aliases":[],"title":"Argo CD: Unauthenticated attacker forges JWTs and gains full Argo CD admin, which in a GitOps cluster","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2022","cvss_score":10,"severity":"critical","kev":false,"impact":"Unauthenticated attacker forges JWTs and gains full Argo CD admin, which in a GitOps cluster means cluster-admin","attack_vector":"Unauthenticated network reaching the Argo CD API","remediation":"Emergency Argo CD upgrade; rotate the signing key and all tokens; audit all Applications for tampering","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-29165"],"status":"curated","published":"2022-05-20"},{"id":"CVE-2022-29226","cve":"CVE-2022-29226","aliases":[],"title":"Envoy: OAuth filter does not validate access tokens, so authentication can be skipped entirely","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy","year":"2022","cvss_score":10,"severity":"critical","kev":false,"impact":"OAuth filter does not validate access tokens, so authentication can be skipped entirely","attack_vector":"Unauthenticated network","remediation":"Emergency Envoy upgrade; sidecar and gateway restart","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-29226"],"status":"curated","published":"2022-06-09"},{"id":"CVE-2022-31481","cve":"CVE-2022-31481","aliases":["CVE-2022-31479","CVE-2022-31483","CVE-2022-31484","CVE-2022-31486","ICSA-22-153-01"],"title":"HID Mercury intelligent controllers sold by Carrier LenelS2 (LNL-X2210/X2220/X3300/X4420/4420","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HID Mercury intelligent controllers sold by Carrier LenelS2","year":"2022","cvss_score":10,"severity":"critical","kev":false,"impact":"A full chain against the access-control panel generation that sits underneath a very large share of enterprise and datacenter badge systems, including Lenel OnGuard and S2 deployments: unauthenticated command execution via a crafted hostname, an unauthenticated buffer overflow in the update path, unauthenticated deletion of web-interface users, and authenticated path traversal and command injection. CISA's own summary is that an attacker gets device access, can monitor all communications to and from the panel, and can modify the onboard relays - the relays are the door strikes. So this is not badge-database tampering, it is direct electrical control of whether a door opens, plus visibility into every badge read at that panel. Someone who walks into your cage on the back of this can pull drives containing model weights and customer data, attach a console to a running node, or plug into the out-of-band switch and reach every BMC in the row. Because HID Mercury panels are OEM'd under multiple brands, many operators do not know they have them - check the board, not the badge software's vendor name.","attack_vector":"Unauthenticated network access to the panel for the worst of the set. Panels sit on the physical-security VLAN, usually in back-of-house electrical or comms closets, sometimes in the same closets tenants and contractors can reach. The multi-brand OEM situation widens exposure: a site can have Mercury boards behind three different vendors' software without a single asset record naming Mercury.","remediation":"Firmware update from the OEM whose badge is on your panel - Carrier LenelS2 published fixed firmware, and other Mercury OEMs issued their own. Identify the actual board model first (LNL-X2210 and friends, S2-LP-*), because the fix tracks the board, not the head-end software. Applying it is a security-integrator engagement with doors in local fallback during the flash, so it needs security staff on site. After patching, assume any panel exposed during the vulnerable window may have had relays or user accounts manipulated: audit the badge database, the panel user list, and the relay configuration against a known-good baseline. Long term, put the physical-security VLAN behind a firewall with an explicit allow-list from the head-end only, and enable port security on the switch ports serving panels.","references":["https://www.cisa.gov/news-events/ics-advisories/icsa-22-153-01","https://nvd.nist.gov/vuln/detail/CVE-2022-31481","https://nvd.nist.gov/vuln/detail/CVE-2022-31479","https://www.corporate.carrier.com/product-security/advisories-resources/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-2825","cve":"CVE-2023-2825","aliases":[],"title":"GitLab: Unauthenticated path traversal reads arbitrary server files when an attachment sits under 5+ nested groups","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"GitLab","year":"2023","cvss_score":10,"severity":"critical","kev":false,"impact":"Unauthenticated path traversal reads arbitrary server files when an attachment sits under 5+ nested groups","attack_vector":"Network (remote)","remediation":"Control-plane: emergency upgrade; assume repository secrets were read","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-2825"],"status":"curated","published":"2023-05-26"},{"id":"CVE-2023-31029","cve":"CVE-2023-31029","aliases":[],"title":"DGX A100 BMC: Full BMC compromise (heap buffer overflow) — worst-case out-of-band takeover","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX A100 BMC","year":"2023","cvss_score":10,"severity":"critical","kev":false,"impact":"Full BMC compromise (heap buffer overflow) — worst-case out-of-band takeover","attack_vector":"Network-adjacent unauthenticated on mgmt LAN","remediation":"Emergency: flash BMC 00.22.05+ out-of-band, audit BMC logs for compromise, rotate all mgmt credentials","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31029","https://github.com/NVIDIA/product-security/tree/main/2024/5510"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-121"],"published":"2024-01-12"},{"id":"CVE-2023-3765","cve":"CVE-2023-3765","aliases":[],"title":"MLflow: Absolute path traversal prior to 2.5.0","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"MLflow","year":"2023","cvss_score":10,"severity":"critical","kev":false,"impact":"Absolute path traversal prior to 2.5.0","attack_vector":"Unauthenticated network","remediation":"Upgrade to 2.5.0+","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-3765"],"status":"curated","published":"2023-07-19"},{"id":"CVE-2023-3939","cve":"CVE-2023-3939","aliases":["CVE-2023-3941","CVE-2023-3940","CVE-2023-3943","CVE-2023-3938","CVE-2023-3942"],"title":"ZKTeco-based OEM biometric access terminals (ZKTeco ProFace X, Smartec ST-FR043/ST-FR041ME and rebadged equivalents)","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"ZKTeco-based OEM biometric access terminals (ZKTeco ProFace X, Smartec ST-FR043/ST-FR041ME and rebadged equivalents)","year":"2023","cvss_score":10,"severity":"critical","kev":false,"impact":"OS command injection where every command runs as root, arbitrary file write as root via path traversal, arbitrary file read, a stack overflow with no stack canaries or PIE, and SQL injection allowing authentication as any user in the device database. This is the reader on the wall next to the door - the device that decides whether a face or a badge opens the hall. Root on it means an attacker opens the door at will, enrols their own biometric template, harvests the biometric templates and badge data of everyone who has ever entered (which is a privacy and regulatory problem on top of a security one), and keeps persistence on a device nobody ever patches. The rebadging is the trap: these terminals ship under many brand names, so an operator's asset list may say Smartec or a local integrator's label with no mention of ZKTeco anywhere. Once someone is through that door and into the cage, they reach drives holding model weights, server console ports, and the out-of-band management switch fronting every BMC in the row.","attack_vector":"Network access to the terminal for the injection and traversal paths - these devices sit on the physical-security VLAN and are frequently given a network address by whoever installed them with no ACL at all. Several of the flaws are also reachable by an attacker with brief physical proximity to the device, since the terminal is by definition mounted on the unsecured side of the door.","remediation":"Firmware from ZKTeco or the OEM that rebadged the device - and this is the problem, because rebadged terminals frequently never receive the upstream fix and the OEM may no longer exist. Start by physically identifying every biometric or badge terminal in the facility and determining the actual manufacturer of the board, not the label. Where a fixed firmware exists, flash it; where it does not, replace the terminal, which is a per-door hardware cost but a small one relative to what is behind the door. Regardless: isolate reader devices onto a segment that cannot reach anything else, never allow a reader to be routable from a tenant or corporate network, and if biometric templates are stored on-device, treat them as already exposed and notify accordingly.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-3939","https://nvd.nist.gov/vuln/detail/CVE-2023-3941","https://nvd.nist.gov/vuln/detail/CVE-2023-3943"],"status":"curated"},{"id":"CVE-2023-43654","cve":"CVE-2023-43654","aliases":["ShellTorch"],"title":"TorchServe: Unauthenticated SSRF","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"TorchServe","year":"2023","cvss_score":10,"severity":"critical","kev":false,"impact":"Unauthenticated SSRF → arbitrary model download → RCE on the serving host","attack_vector":"Unauthenticated network to an exposed TorchServe management port (8081), default config binds broadly","remediation":"Patch to 0.8.2+, restrict `allowed_urls`, and firewall 8080/8081/7070/7071 off any tenant-reachable network. Providers who publish TorchServe images must ship the hardened default","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-43654"],"status":"curated","fleet":{"ubiquity":"Common - TorchServe is the default PyTorch model server on SageMaker-style managed inference and on many neocloud inference offerings","remediation_pain":"`daemon-restart` - upgrade to TorchServe 0.8.2+ and bind `management_address` to 127.0.0.1; restarting the server drops in-flight inference but does not require a node reboot","pain_class":"node-reboot","why_fleet_wide":"The management API listened on 0.0.0.0 with no auth and accepted model URLs from *any* domain, so an unauthenticated attacker uploads a malicious model archive and gets RCE on every exposed inference host at once"},"published":"2023-09-28"},{"id":"CVE-2023-7028","cve":"CVE-2023-7028","aliases":[],"title":"GitLab: Password reset email deliverable to an unverified address","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"GitLab","year":"2023","cvss_score":10,"severity":"critical","kev":true,"impact":"Password reset email deliverable to an unverified address -> full account takeover","attack_vector":"Network (remote)","remediation":"Control-plane: URGENT upgrade + enforce 2FA + audit all reset events","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-7028"],"status":"curated","published":"2024-01-12"},{"id":"CVE-2024-0001","cve":"CVE-2024-0001","aliases":[],"title":"Pure Storage FlashArray Purity (dormant configuration account): A local account intended only for initial array","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Pure Storage FlashArray Purity (dormant configuration account)","year":"2024","cvss_score":10,"severity":"critical","kev":false,"impact":"A local account intended only for initial array configuration stays active, letting an attacker gain elevated privileges on the array. Full control of a storage array serving an AI cluster means access to training data and checkpoints, plus the ability to destroy them.","attack_vector":"Network access to the array management interface, using the leftover account.","remediation":"Apply the Purity update from Pure's security page. Array software upgrade - non-disruptive on a healthy dual-controller array, but schedule it. Verify the configuration account is actually disabled afterwards rather than trusting the version number.","references":["https://purestorage.com/security"],"status":"curated"},{"id":"CVE-2024-0002","cve":"CVE-2024-0002","aliases":[],"title":"Pure Storage FlashArray Purity (privileged remote access account): An attacker uses a privileged account to gain remote","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Pure Storage FlashArray Purity (privileged remote access account)","year":"2024","cvss_score":10,"severity":"critical","kev":false,"impact":"An attacker uses a privileged account to gain remote access to the array, with scope change - complete compromise of the storage tier.","attack_vector":"Remote network access to the array. No prior credentials.","remediation":"Apply the Purity update from Pure's security page as a priority; this one is unauthenticated. Non-disruptive controller upgrade on healthy arrays.","references":["https://purestorage.com/security"],"status":"curated"},{"id":"CVE-2024-11186","cve":"CVE-2024-11186","aliases":[],"title":"Arista CloudVision Portal (on-premise): An authenticated CloudVision user can take actions on managed EOS devices well","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Arista CloudVision Portal (on-premise)","year":"2024","cvss_score":10,"severity":"critical","kev":false,"impact":"An authenticated CloudVision user can take actions on managed EOS devices well beyond what their role should allow. CloudVision is the fabric's configuration and streaming-telemetry brain, so 'broader actions than intended' means pushing configlets to switches you were never granted. In a shared operations model — a neocloud with tenant-facing NOC accounts, or an MSP — this collapses the internal privilege model for the entire fabric.","attack_vector":"Any authenticated CloudVision Portal user on an on-premise deployment.","remediation":"Upgrade CloudVision Portal. Application upgrade on the CVP cluster; the switches keep forwarding. Afterwards review the CVP change log for configlet pushes that did not come from an authorized operator — that audit is the real work.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-11186"],"status":"curated","published":"2025-05-08"},{"id":"CVE-2024-1709","cve":"CVE-2024-1709","aliases":[],"title":"ConnectWise ScreenConnect: Auth bypass via alternate path","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"ConnectWise ScreenConnect","year":"2024","cvss_score":10,"severity":"critical","kev":true,"impact":"Auth bypass via alternate path -> direct access to critical systems; trivially exploited at scale","attack_vector":"Network (remote)","remediation":"Control-plane: patch the RMM - compromise means code execution on every managed node","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-1709"],"status":"curated","published":"2024-02-21"},{"id":"CVE-2024-22216","cve":"CVE-2024-22216","aliases":["CVE-2023-51438"],"title":"Microchip maxView Storage Manager Redfish server (Adaptec SmartRAID / SmartHBA controllers), 3.00.23484 through","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Microchip maxView Storage Manager Redfish server (Adaptec SmartRAID / SmartHBA controllers), 3.00.23484 through…","year":"2024","cvss_score":10,"severity":"critical","kev":false,"impact":"In a default install where the Redfish server is enabled for remote management, the Redfish endpoint accepts unauthorized requests - read and write. The attacker gets the controller's own management API for Adaptec SmartRAID/SmartHBA: enumerate logical drives, change configuration, and reach controller update functions. Wiping or reconfiguring an array under a live training job is a fleet-availability event; the controller-update path is the worse one, because Adaptec controller firmware runs below the host OS with DMA and survives a tenant reimage, so an operator cannot honestly certify tenant handoff on a node that was exposed. CVSS 10.0 with a scope change is the vendor's own rating.","attack_vector":"Any host that can reach the maxView Redfish listener on the management network - no credentials. maxView is commonly installed by OEM server tooling and its Redfish service is on by default, so this is usually reachable from the provisioning VLAN rather than only from a hardened admin subnet.","remediation":"Upgrade maxView Storage Manager to 4.14.00.26068 or later (Microchip also back-patched 3.07.23980 and 4.07.00.25339). Software-only - restart the maxView service, no controller flash, no arrays offline. If you do not actually consume the Redfish interface, disable the maxView Redfish server outright; that is the faster fleet-wide mitigation and costs nothing. Note the OEM lag explicitly: Siemens shipped the same defect as CVE-2023-51438 in its industrial PCs a year later, and Dell/HPE/Supermicro rebadge maxView the same way - check the OEM bundle version, not just Microchip's.","references":["https://www.microchip.com/en-us/solutions/embedded-security/how-to-report-potential-product-security-vulnerabilities/maxview-storage-manager-redfish-server-vulnerability","https://nvd.nist.gov/vuln/detail/CVE-2024-22216","https://nvd.nist.gov/vuln/detail/CVE-2023-51438"],"status":"curated","published":"2024-01-08"},{"id":"CVE-2024-22476","cve":"CVE-2024-22476","aliases":[],"title":"Intel Neural Compressor: An unauthenticated user can reach an input-validation failure in Neural Compressor","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel Neural Compressor","year":"2024","cvss_score":10,"severity":"critical","kev":false,"impact":"An unauthenticated user can reach an input-validation failure in Neural Compressor and escalate. This is scored at the top of the scale, and Neural Compressor is a quantisation and optimisation service that teams commonly stand up as a shared internal endpoint next to their model registry - so an exposed instance is a pre-auth foothold beside your model weights.","attack_vector":"Network-reachable and unauthenticated where the service is exposed. Treat any internal deployment as reachable by anything else on the cluster network.","remediation":"Upgrade Intel Neural Compressor to 2.5.0 or later immediately, and put the service behind authentication and network policy regardless of version. Python package update, restart the service - no node reboot or firmware.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-22476","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01109.html"],"status":"curated","fleet":{"pain_class":"daemon-restart"},"published":"2024-05-16"},{"id":"CVE-2024-23652","cve":"CVE-2024-23652","aliases":[],"title":"BuildKit: \"Leaky Vessels\": RUN --mount empty-file removal can delete arbitrary host files","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"BuildKit","year":"2024","cvss_score":10,"severity":"critical","kev":false,"impact":"\"Leaky Vessels\": RUN --mount empty-file removal can delete arbitrary host files; complete builder compromise","attack_vector":"Malicious Dockerfile or frontend submitted to a shared builder","remediation":"Upgrade BuildKit/buildx everywhere; treat any shared tenant builder as compromised and rebuild it","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-23652"],"status":"curated","published":"2024-01-31"},{"id":"CVE-2024-2912","cve":"CVE-2024-2912","aliases":[],"title":"BentoML: Insecure deserialization","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"BentoML","year":"2024","cvss_score":10,"severity":"critical","kev":false,"impact":"Insecure deserialization → RCE from a crafted POST","attack_vector":"Unauthenticated network to the BentoML serving port","remediation":"Upgrade. A default BentoML service is an unauthenticated RCE endpoint pre-patch","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-2912"],"status":"curated","published":"2024-04-16"},{"id":"CVE-2024-3400","cve":"CVE-2024-3400","aliases":[],"title":"Palo Alto PAN-OS: GlobalProtect arbitrary file creation","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Palo Alto PAN-OS","year":"2024","cvss_score":10,"severity":"critical","kev":true,"impact":"GlobalProtect arbitrary file creation -> command injection, unauthenticated root on the firewall","attack_vector":"Network (remote)","remediation":"Control-plane: emergency hotfix; full forensics and rotate every secret on-box","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-3400"],"status":"curated","published":"2024-04-12"},{"id":"CVE-2024-42479","cve":"CVE-2024-42479","aliases":[],"title":"llama.cpp (RPC backend): Unsafe `data` pointer in `rpc_tensor`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"llama.cpp (RPC backend)","year":"2024","cvss_score":10,"severity":"critical","kev":false,"impact":"Unsafe `data` pointer in `rpc_tensor` → arbitrary address write","attack_vector":"Unauthenticated network to the llama.cpp RPC port on a distributed inference setup","remediation":"Rebuild; the RPC backend has no authentication and must never be tenant-reachable","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-42479"],"status":"curated","published":"2024-08-12"},{"id":"CVE-2024-44984","cve":"CVE-2024-44984","aliases":[],"title":"Linux bnxt_en driver (XDP_REDIRECT double DMA unmap): A double DMA unmap in the XDP_REDIRECT path","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux bnxt_en driver (XDP_REDIRECT double DMA unmap)","year":"2024","cvss_score":10,"severity":"critical","kev":false,"impact":"A double DMA unmap in the XDP_REDIRECT path. Double-unmapping a DMA region is a memory-corruption primitive that touches the IOMMU mapping state, which is precisely the mechanism that keeps a device from reaching memory it should not. Anywhere you run XDP-based load balancing or packet steering in front of inference serving — a common pattern — this is live code.","attack_vector":"Traffic through the driver's XDP_REDIRECT path on a host with an XDP program attached.","remediation":"Kernel/driver upgrade plus host reboot. Interim: detach XDP programs from Broadcom NICs, which is a live change but costs you whatever the XDP program was doing.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-44984"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-04"},{"id":"CVE-2024-45409","cve":"CVE-2024-45409","aliases":[],"title":"GitLab (ruby-saml): Ruby-SAML does not properly verify the SAML Response signature","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"GitLab (ruby-saml)","year":"2024","cvss_score":10,"severity":"critical","kev":false,"impact":"Ruby-SAML does not properly verify the SAML Response signature -> forge an assertion as any user, incl. admin","attack_vector":"Network (remote)","remediation":"Control-plane: URGENT upgrade; treat all SAML sessions as suspect","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-45409"],"status":"curated","published":"2024-09-10"},{"id":"CVE-2024-6886","cve":"CVE-2024-6886","aliases":[],"title":"Gitea: Stored cross-site scripting in Gitea 1.22.0","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Gitea","year":"2024","cvss_score":10,"severity":"critical","kev":false,"impact":"Stored cross-site scripting in Gitea 1.22.0 -> session theft from maintainers","attack_vector":"Network (remote)","remediation":"Control-plane: Gitea upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-6886"],"status":"curated","published":"2024-08-06"},{"id":"CVE-2024-7591","cve":"CVE-2024-7591","aliases":[],"title":"Progress Kemp LoadMaster (including Multi-Tenancy edition): A request handler fails to validate its input before","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Progress Kemp LoadMaster (including Multi-Tenancy edition)","year":"2024","cvss_score":10,"severity":"critical","kev":false,"impact":"A request handler fails to validate its input before passing it to a system call, giving an attacker OS command execution on the LoadMaster with no authentication at all. This explicitly affects the Multi-Tenancy edition — the product line built to let multiple tenants share one LoadMaster — so a single unauthenticated request can compromise the appliance underneath every tenant's virtual services on that instance.","attack_vector":"Fully remote and unauthenticated — a crafted request to the vulnerable API endpoint is sufficient, no login required.","remediation":"Software upgrade to the fixed LoadMaster/Multi-Tenancy release per Kemp's advisory. Patch immediately given the unauthenticated, maximum-severity nature of this bug; if this instance is shared across tenants, treat any exposure window as a potential full-tenant-boundary breach and audit for signs of compromise, not just apply the patch.","references":["https://insinuator.net/2024/11/vulnerability-disclosure-command-injection-in-kemp-loadmaster-load-balancer-cve-2024-7591","https://support.kemptechnologies.com/hc/en-us/articles/29196371689613-LoadMaster-Security-Vulnerability-CVE-2024-7591"],"status":"curated","tags":["tenant-isolation"],"published":"2024-09-05"},{"id":"CVE-2024-8525","cve":"CVE-2024-8525","aliases":["CVE-2024-8526","ICSA-24-326-01"],"title":"Automated Logic WebCTRL 7.0 / WebCTRL Premium Server / Carrier i-Vu building automation server: Unauthenticated file","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Automated Logic WebCTRL 7.0 / WebCTRL Premium Server / Carrier i-Vu building automation server","year":"2024","cvss_score":10,"severity":"critical","kev":false,"impact":"Unauthenticated file upload leading to remote command execution on the BAS server itself - the machine that holds the graphics, the schedules, the trends and, crucially, the write authority over every controller in the building. A 10.0 here means an attacker who reaches the web port gets to be the building operator. They can command CRAH fans to zero, raise chilled-water setpoints, disable economizers, force valves shut, and edit the schedules so the change survives a reboot and looks intentional. Against a hall of 40-140 kW GPU racks that is minutes to thermal shutdown and a hardware-damage risk, and because the same server owns the trend and alarm database, the attacker can also rewrite what the operator sees while it happens. Owning WebCTRL is strictly better than owning any single controller, and it does not require a single credential from the compute network.","attack_vector":"Unauthenticated HTTP POST to the WebCTRL server. WebCTRL is a Windows/Tomcat application, and it is very commonly published beyond the facility VLAN because facilities staff and the controls contractor want browser access - that is the exposure that turns this from 'facility network' to 'internet-exposed via a badly-placed remote-access box or a public DNS entry'. If your site has a WebCTRL login page reachable from anything other than a jump host, treat that as already compromised until proven otherwise.","remediation":"Software upgrade on the BAS server - no controller firmware, no cooling downtime, so this one is genuinely fixable in a normal change window and there is no excuse to defer it. Move to a fixed WebCTRL/i-Vu release per Carrier's advisory. Then do the thing that should have been done first: take the WebCTRL server off any interface reachable from the internet or the corporate LAN, require VPN plus MFA to a jump host, and confirm the Tomcat service account is not a domain administrator. In a leased colo the WebCTRL server is the landlord's and often shared across the whole building - ask for its version and its network position, and treat 'it is behind our VPN' as an unverified claim until you see the rule.","references":["https://www.cisa.gov/news-events/ics-advisories/icsa-24-326-01","https://nvd.nist.gov/vuln/detail/CVE-2024-8525","https://www.corporate.carrier.com/product-security/advisories-resources/"],"status":"curated"},{"id":"CVE-2025-0505","cve":"CVE-2025-0505","aliases":[],"title":"Arista CloudVision (Zero Touch Provisioning): Zero Touch Provisioning can be abused to obtain admin privileges","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Arista CloudVision (Zero Touch Provisioning)","year":"2025","cvss_score":10,"severity":"critical","kev":false,"impact":"Zero Touch Provisioning can be abused to obtain admin privileges on the CloudVision system itself. ZTP is by design an unauthenticated-ish onboarding path — a new switch shows up and asks for its config — so this turns 'plug a device into the provisioning VLAN' into 'own the fabric controller'. For anyone doing rack-and-stack at scale, which is every GPU buildout, the ZTP network is live constantly.","attack_vector":"A device that can participate in ZTP against the CloudVision instance — i.e. anything on the provisioning network. Physical or logical access to that VLAN is the whole requirement.","remediation":"Upgrade CloudVision. Beyond the patch, treat the ZTP/provisioning VLAN as a privileged network: separate it from the general management network, keep it shut down when not actively provisioning, and require MAC/serial allowlisting. Those are config and process changes and they are what actually keeps this closed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0505"],"status":"curated","published":"2025-05-08"},{"id":"CVE-2025-15036","cve":"CVE-2025-15036","aliases":[],"title":"MLflow (`extract_archive_to_dir`): Path traversal in the dbconnect artifact cache","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"MLflow (`extract_archive_to_dir`)","year":"2025","cvss_score":10,"severity":"critical","kev":false,"impact":"Path traversal in the dbconnect artifact cache","attack_vector":"Customer-supplied artifact archive","remediation":"Upgrade immediately — maximum severity","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-15036"],"status":"curated","published":"2026-03-30"},{"id":"CVE-2025-29270","cve":"CVE-2025-29270","aliases":[],"title":"Deep Sea Electronics DSE855 generator communications gateway v1.1.0-v1.1.26 (realtime.cgi): Incorrect access control","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Deep Sea Electronics DSE855 generator communications gateway v1.1.0-v1.1.26 (realtime.cgi)","year":"2025","cvss_score":10,"severity":"critical","kev":false,"impact":"Incorrect access control on the realtime.cgi endpoint hands an attacker the admin panel and complete control of the device with no credentials. The DSE855 is the Ethernet gateway that fronts DSE generator and transfer-switch controllers, exposing them to the BMS and to remote monitoring. Control of it means control of the interface between the building's standby power and everything that watches it. Concretely: an attacker can reconfigure or disable the monitoring path so a genset failure goes unreported, and can reach the controller behind it. For a GPU datacenter this is the second half of the thermal problem - the cooling plant is the thing that stops when utility power drops and the generators do not pick up. A hall that loses chillers on a failed transfer is on the same minutes-to-thermal-shutdown clock as one whose CRAHs were commanded off, except now the UPS is draining under a 40-140 kW-per-rack load that it was probably not sized to ride through for long.","attack_vector":"Unauthenticated HTTP on the facility network - network-adjacent, no credentials. DSE855 units sit in generator yards, switchgear rooms and electrical closets on the building network. They are a classic 'installed by the generator contractor, never inventoried by IT' device, and because they exist to provide remote monitoring they are disproportionately likely to be reachable from outside the building through whatever remote-access arrangement the contractor set up.","remediation":"Firmware update from Deep Sea Electronics for the DSE855. This is a small standalone gateway, so the flash itself is quick and does not require taking generators out of service - but it does need someone with physical or network access to the unit and a maintenance window on the monitoring path, and the generator contractor usually owns that relationship rather than the datacenter operator. Alongside patching: put generator and switchgear network devices on an isolated segment, remove any internet path, and specifically audit the generator contractor's remote-access arrangement, which is frequently a consumer-grade router or a cellular modem nobody in IT knows about. In a leased site the gensets and their gateways are the landlord's - ask who can reach them remotely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-29270","https://blog.byteray.co.uk/shadow-entry-discovery-of-authentication-bypass-vulnerability-in-dse855-communications-device-938e35d4b361"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-32444","cve":"CVE-2025-32444","aliases":[],"title":"vLLM (Mooncake ZMQ/TCP): Unsafe deserialization exposed on all interfaces","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (Mooncake ZMQ/TCP)","year":"2025","cvss_score":10,"severity":"critical","kev":false,"impact":"Unsafe deserialization exposed on all interfaces → RCE","attack_vector":"Unauthenticated network from any host on the cluster fabric","remediation":"Upgrade to 0.8.5+. Highest-severity vLLM issue; assume any pre-0.8.5 disaggregated deployment is compromised","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-32444"],"status":"curated","published":"2025-04-30"},{"id":"CVE-2025-34028","cve":"CVE-2025-34028","aliases":[],"title":"Commvault Command Center: Unauthenticated ZIP upload + path traversal","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Commvault Command Center","year":"2025","cvss_score":10,"severity":"critical","kev":true,"impact":"Unauthenticated ZIP upload + path traversal -> RCE via a malicious JSP","attack_vector":"Network (remote)","remediation":"Control-plane: patch immediately; the backup control plane is a crown-jewel target","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-34028"],"status":"curated","published":"2025-04-22"},{"id":"CVE-2025-37164","cve":"CVE-2025-37164","aliases":[],"title":"HPE OneView (unauthenticated remote code execution): Unauthenticated remote code execution on OneView with scope change","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HPE OneView (unauthenticated remote code execution)","year":"2025","cvss_score":10,"severity":"critical","kev":true,"impact":"Unauthenticated remote code execution on OneView with scope change - a perfect-10 finding. OneView is HPE's fleet management plane: it holds iLO credentials, drives firmware deployment and owns server profiles, so RCE there is effectively root on every managed server. A public Metasploit module exists.","attack_vector":"Anyone who can reach the OneView web interface. No credentials.","remediation":"Patch OneView immediately per HPESBGN04985 - this is the single highest-priority item in this sweep. Appliance update with a service restart. Assume compromise if OneView has been network-reachable and unpatched: rotate every iLO and service-account credential it holds, and review deployed firmware for tampering.","references":["https://support.hpe.com/hpesc/public/docDisplay?docId=hpesbgn04985en_us&docLocale=en_US"],"status":"curated","fleet":{"pain_class":"daemon-restart"}},{"id":"CVE-2026-10520","cve":"CVE-2026-10520","aliases":[],"title":"Ivanti Sentry: OS command injection","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ivanti Sentry","year":"2026","cvss_score":10,"severity":"critical","kev":true,"impact":"OS command injection -> unauthenticated root-level remote code execution","attack_vector":"Network (remote)","remediation":"Control-plane: patch to R10.5.2/R10.6.2/R10.7.1; rebuild the appliance if it was exposed","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-10520"],"status":"curated","published":"2026-06-09"},{"id":"CVE-2026-45829","cve":"CVE-2026-45829","aliases":[],"title":"ChromaDB: Pre-authentication code injection","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"ChromaDB","year":"2026","cvss_score":10,"severity":"critical","kev":false,"impact":"Pre-authentication code injection → arbitrary code execution","attack_vector":"Unauthenticated network to the Chroma server","remediation":"Upgrade. Maximum severity, no auth required — any tenant-reachable Chroma is fully compromised","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-45829"],"status":"curated","published":"2026-05-18"},{"id":"CVE-2026-74280","cve":"CVE-2026-74280","aliases":[],"title":"Linux crypto driver for Marvell OCTEON TX: The scatter-gather cleanup path in the Marvell OCTEON TX crypto driver uses","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux crypto driver for Marvell OCTEON TX","year":"2026","cvss_score":10,"severity":"critical","kev":false,"impact":"The scatter-gather cleanup path in the Marvell OCTEON TX crypto driver uses the wrong loop index, so DMA mappings are torn down against the wrong entries. Getting DMA cleanup wrong on an accelerator is exactly the primitive that undermines IOMMU-based device isolation — the mapping that should have been revoked stays live, or a mapping belonging to something else is revoked. OCTEON is used both as a DPU and as an inline crypto/offload engine in storage and network appliances.","attack_vector":"Local, through the crypto API on a host with an OCTEON TX accelerator.","remediation":"Kernel/driver upgrade plus host reboot. Nothing to flash. If you run OCTEON accelerators in a shared-tenancy role, verify IOMMU is enabled and enforcing on those devices as a standing control.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-74280"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-15"},{"id":"CVE-2026-74475","cve":"CVE-2026-74475","aliases":[],"title":"Linux VXLAN driver (neighbour hardware address read in route_shortcircuit): `route_shortcircuit()` reads a neighbour's","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux VXLAN driver (neighbour hardware address read in route_shortcircuit)","year":"2026","cvss_score":10,"severity":"critical","kev":false,"impact":"`route_shortcircuit()` reads a neighbour's hardware address without taking the seqlock that protects it, so it can observe a torn or partially-updated MAC address while the neighbour subsystem is rewriting it. In an overlay, the destination MAC is what decides which VTEP — and therefore which tenant's segment — a frame is delivered to. A torn read there is not just a memory-safety problem; it is a frame going somewhere the forwarding logic did not intend.","attack_vector":"Concurrent neighbour updates alongside VXLAN forwarding on an affected host or software VTEP. Triggerable by ordinary overlay traffic combined with neighbour churn, which tenants generate routinely.","remediation":"Kernel upgrade plus host reboot, or NOS image upgrade plus switch reload on Linux-based switches. Bundle with the rest of the 2026 VXLAN batch.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-74475"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-08-15"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-284","CWE-918"],"fleet":{"pain_class":"unpatchable / mitigate-only"},"id":"NCVD-2026-042-kubeflow-pipelines-frontend-prox","cve":null,"aliases":["GHSA-gqww-5pj5-8fq7","CVE-2026-54745 (reserved)"],"title":"Kubeflow Pipelines frontend (/_proxy/ route, proxy-middleware.ts): The pipelines frontend hands any unauthenticated","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Kubeflow Pipelines frontend (/_proxy/ route, proxy-middleware.ts)","year":"2026","cvss_score":10,"severity":"critical","kev":false,"impact":"The pipelines frontend hands any unauthenticated caller a general-purpose HTTP client living inside the cluster network. The /_proxy/ route takes a user-supplied URL and forwards method, headers and body verbatim, with no allowlist, no IP filter and no auth gate — loopback, RFC1918, link-local 169.254.0.0/16 and *.svc.cluster.local are all in reach. Concretely an attacker pulls cloud IAM credentials from the instance metadata service, replays stolen service-account or JWT tokens at the Kubernetes API through a forwarded Authorization header, reaches internal services never meant to be exposed (ml-pipeline API on 8888, katib-db-manager on 6789, etcd, kubelet), and spoofs X-Forwarded-For past IP allowlists on internal tooling. Because arbitrary POST bodies and headers pass through, this is not read-only SSRF: it is a write primitive against internal admin endpoints. The sharpest part for an operator is that the documented hardening posture does not help — the auth middleware never gets applied to this route, so ENABLE_AUTHZ=true multi-user mode, the configuration sold as the secure one, is bypassed while every other endpoint on the same instance correctly rejects the same unauthenticated request.","attack_vector":"Network, fully unauthenticated. Anyone who can open a TCP connection to the frontend port qualifies — via Ingress in standalone Kubeflow Pipelines deployments, or any pod in the cluster in Kubeflow Platform installs. A Referer header alone (Referer: http://x/apis/v2beta1/_proxy/http://target/) triggers the rewrite on unrelated paths, so path filtering at a reverse proxy is not sufficient.","remediation":"There is no fixed release as of the advisory (2.16.0 and earlier affected). Gate the route at the edge and inside the cluster: block /apis/v1beta1/_proxy/, /apis/v2beta1/_proxy/ and their /pipeline/-prefixed variants at the ingress, and drop requests carrying a Referer matching the same pattern. Enforce a NetworkPolicy that denies the frontend pod egress to 169.254.169.254 and to cluster services it does not need, and require IMDSv2 so a bare GET cannot lift credentials. Cut the frontend pod's ServiceAccount RBAC to the minimum. Treat the proxy as removable if your deployment does not depend on it.","references":["https://github.com/kubeflow/pipelines/security/advisories/GHSA-gqww-5pj5-8fq7"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2021-21872","cve":"CVE-2021-21872","aliases":["TALOS-2021-1312"],"title":"Lantronix PremierWave 2050 console server (Web Manager): An attacker who can log into the web management console gets","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lantronix PremierWave 2050 console server (Web Manager)","year":"2021","cvss_score":9.9,"severity":"critical","kev":false,"impact":"An attacker who can log into the web management console gets a shell on the console server itself, running with the same privileges as the web daemon. From there they can pivot to every serial-attached device the box terminates, and use the console server as a jump box into the rest of the OOB network.","attack_vector":"Needs an authenticated session to the Web Manager (any role), then sends a crafted HTTP request to the Diagnostics: Traceroute page. The traceroute target field isn't sanitized before being handed to a shell, so shell metacharacters turn it into arbitrary command execution.","remediation":"Firmware upgrade required (Lantronix has patched builds past 8.9.0.0R4) plus a reboot of each unit; no config workaround exists since the flaw is in the diagnostics handler itself. Roll out per-device, one console server at a time — each reboot drops active serial sessions to whatever racks it terminates.","references":["https://talosintelligence.com/vulnerability_reports/TALOS-2021-1312"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2021-12-22"},{"id":"CVE-2021-21883","cve":"CVE-2021-21883","aliases":["TALOS-2021-1327"],"title":"Lantronix PremierWave 2050 console server (Web Manager): Same class of bug as the Traceroute injection on this device","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lantronix PremierWave 2050 console server (Web Manager)","year":"2021","cvss_score":9.9,"severity":"critical","kev":false,"impact":"Same class of bug as the Traceroute injection on this device, reached through the Ping diagnostic instead — full command execution on the console server, with downstream reach into every serial line it terminates.","attack_vector":"Any authenticated Web Manager user submits a crafted host value to the Diagnostics: Ping function; the value flows unsanitized into a shell command.","remediation":"Firmware upgrade past 8.9.0.0R4 and reboot. Same rollout cost as the Traceroute bug — treat both as one patch cycle per unit rather than two separate maintenance windows.","references":["https://talosintelligence.com/vulnerability_reports/TALOS-2021-1327"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2021-12-22"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-22"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2021-25311","cve":"CVE-2021-25311","aliases":["HTCONDOR-2021-0002"],"title":"HTCondor (condor_credd): condor_credd can be told to create or write files as root outside","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HTCondor (condor_credd)","year":"2021","cvss_score":9.9,"severity":"critical","kev":false,"impact":"condor_credd can be told to create or write files as root outside SEC_CREDENTIAL_DIRECTORY_OAUTH. The advisory's own example is planting a file under /etc that later gets executed - so this is a path-traversal-to-root on the credential daemon's host, which is normally the access point of the pool.","attack_vector":"An authenticated pool user who can talk to a running condor_credd.","remediation":"Upgrade to HTCondor 8.9.11 or later and restart condor_credd. If you are not using OAuth credential handling, do not run credd at all. After patching, check /etc and the systemd unit directories on credd hosts for files you did not put there.","references":["https://htcondor.org/security/vulnerabilities/HTCONDOR-2021-0002.html","https://nvd.nist.gov/vuln/detail/CVE-2021-25311"],"status":"curated"},{"id":"CVE-2021-28476","cve":"CVE-2021-28476","aliases":[],"title":"Microsoft Hyper-V: vmswitch fails to validate guest OID requests","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Microsoft Hyper-V","year":"2021","cvss_score":9.9,"severity":"critical","kev":false,"impact":"vmswitch fails to validate guest OID requests - guest reads arbitrary host kernel memory or crashes the host; the highest-rated Hyper-V escape class","attack_vector":"Tenant VM guest","remediation":"Windows update + host reboot with live-migration","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-28476"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-05-11"},{"id":"CVE-2021-36782","cve":"CVE-2021-36782","aliases":[],"title":"Rancher: Cluster owners, members and even base users retrieve plaintext credentials via the Kubernetes API","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Rancher","year":"2021","cvss_score":9.9,"severity":"critical","kev":false,"impact":"Cluster owners, members and even base users retrieve plaintext credentials via the Kubernetes API","attack_vector":"Any authenticated Rancher user","remediation":"Upgrade Rancher; rotate every credential Rancher stores, including cloud and registry credentials","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-36782"],"status":"curated","published":"2022-09-07"},{"id":"CVE-2021-36783","cve":"CVE-2021-36783","aliases":[],"title":"Rancher: Insufficiently protected credentials let project members read passwords and API tokens","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Rancher","year":"2021","cvss_score":9.9,"severity":"critical","kev":false,"impact":"Insufficiently protected credentials let project members read passwords and API tokens","attack_vector":"Any authenticated Rancher user","remediation":"Upgrade Rancher; rotate all stored credentials","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-36783"],"status":"curated","published":"2022-09-07"},{"id":"CVE-2022-24768","cve":"CVE-2022-24768","aliases":[],"title":"Argo CD: Improper access control allows a low-privileged user to escalate to Argo CD admin","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2022","cvss_score":9.9,"severity":"critical","kev":false,"impact":"Improper access control allows a low-privileged user to escalate to Argo CD admin","attack_vector":"Any authenticated Argo CD user","remediation":"Rolling Argo CD upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-24768"],"status":"curated","published":"2022-03-23"},{"id":"CVE-2022-43757","cve":"CVE-2022-43757","aliases":[],"title":"Rancher: Cleartext credential storage lets managed-cluster users read credentials","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Rancher","year":"2022","cvss_score":9.9,"severity":"critical","kev":false,"impact":"Cleartext credential storage lets managed-cluster users read credentials","attack_vector":"Any user on a managed cluster","remediation":"Upgrade Rancher; rotate all credentials","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-43757"],"status":"curated","published":"2023-02-07"},{"id":"CVE-2023-22647","cve":"CVE-2023-22647","aliases":[],"title":"Rancher: Standard users manipulate Kubernetes secrets in the local (management) cluster","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Rancher","year":"2023","cvss_score":9.9,"severity":"critical","kev":false,"impact":"Standard users manipulate Kubernetes secrets in the local (management) cluster","attack_vector":"Any authenticated Rancher user","remediation":"Upgrade Rancher; audit local-cluster secrets","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-22647"],"status":"curated","published":"2023-06-01"},{"id":"CVE-2023-22651","cve":"CVE-2023-22651","aliases":[],"title":"Rancher: Update-logic failure misconfigures Rancher's admission webhook, disabling the validation that enforces","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Rancher","year":"2023","cvss_score":9.9,"severity":"critical","kev":false,"impact":"Update-logic failure misconfigures Rancher's admission webhook, disabling the validation that enforces tenant boundaries","attack_vector":"Any authenticated Rancher user","remediation":"Upgrade Rancher; verify the webhook configuration explicitly after every upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-22651"],"status":"curated","published":"2023-05-04"},{"id":"CVE-2023-32191","cve":"CVE-2023-32191","aliases":[],"title":"RKE / Rancher (k8s control plane): full-cluster-state configmap in kube-system readable by non-admins","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"RKE / Rancher (k8s control plane)","year":"2023","cvss_score":9.9,"severity":"critical","kev":false,"impact":"full-cluster-state configmap in kube-system readable by non-admins -> escalate to cluster-admin","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade RKE; restrict RBAC on kube-system configmaps","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-32191"],"status":"curated","published":"2024-10-16"},{"id":"CVE-2023-34063","cve":"CVE-2023-34063","aliases":[],"title":"VMware Aria Automation (missing access control): An authenticated user reaches remote organizations and workflows they","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"VMware Aria Automation (missing access control)","year":"2023","cvss_score":9.9,"severity":"critical","kev":false,"impact":"An authenticated user reaches remote organizations and workflows they should not see - a tenancy-boundary failure inside the automation platform, with scope change.","attack_vector":"Any authenticated Aria Automation user.","remediation":"Apply the fix per VMSA-2024-0001. Appliance patch; review org/workflow assignments afterwards for signs of cross-org access.","references":["https://www.vmware.com/security/advisories/VMSA-2024-0001.html"],"status":"curated"},{"id":"CVE-2023-40029","cve":"CVE-2023-40029","aliases":[],"title":"Argo CD: Cluster secrets stored in the last-applied-configuration annotation are readable by anyone with get access","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2023","cvss_score":9.9,"severity":"critical","kev":false,"impact":"Cluster secrets stored in the last-applied-configuration annotation are readable by anyone with get access on the secret","attack_vector":"Cluster user with namespace access to argocd","remediation":"Rolling Argo CD upgrade; rotate every cluster credential Argo CD holds","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-40029"],"status":"curated","published":"2023-09-07"},{"id":"CVE-2024-20432","cve":"CVE-2024-20432","aliases":[],"title":"Cisco Nexus Dashboard Fabric Controller (REST API / web UI): A low-privileged NDFC user","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Cisco Nexus Dashboard Fabric Controller (REST API / web UI)","year":"2024","cvss_score":9.9,"severity":"critical","kev":false,"impact":"A low-privileged NDFC user — the sort of read-mostly account you hand to an NOC or a tenant liaison — gets command injection on the fabric controller. NDFC holds the credentials for and pushes config to every switch it manages, so this is a straight path from a minor account to control of the whole leaf/spine build.","attack_vector":"Authenticated but low-privileged, remote. Any valid NDFC login is enough.","remediation":"Upgrade NDFC. Controller-side software upgrade, data plane unaffected. Afterwards rotate the device credentials NDFC stores, because those are what an attacker would have taken.","references":["https://sec.cloudapps.cisco.com/security/center/content/CiscoSecurityAdvisory/cisco-sa-ndfc-raci-T46k3jnN","https://nvd.nist.gov/vuln/detail/CVE-2024-20432"],"status":"curated","published":"2024-10-02"},{"id":"CVE-2024-24594","cve":"CVE-2024-24594","aliases":[],"title":"ClearML web server: XSS","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"ClearML web server","year":"2024","cvss_score":9.9,"severity":"critical","kev":false,"impact":"XSS → code execution in the operator's session","attack_vector":"Attacker-controlled experiment metadata rendered in the UI","remediation":"Upgrade; tenant-supplied experiment names become operator-plane payloads","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-24594"],"status":"curated","published":"2024-02-06"},{"id":"CVE-2024-42327","cve":"CVE-2024-42327","aliases":[],"title":"Zabbix: SQL injection in CUser::addRelatedObjects reachable by ANY non-admin account with API access","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Zabbix","year":"2024","cvss_score":9.9,"severity":"critical","kev":false,"impact":"SQL injection in CUser::addRelatedObjects reachable by ANY non-admin account with API access -> full DB compromise","attack_vector":"Network (remote)","remediation":"Control-plane: URGENT upgrade; Zabbix commonly stores datacenter/IPMI credentials","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-42327"],"status":"curated","published":"2024-11-27"},{"id":"CVE-2024-9264","cve":"CVE-2024-9264","aliases":[],"title":"Grafana: SQL Expressions passes user input to duckdb unsanitized","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Grafana","year":"2024","cvss_score":9.9,"severity":"critical","kev":false,"impact":"SQL Expressions passes user input to duckdb unsanitized -> command injection and local file inclusion (VIEWER+)","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; remove the duckdb binary from the Grafana image","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-9264"],"status":"curated","published":"2024-10-18"},{"id":"CVE-2025-10725","cve":"CVE-2025-10725","aliases":[],"title":"Red Hat OpenShift AI (notebook plane): A low-privileged data-scientist account can escalate to full cluster compromise","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Red Hat OpenShift AI (notebook plane)","year":"2025","cvss_score":9.9,"severity":"critical","kev":false,"impact":"A low-privileged data-scientist account can escalate to full cluster compromise","attack_vector":"Notebook user inside the managed AI platform","remediation":"Patch the platform. Direct tenant→provider escalation on a managed GPU platform","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-10725"],"status":"curated","published":"2025-09-30"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-269"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2025-32445","cve":"CVE-2025-32445","aliases":["GHSA-hmp7-x699-cvhq"],"title":"Argo Events (EventSource / Sensor controller, spec.template.container merge): The controller merges the entire","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo Events (EventSource / Sensor controller, spec.template.container merge)","year":"2025","cvss_score":9.9,"severity":"critical","kev":false,"impact":"The controller merges the entire user-supplied container spec from an EventSource or Sensor custom resource into the pod it creates, so a tenant who can only create those CRs sets securityContext.privileged, adds SYS_ADMIN, and hostPath-mounts the node root. That is node compromise and, through the container runtime socket, cluster compromise - from a namespace-scoped permission that looks harmless. Argo Events is the trigger layer in front of Argo Workflows, so this sits on the same GPU job-submission path.","attack_vector":"Any tenant with create or update rights on EventSource or Sensor CRs in any namespace the controller watches. No cluster-admin, no node access, no direct pod-create permission needed - Pod Security Standards and namespace RBAC are both bypassed.","remediation":"Upgrade Argo Events to v1.9.6, which allow-lists which properties under spec.template.container may be set. Restart the controller. Before and after, audit existing EventSource and Sensor objects for privileged securityContext or hostPath volumes - a resource planted earlier keeps running until you delete it.","references":["https://github.com/argoproj/argo-events/security/advisories/GHSA-hmp7-x699-cvhq","https://nvd.nist.gov/vuln/detail/CVE-2025-32445"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-49844","cve":"CVE-2025-49844","aliases":[],"title":"Redis: \"RediShell\" - authenticated user crafts a Lua script to trigger a use-after-free","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Redis","year":"2025","cvss_score":9.9,"severity":"critical","kev":false,"impact":"\"RediShell\" - authenticated user crafts a Lua script to trigger a use-after-free -> RCE on the host","attack_vector":"Network (remote)","remediation":"Control-plane: URGENT upgrade of every Redis (scheduler/queue/session); restrict EVAL via ACL","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-49844"],"status":"curated","published":"2025-10-03"},{"id":"CVE-2025-54381","cve":"CVE-2025-54381","aliases":[],"title":"BentoML (file upload): SSRF in the file-upload path","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"BentoML (file upload)","year":"2025","cvss_score":9.9,"severity":"critical","kev":false,"impact":"SSRF in the file-upload path","attack_vector":"Authenticated or unauthenticated request depending on deployment","remediation":"Upgrade to 1.4.19+","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-54381"],"status":"curated","published":"2025-07-29"},{"id":"CVE-2026-7374","cve":"CVE-2026-7374","aliases":[],"title":"KubeVirt: Improper symlink validation in virt-handler lets a user with edit rights in one namespace escape to the host","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"KubeVirt","year":"2026","cvss_score":9.9,"severity":"critical","kev":false,"impact":"Improper symlink validation in virt-handler lets a user with edit rights in one namespace escape to the host; the most severe KubeVirt issue","attack_vector":"Cluster user with namespace access","remediation":"Emergency KubeVirt upgrade; virt-handler DaemonSet rollout on every node","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-7374"],"status":"curated","published":"2026-05-26"},{"id":"CVE-2015-8011","cve":"CVE-2015-8011","aliases":[],"title":"lldpd (lldp_decode, management addresses): Buffer overflow in lldpd's LLDP decoder via large management addresses","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"lldpd (lldp_decode, management addresses)","year":"2015","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Buffer overflow in lldpd's LLDP decoder via large management addresses and TLV boundaries, allowing daemon crash and possibly code execution. Old, but included because lldpd is one of those daemons that ships inside embedded switch and appliance images and stays frozen at whatever version the vendor picked years ago — the CVE date tells you nothing about whether your fabric is running it.","attack_vector":"Unauthenticated, adjacent — a crafted LLDP frame.","remediation":"Upgrade lldpd past 0.8.0 and restart. On embedded NOSes and appliances, check the shipped lldpd version explicitly rather than assuming a modern image implies a modern lldpd. Companion crash issue: CVE-2015-8012.","references":["https://nvd.nist.gov/vuln/detail/CVE-2015-8011","https://nvd.nist.gov/vuln/detail/CVE-2015-8012"],"status":"curated","published":"2020-01-28"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2015-8812","cve":"CVE-2015-8812","aliases":[],"title":"Linux kernel iWARP driver drivers/infiniband/hw/cxgb3/iwch_cm.c (Chelsio T3): Unauthenticated remote code execution in","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel iWARP driver drivers/infiniband/hw/cxgb3/iwch_cm.c (Chelsio T3)","year":"2015","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unauthenticated remote code execution in the kernel, reachable from the fabric. The iWARP connection-manager path mis-identifies an error condition and frees a buffer it then keeps using; a crafted packet to the listening RDMA endpoint drives the use-after-free to arbitrary kernel code execution. RDMA connection managers listen before any application-level authentication exists - there is no credential to present - so anything that can put packets on the storage/compute fabric owns the kernel. In a cluster where tenant traffic and RDMA control traffic share the same L2 domain, one compromised tenant VM reaches every node running this HCA.","attack_vector":"Network, pre-auth. Any host able to send packets to the node's iWARP/RDMA-CM listener - which on a flat cluster fabric includes every other tenant. No credentials, no prior foothold on the target.","remediation":"Kernel upgrade to 4.5+ or a vendor backport of commit 67f1aee6f45059fd6b0f5b0ecb2c97ad0451f6b3 (iw_cxgb3: fix incorrectly returning error on success). Rolling reboot required. Compensating control while you schedule it: put the RDMA fabric on its own isolated L2/VLAN with no tenant-reachable path, and firewall the iWARP listener - RDMA-CM has no authentication of its own, so network segmentation is the only pre-patch boundary. The iw_cxgb3 driver was removed from mainline entirely in later kernels; if you still have Chelsio T3 parts racked, that hardware is past end of support and should be scheduled out.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=67f1aee6f45059fd6b0f5b0ecb2c97ad0451f6b3","https://access.redhat.com/security/cve/CVE-2015-8812","https://www.openwall.com/lists/oss-security/2016/02/03/9"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2016-4325","cve":"CVE-2016-4325","aliases":["VU#785823"],"title":"Lantronix xPrintServer: The device ships with a hardcoded root account baked into every unit of a given firmware line","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Lantronix xPrintServer","year":"2016","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The device ships with a hardcoded root account baked into every unit of a given firmware line. Anyone who can reach the management interface gets root on the box without needing to guess or phish a credential — it's the same key for every deployed unit.","attack_vector":"No authentication needed beyond network reachability to the device; the credential is embedded in firmware and identical across all units running the affected build.","remediation":"Firmware flash to 5.0.1-65 or later on every affected unit — this isn't something a config change or password rotation fixes, since the account is compiled into the image. Budget one flash-and-reboot cycle per device; xPrintServer's main job (serial/print bridging) is unavailable during the flash.","references":["http://www.kb.cert.org/vuls/id/785823"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2016-05-14"},{"id":"CVE-2016-4375","cve":"CVE-2016-4375","aliases":[],"title":"HPE iLO3 / iLO4: Multiple unspecified flaws allowing remote information disclosure, data modification and DoS","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HPE iLO3 / iLO4","year":"2016","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Multiple unspecified flaws allowing remote information disclosure, data modification and DoS on the management controller","attack_vector":"Network","remediation":"iLO firmware update (iLO3 <1.88, iLO4 <2.44); the \"unspecified\" advisory style means operators cannot risk-assess individual issues and must patch blind","references":["https://nvd.nist.gov/vuln/detail/CVE-2016-4375"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2016-09-08"},{"id":"CVE-2017-16748","cve":"CVE-2017-16748","aliases":["ICSA-18-191-03","ICSA-19-022-01"],"title":"Tridium Niagara AX (<=3.8) and Niagara 4 (<=4.4) framework: Log into the Niagara platform with a disabled account name","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Tridium Niagara AX (<=3.8) and Niagara 4 (<=4.4) framework","year":"2017","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Log into the Niagara platform with a disabled account name and a blank password and you get administrator. Niagara is the integration layer a very large share of datacenters use to tie CRAC/CRAH, chillers, generators, metering and sometimes access control into one supervisory system, so administrator on the Niagara station is administrator over the whole facility control surface at once. That means arbitrary writes to cooling setpoints and fan commands across every integrated subsystem, the ability to disable alarms and schedules, and access to the credential store Niagara keeps for its downstream drivers - which is how a BMS compromise turns into compromise of the chiller, the ATS controller and the metering network in one step. For a GPU operator the headline is simple: one blank password and the hall's thermal envelope is under attacker control, with minutes of margin before accelerators shut down.","attack_vector":"Unauthenticated login against the Niagara station's web or Fox interface on the facility network. Niagara stations are one of the most consistently internet-exposed classes of building controller in existence - integrators publish them for remote support and forget - so treat internet exposure as likely rather than exceptional until you have checked your own external attack surface for the Niagara Fox port and the station web UI.","remediation":"Patchable via a Niagara framework upgrade (AX 3.8U1 / Niagara 4.4U1 or later), performed by the systems integrator who owns the station. This is a software upgrade on the supervisor plus a JACE controller update, so it needs a maintenance window and a contractor but not a cooling outage. Do it, then audit the account list for disabled-but-present accounts, and pull the station off any internet-facing interface. Because Niagara stations are so often integrator-managed rather than operator-managed, the harder task is organisational: find out who actually holds the platform credentials for your station, whether the integrator has a permanent remote path in, and whether that path is MFA'd. In a leased colo the station belongs to the landlord and typically serves the entire building - demand the framework version and the remote-access architecture in writing.","references":["https://www.cisa.gov/news-events/ics-advisories/icsa-18-191-03","https://nvd.nist.gov/vuln/detail/CVE-2017-16748"],"status":"curated"},{"cvss_vector":"CVSS:3.0/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-121","CWE-119"],"fleet":{"pain_class":"firmware-flash"},"id":"CVE-2017-3774","cve":"CVE-2017-3774","aliases":["LEN-19586"],"title":"Lenovo / IBM Integrated Management Module 2 (IMM2) web administration service: The overflow is inside the","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Lenovo / IBM Integrated Management Module 2 (IMM2) web administration service","year":"2017","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The overflow is inside the authentication routine itself, so a crafted user ID and password pair corrupts the BMC's stack before any credential decision is made - no account needed. Successful exploitation is control of the management controller: power state, virtual media, host firmware, serial console. The pattern is worth naming for anyone building a BMC risk model, because it recurs across vendors and a decade: the code that parses the login attempt is the least-privileged-input, highest-privilege-context code on the device, and it is repeatedly written in C against fixed buffers. Same shape as the Supermicro login.cgi overflow four years earlier.","attack_vector":"Network, pre-auth. Any reachability to the IMM2 web administration service.","remediation":"Flash IMM2 firmware to 4.70+ (Lenovo-branded servers) or 6.60+ (IBM-branded). Out-of-band update, node drain not strictly required but advisable since the IMM restarts. As with every pre-auth BMC bug in this catalogue, the flash is the fix and network isolation is the control that makes the flash schedulable rather than an emergency: if the BMC is only reachable from a bastion, an unpatched pre-auth overflow is a risk you can plan around; if it is reachable from a tenant VLAN or the internet, it is not.","references":["https://support.lenovo.com/us/en/product_security/LEN-19586","https://www.cve.org/CVERecord?id=CVE-2017-3774","https://nvd.nist.gov/vuln/detail/CVE-2017-3774"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2017-5689","cve":"CVE-2017-5689","aliases":["Silent Bob is Silent"],"title":"Intel Active Management Technology / Standard Manageability: An authentication bypass in the AMT web interface: sending","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel Active Management Technology / Standard Manageability","year":"2017","cvss_score":9.8,"severity":"critical","kev":true,"impact":"An authentication bypass in the AMT web interface: sending an empty response hash is accepted as valid, so an unauthenticated network attacker gets full AMT administrative control. AMT provides out-of-band power control, KVM and virtual media below the OS, so this is total control of the machine from the management network, invisible to anything running on the host. Listed in CISA's Known Exploited Vulnerabilities catalog.","attack_vector":"Any attacker who can reach the AMT ports (16992/16993/16994/16995, and 623/664) on a provisioned machine. If your management network is flat or reachable from tenant VLANs, that is everyone.","remediation":"Update Intel CSME/AMT firmware via the OEM. Where firmware is unavailable, unprovision AMT and block the AMT ports at the network layer - that is the mitigation that actually deploys on the same day. Firmware update requires an OEM package, a drain and a reboot. Any machine exposed while vulnerable should be treated as compromised at the firmware level, since AMT access permits persistent implantation below the OS.","references":["https://nvd.nist.gov/vuln/detail/CVE-2017-5689","https://downloadmirror.intel.com/26754/eng/INTEL-SA-00075%20Mitigation%20Guide-Rev%201.1.pdf"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2017-05-02"},{"id":"CVE-2017-8979","cve":"CVE-2017-8979","aliases":[],"title":"HPE iLO2: Authentication bypass and code execution in iLO2 firmware 2.29","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HPE iLO2","year":"2017","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Authentication bypass and code execution in iLO2 firmware 2.29","attack_vector":"Network, unauthenticated","remediation":"iLO2 is EOL — remediation is decommissioning or hard network isolation, not patching","references":["https://nvd.nist.gov/vuln/detail/CVE-2017-8979"],"status":"curated","published":"2018-02-15"},{"id":"CVE-2018-0314","cve":"CVE-2018-0314","aliases":[],"title":"Cisco NX-OS / FXOS (Cisco Fabric Services): Unauthenticated remote code execution as root through Cisco Fabric","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco NX-OS / FXOS (Cisco Fabric Services)","year":"2018","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unauthenticated remote code execution as root through Cisco Fabric Services, the inter-switch distribution protocol. CFS is on by default on many platforms and speaks between switches, so a single compromised switch or a host that can spoof CFS reaches every other switch that trusts the same distribution domain. Part of a 2018 cluster (CVE-2018-0304/0308/0310/0312/0314) that shares this exposure.","attack_vector":"Unauthenticated, remote — the attacker must be able to send CFS messages, which in an unsegmented fabric means any host on a VLAN where CFS-over-IP is enabled.","remediation":"NX-OS upgrade plus reload. Immediately: disable CFS distribution and CFS-over-IP where you do not use it (`no cfs distribute`, `no cfs ipv4 distribute`) — live config, no reload, and it closes the whole 2018 family at once.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-0314","https://nvd.nist.gov/vuln/detail/CVE-2018-0304","https://nvd.nist.gov/vuln/detail/CVE-2022-20624"],"status":"curated","published":"2018-06-20"},{"id":"CVE-2018-12031","cve":"CVE-2018-12031","aliases":[],"title":"Eaton Intelligent Power Manager v1.6 - node_upgrade_srv.js firmware parameter: Local file inclusion through directory","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Eaton Intelligent Power Manager v1.6 - node_upgrade_srv.js firmware parameter","year":"2018","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Local file inclusion through directory traversal on the firmware parameter of the node upgrade service. The affected code path is the one that distributes firmware to managed power devices, so this is an attacker standing in the middle of your UPS firmware supply chain.","attack_vector":"Remote to the IPM server, unauthenticated per the advisory.","remediation":"Upgrade IPM well past 1.6 - anything on 1.6 is also carrying the 2020 and 2021 unauthenticated RCEs. Gate firmware distribution to power devices behind change control regardless of the platform you use.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-12031"],"status":"curated","published":"2018-06-07"},{"id":"CVE-2018-1207","cve":"CVE-2018-1207","aliases":[],"title":"Dell iDRAC7/8: CGI injection giving unauthenticated remote code execution as root on the BMC","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC7/8","year":"2018","cvss_score":9.8,"severity":"critical","kev":false,"impact":"CGI injection giving unauthenticated remote code execution as root on the BMC","attack_vector":"Network, unauthenticated","remediation":"Firmware update to 2.52.52.52+; nodes still on older iDRAC7/8 firmware are the most common legacy hole in a mixed-vintage GPU fleet","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-1207"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2018-03-23"},{"id":"CVE-2018-12171","cve":"CVE-2018-12171","aliases":["INTEL-SA-00149"],"title":"Intel Baseboard Management Controller firmware before 1.43.91f76955 (Intel server boards and systems): An unprivileged","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Baseboard Management Controller firmware before 1.43.91f76955 (Intel server boards and systems)","year":"2018","cvss_score":9.8,"severity":"critical","kev":false,"impact":"An unprivileged user can execute arbitrary code or force denial of service on the BMC. Owning the BMC is owning the node: virtual media boot, host power control, KVM, SOL, firmware update paths for BIOS and ME, and the SMBus. It is below-the-OS persistence in the strongest sense - the BMC has its own flash and its own OS, so nothing you do to the host image removes an implant there, and the node carries it into the next tenant. It also gives an attacker a fleet-wide physical-consequence lever: mass power-off of every node they can reach on the management network.","attack_vector":"Network access to the BMC without valid credentials. Anyone who can route to the out-of-band management network - which in practice includes anything that reaches an internet-exposed or flat-VLAN BMC, and any tenant if the OOB network is not fully separated.","remediation":"BMC firmware update to 1.43.91f76955 or later, from Intel or the board ODM (Quanta, Wiwynn, Supermicro on Intel reference designs). BMC updates do not need a host reboot, so this is one you can roll without draining jobs - do it fleet-wide. In parallel: never expose BMCs to the internet, put them behind a jump host on a dedicated VRF, rotate to unique per-node credentials, and disable the host-side KCS interface.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-12171","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00149.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2018-12327","cve":"CVE-2018-12327","aliases":[],"title":"ntpq / ntpdc (NTP 4.2.8p11 client utilities): Stack buffer overflow in the ntpq and ntpdc command-line tools via a long","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"ntpq / ntpdc (NTP 4.2.8p11 client utilities)","year":"2018","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Stack buffer overflow in the ntpq and ntpdc command-line tools via a long argument, giving code execution or privilege escalation. The interesting case for a cluster operator is automation: monitoring scripts that shell out to ntpq with a hostname taken from inventory turn an inventory-poisoning bug into code execution on the monitoring host.","attack_vector":"Local, via a long argument to ntpq/ntpdc — reachable wherever these tools are invoked with externally influenced arguments.","remediation":"Upgrade the ntp package. No service restart needed for the client tools; nothing to reboot. Audit any monitoring or automation that passes untrusted strings to ntpq.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-12327"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2018-06-20"},{"cvss_vector":"CVSS:3.0/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-77"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2018-14649","cve":"CVE-2018-14649","aliases":[],"title":"Ceph iSCSI gateway (ceph-iscsi-cli / rbd-target-api): rbd-target-api ships with the Werkzeug debug console enabled","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ceph iSCSI gateway (ceph-iscsi-cli / rbd-target-api)","year":"2018","cvss_score":9.8,"severity":"critical","kev":false,"impact":"rbd-target-api ships with the Werkzeug debug console enabled, which is an interactive Python shell exposed over HTTP with no authentication. Anyone who reaches the port executes arbitrary code as root on the iSCSI gateway node and from there controls the RBD images it serves.","attack_vector":"Any host with network reach to the rbd-target-api port on a Ceph iSCSI gateway. Fully pre-authentication.","remediation":"Upgrade ceph-iscsi-cli to the fixed package immediately and restart rbd-target-api. Treat any gateway that was network-reachable as compromised: rebuild it and rotate its CephX keys. Firewall the API to the management network only.","references":["https://access.redhat.com/security/cve/CVE-2018-14649","https://nvd.nist.gov/vuln/detail/CVE-2018-14649"],"status":"curated"},{"id":"CVE-2018-18202","cve":"CVE-2018-18202","aliases":[],"title":"QLogic 4Gb Fibre Channel 5.5.2.6.0 and 4/8Gb SAN 7.10.1.20.0 switch modules for IBM BladeCenter: Three undocumented","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"QLogic 4Gb Fibre Channel 5.5.2.6.0 and 4/8Gb SAN 7.10.1.20.0 switch modules for IBM BladeCenter","year":"2018","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Three undocumented accounts - support, diags and prom - each with a fixed password, baked into the FC switch module firmware. Anyone who knows them owns the switch module: zoning, port state, firmware. The reason this belongs in a modern GPU-datacenter database is not BladeCenter itself but the pattern - embedded FC switch modules and SAS/FC expander boards inherited with second-hand chassis carry vendor service accounts that no amount of operator password hygiene touches, and nothing in a normal build process ever looks for them.","attack_vector":"Anyone who can reach the module's management interface (telnet/SSH/web) on the chassis management network. Credentials are public.","remediation":"Unfixable by configuration - the accounts are in firmware. Either upgrade to a firmware release where they are removed, if one exists for your module, or retire the module. Practically, for inherited or second-hand chassis: treat every embedded switch/expander module as carrying vendor backdoor accounts until proven otherwise, keep chassis management on an isolated segment no tenant or workload VLAN can route to, and make 'scan for vendor service accounts' part of hardware intake rather than something you do after an incident.","references":["http://misteralfa-hack.blogspot.com/2018/10/ibm-bladecenter-qlogic-4g-fibre-channel.html","https://nvd.nist.gov/vuln/detail/CVE-2018-18202"],"status":"curated","published":"2018-10-10"},{"id":"CVE-2018-20687","cve":"CVE-2018-20687","aliases":[],"title":"Raritan CommandCenter Secure Gateway (CC-SG), before 8.0.0: CC-SG is Raritan's single-pane-of-glass gateway that","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Raritan CommandCenter Secure Gateway (CC-SG), before 8.0.0","year":"2018","cvss_score":9.8,"severity":"critical","kev":false,"impact":"CC-SG is Raritan's single-pane-of-glass gateway that brokers KVM and serial console access across a whole fleet of Dominion KX/SX switches. An unauthenticated attacker who reaches its web-services endpoint can send a crafted XML request with a malicious DTD, read arbitrary files off the gateway, or force it to make outbound requests (SSRF) into whatever internal network segments the gateway can reach — which by design is every rack it brokers console access to.","attack_vector":"Fully unauthenticated, remote — a crafted XML request to the CommandCenterWebServices endpoint is enough; no login required.","remediation":"Software upgrade to CC-SG 8.0.0 or later. This is a single appliance (or small HA pair) rather than a per-rack device, so the upgrade footprint is small, but treat it as urgent — the gateway sits in front of console/KVM access to the entire managed fleet, so an unauthenticated file-read/SSRF bug there is a direct path toward every rack it manages.","references":["http://packetstormsecurity.com/files/155359/Raritan-CommandCenter-Secure-Gateway-XML-Injection.html","http://seclists.org/fulldisclosure/2019/Nov/11"],"status":"curated","tags":["tenant-isolation"],"published":"2019-11-18"},{"cvss_vector":"CVSS:3.0/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-732"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2018-20871","cve":"CVE-2018-20871","aliases":["GE-6890"],"title":"Univa Grid Engine (execd spooling with Docker jobs on root_squash): In the specific combination of Docker-based jobs","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Univa Grid Engine (execd spooling with Docker jobs on root_squash)","year":"2018","cvss_score":9.8,"severity":"critical","kev":false,"impact":"In the specific combination of Docker-based jobs plus execd spooling on a root_squash NFS mount, Grid Engine creates spool files that are world-writable. The spool is what the execution daemon reads to decide what to run, so any tenant who can write to it controls job execution on that node.","attack_vector":"Any user who can reach the shared spool directory - on a root_squash NFS export, that is anyone with the mount, i.e. every compute node.","remediation":"Upgrade Univa Grid Engine to 8.6.3 or later (the release notes cover this through 8.6.6). Then audit the modes on the execd spool directory directly - the fix changes what new files get, not what existing files already have. Univa is now Altair, so support for this line runs through Altair.","references":["http://www.univa.com/resources/files/Release_Notes_Univa_Grid_Engine_8.6.6.pdf","https://nvd.nist.gov/vuln/detail/CVE-2018-20871"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.0/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-89"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2018-7033","cve":"CVE-2018-7033","aliases":[],"title":"Slurm (slurmdbd accounting database daemon): SQL injection into SlurmDBD gives an attacker read and write control of","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Slurm (slurmdbd accounting database daemon)","year":"2018","cvss_score":9.8,"severity":"critical","kev":false,"impact":"SQL injection into SlurmDBD gives an attacker read and write control of the cluster's accounting database - the record of which account owns which job, which associations exist, and what fairshare and QOS limits apply. Rewriting associations is how you grant yourself submission rights to another tenant's account, and the same database is what billing and chargeback are computed from.","attack_vector":"Anything that can send RPCs to the slurmdbd port. In most sites slurmdbd is reachable from the login nodes and from slurmctld, so a tenant with a shell on a login node is in position.","remediation":"Upgrade Slurm to 17.02.10 or 17.11.5 and restart slurmdbd. slurmdbd can be restarted independently of slurmctld and running jobs survive it, so this does not need a maintenance window. Audit the assoc and user tables afterwards - the injection leaves no distinctive log line.","references":["https://lists.schedmd.com/pipermail/slurm-announce/2018/000006.html","https://nvd.nist.gov/vuln/detail/CVE-2018-7033","https://www.debian.org/security/2018/dsa-4254"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2018-7243","cve":"CVE-2018-7243","aliases":["SEVD-2018-074-01"],"title":"Schneider Electric MGE Network Management Card Transverse (MGE UPS / MGE STS): The card's integrated web server","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Schneider Electric MGE Network Management Card Transverse (MGE UPS / MGE STS)","year":"2018","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The card's integrated web server has a broken authorization check, letting a remote attacker get full administrative access to the UPS/STS management interface without valid credentials — including whatever load-shedding and shutdown controls that UPS exposes.","attack_vector":"Remote, over the network, to the card's web server on port 80/443 — no valid credentials required to bypass the authorization check.","remediation":"Firmware flash of the Network Management Card required; Schneider's SEVD-2018-074-01 advisory has the fixed build. Roll out per card — each flash briefly drops remote monitoring/management of that UPS (the UPS itself keeps powering its load through the flash).","references":["https://www.schneider-electric.com/en/download/document/SEVD-2018-074-01/"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2018-04-18"},{"id":"CVE-2018-7246","cve":"CVE-2018-7246","aliases":["SEVD-2018-074-01"],"title":"Schneider Electric MGE Network Management Card Transverse (MGE UPS / MGE STS): On default settings without SSL enabled","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Schneider Electric MGE Network Management Card Transverse (MGE UPS / MGE STS)","year":"2018","cvss_score":9.8,"severity":"critical","kev":false,"impact":"On default settings without SSL enabled, repeatedly requesting the card's Access Control page leaks the administrative account credentials in plaintext to anyone who can sniff the traffic — handing over full UPS management access.","attack_vector":"Requires network position to observe traffic to/from the card's web server (or direct access to the unencrypted HTTP endpoint) while an admin session touches the Access Control page.","remediation":"Firmware flash to the fixed build, and as an immediate compensating step, force SSL/TLS on for the card's web interface rather than leaving it on plaintext HTTP. Same per-card rollout as the authorization-bypass companion CVE — do both in the same maintenance pass.","references":["https://www.schneider-electric.com/en/download/document/SEVD-2018-074-01/"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2018-04-18"},{"id":"CVE-2018-7820","cve":"CVE-2018-7820","aliases":[],"title":"APC UPS Network Management Card 2 (AOS 6.5.6): When Remote Monitoring is turned on and then off again, the credentials","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"APC UPS Network Management Card 2 (AOS 6.5.6)","year":"2018","cvss_score":9.8,"severity":"critical","kev":false,"impact":"When Remote Monitoring is turned on and then off again, the credentials used for remote monitoring stay viewable in plaintext on the card. Anyone who gets a look at the card's config (via the web UI, a config export, or a support dump) picks up a working credential to the UPS's remote-monitoring channel.","attack_vector":"Requires some access to the card's configuration or web interface (e.g. a lower-privileged account, an exported config file, or a support bundle) — not a fully unauthenticated remote exploit, but a credential-exposure path.","remediation":"Credential rotation for the affected remote-monitoring account is the immediate fix; pair it with the AOS firmware update from APC that stops persisting the credential in plaintext once monitoring is disabled. Rotate credentials across the whole NMC2 fleet, not just the units you know were exposed.","references":["https://www.apc.com/salestools/CCON-BFQMXC/CCON-BFQMXC_R0_EN.pdf"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2019-09-17"},{"id":"CVE-2018-9091","cve":"CVE-2018-9091","aliases":[],"title":"Kemp LoadMaster (LMOS): A flaw in session management lets a remote, unauthenticated attacker bypass the LoadMaster's","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Kemp LoadMaster (LMOS)","year":"2018","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A flaw in session management lets a remote, unauthenticated attacker bypass the LoadMaster's security protections entirely and run elevated shell commands (ls, ps, cat, etc.) — enough to pull certificates, private keys, and other sensitive data off the load balancer.","attack_vector":"Fully remote and unauthenticated against the LoadMaster's management interface.","remediation":"Software upgrade to LMOS 7.1.35.5 (LTS) or 7.2.41.2+ (mainline) per Kemp's mitigation article. Upgrade and reboot; if this LoadMaster fronts live inference traffic, fail over to a standby instance during the update. Also rotate any certificates/keys that were on the device, since the bug allowed reading them.","references":["https://support.kemptechnologies.com/hc/en-us/articles/360001982452-Mitigation-for-Remote-Access-Execution-Vulnerability"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2018-05-25"},{"id":"CVE-2019-0153","cve":"CVE-2019-0153","aliases":[],"title":"Intel CSME 12.0.0-12.0.34: A buffer overflow in a CSME subsystem reachable over the network by an unauthenticated","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel CSME 12.0.0-12.0.34","year":"2019","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A buffer overflow in a CSME subsystem reachable over the network by an unauthenticated attacker, giving privilege escalation on the management engine. CSME sits below the OS with its own network stack, so a network-reachable overflow there is control of the platform outside anything the host can observe or defend.","attack_vector":"Unauthenticated network access to the affected CSME service. Whether that is reachable depends entirely on how isolated your management network is.","remediation":"Fixed in Intel CSME/SPS firmware, which reaches you as an OEM BIOS or firmware package - not as a microcode or OS update. That means: wait for your server vendor to ship it, drain the node, flash, and reboot. OEM availability is the long pole and routinely lags the Intel advisory by one or more quarters on server platforms. Track it per platform SKU, because vendors ship these unevenly across their own product lines. Until firmware lands, isolate the management network - this class of bug is unreachable if the management plane is not routable from anywhere a tenant can be.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-0153","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00213.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2019-05-17"},{"id":"CVE-2019-1010275","cve":"CVE-2019-1010275","aliases":[],"title":"Helm: Improper certificate validation allows unauthorized clients to connect to Tiller","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Helm","year":"2019","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Improper certificate validation allows unauthorized clients to connect to Tiller","attack_vector":"Unauthenticated network","remediation":"Migrate off Helm 2; there is no Tiller in Helm 3","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-1010275"],"status":"curated","published":"2019-07-17"},{"id":"CVE-2019-11171","cve":"CVE-2019-11171","aliases":["INTEL-SA-00313","CVE-2019-11168","CVE-2019-11170","CVE-2019-11178"],"title":"Intel Baseboard Management Controller firmware (Intel server boards and systems) - web/network services: Heap","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Baseboard Management Controller firmware (Intel server boards and systems) - web/network services","year":"2019","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Heap corruption in the BMC's network-facing code, unauthenticated, yielding information disclosure, privilege escalation and denial of service; the same advisory bundles an outright authentication bypass and a session-validation failure. A compromised BMC is permanent control of the node underneath the OS - virtual media, power, KVM, and the write path to BIOS and ME firmware. It survives tenant reimage by construction, and a fleet-wide compromise of BMCs is a fleet-wide power-off button, which is a hall-level physical event, not a per-node one.","attack_vector":"Unauthenticated network access to the BMC's services. Any host on the OOB management network, anything that reaches a BMC exposed through a misconfigured route or a flat provisioning VLAN, and - where the KCS host interface is enabled - a tenant with root on the node.","remediation":"BMC firmware update from Intel / the board ODM (Quanta, Wiwynn, Supermicro on Intel designs). No host reboot needed, so it can be rolled without draining training jobs, but it must be staged per board family and verified per node. Structurally: dedicated OOB network with no tenant path, jump-host-only access, unique credentials per node, disable KCS on bare-metal SKUs, and re-flash the BMC as part of node reclaim between tenants rather than trusting its current image.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-11171","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00313.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-89"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2019-12838","cve":"CVE-2019-12838","aliases":[],"title":"Slurm (slurmdbd, sacctmgr archive load): A second SQL injection path into SlurmDBD, this one through the 'sacctmgr","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Slurm (slurmdbd, sacctmgr archive load)","year":"2019","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A second SQL injection path into SlurmDBD, this one through the 'sacctmgr archive load' path where strings were not escaped before hitting the database. Same consequence as the 2018 injection - arbitrary read and write of the accounting database that defines account membership, QOS caps and fairshare, which is the data structure the scheduler uses to decide whose jobs get GPUs.","attack_vector":"Anything that can reach the slurmdbd RPC port, typically the login nodes and the controller. The archive-load path is what an operator or an account coordinator invokes to reload archived accounting data.","remediation":"Upgrade to Slurm 18.08.8 or 19.05.1 and restart slurmdbd. SchedMD published fixes only for the then-supported 18.08 and 19.05 lines and warned that earlier versions carry similar flaws with no patch, so anything older has to move forward. Restrict slurmdbd's listener to the controller and admin hosts rather than the whole login-node subnet while you are in there.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-12838","https://lists.schedmd.com/pipermail/slurm-announce/"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2019-13553","cve":"CVE-2019-13553","aliases":["CVE-2019-13549","ICSA-19-297-01"],"title":"Rittal SK 3232-series chiller web interface (built on Carel pCOWeb firmware A1.5.3-B1.2.4): Whoever can reach","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Rittal SK 3232-series chiller web interface (built on Carel pCOWeb firmware A1.5.3-B1.2.4)","year":"2019","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Whoever can reach the chiller's web card owns the chilled-water loop. The hard-coded credentials give direct control over the two operations that matter physically: switching the cooling unit off, and moving the water temperature setpoint. That is a thermal attack, not an IT one. A GB200 or H100 rack drawing 40-140 kW has essentially zero thermal mass at the die; pull chilled water away from the CDUs or rear-door heat exchangers feeding it and inlet temperature crosses the GPU thermal-shutdown threshold in single-digit minutes. Every job in the affected hall dies, unsaved training checkpoints are lost, and repeated thermal cycling degrades VRMs, HBM stacks and pump seals. The subtler and nastier variant is not shutdown but a slow setpoint drift: raise supply water two degrees and the fleet silently throttles, which shows up as unexplained tokens-per-second regression and blown SLA credits long before anyone looks at the chiller.","attack_vector":"Unauthenticated HTTP on whatever network the chiller's pCOWeb card is plugged into. In practice that is the facility/mechanical VLAN, which in a leased colo is the landlord's network, not the tenant's - so a GPU operator may have no visibility into it at all and no idea whether it is flat with the building's office LAN. In owned or built-to-suit sites this card is usually on the same mechanical VLAN as the CRAHs, the BMS front end and the vendor's remote-support jump box. Internet exposure is real but not the common case; the common case is that anyone who lands on any building-systems subnet can reach it, and the credentials are published.","remediation":"There is no meaningful patch path - this is Carel pCOWeb OEM firmware embedded in a Rittal chiller, and the credentials are hard-coded. Realistic fix is network isolation: the chiller card goes on its own VLAN with an allow-list to exactly the BMS front end and nothing else, plus egress deny. If you lease space, you cannot patch this yourself; it is the landlord's mechanical plant. Put it in the contract - demand a network diagram for the mechanical VLAN, an attestation that no chiller/CRAH controller is reachable from any tenant or corporate network, and the right to have a third party validate it. Also demand that thermal-shutdown behaviour be tested: you need to know how many minutes you actually have, not a vendor's brochure number.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-13553","https://www.cisa.gov/news-events/ics-advisories/icsa-19-297-01","https://seclists.org/fulldisclosure/2019/Oct/45"],"status":"curated"},{"id":"CVE-2019-14271","cve":"CVE-2019-14271","aliases":[],"title":"Docker / moby: Code injection into `docker cp` via nsswitch loading a library from the container chroot","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2019","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Code injection into `docker cp` via nsswitch loading a library from the container chroot; host root","attack_vector":"Malicious image","remediation":"Upgrade Docker Engine; daemon restart","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-14271"],"status":"curated","published":"2019-07-29"},{"id":"CVE-2019-18658","cve":"CVE-2019-18658","aliases":[],"title":"Helm: Malicious chart includes sensitive host content such as /etc/passwd, or triggers DoS, when loaded","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Helm","year":"2019","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Malicious chart includes sensitive host content such as /etc/passwd, or triggers DoS, when loaded as a directory","attack_vector":"Malicious chart from a tenant or third-party repo","remediation":"Upgrade Helm; never run `helm package`/`helm lint` on untrusted charts on a privileged host","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-18658"],"status":"curated","published":"2019-11-12"},{"id":"CVE-2019-18801","cve":"CVE-2019-18801","aliases":[],"title":"Envoy: HTTP/2 request writes to the heap outside request buffers when the upstream is HTTP/1","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy","year":"2019","cvss_score":9.8,"severity":"critical","kev":false,"impact":"HTTP/2 request writes to the heap outside request buffers when the upstream is HTTP/1; potential RCE in the proxy","attack_vector":"Unauthenticated network","remediation":"Upgrade Envoy; for a mesh this means restarting every sidecar, which restarts tenant pods","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-18801"],"status":"curated","published":"2019-12-13"},{"id":"CVE-2019-18802","cve":"CVE-2019-18802","aliases":[],"title":"Envoy: Header whitespace handling enables request smuggling and authorization bypass","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy","year":"2019","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Header whitespace handling enables request smuggling and authorization bypass","attack_vector":"Unauthenticated network","remediation":"Upgrade Envoy; sidecar restart","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-18802"],"status":"curated","published":"2019-12-13"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-287"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2019-18823","cve":"CVE-2019-18823","aliases":["HTCONDOR-2020-0001","HTCONDOR-2020-0002","HTCONDOR-2020-0003","HTCONDOR-2020-0004"],"title":"HTCondor (condor_startd, condor_schedd, condor_shadow): One CVE covering four separate authentication failures the","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HTCondor (condor_startd, condor_schedd, condor_shadow)","year":"2019","cvss_score":9.8,"severity":"critical","kev":false,"impact":"One CVE covering four separate authentication failures the HTCondor team disclosed together. The slot claim secret is written in the clear to STARTD_HISTORY and is also sent unencrypted on the wire when daemon-to-daemon encryption is off - either one lets an attacker seize another user's slot and run their own code as that user. Separately, a user with only READ authorization can perform WRITE operations on the job queue, and if CLAIMTOBE is in the READ method list (the default) they can submit and run jobs as any other user. On mixed Windows/Linux pools the shadow will also hand a user's stored Windows password to anyone authenticating as the condor service.","attack_vector":"Depends on the sub-issue: reading a world-readable file from inside your own job on an execute node; passively capturing schedd-to-startd traffic; or simply authenticating to the schedd with any READ-list method. The CLAIMTOBE variant needs nothing but network reach to a default-configured pool.","remediation":"Upgrade to HTCondor 8.8.8 or 8.9.6 and restart all daemons. Independently of the patch, turn on daemon-to-daemon encryption, remove CLAIMTOBE from SEC_READ_AUTHENTICATION_METHODS, and disable match-password authentication in favour of a real method (SSL, Kerberos, TOKEN). Rotate any Windows credentials stored with condor_store_cred.","references":["https://htcondor.org/security/vulnerabilities/HTCONDOR-2020-0001.html","https://htcondor.org/security/vulnerabilities/HTCONDOR-2020-0003.html","https://nvd.nist.gov/vuln/detail/CVE-2019-18823"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2019-18960","cve":"CVE-2019-18960","aliases":[],"title":"Firecracker: vsock buffer overflow producing potentially exploitable crashes","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Firecracker","year":"2019","cvss_score":9.8,"severity":"critical","kev":false,"impact":"vsock buffer overflow producing potentially exploitable crashes","attack_vector":"Any tenant guest VM","remediation":"Upgrade Firecracker; restart microVMs","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-18960"],"status":"curated","published":"2019-12-11"},{"id":"CVE-2019-20427","cve":"CVE-2019-20427","aliases":[],"title":"Lustre ptlrpc module (server-side client packet validation): A Lustre client can send a crafted RPC that overflows a","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Lustre ptlrpc module (server-side client packet validation)","year":"2019","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A Lustre client can send a crafted RPC that overflows a buffer in the server's ptlrpc module, panicking it and possibly achieving remote code execution on the storage server. On a Lustre cluster the client is the tenant's GPU node — Lustre's security model assumes clients are trusted, and this is what that assumption costs. Code execution on an MDS or OSS is access to every tenant's data on the filesystem, and a panic takes the shared filesystem down for every running job.","attack_vector":"Any Lustre client — i.e. any compute node with the filesystem mounted, which in a rented GPU cluster means the tenant's own machine. No privilege escalation needed on the client beyond the ability to send RPCs.","remediation":"Upgrade Lustre servers to 2.12.3 or later. This is a coordinated storage-cluster upgrade: MDS and OSS nodes need the new build and a restart, and while Lustre supports failover pairs, most sites take an I/O pause. Structurally, treat the Lustre network (LNet) as a boundary: put it on a dedicated fabric that tenant workloads cannot address arbitrarily, and use Lustre nodemap/Shared-Secret Key authentication rather than relying on client trust.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-20427"],"status":"curated","tags":["tenant-isolation"],"published":"2020-01-27"},{"id":"CVE-2019-3705","cve":"CVE-2019-3705","aliases":[],"title":"Dell iDRAC7/8: Stack buffer overflow in the iDRAC web server — unauthenticated RCE on the BMC","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC7/8","year":"2019","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Stack buffer overflow in the iDRAC web server — unauthenticated RCE on the BMC","attack_vector":"Network, unauthenticated","remediation":"iDRAC7/8 are EOL on many fleets; remediation may require a chassis refresh rather than a patch","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-3705"],"status":"curated","published":"2019-04-26"},{"id":"CVE-2019-3706","cve":"CVE-2019-3706","aliases":[],"title":"Dell iDRAC9: Authentication bypass in the iDRAC9 web interface — full out-of-band control of the server","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC9","year":"2019","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Authentication bypass in the iDRAC9 web interface — full out-of-band control of the server","attack_vector":"Network, unauthenticated","remediation":"iDRAC firmware update via Lifecycle Controller or racadm; can be scripted fleet-wide but requires a reboot window on older iDRAC lines","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-3706"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2019-04-26"},{"id":"CVE-2019-3707","cve":"CVE-2019-3707","aliases":[],"title":"Dell iDRAC9: Authentication bypass via the WS-MAN interface","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC9","year":"2019","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Authentication bypass via the WS-MAN interface","attack_vector":"Network, unauthenticated","remediation":"Same iDRAC firmware update; also disable WS-MAN if unused","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-3707"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2019-04-26"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-306"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2019-5504","cve":"CVE-2019-5504","aliases":[],"title":"NetApp ONTAP Select Deploy administration utility (HTTP service): An unauthenticated attacker performs administrative","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"NetApp ONTAP Select Deploy administration utility (HTTP service)","year":"2019","cvss_score":9.8,"severity":"critical","kev":false,"impact":"An unauthenticated attacker performs administrative actions on the utility that deploys and manages ONTAP Select clusters. That is control over the storage serving the GPU fleet, with no credential at all.","attack_vector":"Any network path to the Deploy appliance's HTTP service. The service binds to the network and does not require authentication, so a foothold on the management VLAN is enough.","remediation":"Upgrade ONTAP Select Deploy 2.12/2.12.1 to a fixed release. Until then, firewall the Deploy appliance so only the storage admin jump host can reach it, and check the Deploy audit trail for actions you did not initiate.","references":["https://security.netapp.com/advisory/ntap-20190923-0001/","https://nvd.nist.gov/vuln/detail/CVE-2019-5504"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-319","CWE-522"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2019-5505","cve":"CVE-2019-5505","aliases":[],"title":"NetApp ONTAP Select Deploy administration utility (credential transport): Deploy sends its credentials in plaintext, so","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"NetApp ONTAP Select Deploy administration utility (credential transport)","year":"2019","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Deploy sends its credentials in plaintext, so anyone able to observe management traffic recovers the account that provisions and controls ONTAP Select storage clusters.","attack_vector":"A passive position on any network segment carrying Deploy management traffic - a mirrored port, a compromised switch, or a shared management VLAN.","remediation":"Upgrade ONTAP Select Deploy 2.2 through 2.12.1 to a fixed release, then rotate every credential that was ever used with the affected versions. Assume anything on that wire is already known.","references":["https://security.netapp.com/advisory/ntap-20190923-0002/","https://nvd.nist.gov/vuln/detail/CVE-2019-5505"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2019-5509","cve":"CVE-2019-5509","aliases":[],"title":"NetApp ONTAP Select Deploy administration utility (code injection): An unauthenticated remote attacker injects code and","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"NetApp ONTAP Select Deploy administration utility (code injection)","year":"2019","cvss_score":9.8,"severity":"critical","kev":false,"impact":"An unauthenticated remote attacker injects code and enables a privileged account on the Deploy appliance, taking over provisioning for every ONTAP Select cluster it manages.","attack_vector":"Network access to ONTAP Select Deploy 2.11.2 through 2.12.2. No prior account needed.","remediation":"Upgrade to a fixed Deploy release, then enumerate local accounts on the appliance and delete any privileged account you did not create. Rotate the credentials of the ones you keep.","references":["https://security.netapp.com/advisory/ntap-20191121-0001/","https://nvd.nist.gov/vuln/detail/CVE-2019-5509"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2019-5685","cve":"CVE-2019-5685","aliases":[],"title":"NVIDIA Windows GPU Display Driver, DirectX driver (shader compiler/runtime): A crafted shader overruns a shader-local","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver, DirectX driver (shader compiler/runtime)","year":"2019","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A crafted shader overruns a shader-local temporary array, giving code execution in the graphics driver. Same delivery surface as the texture-array bug: wherever attacker-authored shaders reach your GPUs, this is remote code execution against the host driver.","attack_vector":"Anyone able to submit shaders - remote/VDI session users, tenant VMs, or local processes.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://support.lenovo.com/us/en/product_security/LEN-28096","https://www.talosintelligence.com/vulnerability_reports/TALOS-2019-0812","https://nvd.nist.gov/vuln/detail/CVE-2019-5685"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2019-08-06"},{"id":"CVE-2019-6260","cve":"CVE-2019-6260","aliases":["Pantsdown"],"title":"ASPEED AST2400 / AST2500 BMC SoC: Arbitrary read/write of the BMC's entire physical address space **from the host CPU**","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ASPEED AST2400 / AST2500 BMC SoC","year":"2019","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Arbitrary read/write of the BMC's entire physical address space **from the host CPU** — host-to-BMC boundary collapse. A tenant with host root can implant the BMC; survives reimaging and node reallocation","attack_vector":"Local, from the host OS via iLPC2AHB / PCIe VGA / X-DMA bridges","remediation":"Fix is a BMC firmware build that disables the AHB bridges (OpenBMC has it; many ODM builds do not). On multi-tenant bare metal this is the single most important control — otherwise every tenant handoff is a potential persistent implant","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-6260"],"status":"curated","fleet":{"ubiquity":"universal - ASPEED is effectively the sole-source BMC SoC in x86 server boards, including GPU servers","remediation_pain":"firmware-flash per node; on some boards the only mitigation is disabling the LPC/PCIe P2A bridges in an OEM image respin, and several SKUs remain unpatchable-mitigate-only","pain_class":"unpatchable / mitigate-only","why_fleet_wide":"Host-side root can read/write the BMC's entire physical address space over LPC/PCIe, so any tenant that gets host root pivots into the always-on management processor - below the hypervisor, persistent across reimaging, on every node of an ASPEED-based fleet."},"published":"2019-01-22"},{"cvss_vector":"CVSS:3.0/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","fleet":{"pain_class":"daemon-restart"},"id":"CVE-2019-6438","cve":"CVE-2019-6438","aliases":[],"title":"Slurm (32-bit RPC handling): Memory corruption on 32-bit Slurm builds reachable from a crafted RPC, up to control of","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Slurm (32-bit RPC handling)","year":"2019","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Memory corruption on 32-bit Slurm builds reachable from a crafted RPC, up to control of the daemon process. SchedMD states 64-bit builds - the overwhelming majority - are not affected, so this only matters if you still run 32-bit management or login hosts.","attack_vector":"Network reach to a 32-bit Slurm daemon. No credentials described as required.","remediation":"Upgrade to Slurm 17.11.13 or 18.08.5. SchedMD published fixes only for the then-supported 17.11 and 18.08 lines and states that similar flaws affect earlier 32-bit builds with no fix available, so on anything older the only resolution is upgrading. If you have 32-bit Slurm hosts left in the estate, retire them.","references":["https://lists.schedmd.com/pipermail/slurm-announce/2019/000018.html","https://nvd.nist.gov/vuln/detail/CVE-2019-6438"],"status":"curated"},{"id":"CVE-2019-7276","cve":"CVE-2019-7276","aliases":["CVE-2019-7279","CVE-2019-7274","CVE-2019-7273"],"title":"Optergy Proton / Enterprise building management platform: A backdoor console giving remote root code execution","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Optergy Proton / Enterprise building management platform","year":"2019","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A backdoor console giving remote root code execution, alongside hard-coded credentials, CSRF, and authenticated file upload that also runs as root. Optergy is a full BMS platform - it aggregates HVAC, metering and often access control for a site - so root on it is root over the facility's control surface and its credential store. An attacker gets to command cooling equipment directly, alter schedules and setpoints, silence alarms, and read out whatever downstream device passwords the platform holds. For a GPU operator the physical consequence is the familiar one: cooling commanded away from a hall of 40 kW+ racks means thermal shutdown in minutes, lost checkpoints and thermal-cycling damage. The backdoor is the part that should decide the response - a deliberate hidden access path means you cannot reason about who has been in the system historically.","attack_vector":"Remote, unauthenticated, over the platform's web interface on the facility network. Optergy deployments are frequently published for remote access because the product is sold on browser-based management, so internet exposure is a realistic assumption rather than an edge case. Hard-coded credentials mean even a 'secured' instance is open to anyone who read the advisory.","remediation":"Vendor firmware/software update - a platform upgrade rather than controller flashing, so no cooling downtime, but it requires the integrator and a version that removes the backdoor console. Given a deliberate backdoor was present, patching alone is not sufficient: rebuild or re-image the platform, rotate every credential it ever stored, and rotate credentials on every downstream device it integrated. Then remove all internet exposure and put it behind a jump host with MFA. If the platform is under an integrator's remote-support contract, audit that path specifically.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-7276","https://nvd.nist.gov/vuln/detail/CVE-2019-7279","https://applied-risk.com/resources/advisories"],"status":"curated"},{"id":"CVE-2019-9569","cve":"CVE-2019-9569","aliases":["McAfee ATR HVACking"],"title":"Delta Controls enteliBUS Manager (eBMGR) V3.40_B-571848, dactetra service: Unauthenticated remote code execution","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Delta Controls enteliBUS Manager (eBMGR) V3.40_B-571848, dactetra service","year":"2019","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unauthenticated remote code execution on a central plant controller. The enteliBUS Manager is the box that sequences chillers, pumps, boilers and air handling for a whole building; the research that produced this bug demonstrated full control over HVAC and, on the same hardware family, over the access-control and life-safety points wired into it. Owning it gives an attacker root-equivalent control over the mechanical plant with no credentials and no user interaction, plus the ability to lie to the supervisory system above it. Against a GPU hall this is the textbook scenario: command the plant off or drive setpoints, and a 40 kW+ rack crosses thermal shutdown in minutes while the BMS graphic still shows normal operation. Because the compromise is code execution rather than a config change, it also persists across reboots and survives the operator's instinctive 'restart the controller' response.","attack_vector":"Unauthenticated network access to the controller's service port on the facility network - no credentials, no interaction. enteliBUS controllers are commonly reachable from anywhere on the building VLAN and, in sites where the controls contractor set up remote support, from the internet through a poorly-placed remote-access appliance.","remediation":"Firmware update from Delta Controls, applied through the certified Delta dealer who owns the site - operators generally cannot obtain or apply Delta firmware directly, which is itself the problem. That means a scheduled contractor visit and a plant controller offline during the flash: real cooling risk, real cost, and a window most operators will not take without an incident to justify it. Treat segmentation as the primary control: dedicated VLAN, deny-by-default with an allow-list from the supervisor only, and no path from the internet. Verify by scanning your own facility VLAN for the controller's service port rather than trusting the dealer's assurance.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-9569","https://www.mcafee.com/blogs/other-blogs/mcafee-labs/hvacking-understanding-the-delta-between-security-and-reality/","https://www.deltacontrols.com/products/hvac-controls/central-plant-controllers/entelibus"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2020-11483","cve":"CVE-2020-11483","aliases":[],"title":"NVIDIA DGX BMC (AMI firmware): Hard-coded credentials in the DGX BMC firmware","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVIDIA DGX BMC (AMI firmware)","year":"2020","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Hard-coded credentials in the DGX BMC firmware. Anyone who can reach the BMC's management interface authenticates with credentials baked into a firmware image that is publicly downloadable - meaning no credential rotation on your side ever helped. That is full out-of-band control of a DGX-1 or DGX-2: power, console, virtual media, and a path to reflashing the host. Affects DGX-1 before BMC 3.38.30 and DGX-2 before 1.06.06.","attack_vector":"Anyone with network reach to the BMC. If the management network is flat, shared with tenants, or accidentally routable, that is effectively anyone inside the datacenter - and internet-exposed BMCs make it anyone at all.","remediation":"Flash the DGX BMC firmware from NVIDIA's DGX firmware update container (DGX-1 to 3.38.30 or later, DGX-2 to 1.06.06 or later; DGX A100 per the bulletin's table). A BMC flash does not require the host OS to reboot but drops out-of-band management for several minutes and NVIDIA recommends a host power cycle afterwards, so treat it as a per-node maintenance window. Rotate every BMC and IPMI credential after the flash - flashing does not invalidate secrets an attacker already pulled. Keep BMCs on an isolated management VLAN with no route from tenant or job networks.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-11483"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2020-10-29"},{"id":"CVE-2020-11486","cve":"CVE-2020-11486","aliases":[],"title":"NVIDIA DGX BMC (AMI firmware): File upload into the BMC that gets automatically processed, yielding remote code","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVIDIA DGX BMC (AMI firmware)","year":"2020","cvss_score":9.8,"severity":"critical","kev":false,"impact":"File upload into the BMC that gets automatically processed, yielding remote code execution on the baseboard management controller of a DGX-1. Code on the BMC is below the host OS: it survives host reinstall, sees the host's memory and storage paths, controls power and firmware, and is invisible to everything running on the node. This is the worst outcome in the DGX-1 BMC set. DGX-1 before BMC 3.38.30.","attack_vector":"Anyone with network reach to the BMC's management interface. Combined with the hard-coded credentials in the same bulletin, that is unauthenticated in practice.","remediation":"Flash the DGX BMC firmware from NVIDIA's DGX firmware update container (DGX-1 to 3.38.30 or later, DGX-2 to 1.06.06 or later; DGX A100 per the bulletin's table). A BMC flash does not require the host OS to reboot but drops out-of-band management for several minutes and NVIDIA recommends a host power cycle afterwards, so treat it as a per-node maintenance window. Rotate every BMC and IPMI credential after the flash - flashing does not invalidate secrets an attacker already pulled. Keep BMCs on an isolated management VLAN with no route from tenant or job networks.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-11486"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2020-10-29"},{"id":"CVE-2020-11951","cve":"CVE-2020-11951","aliases":[],"title":"Rittal PDU-3C002DEC rack PDU firmware (through 5.17.10): A backdoor root account in the PDU firmware. Not a weak","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Rittal PDU-3C002DEC rack PDU firmware (through 5.17.10)","year":"2020","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A backdoor root account in the PDU firmware. Not a weak default that an operator could change - an undocumented account shipped in the image. Anyone who knows it owns the device that switches power to the rack, and nothing in your provisioning process would ever have noticed it.","attack_vector":"Network access to the PDU management interface, with credentials that are public knowledge once the advisory is out.","remediation":"Firmware update to a build that removes the account. This is a per-PDU flash, two per rack, and the update must be verified rather than assumed - check that the account is actually gone on a sample. Treat undocumented accounts as a procurement question going forward: require vendors to attest that no non-configurable accounts exist before you buy a PDU SKU at fleet scale.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-11951"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["physical-impact"],"published":"2020-07-14"},{"id":"CVE-2020-11956","cve":"CVE-2020-11956","aliases":[],"title":"Rittal PDU-3C002DEC rack PDU firmware (through 5.17.10): Least-privilege violation: low-privilege users on the PDU get","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Rittal PDU-3C002DEC rack PDU firmware (through 5.17.10)","year":"2020","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Least-privilege violation: low-privilege users on the PDU get far more capability than the role implies, up to and including control of the device. Read-only monitoring accounts handed to a DCIM system or an NOC contractor become power control.","attack_vector":"Any authenticated user on the PDU, including monitoring accounts.","remediation":"Firmware update per PDU. Audit third-party accounts on the power estate at the same time - the monitoring integration is usually how the low-privilege credential got into someone else's hands.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-11956"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2020-07-14"},{"id":"CVE-2020-12812","cve":"CVE-2020-12812","aliases":["FG-IR-19-283"],"title":"Fortinet FortiOS SSL-VPN: A logic flaw lets a user who changes their login case (e.g","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Fortinet FortiOS SSL-VPN","year":"2020","cvss_score":9.8,"severity":"critical","kev":true,"impact":"A logic flaw lets a user who changes their login case (e.g. 'User' vs 'user') complete SSL-VPN authentication without ever being prompted for their second factor — a clean bypass of FortiToken MFA on the VPN gateway. Confirmed in CISA's KEV catalog as actively exploited; if this FortiGate is the VPN entry point into a cluster's management network, MFA was supposed to be the thing stopping a stolen password from being enough.","attack_vector":"Requires a valid username/password (e.g. phished or reused) but no second factor — the attacker just varies the case of the username at login to skip the FortiToken prompt.","remediation":"Firmware upgrade to the fixed FortiOS release per Fortinet PSIRT FG-IR-19-283. Given confirmed active exploitation, patch ahead of routine cycles, and afterward force a credential rotation for any accounts that authenticated to SSL-VPN during the vulnerable window in case MFA was bypassed on them already.","references":["https://fortiguard.com/psirt/FG-IR-19-283","https://www.cisa.gov/known-exploited-vulnerabilities-catalog?field_cve=CVE-2020-12812"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2020-07-24"},{"id":"CVE-2020-13092","cve":"CVE-2020-13092","aliases":[],"title":"scikit-learn / joblib: `joblib.load()` executes commands from an untrusted file via `__reduce__`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"scikit-learn / joblib","year":"2020","cvss_score":9.8,"severity":"critical","kev":false,"impact":"`joblib.load()` executes commands from an untrusted file via `__reduce__`","attack_vector":"Customer-supplied `.joblib`/`.pkl` model","remediation":"No fix — this is pickle semantics. Reject joblib artifacts from untrusted sources; use skops or ONNX","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-13092"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"published":"2020-05-15"},{"id":"CVE-2020-15373","cve":"CVE-2020-15373","aliases":[],"title":"Brocade Fabric OS REST API: Multiple buffer overflows in the Fabric OS REST API reachable by an unauthenticated remote","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Brocade Fabric OS REST API","year":"2020","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Multiple buffer overflows in the Fabric OS REST API reachable by an unauthenticated remote attacker. The REST API is what modern SAN automation drives, so it is enabled in exactly the environments that also automate zoning — meaning the vulnerable interface is the one wired into your provisioning pipeline.","attack_vector":"Unauthenticated, remote to the FOS REST API on v8.2.1 through v8.2.1d, and 8.2.2 before v8.2.2c.","remediation":"Fabric OS upgrade plus reboot, per fabric. If you do not use the REST API, disabling it is a live config change that removes the exposure without a maintenance window. Related: CVE-2020-15374, CVE-2020-15371.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-15373","https://nvd.nist.gov/vuln/detail/CVE-2020-15371"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2020-09-25"},{"id":"CVE-2020-15639","cve":"CVE-2020-15639","aliases":["CVE-2020-15640","CVE-2020-15641","CVE-2020-15642","CVE-2020-15643","CVE-2020-15644","CVE-2020-15645","CVE-2020-17387","CVE-2020-17388","CVE-2020-17389","CVE-2020-5803","CVE-2020-5804","CVE-2020-5805"],"title":"Marvell QConvergeConsole GUI 5.5.0.64 - 5.5.0.74 (QLogic HBA management): The earlier cluster on the same console","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Marvell QConvergeConsole GUI 5.5.0.64 - 5.5.0.74 (QLogic HBA management)","year":"2020","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The earlier cluster on the same console: unauthenticated RCE via decryptFile, unauthenticated file disclosure via getFileUploadBytes, several RCE paths whose authentication requirement is defeated by a bypass in the console's own auth mechanism, path traversal that deletes arbitrary files as SYSTEM or root, and Tomcat credentials stored in cleartext in tomcat-users.xml. Included here because operators routinely find 5.5.0.6x/7x consoles still running on legacy management servers that were never inventoried - the same fleet-wide HBA firmware control as the 2025 cluster, on hosts nobody is patching.","attack_vector":"Any host reachable to the QConvergeConsole web port. Unauthenticated for the RCE and disclosure primitives; the cleartext tomcat-users.xml additionally gives any local OS user on the console host a working login.","remediation":"Do not patch this generation - retire it. Inventory for QConvergeConsole installs by port and by package, uninstall from all hosts, and replace with CLI-driven HBA management. If a console must stay, upgrade to the current release, rotate the Tomcat credentials and put the port behind an admin-only ACL. No HBA firmware flash and no storage downtime is required to remove the console.","references":["https://www.marvell.com/content/dam/marvell/en/public-collateral/fibre-channel/marvell-fibre-channel-security-advisory-2020-07.pdf","https://www.zerodayinitiative.com/advisories/ZDI-20-967/","https://nvd.nist.gov/vuln/detail/CVE-2020-15639"],"status":"curated","published":"2020-08-25"},{"id":"CVE-2020-27745","cve":"CVE-2020-27745","aliases":[],"title":"Slurm: RPC buffer overflow in the PMIx MPI plugin","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Slurm","year":"2020","cvss_score":9.8,"severity":"critical","kev":false,"impact":"RPC buffer overflow in the PMIx MPI plugin","attack_vector":"Any user who can submit a job","remediation":"Upgrade Slurm; restart slurmctld and slurmd. Draining a Slurm partition means killing running training jobs","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-27745"],"status":"curated","published":"2020-11-27"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-732"],"fleet":{"pain_class":"hot-patch"},"id":"CVE-2020-36770","cve":"CVE-2020-36770","aliases":[],"title":"Slurm (Gentoo ebuild pkg_postinst): The Gentoo packaging runs chown across paths on the live root filesystem during","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Slurm (Gentoo ebuild pkg_postinst)","year":"2020","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The Gentoo packaging runs chown across paths on the live root filesystem during install, so a local user who can pre-create or symlink those paths gets root-owned files planted where they choose. This is a packaging defect, not a Slurm code defect, but it lands on the controller host with root.","attack_vector":"A local user on a Gentoo host at the moment the slurm package is installed or upgraded.","remediation":"Only relevant if you deploy Slurm from Gentoo ebuilds - most GPU sites do not. Update to a fixed ebuild, or install Slurm from SchedMD tarballs or distro packages you control. Verify ownership under the Slurm state and spool directories after any install.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-36770","https://bugs.gentoo.org/631552"],"status":"curated"},{"id":"CVE-2020-3992","cve":"CVE-2020-3992","aliases":[],"title":"VMware ESXi (OpenSLP): Use-after-free in OpenSLP on port 427 - unauthenticated remote code execution on the hypervisor","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"VMware ESXi (OpenSLP)","year":"2020","cvss_score":9.8,"severity":"critical","kev":true,"impact":"Use-after-free in OpenSLP on port 427 - unauthenticated remote code execution on the hypervisor [KEV]","attack_vector":"Unauthenticated network on the management segment","remediation":"Patch and disable the SLP service entirely (VMware's own recommendation). Any ESXi with 427 reachable should be treated as already compromised","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-3992"],"status":"curated","published":"2020-10-20"},{"id":"CVE-2020-5344","cve":"CVE-2020-5344","aliases":[],"title":"Dell iDRAC9: Stack-based buffer overflow via crafted remote input — pre-auth code execution on the BMC","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC9","year":"2020","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Stack-based buffer overflow via crafted remote input — pre-auth code execution on the BMC","attack_vector":"Network, unauthenticated","remediation":"iDRAC firmware update; requires a rolling out-of-band update campaign across the fleet","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5344"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2020-03-31"},{"id":"CVE-2020-7521","cve":"CVE-2020-7521","aliases":["SEVD-2020-224-01"],"title":"APC Easy UPS On-Line Software (SFAPV9601) FileUploadServlet: Path traversal in a file upload servlet allows writing","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"APC Easy UPS On-Line Software (SFAPV9601) FileUploadServlet","year":"2020","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Path traversal in a file upload servlet allows writing an executable anywhere on the host, giving code execution on the machine that manages UPS shutdown across the site. Same PHYSICAL end state as the newer Easy UPS bugs - an attacker gains the ability to command an orderly power-down of everything the software manages.","attack_vector":"Unauthenticated network access to the software's web servlet.","remediation":"Upgrade past v2.0. Server-side upgrade, cheap. If the host was exposed, rebuild rather than patch - and rotate every UPS and host credential it stored.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-7521"],"status":"curated","published":"2020-08-31"},{"id":"CVE-2021-1104","cve":"CVE-2021-1104","aliases":[],"title":"RISC-V ISA (MTVEC register) as used in NVIDIA GPU microcontrollers: A documented ambiguity in the RISC-V specification","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"RISC-V ISA (MTVEC register) as used in NVIDIA GPU microcontrollers","year":"2021","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A documented ambiguity in the RISC-V specification leaves the machine trap vector base address register in an undefined state at reset, which fault injection can exploit to redirect trap handling and break the secure boot chain of an embedded microcontroller. NVIDIA's own offensive-security team assigned this; it is relevant because modern NVIDIA GPU and platform microcontrollers are RISC-V based, so it describes a weakness in the root of trust below the GPU driver. Exploitation needs physical glitching, not remote access - so it belongs in your supply-chain and physical-security threat model, not your patch queue.","attack_vector":"An attacker with physical access to the board who can perform voltage or clock glitching. Not reachable from software, local or remote.","remediation":"Architectural ambiguity in the RISC-V ISA specification as implemented in embedded microcontrollers, including NVIDIA's. There is no operator-installable patch: the fix is in silicon and in hardened boot firmware from the chip vendor. For an operator this is unpatchable in the field - mitigate by treating physical access to a GPU as game over, keeping fault-injection-capable access (open chassis, exposed board) out of shared-tenancy racks, and preferring hardware generations that NVIDIA states carry the hardened boot ROM.","references":["https://riscv.org/news/2021/08/video-glitching-risc-v-chips-mtvec-corruption-for-hardening-isa-adam-zabrocki-and-alex-matrosov-def-con-29/","https://nvd.nist.gov/vuln/detail/CVE-2021-1104"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"published":"2021-08-13"},{"id":"CVE-2021-1361","cve":"CVE-2021-1361","aliases":[],"title":"Cisco Nexus 3000/9000 (internal file management service): Unauthenticated remote file write, read and delete as root","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Cisco Nexus 3000/9000 (internal file management service)","year":"2021","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unauthenticated remote file write, read and delete as root on the switch, over a service that listens on the management interface by default. An attacker on the management network can drop a file, replace a config, or wipe the box without ever having a credential. This is the single worst pre-auth exposure in the Nexus line and it is a good argument for treating the switch management VLAN as production, not as 'internal'.","attack_vector":"Unauthenticated, remote — anything that can reach TCP/9075 on the switch's mgmt0 interface. No credentials required.","remediation":"NX-OS image upgrade and switch reload. As a stopgap, an interface ACL on mgmt0 restricting the affected port materially reduces exposure and can be applied live with no reload. Rollout: one reload per switch, drain-and-patch per MLAG pair.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1361"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2021-02-24"},{"id":"CVE-2021-27797","cve":"CVE-2021-27797","aliases":[],"title":"Brocade Fabric OS (hard-coded credentials): Documented hard-coded credentials in Brocade Fabric OS","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Brocade Fabric OS (hard-coded credentials)","year":"2021","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Documented hard-coded credentials in Brocade Fabric OS — the operating system on the Fibre Channel directors and switches that carry NVMe-over-FC and FC-attached storage into AI clusters. Hard-coded credentials on a SAN switch mean anyone who can reach its management interface owns the zoning configuration, and zoning is the only thing keeping one host's initiators from seeing another's LUNs.","attack_vector":"Anyone with network reachability to the FOS management interface. The credentials are public.","remediation":"Upgrade Fabric OS to 8.2.1c / 8.1.2h or later — all 8.0.x and 7.x versions are affected with no fixed release, so those switches need replacing or hard isolation. A FOS upgrade is a firmware install plus switch reboot; on a redundant dual-fabric SAN you do one fabric at a time and multipathing covers it. Until then, put FOS management interfaces on an isolated OOB network.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-27797"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2022-02-21"},{"id":"CVE-2021-28235","cve":"CVE-2021-28235","aliases":[],"title":"etcd: Authentication flaw via the debug function","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"etcd","year":"2021","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Authentication flaw via the debug function -> remote privilege escalation","attack_vector":"Network (remote)","remediation":"Control-plane: CRITICAL - etcd holds every k8s secret; upgrade + rotate all stored secrets","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-28235"],"status":"curated","published":"2023-04-04"},{"id":"CVE-2021-28503","cve":"CVE-2021-28503","aliases":[],"title":"Arista EOS (eAPI certificate auth): Certificate-based eAPI authentication skips credential re-evaluation","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (eAPI certificate auth)","year":"2021","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Certificate-based eAPI authentication skips credential re-evaluation — authentication bypass on the switch's programmatic API, which is exactly the interface a neocloud's fabric automation uses","attack_vector":"Network","remediation":"EOS upgrade with fabric failover; also rotate any eAPI client certificates issued while vulnerable","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-28503"],"status":"curated","published":"2022-02-04"},{"id":"CVE-2021-29212","cve":"CVE-2021-29212","aliases":["HPESBGN04189","ZDI-21-1278"],"title":"HPE iLO Amplifier Pack (unauthenticated directory traversal): Unauthenticated directory traversal on the iLO Amplifier","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HPE iLO Amplifier Pack (unauthenticated directory traversal)","year":"2021","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unauthenticated directory traversal on the iLO Amplifier Pack appliance, letting an attacker read arbitrary files off it. Amplifier Pack is the fleet-scale iLO management appliance - it inventories every iLO it manages and holds the credentials it uses to reach them. Reading its filesystem without authentication is therefore a credential harvest for the entire out-of-band plane, not a single-host information leak. One unauthenticated request against one appliance can yield the keys to every BMC behind it. Affects versions 1.80, 1.81, 1.90 and 1.95.","attack_vector":"Anything routable to the Amplifier Pack appliance, unauthenticated. In most deployments that means the management network - so the question is which hosts, VPN pools and monitoring systems share a segment with your fleet management appliance.","remediation":"Upgrade iLO Amplifier Pack past the affected 1.80-1.95 range - a single appliance upgrade, no per-node work, no reboots and no job drain, so rollout cost is near zero. What is not near zero: if this appliance was reachable while unpatched, assume the iLO credentials it stored are burned and rotate them fleet-wide. That credential rotation is the expensive part, and skipping it leaves the actual exposure open after the appliance is patched.","references":["https://support.hpe.com/hpsc/doc/public/display?docLocale=en_US&docId=emr_na-hpesbgn04189en_us","https://www.zerodayinitiative.com/advisories/ZDI-21-1278/","https://nvd.nist.gov/vuln/detail/CVE-2021-29212"],"status":"curated","published":"2021-11-01"},{"id":"CVE-2021-31884","cve":"CVE-2021-31884","aliases":["CVE-2021-31886","CVE-2021-31887","CVE-2021-31888","CVE-2017-9946","SSA-044112","SSA-114589"],"title":"Siemens APOGEE PXC / MEC / MBC and TALON TC BACnet and P2 automation controllers: A cluster of critical flaws","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Siemens APOGEE PXC / MEC / MBC and TALON TC BACnet and P2 automation controllers","year":"2021","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A cluster of critical flaws in the Siemens field controllers that actually drive air handling units, VAV boxes and chilled-water valves, including unauthenticated paths to the integrated web server and memory-safety issues reachable over the network. Compromise of an APOGEE or TALON controller is compromise of the physical actuator layer: fan enable, damper position, valve position, setpoint. There is no supervisory system to overrule it because the controller is the thing that closes the loop. The older companion issue (CVE-2017-9946) lets an attacker bypass authentication on the embedded web server and pull configuration off the device, which hands over the point map needed to make the attack surgical. In a GPU hall, an attacker who owns the AHU controllers can drive inlet temperatures past the accelerator shutdown threshold in minutes and, by holding valve position, keep them there.","attack_vector":"Network access to the controller's HTTP/HTTPS ports (80/443) on the facility network. These are wall-mounted controllers in mechanical rooms and ceiling spaces, on a flat building VLAN with no port security in the overwhelming majority of sites. Physical access to the panel is also a real vector because these are often in unlocked or shared mechanical spaces that tenant security policy does not cover.","remediation":"Firmware upgrade per Siemens SSAs - APOGEE PXC Compact/Modular to V3.5.4+ (BACnet) or V2.8.19+ (P2), TALON TC accordingly. That is a per-controller flash by a Siemens-certified technician, with each AHU losing automatic control during its flash, so it is a genuine maintenance-window project measured in technician-days across a large site and one that most operators will keep deferring. Because of that, the load-bearing control is network isolation: controllers on a dedicated VLAN, HTTP/HTTPS permitted only from the supervisory station, and physical locks on mechanical rooms and control panels. If you lease, these are the landlord's controllers - ask for the firmware baseline, and if the answer is 'all versions' assume the cluster applies.","references":["https://cert-portal.siemens.com/productcert/pdf/ssa-044112.pdf","https://cert-portal.siemens.com/productcert/pdf/ssa-148078.pdf","https://nvd.nist.gov/vuln/detail/CVE-2021-31884","https://nvd.nist.gov/vuln/detail/CVE-2017-9946"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-31921","cve":"CVE-2021-31921","aliases":[],"title":"Istio: With AUTO_PASSTHROUGH gateways, an external client reaches arbitrary in-cluster services, bypassing","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Istio","year":"2021","cvss_score":9.8,"severity":"critical","kev":false,"impact":"With AUTO_PASSTHROUGH gateways, an external client reaches arbitrary in-cluster services, bypassing all authorization","attack_vector":"Unauthenticated network via the ingress gateway","remediation":"Emergency istiod and gateway upgrade; audit for AUTO_PASSTHROUGH gateways","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-31921"],"status":"curated","published":"2021-06-02"},{"id":"CVE-2021-32974","cve":"CVE-2021-32974","aliases":["ICSA-21-187-01"],"title":"Moxa NPort IAW5000A-I/O serial device server: The built-in web server doesn't validate input properly, letting a remote","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Moxa NPort IAW5000A-I/O serial device server","year":"2021","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The built-in web server doesn't validate input properly, letting a remote unauthenticated attacker run arbitrary commands on the device server — full takeover of the box bridging serial equipment onto the network.","attack_vector":"Remote, unauthenticated — a crafted HTTP request to the web management interface is enough.","remediation":"Firmware upgrade to the version Moxa published in its security advisory; requires a flash and reboot on every affected unit, which briefly drops the serial sessions it's carrying. No compensating config change exists since the flaw is in input handling, not a feature you can disable.","references":["https://www.cisa.gov/uscert/ics/advisories/icsa-21-187-01","https://www.moxa.com/en/support/product-support/security-advisory/nport-iaw5000a-io-serial-device-server-vulnerabilities"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-04-01"},{"id":"CVE-2021-39226","cve":"CVE-2021-39226","aliases":[],"title":"Grafana: Unauthenticated access to snapshots via /api/snapshots/:key","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Grafana","year":"2021","cvss_score":9.8,"severity":"critical","kev":true,"impact":"Unauthenticated access to snapshots via /api/snapshots/:key -> view and delete snapshot data","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade the monitoring tier","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-39226"],"status":"curated","published":"2021-10-05"},{"id":"CVE-2021-39815","cve":"CVE-2021-39815","aliases":["PowerVR pinned-memory UAF"],"title":"Imagination PowerVR GPU driver - pinned memory lifecycle: An unprivileged app allocates pinned GPU memory, unpins it so","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Imagination PowerVR GPU driver - pinned memory lifecycle","year":"2021","cvss_score":9.8,"severity":"critical","kev":false,"impact":"An unprivileged app allocates pinned GPU memory, unpins it so the page can be freed and reallocated elsewhere, and keeps using it in GPU calls - a use-after-free that reads and writes pages now belonging to something else. Scored 9.8. This is the GPU-memory-lifecycle bug class that matters most for shared accelerators: the GPU keeps a mapping the OS thinks it revoked.","attack_vector":"Unprivileged local application with GPU access. Re-filed as CVE-2022-20122 in a later Android bulletin.","remediation":"Update the vendor GPU driver. The transferable lesson for a GPU fleet is to ask, for whichever accelerator you run, whether unpinning actually tears down the device-side mapping - the same design mistake is what makes GPU memory reuse dangerous on any vendor.","references":["https://source.android.com/security/bulletin/2022-04-01","https://nvd.nist.gov/vuln/detail/CVE-2021-39815"],"status":"curated","tags":["tenant-isolation"],"published":"2022-08-24"},{"id":"CVE-2021-41842","cve":"CVE-2021-41842","aliases":["VU#796611"],"title":"Insyde InsydeH2O (AtaLegacySmm SMM driver): The SMI handler in the legacy ATA driver does not validate the CommBuffer","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (AtaLegacySmm SMM driver)","year":"2021","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The SMI handler in the legacy ATA driver does not validate the CommBuffer it is handed, so a caller from the OS can steer SMM into executing attacker-supplied code. That is full ring -2 control: persistence that survives OS reinstall and disk wipe, the ability to disable or forge Secure Boot and TPM measurements, and a vantage point beneath any hypervisor. The highest-scored entry in the whole 2021-2022 Insyde/Binarly batch.","attack_vector":"Local admin or root on the host OS, then a software SMI invoking the vulnerable handler. On bare-metal GPU rental this is exactly the privilege a tenant already has on their leased node.","remediation":"InsydeH2O kernel fix (5.0 / 05.08.46, 5.1 / 05.16.46, 5.2 / 05.26.46, 5.3 / 05.35.46, 5.4 / 05.43.46, 5.5 / 05.51.45) - but you cannot apply that. You need the BIOS image your server OEM built on top of it, and the rebase lag from Insyde's kernel drop to a shipping Dell/HPE/Lenovo/Supermicro payload ran into many months for this batch. Firmware flash plus one reboot per node. No config workaround: SMM cannot be turned off. If you rent bare metal to untrusted tenants on unpatched firmware, treat every returned node as compromised and re-flash rather than reimage.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-41842","https://kb.cert.org/vuls/id/796611","https://www.insyde.com/security-pledge"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-01-06"},{"id":"CVE-2021-42774","cve":"CVE-2021-42774","aliases":["CVE-2021-42772","CVE-2021-42773","CVE-2021-42775"],"title":"Broadcom Emulex HBA Manager / OneCommand Manager (Fibre Channel and FC-NVMe HBAs), before 11.4.425.0 and 12.8.542.31","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Broadcom Emulex HBA Manager / OneCommand Manager (Fibre Channel and FC-NVMe HBAs), before 11.4.425.0 and 12.8.542.31","year":"2021","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unless the agent was installed in Strictly Local Management mode, the Emulex management daemon accepts unauthenticated remote commands. Two of them matter: the remote firmware-download path has a buffer overflow reachable pre-auth, and the same path lets an attacker place or replace an arbitrary file on the host. That is remote, unauthenticated code execution on the host plus a direct route to writing HBA firmware. An Emulex HBA is a PCIe device with DMA and its own processor that owns the node's path to shared storage - firmware planted there survives an OS reimage and can silently mirror or corrupt a later tenant's FC traffic. The companion GetDumpFile issues additionally let an unauthenticated caller pull arbitrary files off the host.","attack_vector":"Any host that can reach the HBA Manager remote management listener on the host's management interface. No credentials in non-secure (default remote) mode. This listener frequently ends up on the same flat provisioning VLAN as the BMCs.","remediation":"Upgrade HBA Manager to 11.4.425.0 or 12.8.542.31 and above on every host with an Emulex HBA. The stronger and faster fix is to reinstall the agent in Strictly Local Management mode, which removes the remote listener entirely and leaves only local hbacmd - do this by default on bare-metal tenant nodes, since remote HBA management is rarely worth the exposure. Neither step requires an HBA firmware flash or an array outage; it is an agent reinstall and service restart. If a node was exposed, treat the HBA firmware as untrusted and reflash from the vendor image before returning it to the pool.","references":["https://docs.broadcom.com/doc/elx_HBAManager-Lin-RN12811-101.pdf","https://www.broadcom.com/products/storage/fibre-channel-host-bus-adapters/emulex-hba-manager","https://nvd.nist.gov/vuln/detail/CVE-2021-42774"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2021-11-12"},{"id":"CVE-2021-43215","cve":"CVE-2021-43215","aliases":["CVE-2017-0104"],"title":"Microsoft iSNS Server service (Internet Storage Name Service for iSCSI discovery): Memory corruption in the iSNS Server","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Microsoft iSNS Server service (Internet Storage Name Service for iSCSI discovery)","year":"2021","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Memory corruption in the iSNS Server service leading to remote code execution, with an earlier integer-overflow variant in the same service. iSNS is the discovery registry for an iSCSI fabric: it tells initiators which targets and portals exist and which discovery domains they belong to. Code execution there gives an attacker the ability to re-point initiators at targets of their choosing, which is a fabric-wide man-in-the-middle on block storage without ever touching a target or a host. Discovery domains are also the iSCSI analogue of FC zoning, so controlling iSNS is controlling who can see whose LUNs.","attack_vector":"Any host that can reach the iSNS server's registration port (TCP 3205) on the storage or management network. No authentication - iSNS has essentially none in common deployments.","remediation":"Patch the Windows host running the iSNS Server role and reboot. Better: most iSCSI deployments in a GPU datacenter do not need iSNS at all - initiators are configured with explicit target portals by the provisioning system. If that describes you, remove the iSNS Server role and strip the iSNS configuration from initiators; that removes an unauthenticated fabric-control service permanently and costs nothing operationally. If you do need it, put port 3205 behind an ACL that only initiator hosts can traverse.","references":["https://msrc.microsoft.com/update-guide/vulnerability/CVE-2021-43215","https://nvd.nist.gov/vuln/detail/CVE-2021-43215"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-12-15"},{"id":"CVE-2021-47215","cve":"CVE-2021-47215","aliases":["net/mlx5e kTLS crash in RX resync flow"],"title":"Linux kernel mlx5_core kTLS RX offload: TLS RX resync list corruption: entries are moved by the resync handler","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_core kTLS RX offload","year":"2021","cvss_score":9.8,"severity":"critical","kev":false,"impact":"TLS RX resync list corruption: entries are moved by the resync handler while still in use in NAPI, corrupting the list from the receive softirq. Remote and unauthenticated on any node using mlx5 hardware kTLS RX offload.","attack_vector":"Remote sender over a TLS connection handled by mlx5 kTLS RX offload, able to trigger resync conditions (out-of-order or retransmitted TLS records).","remediation":"Upgrade the host kernel to 5.16 or the 5.15.5 stable backport. Rolling reboot. Interim: disable kTLS RX offload (ethtool -K <dev> tls-hw-rx-offload off) as a live config change.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-47215","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2021/CVE-2021-47215.json"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-04-10"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2021-47378","cve":"CVE-2021-47378","aliases":[],"title":"Linux kernel (drivers/nvme/host): The NVMe/RDMA initiator destroys the queue pair before the connection manager ID, so","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/nvme/host)","year":"2021","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The NVMe/RDMA initiator destroys the queue pair before the connection manager ID, so RDMA CM events that arrive after the QP is gone land on freed memory. A peer that manipulates connection establishment gets a use-after-free inside the host's RDMA connection path - kernel memory corruption on a compute node driven from the fabric.","attack_vector":"This is initiator-side, driven by the remote end of the RDMA connection: a target (or anything that can answer/reject/stall CM traffic on the fabric) that forces errors during connection establishment causes CM events to be delivered after the QP teardown. It matters wherever your nodes connect NVMe/RDMA to a target you do not fully control, or where a tenant sits on the same RDMA fabric and can inject CM traffic. Requires nvme-rdma in use; no local privilege on the initiator is needed.","remediation":"No fixed version is listed on this record - boot a kernel carrying the linked stable commits. Interim: restrict which targets initiators may connect to, and segment the RDMA fabric so tenant workloads cannot reach the CM path used for storage connections.","references":["https://git.kernel.org/stable/c/ecf0dc5a904830c926a64feffd8e01141f89822f","https://git.kernel.org/stable/c/d268a182c56e8361e19fb781137411643312b994","https://nvd.nist.gov/vuln/detail/CVE-2021-47378"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-191","CWE-787"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2021-47496","cve":"CVE-2021-47496","aliases":[],"title":"Linux kernel (net/tls): KTLS stored a negative errno into the socket error field where a positive value is expected. A","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/tls)","year":"2021","cvss_score":9.8,"severity":"critical","kev":false,"impact":"KTLS stored a negative errno into the socket error field where a positive value is expected. A splice or sendfile consumer reads it back and negates it, so an error is interpreted as a positive count of bytes written. The pipe buffer length and the descriptor length both underflow, producing an enormous buffer offset and bogus addresses passed to the copy actor - memory corruption in a shared code path, triggered by nothing more than a TLS crypto request failing.","attack_vector":"Any unprivileged local process using kTLS with splice/sendfile-style zero-copy transmission - available to every tenant container, since kTLS needs no capability. The trigger is an encrypt request returning an error, which a tenant can precipitate by saturating the crypto engine or by using a crypto backend under memory pressure; a co-tenant supplying that pressure works just as well. No device node, no fabric peer required, though a peer driving the connection helps keep the splice loop active.","remediation":"Boot a kernel carrying the fix commits below (no fixed stable version published). This one is old enough that most maintained kernels already carry it - verify by commit rather than assume, particularly on long-lived vendor kernels. Interim control: avoid splice/sendfile on kTLS sockets, or blacklist the tls ULP.","references":["https://git.kernel.org/stable/c/e0cfd5159f314d6b304d030363650b06a2299cbb","https://git.kernel.org/stable/c/f3dec7e7ace38224f82cf83f0049159d067c2e19","https://nvd.nist.gov/vuln/detail/CVE-2021-47496"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787","CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2021-47536","cve":"CVE-2021-47536","aliases":[],"title":"Linux kernel (net/smc): The early link-group cleanup path deletes the list head instead of the link group, so the group","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/smc)","year":"2021","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The early link-group cleanup path deletes the list head instead of the link group, so the group stays on the global list and is then memset while still linked. Subsequent list operations write through a poisoned pointer - the published failure is a list-corruption BUG in the link-down worker, and the same defect gives an attacker who can force early link-group teardown a controlled-ish write into freed memory.","attack_vector":"Driven from the fabric: the cleanup path runs when link-group setup aborts early, which a peer can force by failing or aborting the CLC handshake, and the crash was observed from the smc_link_down worker (a fabric link event). Local tenants reach the same path through repeated AF_SMC connect attempts; socket(AF_SMC, ...) is unprivileged and autoloads the module.","remediation":"Boot a kernel carrying the fix commits. Interim: blacklist the smc module or block socket family 43 for tenants on nodes that do not intentionally run SMC-R.","references":["https://git.kernel.org/stable/c/77731fede297a23d26f2d169b4269466b2c82529","https://git.kernel.org/stable/c/789b6cc2a5f9123b9c549b886fdc47c865cfe0ba","https://nvd.nist.gov/vuln/detail/CVE-2021-47536"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2022-1388","cve":"CVE-2022-1388","aliases":["K23605346"],"title":"F5 BIG-IP (iControl REST): An unauthenticated attacker can send undisclosed requests to the iControl REST management","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"F5 BIG-IP (iControl REST)","year":"2022","cvss_score":9.8,"severity":"critical","kev":true,"impact":"An unauthenticated attacker can send undisclosed requests to the iControl REST management API and bypass authentication entirely, then run arbitrary system commands, create/delete files, or disable services — full compromise of the load balancer. This is one of the most widely exploited F5 bugs on record and is confirmed in CISA's KEV catalog; mass scanning for it started within days of disclosure.","attack_vector":"Remote and unauthenticated, but only if the iControl REST management interface is reachable from the attacker's network position — the standard defense-in-depth advice from F5 is that this interface should never be internet-facing, only reachable from a management network.","remediation":"Software upgrade to the fixed BIG-IP version per F5 K23605346, then reboot; if immediate patching isn't possible, restricting iControl REST access to a trusted management network (or disabling it on the self-IP/external interfaces) is the documented interim mitigation. Given confirmed mass exploitation, treat any unpatched device with an internet-reachable management plane as likely already compromised, not just vulnerable.","references":["https://support.f5.com/csp/article/K23605346","https://www.cisa.gov/known-exploited-vulnerabilities-catalog?field_cve=CVE-2022-1388"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2022-05-05"},{"id":"CVE-2022-20861","cve":"CVE-2022-20861","aliases":[],"title":"Cisco Nexus Dashboard (web UI / CSRF): One of a batch of unauthenticated flaws in Nexus Dashboard that together allow","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Cisco Nexus Dashboard (web UI / CSRF)","year":"2022","cvss_score":9.8,"severity":"critical","kev":false,"impact":"One of a batch of unauthenticated flaws in Nexus Dashboard that together allow remote command execution, reading and uploading container images, and CSRF. Nexus Dashboard is the pane of glass over your whole data-center fabric, so an attacker who lands here can push configuration and container workloads to every managed switch. Uploading a container image is the persistence path — it survives the dashboard being patched.","attack_vector":"Unauthenticated, remote to the Nexus Dashboard web interface. CSRF variant needs an admin to visit an attacker page while logged in.","remediation":"Upgrade the Nexus Dashboard cluster software. Application upgrade, no switch reload; the managed fabric keeps forwarding throughout. Re-verify every managed device's running image afterwards, since image upload was in scope.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-20861","https://nvd.nist.gov/vuln/detail/CVE-2022-20858","https://nvd.nist.gov/vuln/detail/CVE-2022-20857"],"status":"curated","published":"2022-07-21"},{"id":"CVE-2022-23676","cve":"CVE-2022-23676","aliases":[],"title":"ArubaOS-Switch (HPE Aruba wired switches): Remote arbitrary code execution on ArubaOS-Switch devices, affecting","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ArubaOS-Switch (HPE Aruba wired switches)","year":"2022","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Remote arbitrary code execution on ArubaOS-Switch devices, affecting every version of the 15.xx and early 16.xx trains. Aruba wired switches show up in GPU-cluster builds as the out-of-band management and provisioning fabric rather than the GPU data path — which makes RCE here a foothold on the network that reaches every BMC in the building.","attack_vector":"Remote to the switch's services. No credentials in the affected paths.","remediation":"ArubaOS-Switch firmware upgrade plus switch reload. Management switches are usually not redundant the way the GPU fabric is, so a reload means the OOB network drops — schedule it when you do not need remote console. Restrict management-plane reachability first.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-23676"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-05-10"},{"id":"CVE-2022-27518","cve":"CVE-2022-27518","aliases":[],"title":"Citrix ADC/Gateway: SAML SP/IdP config","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Citrix ADC/Gateway","year":"2022","cvss_score":9.8,"severity":"critical","kev":true,"impact":"SAML SP/IdP config -> unauthenticated remote arbitrary code execution (nation-state exploited)","attack_vector":"Network (remote)","remediation":"Control-plane: patch immediately","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-27518"],"status":"curated","published":"2022-12-13"},{"id":"CVE-2022-28163","cve":"CVE-2022-28163","aliases":[],"title":"Brocade SANnav Management Portal - Zone management endpoints, before SANnav 2.2.0: SQL injection in multiple endpoints","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Brocade SANnav Management Portal - Zone management endpoints, before SANnav 2.2.0","year":"2022","cvss_score":9.8,"severity":"critical","kev":false,"impact":"SQL injection in multiple endpoints associated with zone management, allowing arbitrary SQL against the SANnav database. SANnav is the single management plane for an entire Fibre Channel estate; its database holds the fabric inventory, the zoning configuration and the stored switch credentials. Arbitrary SQL there means reading every switch password SANnav holds and manipulating the zoning configuration it pushes - so one flaw in a management appliance converts into cross-tenant LUN exposure across every fabric it manages, without ever touching a switch directly.","attack_vector":"A user who can reach the SANnav web application. The zone-management endpoints sit behind the portal login, so realistically this is a low-privilege operator account, a stolen session, or an attacker who first used one of the SANnav authentication-bypass defects.","remediation":"Upgrade the SANnav Management Portal to 2.2.0 or later. This is an appliance/VM upgrade, not a switch firmware flash - no fabric downtime and no arrays offline, but it does take the management plane out for the duration. Afterwards, rotate every switch credential stored in SANnav, because the point of the bug is that they were readable. Keep SANnav off any network a tenant or a BMC can reach.","references":["https://www.broadcom.com/support/fibre-channel-networking/security-advisories/brocade-security-advisory-2022-1842","https://nvd.nist.gov/vuln/detail/CVE-2022-28163"],"status":"curated","published":"2022-05-06"},{"id":"CVE-2022-29264","cve":"CVE-2022-29264","aliases":[],"title":"coreboot 4.13-4.16 (SMM handling on application processors): Arbitrary code execution in System Management Mode","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"coreboot 4.13-4.16 (SMM handling on application processors)","year":"2022","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Arbitrary code execution in System Management Mode on application processors - every core other than the bootstrap processor. Scored critical. On a many-core GPU host that is a large number of entry points into ring -2, and SMM compromise means firmware persistence beneath the OS and the hypervisor, forged or suppressed boot measurements, and an implant that a reimage between tenants does not touch. Relevant to operators running open firmware stacks (OCP-style, Open System Firmware, or coreboot-based management and storage nodes) rather than vendor BIOS.","attack_vector":"Local attacker on the host able to reach SMM on a non-bootstrap core. Requires code execution on the node, not remote access.","remediation":"Rebuild and reflash coreboot at 4.17 or later - which for coreboot-based fleets is your own build pipeline rather than an OEM download, so the rebase lag is yours to control and can be much shorter than the IBV-to-OEM path. Firmware flash, one reboot per node. No config workaround. Note that coreboot does not publish a CVE-indexed advisory page and has no GitHub security advisories, so tracking its security fixes means watching commits and release notes directly rather than waiting for an advisory feed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-29264","https://github.com/advisories/GHSA-3295-v7jw-7p5g"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-04-25"},{"id":"CVE-2022-29502","cve":"CVE-2022-29502","aliases":[],"title":"Slurm: Incorrect access control leading to privilege escalation","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Slurm","year":"2022","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Incorrect access control leading to privilege escalation","attack_vector":"Any user who can submit a job","remediation":"Upgrade Slurm; restart the controller and all node daemons","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-29502"],"status":"curated","published":"2022-05-05"},{"id":"CVE-2022-32295","cve":"CVE-2022-32295","aliases":["AMP-SB-0002","Altra SPI-NOR SMC protection"],"title":"Ampere Altra and Altra Max UEFI reference design before SRP 1.09 - SMC interface exposing SPI-NOR flash: The OS","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Ampere Altra and Altra Max UEFI reference design before SRP 1.09 - SMC interface exposing SPI-NOR flash","year":"2022","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The OS or hypervisor can reach the SPI-NOR boot flash through an insufficiently protected SMC. That means whoever owns the kernel on an Altra box can rewrite platform firmware - a persistent, below-the-OS implant that survives reimaging, disk wipe and tenant handoff. For a bare-metal Arm GPU provider this is the canonical tenant-persistence failure: rent a node for an hour, own it for its service life. It also destroys any attestation story you have told customers.","attack_vector":"Any code at host kernel or hypervisor level on an Altra / Altra Max system - which, on bare-metal rental, means the tenant by design. No physical access needed.","remediation":"Update to Altra SRP 1.09 or later from the board OEM (the fix hardens the SMC so the non-secure world can no longer drive SPI-NOR). Flash + reboot + drain per node, and the OEM has to ship an SRP build for your specific board - Ampere publishes the reference, your ODM integrates it, so the lag is on them. Independently and more importantly: on bare-metal, verify boot flash contents against a golden image at every tenant handoff. Assume any node rented before the SRP update may already carry an implant and reflash it from an out-of-band path rather than trusting in-band verification.","references":["https://amperecomputing.com/products/security-bulletins/altra-spi-nor-smc.html","https://nvd.nist.gov/vuln/detail/CVE-2022-32295","https://amperecomputing.com/products/product-security"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-07-01"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-32554","cve":"CVE-2022-32554","aliases":[],"title":"Pure Storage Purity//FA and Purity//FB management interface (exposed credential): A password for the array's management","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Pure Storage Purity//FA and Purity//FB management interface (exposed credential)","year":"2022","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A password for the array's management interface may be known outside Pure Storage. Anyone holding it executes arbitrary instructions on the array as root, which is total control of the storage backing the cluster.","attack_vector":"Network reach to the management interface of an affected FlashArray or FlashBlade. No legitimate account is needed, because the credential is the vulnerability.","remediation":"Take the opt-in patch, apply the manual patch, or upgrade Purity//FA and Purity//FB to an unaffected release - all three routes are offered by Pure. Rotating your own admin passwords does not close this; the shipped credential has to be removed by the patch.","references":["https://support.purestorage.com/Pure_Security/Security_Bundle_2022-04-04/Security_Advisory_for_%E2%80%9Csecurity-bundle-2022-04-04","https://nvd.nist.gov/vuln/detail/CVE-2022-32554"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2022-33186","cve":"CVE-2022-33186","aliases":[],"title":"Brocade Fabric OS (unauthenticated remote code execution): Unauthenticated remote code execution on a Fibre Channel","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Brocade Fabric OS (unauthenticated remote code execution)","year":"2022","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unauthenticated remote code execution on a Fibre Channel switch running Fabric OS. Code execution on a SAN switch is total control of the storage fabric: zoning, LUN masking enforcement, and the path every host takes to its data. Affects v9.1.1, v9.0.1e, v8.2.3c, v7.4.2j and earlier.","attack_vector":"Unauthenticated, remote to the switch's management services.","remediation":"Fabric OS upgrade plus switch reboot, one fabric at a time so multipathing keeps hosts online. Restrict FOS management reachability to a dedicated OOB network as an immediate config control.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-33186"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2022-12-08"},{"id":"CVE-2022-40684","cve":"CVE-2022-40684","aliases":[],"title":"Fortinet FortiOS/FortiProxy: Auth bypass via an alternate path","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Fortinet FortiOS/FortiProxy","year":"2022","cvss_score":9.8,"severity":"critical","kev":true,"impact":"Auth bypass via an alternate path -> unauthenticated operations on the admin interface","attack_vector":"Network (remote)","remediation":"Control-plane: firmware + audit for attacker-added admin accounts and SSH keys","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-40684"],"status":"curated","published":"2022-10-18"},{"id":"CVE-2022-42475","cve":"CVE-2022-42475","aliases":[],"title":"Fortinet FortiOS: SSL-VPN heap-based buffer overflow","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Fortinet FortiOS","year":"2022","cvss_score":9.8,"severity":"critical","kev":true,"impact":"SSL-VPN heap-based buffer overflow -> unauthenticated RCE, exploited as a zero-day","attack_vector":"Network (remote)","remediation":"Control-plane: firmware upgrade; full IOC sweep","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42475"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-01-02"},{"id":"CVE-2022-42970","cve":"CVE-2022-42970","aliases":["SEVD-2022-256-01"],"title":"APC Easy UPS Online Monitoring Software (Windows and Windows Server): Critical functions in the UPS monitoring server","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"APC Easy UPS Online Monitoring Software (Windows and Windows Server)","year":"2022","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Critical functions in the UPS monitoring server are exposed with no authentication at all. This software is what issues graceful-shutdown commands to hosts and controls UPS behaviour - so an attacker who reaches it can trigger a coordinated shutdown of everything the software manages. On a GPU fleet that is a clean, deniable way to kill every running job at once, no memory corruption required.","attack_vector":"Unauthenticated, over the network to the monitoring server. This software typically runs on a Windows box on the facility or management network with wide reachability, because it needs to talk to every UPS and every managed host.","remediation":"Upgrade the monitoring software (SEVD-2022-256-01). This is a server-side upgrade rather than device firmware, so it is cheap - a reboot of one Windows host, not a maintenance window on the power train. The harder work is the network position: this box should not be reachable from the compute network or from tenant-facing subnets, and its shutdown-command channel should be authenticated.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42970"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["physical-impact"],"published":"2023-02-01"},{"id":"CVE-2022-42971","cve":"CVE-2022-42971","aliases":["SEVD-2022-256-01"],"title":"APC Easy UPS Online Monitoring Software (Windows and Windows Server): Unrestricted file upload leads to remote code","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"APC Easy UPS Online Monitoring Software (Windows and Windows Server)","year":"2022","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unrestricted file upload leads to remote code execution by dropping a JSP payload. Full control of the host that orchestrates UPS shutdowns across the site - which is the same PHYSICAL outcome as above, plus a persistent Windows foothold sitting on the facility network.","attack_vector":"Unauthenticated network access to the monitoring server's web component.","remediation":"Software upgrade per SEVD-2022-256-01, then treat the host as potentially compromised and rebuild it rather than patching in place if it was ever internet-reachable. Verify what the box could reach - it usually holds credentials for every UPS and every managed host it can shut down.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42971"],"status":"curated","published":"2023-02-01"},{"id":"CVE-2022-45907","cve":"CVE-2022-45907","aliases":[],"title":"PyTorch (`torch.jit.annotations.parse_type_line`): Arbitrary code execution via unsafe `eval` in TorchScript type","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"PyTorch (`torch.jit.annotations.parse_type_line`)","year":"2022","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Arbitrary code execution via unsafe `eval` in TorchScript type parsing","attack_vector":"Customer-supplied TorchScript model file","remediation":"Patch torch in base images; TorchScript ingest of untrusted models should be sandboxed regardless","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-45907"],"status":"curated","published":"2022-11-26"},{"id":"CVE-2022-46892","cve":"CVE-2022-46892","aliases":["AMP-SB-0006","root complex OS re-enable"],"title":"Ampere Altra and Altra Max before firmware 2.10c - PCIe root complex access control: The OS can re-initialise a PCIe","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Ampere Altra and Altra Max before firmware 2.10c - PCIe root complex access control","year":"2022","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The OS can re-initialise a PCIe root complex the platform firmware deliberately disabled. Operators disable root complexes to fence off slots - a management NIC, a device belonging to another tenant partition, a port that should not exist in this SKU. Re-enabling it hands the tenant a live PCIe path to hardware they were never allocated, and a PCIe device is a DMA master. In a GPU-passthrough fleet this is the difference between 'the tenant has the two GPUs we gave them' and 'the tenant can talk to whatever else is on the fabric'.","attack_vector":"Host kernel or hypervisor code on an Altra / Altra Max node. On bare-metal rental the tenant already has this. On a virtualised Arm host it needs a prior host-kernel compromise.","remediation":"Update Altra / Altra Max platform firmware to 2.10c or later via the board OEM. Flash + reboot + drain. There is no software workaround - the access control lives in firmware. Compensating control for anyone who cannot patch quickly: stop relying on 'firmware disabled that root complex' as an isolation boundary and physically depopulate or electrically isolate slots that must not be reachable.","references":["https://amperecomputing.com/products/security-bulletins/root-complex-OS-re-enable","https://nvd.nist.gov/vuln/detail/CVE-2022-46892"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2023-02-15"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-48673","cve":"CVE-2022-48673","aliases":[],"title":"Linux kernel (net/smc): When an SMC-R link is torn down, the kernel moves the QP to Error state and then destroys the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/smc)","year":"2022","cvss_score":9.8,"severity":"critical","kev":false,"impact":"When an SMC-R link is torn down, the kernel moves the QP to Error state and then destroys the QP and frees the link group without waiting for the outstanding receive work requests to flush. The RDMA completion tasklet then writes into freed link-group memory, giving a fabric peer that can force a link teardown a write-after-free in softirq context on the host - observed as a page fault inside a spinlock acquire, i.e. a node panic, with corruption of adjacent slab objects as the worse case.","attack_vector":"Reachable by any RDMA fabric peer holding an SMC-R link group with the node: the peer simply drives the link down (LLC delete-link, port flap, abrupt teardown) while traffic is in flight. On the victim side all that is required is that SMC-R is in use over an RoCE/IB device; socket(AF_SMC, ...) is unprivileged and autoloads the smc module through the net-pf-43 alias in default distro configs, so a tenant container needs no device node and no capability to bring the code path online.","remediation":"Boot a kernel carrying the fix commits (the record publishes no fixed version - check your distro's mapping for this CVE). Interim: prevent the family from loading at all with a modprobe.d entry (`blacklist smc` plus `install smc /bin/false`), or block socket family 43 in the tenant seccomp profile, so tenants cannot instantiate SMC-R link groups.","references":["https://git.kernel.org/stable/c/89fcb70f1acd6b0bbf2f7bfbf45d7aa75a9bdcde","https://git.kernel.org/stable/c/e9b1a4f867ae9c1dbd1d71cd09cbdb3239fb4968","https://nvd.nist.gov/vuln/detail/CVE-2022-48673"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-48686","cve":"CVE-2022-48686","aliases":[],"title":"Linux kernel NVMe-oF TCP host (nvme-tcp, digest error handling in io_work): The initiator kept reading from the socket","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel NVMe-oF TCP host (nvme-tcp, digest error handling in io_work)","year":"2022","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The initiator kept reading from the socket after deciding the TCP stream was already out of sync, producing a use-after-free on the compute node. This is the direction operators usually forget to threat-model: the storage target attacking its clients. A compromised or spoofed NVMe/TCP target - or anyone who can occupy that address on the storage network - gets kernel memory corruption on every GPU node that mounts from it, which is a fleet-wide blast radius from a single storage endpoint.","attack_vector":"Remote, from the target side. Requires being (or impersonating) the NVMe/TCP target the compute node connects to.","remediation":"Kernel update on compute nodes bailing out of the io_work loop once the stream is known bad. Structurally: authenticate the target, not just the initiator - most clusters configure host-NQN allow lists in one direction only and leave the initiator trusting whatever answers on the storage IP.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=13c80a6c112467bab5e44d090767930555fc17a5","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2022/CVE-2022-48686.json"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-48697","cve":"CVE-2022-48697","aliases":[],"title":"Linux kernel NVMe target core (nvmet, request completion during IO connect): KASAN-confirmed use-after-free reached","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel NVMe target core (nvmet, request completion during IO connect)","year":"2022","cvss_score":9.8,"severity":"critical","kev":false,"impact":"KASAN-confirmed use-after-free reached through nvmet_execute_io_connect - the I/O queue connect command. The request is completed and freed by the transport's queue_response callback, and nvmet_req_complete() then dereferences it. Connect is the command every initiator sends first, so the freed object is manipulated on the path that establishes a tenant's connection to shared namespaces. The kernel CNA scores it network, unauthenticated, full CIA.","attack_vector":"Remote and unauthenticated - the I/O connect path is exercised before any in-band authentication completes.","remediation":"Kernel update on target nodes. Restrict which initiators can reach the target and enforce host-NQN allow lists; neither closes the pre-auth window, but both shrink who can enter it.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=17f121ca3ec6be0fb32d77c7f65362934a38cc8e","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2022/CVE-2022-48697.json"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-48788","cve":"CVE-2022-48788","aliases":[],"title":"Linux kernel (drivers/nvme/host): On the NVMe/RDMA initiator, an async-event command can be submitted against an admin","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/nvme/host)","year":"2022","cvss_score":9.8,"severity":"critical","kev":false,"impact":"On the NVMe/RDMA initiator, an async-event command can be submitted against an admin queue that error recovery is concurrently destroying, giving a use-after-free in the host kernel. The node corrupts freed transport state and typically panics, dropping every tenant workload on it.","attack_vector":"Initiator-side, but the trigger is remote: error recovery only runs when the RDMA connection breaks, and the target (or a peer that can disrupt the fabric path) decides when that happens. If your compute nodes mount NVMe/RDMA volumes from targets a tenant controls or can influence, that peer can time connection drops against the controller's own AER traffic. Requires nvme-rdma in use; no local privilege needed on the initiator.","remediation":"No fixed version is listed on this record - boot a kernel carrying the linked stable commits. Interim: keep NVMe/RDMA initiator traffic on a fabric segment tenants cannot disturb, and drain nodes that show repeated transport error recovery.","references":["https://git.kernel.org/stable/c/5593f72d1922403c11749532e3a0aa4cf61414e9","https://git.kernel.org/stable/c/d411b2a5da68b8a130c23097014434ac140a2ace","https://nvd.nist.gov/vuln/detail/CVE-2022-48788"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-48789","cve":"CVE-2022-48789","aliases":[],"title":"Linux kernel (drivers/nvme/host): Same race as the RDMA variant but on NVMe/TCP, which is the far more common fabric in","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/nvme/host)","year":"2022","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Same race as the RDMA variant but on NVMe/TCP, which is the far more common fabric in a mixed cluster - an async-event submission can reach an admin queue whose socket error recovery has already released, giving a use-after-free and a kernel panic on the compute node.","attack_vector":"Initiator-side, triggered whenever the TCP connection to the target drops while the host has an AER outstanding. The remote target controls both halves: it sends the async event notification and it can close or stall the connection. This is a live concern if the operator connects initiators to tenant-controlled or shared NVMe/TCP targets, or if a tenant can interfere with the storage path on the network. Requires nvme-tcp in use.","remediation":"No fixed version is listed on this record - boot a kernel carrying the linked stable commits. Interim: pin NVMe/TCP storage traffic to a network tenants cannot reach or disrupt, and only connect initiators to targets you operate.","references":["https://git.kernel.org/stable/c/61a26ffd5ad3ece456d74c4c79f7b5e3f440a141","https://git.kernel.org/stable/c/e192184cf8bce8dd55d619f5611a2eaba996fa05","https://nvd.nist.gov/vuln/detail/CVE-2022-48789"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-49003","cve":"CVE-2022-49003","aliases":[],"title":"Linux kernel (drivers/nvme/host): The multipath sibling list is walked without SRCU protection during path revalidation","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/nvme/host)","year":"2022","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The multipath sibling list is walked without SRCU protection during path revalidation while concurrent scan work frees namespaces out from under it, giving a use-after-free on the initiator. The reported panic came from a live NVMe/RDMA multipath setup - the node dies and takes every tenant on it.","attack_vector":"Initiator-side, driven by the remote target: namespace-change async event notifications are what schedule scan work, and the target decides when and how often to send them. A target that a tenant controls, or one an attacker has compromised, can spray namespace-change AENs and capacity changes to force the concurrent revalidate/remove race. Requires native NVMe multipath enabled (the default) with a fabric transport.","remediation":"No fixed version is listed on this record - boot a kernel carrying the linked stable commits. Interim: only connect initiators to targets you operate, and treat a node emitting repeated capacity-change/rescan events as suspect - drain it.","references":["https://git.kernel.org/stable/c/787d81d4eb150e443e5d1276c6e8f03cfecc2302","https://git.kernel.org/stable/c/5b566d09ab1b975566a53f9c5466ee260d087582","https://nvd.nist.gov/vuln/detail/CVE-2022-49003"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-125"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-49094","cve":"CVE-2022-49094","aliases":[],"title":"Linux kernel (net/tls): KTLS allocates a 12-byte IV buffer for AES-128-CCM but the decrypt path copies 16 bytes out of","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/tls)","year":"2022","cvss_score":9.8,"severity":"critical","kev":false,"impact":"KTLS allocates a 12-byte IV buffer for AES-128-CCM but the decrypt path copies 16 bytes out of it, so every inbound record over-reads four bytes of adjacent slab memory into the AEAD input. Neighbouring heap contents feed the crypto operation and a remote peer holds a repeatable slab-overread primitive against the node terminating its connection.","attack_vector":"Reachable by any unprivileged process on the node that owns a TCP socket - no device node required. setsockopt(SOL_TLS, TLS_RX, TLS_CIPHER_AES_CCM_128) with TLS 1.3, then receive records from the peer. Applies to tenant workloads using kTLS and to any storage or control-plane daemon on the node configured for CCM; the peer supplying the records drives the overread.","remediation":"Boot a kernel carrying the linked stable commits. Interim: restrict kTLS ciphersuites to AES-GCM in the daemons you control and disable AES-CCM-128 offload, or disable the tls ULP (blacklist the tls module) where it is not required.","references":["https://git.kernel.org/stable/c/2b7d14c105dd8f6412eda5a91e1e6154653731e3","https://git.kernel.org/stable/c/589154d0f18945f41d138a5b4e49e518d294474b","https://nvd.nist.gov/vuln/detail/CVE-2022-49094"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-191"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-49280","cve":"CVE-2022-49280","aliases":[],"title":"Linux NFS server (nfsd, nfssvc_decode_writeargs): The NFSv2/v3 write argument decoder has no lower bound on the length","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux NFS server (nfsd, nfssvc_decode_writeargs)","year":"2022","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The NFSv2/v3 write argument decoder has no lower bound on the length field, so a negative value underflows and the server reads or writes outside the intended buffer. Unauthenticated remote memory corruption in the kernel on the file server.","attack_vector":"Any host able to send NFS RPC to the server. No authentication needed.","remediation":"Update the storage server kernel and reboot. If your workloads only need NFSv4, disable v2/v3 in /etc/nfs.conf (vers2=n, vers3=n) to remove this decoder from the reachable surface.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49280","https://git.kernel.org/pub/scm/linux/security/vulns.git"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-49356","cve":"CVE-2022-49356","aliases":[],"title":"Linux SUNRPC / NFS-over-RDMA server (svc_rdma_build_writes): svc_rdma_build_writes can walk off the end of a Write","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux SUNRPC / NFS-over-RDMA server (svc_rdma_build_writes)","year":"2022","cvss_score":9.8,"severity":"critical","kev":false,"impact":"svc_rdma_build_writes can walk off the end of a Write chunk's segment array, caught by KASAN as an out-of-bounds access. A client that supplies a crafted RDMA Write chunk corrupts kernel memory on the NFS server.","attack_vector":"Any NFS/RDMA client on the storage fabric.","remediation":"Update the storage server kernel and reboot. Falling back to NFS over TCP removes the RDMA chunk path entirely while you schedule the reboot.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git","https://nvd.nist.gov/vuln/detail/CVE-2022-49356"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-50666","cve":"CVE-2022-50666","aliases":[],"title":"Linux kernel (drivers/infiniband/sw/siw): A remote peer turns a connection drop into a kernel use-after-free. The","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/sw/siw)","year":"2022","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A remote peer turns a connection drop into a kernel use-after-free. The driver lets queue-pair destruction return while its own connection-manager work still holds references, so the RDMA core frees the QP and the late work handler then operates on freed memory. Rated critical and network-reachable by the kernel CNA.","attack_vector":"A peer on the fabric resets or drops the TCP connection at the moment the local side is tearing the QP down - the upstream report comes from ordinary NFS-over-RDMA testing, so no exotic crafting is needed, and no local credentials are involved. Requires the siw (soft-iWARP) module.","remediation":"No fixed release is published in this record - apply the listed stable fix commits or run a current stable kernel. Interim: unload/blacklist siw unless soft-iWARP is deliberately in use, and limit which peers can open iWARP connections to the node.","references":["https://git.kernel.org/stable/c/5c75d608fad58301b63e7d69200c13c3a1d411da","https://git.kernel.org/stable/c/74ad141e995a730760b1bcfa14854b7f1057d6bc","https://nvd.nist.gov/vuln/detail/CVE-2022-50666"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2022-50717","cve":"CVE-2022-50717","aliases":["nvmet-tcp Transfer Tag bounds check","NVMe/TCP H2C ttag out-of-bounds"],"title":"Linux kernel - NVMe-oF TCP target, drivers/nvme/target/tcp.c: The NVMe/TCP target used the host-supplied Transfer Tag","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel - NVMe-oF TCP target, drivers/nvme/target/tcp.c","year":"2022","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The NVMe/TCP target used the host-supplied Transfer Tag directly as an array index to look up the command structure, with no bounds check. A connected initiator sets an arbitrary ttag in an H2C Data PDU and the target reads and operates on memory outside the command array - remote out-of-bounds access on the storage node with full confidentiality, integrity and availability impact. Because NVMe/TCP requires no authentication by default, the 'connected initiator' bar is effectively 'anyone who can reach port 4420'. This one has been in shipping kernels since NVMe/TCP target support landed and was only assigned a CVE retroactively, so long-lived storage nodes are the ones to check.","attack_vector":"Open an NVMe/TCP connection to the target and send an H2CData PDU with an out-of-range Transfer Tag. No authentication needed unless DH-HMAC-CHAP has been explicitly configured. Reachable across any routed path to the target port.","remediation":"Host reboot / kernel upgrade - and specifically check long-lived storage nodes, since the fix was backported late and a node that has not been rebooted in a year may still be exposed. Interim: firewall NVMe/TCP 4420 to known initiators and enable in-band DH-HMAC-CHAP on a patched kernel so the connection itself requires credentials. Roll targets in waves behind multipath so tenants see no outage.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2022/CVE-2022-50717.json","https://nvd.nist.gov/vuln/detail/CVE-2022-50717"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-12-24"},{"id":"CVE-2023-0052","cve":"CVE-2023-0052","aliases":["CVE-2023-0053","ICSA-23-012-05"],"title":"SAUTER Controls Nova 200-220 series (firmware <=3.3-006) with BACnetstac <=4.2.1: Commands execute with no credentials","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"SAUTER Controls Nova 200-220 series (firmware <=3.3-006) with BACnetstac <=4.2.1","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Commands execute with no credentials at all, and the only management protocols the device offers are Telnet and FTP - both cleartext. An unauthorized user can log in, change the device configuration and modify the control program. These are HVAC automation stations; whoever holds them holds the air handling and chilled-water sequences they run. The cleartext management path compounds it: any credentials that do exist elsewhere in the building, entered through these devices, are recoverable by passive sniffing on the facility VLAN, which turns one weak controller into a credential source for the rest of the BMS. In a GPU hall the direct consequence is loss of thermal control with a hardware-damage tail; the indirect one is that your entire building-controls credential set should be considered compromised if these devices are present and the network is shared.","attack_vector":"Unauthenticated Telnet/FTP from anywhere on the facility network. There is nothing to bypass. Passive sniffing on the same segment additionally yields any credentials in flight. Internet exposure of Telnet on building controllers is a recurring Shodan finding, so check your external surface for port 23 as well.","remediation":"SAUTER's guidance is upgrade to fixed firmware where available and otherwise disable the affected services - but on this generation, disabling Telnet and FTP removes the only management path the device has, which is why sites leave them on. Treat this as effectively unpatchable in place: the durable fix is replacing the controller generation, a capital project with a contractor and per-device downtime. Interim controls: strict VLAN isolation with an allow-list from the supervisor only, switch ACLs blocking 21/23 from everything else, and physical security on the panels. Leased colo: this is landlord equipment, so the honest remediation is contractual - require disclosure of controller make/model/firmware across the mechanical plant and the right to audit the segment.","references":["https://www.cisa.gov/news-events/ics-advisories/icsa-23-012-05","https://nvd.nist.gov/vuln/detail/CVE-2023-0052","https://nvd.nist.gov/vuln/detail/CVE-2023-0053"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2023-20596","cve":"CVE-2023-20596","aliases":[],"title":"AMD SMM Supervisor (AMD-SB-7011): The highest-scored AMD platform CVE in this database at 9.8 critical. A flaw in the","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD SMM Supervisor (AMD-SB-7011)","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The highest-scored AMD platform CVE in this database at 9.8 critical. A flaw in the AMD SMM Supervisor - the component that is supposed to *constrain* what SMM code can do - yields full compromise of confidentiality, integrity and availability. SMM sits above the hypervisor and can reach all physical memory; owning it means owning every VM, container and confidential guest on the node, persistently and invisibly to anything running above.","attack_vector":"NVD scores this as network-reachable with no privileges required, which is unusually severe for an SMM issue and worth treating at face value until you can prove otherwise for your platform. AMD's own framing is narrower. Given the disagreement, patch first and reconcile the vector later.","remediation":"Fixed in AMD PI/AGESA firmware and delivered only as an OEM SBIOS package - AMD ships the PI drop to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before releasing BIOS. **Budget one to six months of OEM lag**, and note that several CVEs in this batch are marked 'no fix planned' on Naples (EPYC 7001) - for those the only remediation is retiring the hardware. Applying it means cordon, drain and a full power cycle per node; there is no driver reload, no live patch and no VBIOS step. Treat as the top of your AMD firmware queue purely on score and blast radius. If your OEM has not shipped a BIOS carrying it, escalate with them rather than waiting on the normal cycle.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-20596","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-7011.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"tags":["tenant-isolation"],"published":"2023-11-14"},{"id":"CVE-2023-2780","cve":"CVE-2023-2780","aliases":[],"title":"MLflow: Path traversal prior to 2.3.1","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"MLflow","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Path traversal prior to 2.3.1","attack_vector":"Unauthenticated network","remediation":"Upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-2780"],"status":"curated","published":"2023-05-17"},{"id":"CVE-2023-27997","cve":"CVE-2023-27997","aliases":["FG-IR-23-097","XORtigate"],"title":"Fortinet FortiOS / FortiProxy SSL-VPN: A heap-based buffer overflow in the SSL-VPN daemon lets a remote","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Fortinet FortiOS / FortiProxy SSL-VPN","year":"2023","cvss_score":9.8,"severity":"critical","kev":true,"impact":"A heap-based buffer overflow in the SSL-VPN daemon lets a remote, unauthenticated attacker run arbitrary code on the FortiGate — full device compromise. This is the 'XORtigate' bug, confirmed in CISA's KEV catalog as actively exploited; if this FortiGate is the VPN gateway into your cluster's management network, an attacker doesn't need any credentials to get a foothold there.","attack_vector":"Remote, unauthenticated — a specifically crafted request to the SSL-VPN service is sufficient, no login required.","remediation":"Firmware upgrade of FortiOS/FortiProxy to the fixed release per Fortinet PSIRT FG-IR-23-097. Given confirmed active exploitation, patch immediately rather than waiting for a scheduled window, and assume compromise on any internet-facing unit that was unpatched during the exploitation window — a reboot alone doesn't remediate a box that was already popped.","references":["https://fortiguard.com/psirt/FG-IR-23-097","https://www.cisa.gov/known-exploited-vulnerabilities-catalog?field_cve=CVE-2023-27997"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-06-13"},{"id":"CVE-2023-29374","cve":"CVE-2023-29374","aliases":[],"title":"LangChain (`LLMMathChain`): Prompt injection","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LangChain (`LLMMathChain`)","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Prompt injection → arbitrary Python execution","attack_vector":"Untrusted prompt or retrieved document reaching an agent running on the GPU node","remediation":"No patch for the pattern — LLM output feeding `exec` is the design. Sandbox agent execution; provider isolates the node's credentials","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-29374"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"published":"2023-04-05"},{"id":"CVE-2023-31024","cve":"CVE-2023-31024","aliases":[],"title":"DGX A100 BMC: RCE on BMC (stack buffer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX A100 BMC","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"RCE on BMC (stack buffer overflow)","attack_vector":"Network-adjacent mgmt-LAN attacker","remediation":"Flash BMC 00.22.05+ out-of-band immediately; isolate BMC VLAN","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31024","https://github.com/NVIDIA/product-security/tree/main/2024/5510"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-121"],"published":"2024-01-12"},{"id":"CVE-2023-31030","cve":"CVE-2023-31030","aliases":[],"title":"DGX A100 BMC: RCE on BMC (stack buffer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX A100 BMC","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"RCE on BMC (stack buffer overflow)","attack_vector":"Network-adjacent mgmt-LAN attacker","remediation":"Flash BMC 00.22.05+ out-of-band","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31030","https://github.com/NVIDIA/product-security/tree/main/2024/5510"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-121"],"published":"2024-01-12"},{"id":"CVE-2023-32484","cve":"CVE-2023-32484","aliases":[],"title":"Dell Enterprise SONiC OS (input validation): Improper input validation on Dell Networking switches running Enterprise","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell Enterprise SONiC OS (input validation)","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Improper input validation on Dell Networking switches running Enterprise SONiC, exploitable by a remote unauthenticated attacker. Affects 4.1.0, 4.0.5, 3.5.4 and below — i.e. essentially every SONiC release before the 2024 hardening pass. If you bought Dell switches with SONiC for a cost-optimised GPU buildout in 2022-2023 and have not touched the NOS since, this is live.","attack_vector":"Unauthenticated, remote to the switch.","remediation":"NOS image upgrade and switch reboot. On SONiC that is a full image install, so the switch is down for minutes — stage across MLAG pairs. Put SONiC management interfaces on an isolated OOB network as the standing control.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-32484"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-02-15"},{"id":"CVE-2023-32485","cve":"CVE-2023-32485","aliases":[],"title":"Dell SmartFabric Storage Software: Improper input validation in Dell SmartFabric Storage Software 1.3 and lower","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell SmartFabric Storage Software","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Improper input validation in Dell SmartFabric Storage Software 1.3 and lower, exploitable by a remote unauthenticated attacker. SmartFabric Storage Software is the NVMe-over-TCP fabric controller — it performs the discovery and zoning that decides which host initiators can see which NVMe subsystems. Compromising it is compromising the storage access-control layer for the whole cluster. Companion unauthenticated command injection: CVE-2022-31232.","attack_vector":"Unauthenticated, remote to the SmartFabric Storage Software service.","remediation":"Upgrade the SmartFabric Storage Software appliance/VM past 1.3 (and past 1.4 for the CVE-2023-4306x set). Application upgrade with a service restart; NVMe-oF sessions reconnect. Afterwards, re-verify the zoning database against intent, because an attacker with control here would change exactly that.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-32485","https://nvd.nist.gov/vuln/detail/CVE-2022-31232"],"status":"curated","fleet":{"pain_class":"daemon-restart"},"published":"2023-10-05"},{"id":"CVE-2023-3265","cve":"CVE-2023-3265","aliases":["ZDI-23-1147"],"title":"CyberPower PowerPanel Enterprise DCIM - username handling: Authentication bypass: appending a non-printable character","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"CyberPower PowerPanel Enterprise DCIM - username handling","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Authentication bypass: appending a non-printable character to the built-in 'cyberpower' username logs an attacker straight in. Unauthenticated to full DCIM administrator, with no exploit development required. PowerPanel Enterprise manages UPS and PDU estates, so this is direct PHYSICAL exposure of the power layer.","attack_vector":"Unauthenticated, remote, against the PowerPanel Enterprise login. Anyone who can reach the web interface.","remediation":"Upgrade PowerPanel Enterprise. Software upgrade on one host - genuinely cheap. Then check whether the default 'cyberpower' account exists at all and remove it. Get the DCIM off any network a tenant workload can route to.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-3265"],"status":"curated","published":"2023-08-14"},{"id":"CVE-2023-3266","cve":"CVE-2023-3266","aliases":["ZDI-23-1148"],"title":"CyberPower PowerPanel Enterprise DCIM - LDAP authentication path: If LDAP authentication is selected","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"CyberPower PowerPanel Enterprise DCIM - LDAP authentication path","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"If LDAP authentication is selected, the authentication mechanism is incomplete and every check can be bypassed. The operators most likely to be affected are the mature ones - the ones who wired their DCIM into corporate directory rather than using local accounts. Doing the responsible thing is what turns the bug on.","attack_vector":"Unauthenticated, remote, on any PowerPanel Enterprise instance configured for LDAP.","remediation":"Upgrade PowerPanel Enterprise immediately. As an interim, switching off LDAP mode removes the vulnerable path but costs you central account control - a genuinely unpleasant trade, so prioritise the upgrade.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-3266"],"status":"curated","published":"2023-08-14"},{"id":"CVE-2023-34048","cve":"CVE-2023-34048","aliases":[],"title":"VMware vCenter: Out-of-bounds write in the DCERPC implementation - unauthenticated remote code execution","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"VMware vCenter","year":"2023","cvss_score":9.8,"severity":"critical","kev":true,"impact":"Out-of-bounds write in the DCERPC implementation - unauthenticated remote code execution; exploited as a zero-day by UNC3886 since at least 2021 [KEV]","attack_vector":"Unauthenticated network to the management plane","remediation":"vCenter patch + service restart. Assume compromise on any vCenter that was internet- or tenant-reachable before Oct 2023","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-34048"],"status":"curated","fleet":{"pain_class":"daemon-restart"},"published":"2023-10-25"},{"id":"CVE-2023-34362","cve":"CVE-2023-34362","aliases":[],"title":"Progress MOVEit Transfer: Unauthenticated SQL injection into the web app","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Progress MOVEit Transfer","year":"2023","cvss_score":9.8,"severity":"critical","kev":true,"impact":"Unauthenticated SQL injection into the web app -> DB access and RCE (Cl0p mass exploitation)","attack_vector":"Network (remote)","remediation":"Control-plane: patch or decommission; assume data exfiltration if it was internet-facing","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-34362"],"status":"curated","published":"2023-06-02"},{"id":"CVE-2023-3519","cve":"CVE-2023-3519","aliases":[],"title":"Citrix NetScaler ADC/Gateway: Unauthenticated remote code execution on the gateway appliance","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Citrix NetScaler ADC/Gateway","year":"2023","cvss_score":9.8,"severity":"critical","kev":true,"impact":"Unauthenticated remote code execution on the gateway appliance","attack_vector":"Network (remote)","remediation":"Control-plane: emergency patch plus a webshell hunt","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-3519"],"status":"curated","published":"2023-07-19"},{"id":"CVE-2023-35861","cve":"CVE-2023-35861","aliases":[],"title":"Supermicro BMC email/SMTP alert notification handler (H12DST-B): Command execution as root on the BMC, reached through","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC email/SMTP alert notification handler (H12DST-B)","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Command execution as root on the BMC, reached through a feature nearly every operator turns on because they want temperature and PSU alerts. Root on the BMC is total control of the node's out-of-band plane: power, boot device, serial and graphical console, and the ability to write persistent code that outlives any host reinstall. On a rented bare-metal GPU node it also means the previous tenant's implant can be watching the next tenant's console. Firmware 03.10.35 and present across the same BMC codebase on other Supermicro boards. User-controlled notification fields reach a shell without sanitisation.","attack_vector":"Reachable over the network to the BMC. The alerting configuration surface is exactly the kind of thing left enabled and reachable from the monitoring VLAN, so an attacker who compromises a monitoring or DCIM host is already in position.","remediation":"Firmware flash to 03.10.35 or later per board SKU, from Supermicro's June 2023 SMTP advisory. As an immediate config-only stopgap you can disable BMC email alerting entirely, which removes the vulnerable path at the cost of losing hardware alerts - acceptable for a few days, not as a permanent posture. Because this bug is in shared Supermicro BMC code rather than one board's, audit the whole fleet rather than only H12DST-B.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-35861","https://blog.freax13.de/cve/cve-2023-35861","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2023/35xxx/CVE-2023-35861.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-36258","cve":"CVE-2023-36258","aliases":[],"title":"LangChain (PALChain): Arbitrary code execution via `os.system`/`exec` in generated code","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LangChain (PALChain)","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Arbitrary code execution via `os.system`/`exec` in generated code","attack_vector":"Untrusted prompt input","remediation":"Upgrade past 0.0.236; the fix was bypassed twice (see below)","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-36258"],"status":"curated","published":"2023-07-03"},{"id":"CVE-2023-36281","cve":"CVE-2023-36281","aliases":[],"title":"LangChain (`load_prompt`): Arbitrary code execution from a JSON prompt file","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LangChain (`load_prompt`)","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Arbitrary code execution from a JSON prompt file","attack_vector":"Customer-supplied prompt file","remediation":"Upgrade; prompt files are executable content","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-36281"],"status":"curated","published":"2023-08-22"},{"id":"CVE-2023-36845","cve":"CVE-2023-36845","aliases":[],"title":"Juniper Junos OS J-Web (EX/SRX): Unauthenticated remote code execution by setting `PHPRC` through a crafted J-Web","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Juniper Junos OS J-Web (EX/SRX)","year":"2023","cvss_score":9.8,"severity":"critical","kev":true,"impact":"Unauthenticated remote code execution by setting `PHPRC` through a crafted J-Web request; chained with the other J-Web bugs for full device takeover","attack_vector":"Network, unauthenticated","remediation":"Junos upgrade or disabling J-Web entirely; on a management-plane switch the fastest mitigation is turning J-Web off, which removes the GUI ops staff depend on","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-36845"],"status":"curated","published":"2023-08-17"},{"id":"CVE-2023-38408","cve":"CVE-2023-38408","aliases":[],"title":"OpenSSH (ssh-agent): Remote code execution in ssh-agent PKCS#11 support when agent forwarding reaches a hostile host","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"OpenSSH (ssh-agent)","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Remote code execution in ssh-agent PKCS#11 support when agent forwarding reaches a hostile host","attack_vector":"Unauthenticated network (against a forwarded agent)","remediation":"Package update; restart agents. Policy fix: forbid agent forwarding into tenant-reachable bastions","references":["https://access.redhat.com/security/cve/CVE-2023-38408"],"status":"curated","published":"2023-07-20"},{"id":"CVE-2023-39281","cve":"CVE-2023-39281","aliases":["INSYDE-SA-2023054"],"title":"Insyde InsydeH2O (AsfSecureBootDxe): Stack buffer overflow leading to arbitrary code execution during the DXE phase","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (AsfSecureBootDxe)","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Stack buffer overflow leading to arbitrary code execution during the DXE phase - and it lives in the driver responsible for Secure Boot handling for ASF (Alert Standard Format, the out-of-band manageability path). Code execution in DXE means running before the OS with firmware privileges and with Secure Boot policy still under the attacker's influence. Scored critical, and the placement inside the Secure Boot path is what makes it worse than the raw score suggests.","attack_vector":"Attacker able to supply the oversized input the DXE driver parses during boot. Given the ASF/manageability association, treat anything that can reach the platform's out-of-band alerting path as in scope alongside local OS-level access.","remediation":"OEM BIOS update on the fixed Insyde kernel (5.0-5.5 affected). Firmware flash, reboot per node. No config workaround inside firmware; on the network side, keep the BMC and manageability interfaces on an isolated management VLAN with no route from tenant or provisioning networks - which is good practice independent of this CVE and materially reduces who can reach the ASF path.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-39281","https://www.insyde.com/security-pledge/SA-2023054"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-11-01"},{"id":"CVE-2023-41910","cve":"CVE-2023-41910","aliases":[],"title":"lldpd (CDP PDU parser, cdp_decode): A crafted CDP PDU with specific CDP_TLV_ADDRESSES TLVs forces lldpd","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"lldpd (CDP PDU parser, cdp_decode)","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A crafted CDP PDU with specific CDP_TLV_ADDRESSES TLVs forces lldpd into an out-of-bounds heap read. lldpd is what runs on Linux-based switch OSes and on servers that advertise link topology, it runs as root, and it accepts input from any directly attached device with no authentication whatsoever. In a GPU cluster where LLDP is used to verify rail-optimized cabling, lldpd is running on every node and every switch.","attack_vector":"Unauthenticated, adjacent — a single crafted frame from a directly connected device. Any tenant bare-metal node can attack the switch or the neighbours it is cabled to.","remediation":"Upgrade lldpd to 1.0.17 or later and restart the daemon — a package upgrade with a service restart, no reboot, no switch reload. On appliance NOSes this arrives as a NOS image update instead. Cheap fix; the reason it lingers is that nobody inventories lldpd versions.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-41910"],"status":"curated","fleet":{"pain_class":"daemon-restart"},"published":"2023-09-05"},{"id":"CVE-2023-42793","cve":"CVE-2023-42793","aliases":[],"title":"JetBrains TeamCity: Authentication bypass leading to remote code execution on TeamCity Server","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"JetBrains TeamCity","year":"2023","cvss_score":9.8,"severity":"critical","kev":true,"impact":"Authentication bypass leading to remote code execution on TeamCity Server","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; treat previously built artifacts as suspect","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-42793"],"status":"curated","published":"2023-09-19"},{"id":"CVE-2023-4323","cve":"CVE-2023-4323","aliases":["CVE-2023-4324","CVE-2023-4325","CVE-2023-4326","CVE-2023-4329","CVE-2023-4331","CVE-2023-4332","CVE-2023-4333","CVE-2023-4345"],"title":"Broadcom LSI Storage Authority (LSA) / Intel RAID Web Console 3 (RWC3)","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Broadcom LSI Storage Authority (LSA) / Intel RAID Web Console 3 (RWC3) - management service for MegaRAID and LSI HBA…","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"LSA is the agent that fronts the MegaRAID/LSI controller on every server that has one, and it runs as root/SYSTEM with the ability to create, delete and re-initialize virtual drives and to push controller firmware. Sessions are not invalidated properly in Gateway (multi-node) setup, so an attacker who reaches the LSA web service can ride an admin session and drive the RAID controller on any managed node: wipe or re-create arrays under a running tenant, or stage a controller firmware update. Because the controller sits below the OS with DMA to host memory, a firmware push through this path is below-the-OS persistence that a tenant reimage does not remove - it breaks tenant handoff. The sibling issues in the same disclosure are the supporting weaknesses: a bundled vulnerable libcurl, no CSP, no SameSite on the session cookie, SHA-1 ciphersuites and obsolete TLS versions on the management listener, and world-readable log files.","attack_vector":"Any host that can reach the LSA HTTPS listener (default TCP 2463) on the management or provisioning network. In Gateway mode one LSA instance manages many nodes, so one reachable management endpoint fans out to the whole fleet. No valid credentials are needed to abuse the session-handling flaw; the weak TLS and cookie defaults widen it to on-path and browser-side attackers.","remediation":"Software-only: upgrade LSA / Intel RWC3 to 7.017.011.000 or later on every managed node and on the Gateway. No controller firmware flash and no array downtime - restarting the LSA service is enough. The real cost is that this agent is installed by OEM tooling on every server with a Broadcom controller, and Dell/HPE/Supermicro/Intel ship their own rebadged LSA build months behind Broadcom, so you often cannot take the upstream package. Interim control: bind LSA to loopback or a dedicated management VLAN and firewall 2463 off the tenant and provisioning networks; if you do not use the web UI, uninstall LSA and drive the controller with StorCLI from a config-managed path instead.","references":["https://www.broadcom.com/support/resources/product-security-center","https://nvd.nist.gov/vuln/detail/CVE-2023-4323","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00926.html"],"status":"curated","published":"2023-08-15"},{"id":"CVE-2023-43845","cve":"CVE-2023-43845","aliases":[],"title":"ATEN PE6208 switched PDU: The PDU ships with a default telnet account and never forces the operator to change","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ATEN PE6208 switched PDU","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The PDU ships with a default telnet account and never forces the operator to change it on first login. Anyone who finds an un-rotated unit gets an administrator telnet session — full outlet control, including turning power off to whatever racks that PDU feeds.","attack_vector":"Network reachability to the PDU's telnet service plus knowledge of the published default credential; no exploit development needed.","remediation":"Credential rotation is the immediate fix — audit every deployed PE6208 for the default telnet account and change it now. ATEN's firmware update additionally forces a credential change on first login for new deployments, so pair the audit with a firmware upgrade where feasible.","references":["https://github.com/setersora/pe6208"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-05-28"},{"id":"CVE-2023-44467","cve":"CVE-2023-44467","aliases":[],"title":"langchain-experimental (PALChain): Bypass of the CVE-2023-36258 fix","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"langchain-experimental (PALChain)","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Bypass of the CVE-2023-36258 fix","attack_vector":"Untrusted prompt input","remediation":"Upgrade past 0.0.306","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-44467"],"status":"curated","published":"2023-10-09"},{"id":"CVE-2023-45249","cve":"CVE-2023-45249","aliases":[],"title":"Acronis Cyber Infrastructure: Default passwords","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Acronis Cyber Infrastructure","year":"2023","cvss_score":9.8,"severity":"critical","kev":true,"impact":"Default passwords -> unauthenticated remote command execution","attack_vector":"Network (remote)","remediation":"Control-plane: patch + change all default credentials on the storage/backup appliance","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-45249"],"status":"curated","published":"2024-07-24"},{"id":"CVE-2023-48022","cve":"CVE-2023-48022","aliases":["ShadowRay"],"title":"Ray (job submission API): Unauthenticated RCE — the Jobs API accepts arbitrary code by design","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ray (job submission API)","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unauthenticated RCE — the Jobs API accepts arbitrary code by design","attack_vector":"Unauthenticated network to the Ray dashboard/Jobs API (default 8265)","remediation":"**No patch — vendor disputes it.** The only remediation is network isolation and an auth proxy. Exploited in the wild against GPU clusters; a neocloud must never let 8265 reach a tenant or public network","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-48022"],"status":"curated","fleet":{"ubiquity":"Very common - Ray is the default distributed-compute layer for RL, tuning and batch inference on GPU clusters; ~230k exposed instances observed","remediation_pain":"`node-drain` + network re-architecture - there is no vendor patch (Anyscale calls it intended behavior), so remediation means putting auth in front of every dashboard/Jobs API and restarting every Ray cluster in the fleet","pain_class":"node-drain","why_fleet_wide":"The Jobs API has no authorization: anyone who can reach the dashboard submits arbitrary code to the whole cluster, so one exposed head node hands over every GPU and every dataset attached to it"},"published":"2023-11-28"},{"id":"CVE-2023-48788","cve":"CVE-2023-48788","aliases":[],"title":"Fortinet FortiClient EMS: Unauthenticated SQL injection","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Fortinet FortiClient EMS","year":"2023","cvss_score":9.8,"severity":"critical","kev":true,"impact":"Unauthenticated SQL injection -> command execution as SYSTEM","attack_vector":"Network (remote)","remediation":"Control-plane: patch the endpoint-management server","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-48788"],"status":"curated","published":"2024-03-12"},{"id":"CVE-2023-49934","cve":"CVE-2023-49934","aliases":[],"title":"Slurm: SQL injection against the SlurmDBD accounting database","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Slurm","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"SQL injection against the SlurmDBD accounting database","attack_vector":"Any user who can submit a job","remediation":"Upgrade Slurm to 23.11.1+; audit accounting data","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-49934"],"status":"curated","published":"2023-12-14"},{"id":"CVE-2023-49937","cve":"CVE-2023-49937","aliases":[],"title":"Slurm: Double free allowing denial of service or possibly arbitrary code execution","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Slurm","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Double free allowing denial of service or possibly arbitrary code execution","attack_vector":"Any user who can submit a job","remediation":"Upgrade Slurm","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-49937"],"status":"curated","published":"2023-12-14"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-52515","cve":"CVE-2023-52515","aliases":[],"title":"Linux kernel (drivers/infiniband/ulp/srp): The SRP abort handler completes the SCSI command itself, after which the","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/ulp/srp)","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The SRP abort handler completes the SCSI command itself, after which the SCSI error handler re-queues or re-finishes the same command - a use-after-free of the command structure. This fires on any aborted SRP command, so a storage timeout caused by fabric congestion (including congestion a noisy tenant creates) corrupts kernel memory on the node that mounts the SRP storage.","attack_vector":"Triggered on the initiator side by any SCSI abort, which means an ordinary command timeout - not only a hostile target. In a shared cluster a tenant that saturates the RDMA fabric can induce those timeouts on nodes using SRP-over-IB storage. Conditional on the ib_srp module being loaded and SRP targets being mounted.","remediation":"Update to a stable kernel carrying the srp_abort fix (commits 26788a5b48d9 / b9bdffb3f9aa); the record lists the affected series as 3.1 through 3.7 baselines, so verify your distro backport. Interim: move affected nodes off SRP-over-IB storage, or drain them if SRP timeouts are already being observed.","references":["https://git.kernel.org/stable/c/26788a5b48d9d5cd3283d777d238631c8cd7495a","https://git.kernel.org/stable/c/b9bdffb3f9aaeff8379c83f5449c6b42cb71c2b5","https://nvd.nist.gov/vuln/detail/CVE-2023-52515"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-53116","cve":"CVE-2023-53116","aliases":[],"title":"Linux kernel NVMe target core (nvmet_req_complete submission-queue dereference): Nvmet_req_complete() dereferenced req","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel NVMe target core (nvmet_req_complete submission-queue dereference)","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Nvmet_req_complete() dereferenced req to reach the submission queue after calling __nvmet_req_complete(), which a transport's queue_response implementation is allowed to free. Every completed command on the target passes through this function, so the use-after-free is on the hottest path in the process serving the cluster's block storage - and command completion is driven by whatever the remote initiator submitted.","attack_vector":"Remote, unauthenticated. Any initiator that can submit commands to the target reaches the completion path.","remediation":"Kernel update caching the sq pointer before completion. Contain by network segmentation of the storage fabric.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=04c394208831d5e0d5cfee46722eb0f033cd4083","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2023/CVE-2023-53116.json"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-476","CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-53382","cve":"CVE-2023-53382","aliases":[],"title":"Linux kernel (net/smc): When an incoming connection tries SMC-Rv2 and device setup fails, the listener does not reset","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/smc)","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"When an incoming connection tries SMC-Rv2 and device setup fails, the listener does not reset the connection and goes on to build the CLC ACCEPT/CONFIRM message from half-initialised state, dereferencing a NULL link pointer inside the handshake worker. A remote client can crash the server node during connection setup, before any authentication, and the reproducer is nothing more exotic than nginx plus a load generator over SMC-R.","attack_vector":"Remote and pre-authentication: the fault is in smc_listen_work / smc_clc_send_confirm_accept, i.e. the server side of the CLC handshake, reachable by any fabric peer that connects to an SMC-enabled listening service. Requires SMC-Rv2 negotiation over RoCE (the report used two Mellanox ConnectX-4 adapters); the crash is in a kworker, so it takes the node with it.","remediation":"Boot a kernel carrying the fix commits (resets the connection when SMC-Rv2 setup fails). Interim: disable SMC-Rv2 negotiation, keep tenant-reachable services off SMC listeners, and blacklist the smc module on nodes that do not need it.","references":["https://git.kernel.org/stable/c/9540765d1882d15497d880096de99fafabcfa08c","https://git.kernel.org/stable/c/35112271672ae98f45df7875244a4e33aa215e31","https://nvd.nist.gov/vuln/detail/CVE-2023-53382"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-415","CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-54223","cve":"CVE-2023-54223","aliases":[],"title":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core/en/xsk): An RX buffer on the legacy receive queue is released","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core/en/xsk)","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"An RX buffer on the legacy receive queue is released twice - once by the XDP_REDIRECT path and again by the mlx5 driver - so a page that has already gone back to the allocator is still owned by the NIC receive ring. That is a classic double-free/use-after-free on the packet path: the freed page can be handed to another workload while the device keeps DMAing peer-supplied packet bytes into it, and the observed symptom is a general protection fault that panics the whole node.","attack_vector":"Reachable from the RX packet path, so any peer able to send frames to the node's Ethernet interface drives the code. Conditional on AF_XDP zero-copy being in use on an mlx5 interface configured with legacy (non-striding) RQ - typical for a tenant or host agent running an XDP/AF_XDP dataplane. No tenant device node is needed; the corruption happens in shared kernel memory on the host, so a single affected node is a blast radius covering every tenant on it.","remediation":"Boot a kernel with the fix (6.4.10 or later on that stream, plus the corresponding backports). Interim: stop running AF_XDP zero-copy sockets on mlx5 interfaces, or switch the affected interfaces to striding RQ so the legacy-RQ path is not used.","references":["https://git.kernel.org/stable/c/58a113a35846d9a5bd759beb332e551e28451f09","https://git.kernel.org/stable/c/e0f52298fee449fec37e3e3c32df60008b509b16","https://nvd.nist.gov/vuln/detail/CVE-2023-54223"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-362","CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-54237","cve":"CVE-2023-54237","aliases":[],"title":"Linux kernel (net/smc): On the server side of the SMC-R LLC handshake, adding a second link to a link group runs","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/smc)","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"On the server side of the SMC-R LLC handshake, adding a second link to a link group runs without the link-configuration mutex, so a remote client can race link addition against buffer registration and drive the host into ib_alloc_mr with a torn link structure. The published result is a kernel page fault during memory-region allocation - a remote, pre-authentication panic of the node from an unauthenticated connecting peer.","attack_vector":"Pre-authentication and driven entirely by the connecting peer: the LLC ADD LINK exchange happens during SMC-R connection setup, before any application-level authentication. Any host on the RDMA fabric that can complete a TCP connection to an SMC-enabled listener and negotiate SMC-R can drive it. Requires SMC-R to be in use (an RoCE/IB device plus a listener whose sockets fall into SMC), which is the normal case once the smc module is loaded.","remediation":"Boot a kernel carrying the fix commits (serialises smc_llc_srv_add_link under llc_conf_mutex). Interim: stop terminating tenant or east-west traffic on SMC-R listeners, blacklist the smc module on nodes not using it, and restrict which fabric peers can open TCP connections to SMC-capable services.","references":["https://git.kernel.org/stable/c/f2f46de98c11d41ac8d22765f47ba54ce5480a5b","https://git.kernel.org/stable/c/e40b801b3603a8f90b46acbacdea3505c27f01c0","https://nvd.nist.gov/vuln/detail/CVE-2023-54237"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2023-6014","cve":"CVE-2023-6014","aliases":[],"title":"MLflow: Arbitrary account creation bypassing authentication","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"MLflow","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Arbitrary account creation bypassing authentication","attack_vector":"Unauthenticated network to the tracking server","remediation":"Upgrade; the basic-auth plugin is not a boundary","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-6014"],"status":"curated","fleet":{"ubiquity":"Common - same MLflow footprint; this is the incomplete-fix follow-on","remediation_pain":"`daemon-restart` - server upgrade, plus rotation of everything the server could reach","pain_class":"daemon-restart","why_fleet_wide":"Basic-auth bypass on the tracking server, so the one control that was supposed to contain the previous LFI does not hold; full model-registry and artifact-store compromise"},"published":"2023-11-16"},{"id":"CVE-2023-6018","cve":"CVE-2023-6018","aliases":[],"title":"MLflow: Overwrite any file on the MLflow host without authentication","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"MLflow","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Overwrite any file on the MLflow host without authentication","attack_vector":"Unauthenticated network","remediation":"Upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-6018"],"status":"curated","published":"2023-11-16"},{"id":"CVE-2023-6019","cve":"CVE-2023-6019","aliases":[],"title":"Ray (dashboard `cpu_profile`): Command injection","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ray (dashboard `cpu_profile`)","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Command injection → OS command execution, unauthenticated","attack_vector":"Unauthenticated network to the Ray dashboard","remediation":"Upgrade to 2.8.1+ and isolate the dashboard","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-6019"],"status":"curated","published":"2023-11-16"},{"id":"CVE-2024-0012","cve":"CVE-2024-0012","aliases":[],"title":"Palo Alto PAN-OS: Management web interface authentication bypass","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Palo Alto PAN-OS","year":"2024","cvss_score":9.8,"severity":"critical","kev":true,"impact":"Management web interface authentication bypass -> PAN-OS administrator privileges","attack_vector":"Network (remote)","remediation":"Control-plane: patch + remove the management interface from the internet","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0012"],"status":"curated","published":"2024-11-18"},{"id":"CVE-2024-0138","cve":"CVE-2024-0138","aliases":[],"title":"Base Command Manager (CMDaemon): Unauthenticated RCE on the cluster manager","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Base Command Manager (CMDaemon)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unauthenticated RCE on the cluster manager -> full cluster takeover","attack_vector":"Network-adjacent unauthenticated attacker reaching CMDaemon","remediation":"Emergency: patch Base Command Manager, restrict CMDaemon to the mgmt network, audit for compromise; cluster-wide credential rotation","references":["https://github.com/NVIDIA/product-security/tree/main/2024/5595","https://nvd.nist.gov/vuln/detail/CVE-2024-0138"],"status":"curated","fleet":{"ubiquity":"Common - Base Command / Bright Cluster Manager is the control plane on many enterprise and neocloud GPU clusters","remediation_pain":"`daemon-restart` of the cluster control plane, which is itself a scheduling outage for the whole cluster","pain_class":"daemon-restart","why_fleet_wide":"Missing authentication in CMDaemon, remotely exploitable with no user interaction or privileges: compromising the cluster manager means owning provisioning for every node in the cluster at once"},"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-862"],"published":"2024-11-23"},{"id":"CVE-2024-11041","cve":"CVE-2024-11041","aliases":[],"title":"vLLM (MessageQueue / ZMQ): `pickle.loads` on socket data","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (MessageQueue / ZMQ)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"`pickle.loads` on socket data → unauthenticated RCE","attack_vector":"Unauthenticated network to the vLLM internal ZMQ socket, reachable by a co-tenant","remediation":"Upgrade; bind ZMQ to loopback and enforce per-tenant network policy. Internal IPC sockets must never cross the tenant boundary","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-11041"],"status":"curated","published":"2025-03-20"},{"id":"CVE-2024-1305","cve":"CVE-2024-1305","aliases":[],"title":"OpenVPN (tap-windows6): Unchecked write size","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"OpenVPN (tap-windows6)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unchecked write size -> memory buffer overflow and potential arbitrary code execution in kernel space","attack_vector":"Network (remote)","remediation":"Control-plane: operator endpoint driver update","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-1305"],"status":"curated","published":"2024-07-08"},{"id":"CVE-2024-21652","cve":"CVE-2024-21652","aliases":[],"title":"Argo CD: Chained DoS plus in-memory data manipulation lets an attacker bypass authentication","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Chained DoS plus in-memory data manipulation lets an attacker bypass authentication","attack_vector":"Unauthenticated network reaching the Argo CD API","remediation":"Rolling Argo CD upgrade to 2.8.13/2.9.9/2.10.4+","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21652"],"status":"curated","published":"2024-03-18"},{"id":"CVE-2024-21762","cve":"CVE-2024-21762","aliases":[],"title":"Fortinet FortiOS: SSL-VPN out-of-bounds write","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Fortinet FortiOS","year":"2024","cvss_score":9.8,"severity":"critical","kev":true,"impact":"SSL-VPN out-of-bounds write -> unauthenticated remote code execution","attack_vector":"Network (remote)","remediation":"Control-plane: emergency firmware; rotate all VPN user credentials","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21762"],"status":"curated","published":"2024-02-09"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-259","CWE-798"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2024-21990","cve":"CVE-2024-21990","aliases":[],"title":"NetApp ONTAP Select Deploy administration utility (hard-coded credentials): Baked-in credentials let an attacker read","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"NetApp ONTAP Select Deploy administration utility (hard-coded credentials)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Baked-in credentials let an attacker read Deploy configuration and change account credentials, which hands over the management plane for every ONTAP Select cluster the appliance controls.","attack_vector":"Network reach to ONTAP Select Deploy 9.12.1.x, 9.13.1.x or 9.14.1.x. The credential ships with the product, so it is the same everywhere and is not something an operator can rotate away.","remediation":"Upgrade Deploy to 9.15.1 or the fixed patch level NetApp names. Rotating passwords does not help while the hard-coded pair is present, so treat network isolation of the appliance as the only interim control.","references":["https://security.netapp.com/advisory/ntap-20240411-0002/","https://nvd.nist.gov/vuln/detail/CVE-2024-21990"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-2221","cve":"CVE-2024-2221","aliases":[],"title":"Qdrant (snapshot upload): Path traversal + arbitrary file upload via `/collections/{c}/snapshots/upload`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Qdrant (snapshot upload)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Path traversal + arbitrary file upload via `/collections/{c}/snapshots/upload`","attack_vector":"Network user able to upload a snapshot","remediation":"Upgrade; snapshot upload is an arbitrary-write primitive","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-2221"],"status":"curated","published":"2024-04-10"},{"id":"CVE-2024-22441","cve":"CVE-2024-22441","aliases":[],"title":"HPE Cray Parallel Application Launch Service (PALS) authentication bypass: Authentication bypass in the service","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HPE Cray Parallel Application Launch Service (PALS) authentication bypass","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Authentication bypass in the service that launches parallel jobs on Cray EX supercomputers. Bypassing PALS auth means launching arbitrary work on the cluster as another user - direct compromise of the HPC job-execution path.","attack_vector":"Unauthenticated network access to the PALS service on a Cray EX system.","remediation":"Apply the HPE Cray fix per HPESBCR04653. This is a system-management-stack update on the Cray EX; coordinate with your Cray support contact because the update path is tied to the CSM release train, not a simple package bump.","references":["https://support.hpe.com/hpesc/public/docDisplay?docId=hpesbcr04653en_us&docLocale=en_US"],"status":"curated"},{"id":"CVE-2024-23653","cve":"CVE-2024-23653","aliases":[],"title":"BuildKit: Interactive-container API lacks entitlement checks, so a build can run a privileged container and escape","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"BuildKit","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Interactive-container API lacks entitlement checks, so a build can run a privileged container and escape the builder","attack_vector":"Anyone who can submit a build to a shared builder","remediation":"Upgrade BuildKit; rebuild builder nodes; never share one BuildKit daemon across tenants","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-23653"],"status":"curated","published":"2024-01-31"},{"id":"CVE-2024-23897","cve":"CVE-2024-23897","aliases":[],"title":"Jenkins: CLI parser expands `@file` into argument contents","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Jenkins","year":"2024","cvss_score":9.8,"severity":"critical","kev":true,"impact":"CLI parser expands `@file` into argument contents -> unauthenticated arbitrary file read, chains to RCE","attack_vector":"Network (remote)","remediation":"Control-plane: URGENT patch; disable the CLI; rotate every credential in the Jenkins store","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-23897"],"status":"curated","published":"2024-01-24"},{"id":"CVE-2024-2421","cve":"CVE-2024-2421","aliases":["CVE-2024-2420","CVE-2024-2422","ICSA-24-151-01"],"title":"LenelS2 NetBox access control and event monitoring system (<=5.6.1): Unauthenticated remote code execution","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"LenelS2 NetBox access control and event monitoring system (<=5.6.1)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unauthenticated remote code execution with elevated privileges, plus hardcoded credentials that bypass authentication outright, plus a second authenticated RCE. NetBox is the head-end - the system that holds the cardholder database, the access rules, the door schedules and the event log for the whole site. Owning it is strictly more powerful than owning a single panel: an attacker can grant themselves a credential that works on every door, schedule doors to unlock, and erase the events showing they did. Physical entry to the hall and the cage follows, and from inside the cage the attacker reaches drives holding customer data and model weights, server console ports, and the out-of-band management switch that fronts every BMC in the row. For a bare-metal GPU provider, an attacker with head-end control can also target one specific tenant's cage on demand, which turns a security incident into a customer-trust and contractual event. Hardcoded credentials mean the exposure predates any breach you can detect.","attack_vector":"Unauthenticated over the network for the RCE and the hardcoded-credential bypass. NetBox is a web-managed appliance; sites routinely make it reachable from the corporate network so security staff can administer badges, and internet-exposed NetBox instances have been observed. Any of those makes this a direct, no-credential path from outside to physical door control.","remediation":"Upgrade NetBox past 5.6.1 to the fixed release per Carrier's advisory - a head-end software update in a normal change window, no door hardware touched, so there is no operational excuse to defer. Because hardcoded credentials were present, patching alone does not establish trust: rotate all system and integration credentials, audit the cardholder database and access rules against a known-good export, review the event log for gaps, and check whether any credentials were added during the exposure window. Then remove NetBox from any internet-facing or general corporate reachability and put it behind a jump host with MFA.","references":["https://www.cisa.gov/news-events/ics-advisories/icsa-24-151-01","https://nvd.nist.gov/vuln/detail/CVE-2024-2421","https://nvd.nist.gov/vuln/detail/CVE-2024-2420","https://www.corporate.carrier.com/product-security/advisories-resources/"],"status":"curated"},{"id":"CVE-2024-24592","cve":"CVE-2024-24592","aliases":[],"title":"ClearML fileserver: No authentication — arbitrary read/write/delete of all stored artifacts","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"ClearML fileserver","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"No authentication — arbitrary read/write/delete of all stored artifacts","attack_vector":"Unauthenticated network to the fileserver","remediation":"Upgrade and front with auth. All experiment data and models are exposed","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-24592"],"status":"curated","published":"2024-02-06"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-26582","cve":"CVE-2024-26582","aliases":[],"title":"Linux kernel (net/tls): The async decrypt completion released pages that the decrypt path never took a reference on, so","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/tls)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The async decrypt completion released pages that the decrypt path never took a reference on, so a partially-read record left the receive list pointing at freed pages. The next read walks freed memory - a network-driven use-after-free in the kTLS receive path.","attack_vector":"Remote: a peer sends records to a kTLS RX socket while the local reader does partial reads (a short recv() buffer is enough). Any kTLS connection on the node qualifies; no local privilege or device node required. Async decrypt must be in play, i.e. an async-capable AEAD driver such as cryptd-backed AES-NI.","remediation":"Boot a kernel carrying the linked stable commits, along with the rest of the tls async-decrypt series. Interim: disable async crypto offload for kTLS.","references":["https://git.kernel.org/stable/c/20b4ed034872b4d024b26e2bc1092c3f80e5db96","https://git.kernel.org/stable/c/d684763534b969cca1022e2a28645c7cc91f7fa5","https://nvd.nist.gov/vuln/detail/CVE-2024-26582"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-26583","cve":"CVE-2024-26583","aliases":[],"title":"Linux kernel (net/tls): The thread in recvmsg/sendmsg can exit as soon as the async crypto callback signals completion","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/tls)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The thread in recvmsg/sendmsg can exit as soon as the async crypto callback signals completion, so everything the callback touches afterwards is already-freed socket and context memory. A peer that times its records against a closing socket gets a use-after-free on the node.","attack_vector":"Remote plus a local close: the peer supplies records to a kTLS socket while the owning process closes or returns from the syscall - a normal pattern for short-lived tenant connections, and one an attacker can encourage by resetting connections. No privilege or device node needed; requires an async-capable AEAD driver.","remediation":"Boot a kernel carrying the linked stable commits, along with the rest of the tls async-decrypt series. Interim: disable async crypto offload for kTLS.","references":["https://git.kernel.org/stable/c/f17d21ea73918ace8afb9c2d8e734dbf71c2c9d7","https://git.kernel.org/stable/c/7a3ca06d04d589deec81f56229a9a9d62352ce01","https://nvd.nist.gov/vuln/detail/CVE-2024-26583"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-754","CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-26584","cve":"CVE-2024-26584","aliases":[],"title":"Linux kernel (net/tls): When the crypto queue is full the AEAD call returns -EBUSY instead of -EINPROGRESS and the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/tls)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"When the crypto queue is full the AEAD call returns -EBUSY instead of -EINPROGRESS and the async callback fires twice. kTLS treated that as an error and unwound buffers the callback still owns, giving a double-release of record memory on a node under crypto load.","attack_vector":"Remote peers supply the record volume; the trigger is a saturated cryptd queue, which a co-tenant driving heavy kTLS or dm-crypt traffic produces on the same node. Any kTLS socket is in scope, no privilege needed. Requires a backlog-capable async AEAD (cryptd/AES-NI).","remediation":"Boot a kernel carrying the linked stable commits, along with the rest of the tls async-decrypt series. Interim: disable async crypto offload for kTLS, or raise cryptd queue limits so the backlog path is not hit.","references":["https://git.kernel.org/stable/c/3ade391adc584f17b5570fd205de3ad029090368","https://git.kernel.org/stable/c/cd1bbca03f3c1d845ce274c0d0a66de8e5929f72","https://nvd.nist.gov/vuln/detail/CVE-2024-26584"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-26585","cve":"CVE-2024-26585","aliases":[],"title":"Linux kernel (net/tls): The async crypto callback signalled completion before scheduling the transmit work, so the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/tls)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The async crypto callback signalled completion before scheduling the transmit work, so the submitting thread could exit and free the socket context while the callback was still queueing work against it - a use-after-free on the kTLS transmit path.","attack_vector":"Remote peers drive the record flow; the race closes on a socket close or syscall return, which an attacker can encourage by resetting connections against a kTLS sender. Any kTLS TX socket on the node, no privilege or device node needed. Requires an async-capable AEAD driver.","remediation":"Boot a kernel carrying the linked stable commits, along with the rest of the tls async series. Interim: disable async crypto offload for kTLS.","references":["https://git.kernel.org/stable/c/dd32621f19243f89ce830919496a5dcc2158aa33","https://git.kernel.org/stable/c/196f198ca6fce04ba6ce262f5a0e4d567d7d219d","https://nvd.nist.gov/vuln/detail/CVE-2024-26585"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-415"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-26800","cve":"CVE-2024-26800","aliases":[],"title":"Linux kernel (net/tls): When a decrypt goes to the crypto backlog and a sibling decrypt fails, the error path releases","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/tls)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"When a decrypt goes to the crypto backlog and a sibling decrypt fails, the error path releases pages that the async callback has already freed. That is a double free of record pages in the kTLS receive path, reachable from the network.","attack_vector":"Remote records plus a saturated crypto queue, which a co-tenant running heavy encrypted I/O on the same node produces. Any kTLS RX socket qualifies; no local privilege or device node. Requires a backlog-capable async AEAD (cryptd/AES-NI).","remediation":"Update to 6.6.21 / 6.7.9 or later, or a kernel carrying the linked stable commits. Interim: disable async crypto offload for kTLS.","references":["https://git.kernel.org/stable/c/f2b85a4cc763841843de693bbd7308fe9a2c4c89","https://git.kernel.org/stable/c/81be85353b0f5a7b660635634b655329b429eefe","https://nvd.nist.gov/vuln/detail/CVE-2024-26800"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-27198","cve":"CVE-2024-27198","aliases":[],"title":"JetBrains TeamCity: Alternative-path authentication bypass","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"JetBrains TeamCity","year":"2024","cvss_score":9.8,"severity":"critical","kev":true,"impact":"Alternative-path authentication bypass -> unauthenticated admin actions on the CI server","attack_vector":"Network (remote)","remediation":"Control-plane: URGENT upgrade; rotate all build secrets and signing keys","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-27198"],"status":"curated","published":"2024-03-04"},{"id":"CVE-2024-27444","cve":"CVE-2024-27444","aliases":[],"title":"langchain-experimental: Second bypass of CVE-2023-44467","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"langchain-experimental","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Second bypass of CVE-2023-44467","attack_vector":"Untrusted prompt input","remediation":"Upgrade past 0.1.8. Three CVEs for one sandbox — code-generating chains have no safe configuration","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-27444"],"status":"curated","published":"2024-02-26"},{"id":"CVE-2024-29849","cve":"CVE-2024-29849","aliases":[],"title":"Veeam Backup Enterprise Manager: Unauthenticated users can log in as any user to the Enterprise Manager web interface","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Veeam Backup Enterprise Manager","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unauthenticated users can log in as any user to the Enterprise Manager web interface","attack_vector":"Network (remote)","remediation":"Control-plane: patch or decommission Enterprise Manager","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-29849"],"status":"curated","published":"2024-05-22"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-269"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-3057","cve":"CVE-2024-3057","aliases":[],"title":"Pure Storage FlashArray Purity API endpoint: A specific call to a FlashArray endpoint escalates the caller's privileges","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Pure Storage FlashArray Purity API endpoint","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A specific call to a FlashArray endpoint escalates the caller's privileges on the array, which puts every volume, snapshot and host mapping under attacker control.","attack_vector":"Network access to the FlashArray management endpoint. The CVSS vector records no privileges required, so treat any route to the management interface as sufficient.","remediation":"Upgrade Purity//FA to the fixed release Pure names in its security bulletin. Confirm the array's management interface is not reachable from tenant or compute networks, and audit recently created accounts and host mappings.","references":["https://support.purestorage.com/category/m_pure_storage_product_security","https://nvd.nist.gov/vuln/detail/CVE-2024-3057"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-32053","cve":"CVE-2024-32053","aliases":[],"title":"CyberPower PowerPanel platform - hardcoded database, service and cloud credentials: Hardcoded credentials used","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"CyberPower PowerPanel platform - hardcoded database, service and cloud credentials","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Hardcoded credentials used by the platform to authenticate to its database, to other services and to CyberPower's cloud. The cloud element is what makes this worth calling out separately: the credential is shared across deployments, so an attacker who extracts it once has a position against many operators' installations at the vendor's back end, not just yours.","attack_vector":"Anyone with the shipped software. The cloud credential in particular is reachable from the internet by design.","remediation":"Vendor upgrade is the only fix. Ask the vendor directly whether the cloud-side credential was rotated on their end - your upgrade does not do that. If the answer is unsatisfying, disable the cloud integration.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-32053"],"status":"curated","published":"2024-05-15"},{"id":"CVE-2024-32735","cve":"CVE-2024-32735","aliases":[],"title":"CyberPower PowerPanel Enterprise prior to v2.8.3 - PDNU REST APIs: Certain utility REST APIs have no authentication","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"CyberPower PowerPanel Enterprise prior to v2.8.3 - PDNU REST APIs","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Certain utility REST APIs have no authentication at all, giving an unauthenticated remote attacker a direct route into the application. Yet another unauthenticated path into a system that controls power distribution - the pattern across this product line is that authentication is applied per-endpoint and repeatedly missed.","attack_vector":"Unauthenticated, remote, to the PowerPanel Enterprise API surface.","remediation":"Upgrade to v2.8.3 or later. Given the density of unauthenticated findings in this product across 2023 and 2024, an operator should also decide whether it belongs in the design at all, or whether the power estate should be monitored through something with a better track record.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-32735"],"status":"curated","published":"2024-05-14"},{"id":"CVE-2024-33625","cve":"CVE-2024-33625","aliases":[],"title":"CyberPower PowerPanel business application - JWT signing key: The JWT signing key is hardcoded in the application, so","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"CyberPower PowerPanel business application - JWT signing key","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The JWT signing key is hardcoded in the application, so an attacker forges any token they like and becomes any user. Same shape as the hardcoded credentials: a secret that is not secret and cannot be rotated by the operator.","attack_vector":"Unauthenticated, remote. Requires only the shipped software to extract the key.","remediation":"Vendor upgrade. Nothing an operator can configure fixes a hardcoded signing key. Isolate the host until patched, and treat any PowerPanel instance that was internet-reachable as compromised.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-33625"],"status":"curated","published":"2024-05-15"},{"id":"CVE-2024-34025","cve":"CVE-2024-34025","aliases":[],"title":"CyberPower PowerPanel business application - hardcoded authentication credentials: A hardcoded credential set compiled","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"CyberPower PowerPanel business application - hardcoded authentication credentials","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A hardcoded credential set compiled into the application gives administrator access to anyone who reads the binary. There is no configuration that removes it and no password rotation that helps. On a platform that controls power distribution, this is a permanent unauthenticated back door until the vendor ships a build without it.","attack_vector":"Unauthenticated, remote, using credentials extractable from the shipped software by anyone.","remediation":"Upgrade to a build that removes the credentials - configuration changes cannot help. Until upgraded, the only real control is network isolation: the PowerPanel host must be unreachable from anything but a management jump box.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-34025"],"status":"curated","published":"2024-05-15"},{"id":"CVE-2024-35198","cve":"CVE-2024-35198","aliases":[],"title":"TorchServe: `allowed_urls` bypass","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"TorchServe","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"`allowed_urls` bypass → arbitrary model load → RCE","attack_vector":"Unauthenticated network to the management API","remediation":"Patch; the earlier ShellTorch fix is insufficient. Treat the management API as never tenant-reachable","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-35198"],"status":"curated","published":"2024-07-19"},{"id":"CVE-2024-36435","cve":"CVE-2024-36435","aliases":[],"title":"Supermicro BMC firmware web/management service (X11/X12/X13/H12/H13/B12/B13, CMM6): An attacker who never authenticates","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC firmware web/management service (X11/X12/X13/H12/H13/B12/B13, CMM6)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"An attacker who never authenticates gets code execution inside the BMC's own firmware OS. On a GPU fleet that means out-of-band power control over every affected node, the ability to mount an attacker-supplied ISO as virtual media and reboot a node into it, KVM access to whatever a tenant has on the console, and a persistent implant living beneath the hypervisor and the host OS that survives every reimage, every OS patch and every tenant handoff. Because the BMC also drives the CMM6 in blade chassis, a single compromised management module gives leverage over an entire enclosure rather than one node. H13, B12 and B13 motherboards plus CMM6 blade chassis management modules - i.e. essentially the whole current Supermicro server line including the GPU chassis and SuperBlade enclosures.","attack_vector":"Anything that can open a TCP connection to the BMC's management interface, with no credentials at all. In practice that is anyone with a route to the out-of-band management VLAN - a jump host, a misconfigured L3 leaf, a compromised DCIM or monitoring box, or a BMC that ended up with a public address. No host-side foothold and no tenant workload access is required.","remediation":"Firmware flash, per node, out of band. Supermicro shipped fixed BMC images in its July 2024 BMC/IPMI advisory batch, but the fixed version differs per motherboard SKU, so an operator with a mixed X11/X12/X13 fleet has to build a board-to-image matrix before flashing anything. Budget a BMC reboot per node - the host stays up during a BMC flash on most Supermicro boards but a failed flash can leave the BMC unresponsive and requires a physical recovery, so stage it rack by rack. Until every node is flashed, the only mitigation that actually holds is hard network isolation: BMCs on a dedicated VLAN with no route from tenant networks and an explicit allowlist for the handful of management hosts that need them.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36435","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2024/36xxx/CVE-2024-36435.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-476","CWE-562"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-36476","cve":"CVE-2024-36476","aliases":[],"title":"Linux kernel (drivers/infiniband/ulp/rtrs): The RTRS server builds an RDMA work request around a scatter-gather list","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/ulp/rtrs)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The RTRS server builds an RDMA work request around a scatter-gather list whose storage has already gone out of scope, so the transport copies data through a dangling descriptor. Upstream saw it as a kernel NULL-pointer fault inside the memory-copy path on the server - a remote client can crash the storage-serving node, and the underlying stale descriptor is a corruption primitive, not just a panic.","attack_vector":"Server-side and driven by the fabric: a client that establishes an RTRS/RNBD session and issues I/O drives the affected path on the target node, before any application-level trust decision. Conditional on the rtrs-srv module being loaded and exporting block devices (RNBD storage backend); the reported trace runs over soft-RoCE (rdma_rxe), which makes it reachable without special hardware.","remediation":"No fixed version is recorded in this entry; boot a stable kernel carrying the ib_sge scope fix (commits 7eaa71f56a6f / 143378075904). Interim: stop exporting RNBD/RTRS targets from shared nodes, or restrict which fabric addresses may open RTRS sessions.","references":["https://git.kernel.org/stable/c/7eaa71f56a6f7ab87957213472dc6d4055862722","https://git.kernel.org/stable/c/143378075904e78b3b2a810099bcc3b3d82d762f","https://nvd.nist.gov/vuln/detail/CVE-2024-36476"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-1259"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2024-36533","cve":"CVE-2024-36533","aliases":["GHSA-5g3x-8g2v-r8x8"],"title":"Volcano (v1.8.2 and earlier, service account token permissions): Volcano 1.8.2 ships over-permissive settings that let","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Volcano (v1.8.2 and earlier, service account token permissions)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Volcano 1.8.2 ships over-permissive settings that let an attacker obtain the scheduler's service account token. Volcano is the component that decides which tenant's job lands on which GPU, so holding its token means reading and rewriting placement for the entire cluster, plus whatever cluster-wide objects that account can touch.","attack_vector":"An attacker who reaches the Volcano components in-cluster. The advisory reports it as network-reachable with no privileges required.","remediation":"Upgrade Volcano to 1.10.0-alpha.0 or later and restart the scheduler, controller and webhook deployments. Rotate the Volcano service account token after upgrading and review RBAC bindings for the account, since the pre-upgrade token may already be out.","references":["https://github.com/advisories/GHSA-5g3x-8g2v-r8x8","https://nvd.nist.gov/vuln/detail/CVE-2024-36533"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-3660","cve":"CVE-2024-3660","aliases":[],"title":"Keras / TensorFlow: Arbitrary code injection in Keras < 2.13 via Lambda-layer model loading","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Keras / TensorFlow","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Arbitrary code injection in Keras < 2.13 via Lambda-layer model loading","attack_vector":"Customer-supplied `.h5` model","remediation":"Rebuild images off TF/Keras < 2.13; no runtime mitigation","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-3660"],"status":"curated","published":"2024-04-16"},{"id":"CVE-2024-37079","cve":"CVE-2024-37079","aliases":[],"title":"VMware vCenter: Heap overflow in the DCERPC implementation - unauthenticated remote code execution on vCenter","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"VMware vCenter","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Heap overflow in the DCERPC implementation - unauthenticated remote code execution on vCenter","attack_vector":"Unauthenticated network to the management plane","remediation":"vCenter patch + service restart. Never expose vCenter to a tenant-reachable network segment","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-37079"],"status":"curated","fleet":{"pain_class":"daemon-restart"},"published":"2024-06-18"},{"id":"CVE-2024-37080","cve":"CVE-2024-37080","aliases":[],"title":"VMware vCenter Server (DCERPC heap overflow): A heap overflow in the DCERPC implementation lets an unauthenticated","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"VMware vCenter Server (DCERPC heap overflow)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A heap overflow in the DCERPC implementation lets an unauthenticated network attacker reach remote code execution on vCenter with a single crafted packet.","attack_vector":"Network access to vCenter Server. No authentication.","remediation":"Apply the VMSA fix per Broadcom advisory 24453. vCenter appliance patch and restart. vCenter should never be reachable from tenant or general-purpose networks.","references":["https://support.broadcom.com/web/ecx/support-content-notification/-/external/content/SecurityAdvisories/0/24453"],"status":"curated"},{"id":"CVE-2024-3817","cve":"CVE-2024-3817","aliases":[],"title":"Terraform (go-getter): Argument injection when go-getter shells out to Git for remote branch discovery","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Terraform (go-getter)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Argument injection when go-getter shells out to Git for remote branch discovery -> RCE in the IaC runner","attack_vector":"Network (remote)","remediation":"Control-plane: bump go-getter/Terraform in the CI runner image","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-3817"],"status":"curated","published":"2024-04-17"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-38544","cve":"CVE-2024-38544","aliases":[],"title":"Linux kernel Soft-RoCE completer (rdma_rxe, rxe_comp_queue_pkt): An inbound response packet is queued to the completer","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel Soft-RoCE completer (rdma_rxe, rxe_comp_queue_pkt)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"An inbound response packet is queued to the completer before the code dereferences the same skb to bump a counter. If the completer task is already running on another CPU it can free the skb first, so the counter update reads freed memory. The trigger is a remote packet arriving at the right moment, which means an attacker who can send RoCE traffic at a node - a co-tenant on the same L2 fabric, since RoCEv2 is just UDP/4791 - can drive a kernel use-after-free with no credential. The kernel CNA rates it network, unauthenticated, full CIA.","attack_vector":"Remote and unauthenticated. Any host that can put RoCEv2 packets onto the node's fabric interface, including a container on another node in the same tenant network if RoCE is not segmented.","remediation":"Kernel update reordering the counter access ahead of the enqueue. Practical short-term control: unload rdma_rxe on nodes that have real RDMA NICs and do not need software RoCE, and enforce a fabric ACL so UDP/4791 is only accepted from the cluster's own RDMA subnet rather than from any tenant-routable network.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=21b4c6d4d89030fd4657a8e7c8110fd941049794","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2024/CVE-2024-38544.json"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-38812","cve":"CVE-2024-38812","aliases":[],"title":"VMware vCenter: Heap overflow in DCERPC - unauthenticated remote code execution on vCenter","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"VMware vCenter","year":"2024","cvss_score":9.8,"severity":"critical","kev":true,"impact":"Heap overflow in DCERPC - unauthenticated remote code execution on vCenter [KEV]","attack_vector":"Unauthenticated network to the management plane","remediation":"vCenter patch + service restart; the first patch was incomplete, so verify the build number rather than trusting the advisory date","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-38812"],"status":"curated","fleet":{"pain_class":"daemon-restart"},"published":"2024-09-17"},{"id":"CVE-2024-39236","cve":"CVE-2024-39236","aliases":[],"title":"Gradio: Code injection via `gradio/component_meta.py`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Gradio","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Code injection via `gradio/component_meta.py`","attack_vector":"Attacker-influenced component definition","remediation":"Upgrade past 4.36.1","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-39236"],"status":"curated","published":"2024-07-01"},{"id":"CVE-2024-40711","cve":"CVE-2024-40711","aliases":[],"title":"Veeam Backup & Replication: Deserialization of untrusted data","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Veeam Backup & Replication","year":"2024","cvss_score":9.8,"severity":"critical","kev":true,"impact":"Deserialization of untrusted data -> unauthenticated remote code execution","attack_vector":"Network (remote)","remediation":"Control-plane: URGENT - the backup server owns the restore path; patch, isolate, rotate","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-40711"],"status":"curated","published":"2024-09-07"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-415"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-41073","cve":"CVE-2024-41073","aliases":[],"title":"Linux kernel (drivers/nvme/host): A discard (TRIM) request that is retried and fails again before a fresh payload is","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/nvme/host)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A discard (TRIM) request that is retried and fails again before a fresh payload is attached frees the same special payload twice, corrupting the kernel heap on a node shared by many tenants. Double-free of a slab object is the classic starting point for privilege escalation, not just a crash.","attack_vector":"Driven by ordinary discard traffic from an unprivileged tenant - fstrim, a filesystem's online discard, or a thin-provisioned volume - so no device passthrough is needed. Turning it into a reliable double free requires the discard to fail and be retried, which is much easier to arrange when the namespace lives on a fabric target rather than a local SSD: the target decides which commands error. Say so plainly - on operator-run local NVMe this is opportunistic, on a tenant-influenced NVMe-oF target it is drivable.","remediation":"No fixed version is listed on this record - boot a kernel carrying the linked stable commits. Interim: disable online discard on tenant filesystems (mount without `discard`, batch with scheduled fstrim on the host) to shrink the exposed path.","references":["https://git.kernel.org/stable/c/882574942a9be8b9d70d13462ddacc80c4b385ba","https://git.kernel.org/stable/c/c5942a14f795de957ae9d66027aac8ff4fe70057","https://nvd.nist.gov/vuln/detail/CVE-2024-41073"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-41660","cve":"CVE-2024-41660","aliases":["GHSA-wmgv-jffg-v3xr"],"title":"OpenBMC slpd-lite (Service Location Protocol daemon, UDP 427): slpd-lite is a small SLP responder that OpenBMC installs","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"OpenBMC slpd-lite (Service Location Protocol daemon, UDP 427)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"slpd-lite is a small SLP responder that OpenBMC installs by default, and it overflows memory on crafted UDP packets. Because it is in the default build, this is not a niche configuration - if your image came from an OpenBMC tree and nobody explicitly removed the package, the daemon is listening. Unauthenticated remote memory corruption in a BMC-resident network daemon is the entry point that everything else in this cluster builds on: get code running on the BMC, then use the LPC-control or crypto kernel bugs to reach BMC root, then write flash and persist across tenant handover.","attack_vector":"Unauthenticated, network, UDP port 427 on the BMC's management interface. Nothing on the host and no credentials required. SLP is a discovery protocol nobody in a modern GPU fleet actually uses, which makes the exposure pure cost.","remediation":"Two moves and the cheap one is very cheap. Config-only: block UDP 427 at the management-VLAN boundary and, better, remove slpd-lite from the image or stop and mask the service. It provides nothing an operator needs - Redfish discovery does not depend on it. The durable fix is upstream in the slpd-lite repository and reaches nodes through a BMC firmware flash: per node, out-of-band, ODM-lagged. Do the ACL and service disable now; batch the flash.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-41660","https://github.com/openbmc/slpd-lite/security/advisories/GHSA-wmgv-jffg-v3xr"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-07-31"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-42285","cve":"CVE-2024-42285","aliases":[],"title":"Linux kernel (drivers/infiniband/core): Tearing down an iWARP connection frees the rdma_id_private while","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/core)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Tearing down an iWARP connection frees the rdma_id_private while connection-manager work is still queued against it, so the CM worker runs on freed memory. That is a use-after-free in kernel connection state driven by ordinary connect/disconnect timing, which gives an attacker a write primitive into reclaimed slab memory shared with every other tenant on the node.","attack_vector":"Reached through the iWARP connection manager, which handles incoming connection requests before any application-level authentication. A peer on the IP/RDMA fabric that connects to a listening rdma_cm endpoint drives iw_conn_req_handler, and racing the local destroy path wins the free. Requires an iWARP-capable path in use (irdma, cxgb4, or the software siw driver); a tenant holding /dev/infiniband/rdma_cm can drive both sides of the race locally.","remediation":"No fixed version is recorded in this entry; boot a stable kernel carrying the iwcm lifetime fix (commits d91d253c87fd / 7f25f296fc9b). Interim: blacklist siw if soft-iWARP is not needed, do not expose /dev/infiniband/rdma_cm to tenants, and restrict which peers may open iWARP connections to the node.","references":["https://git.kernel.org/stable/c/d91d253c87fd1efece521ff2612078a35af673c6","https://git.kernel.org/stable/c/7f25f296fc9bd0435be14e89bf657cd615a23574","https://nvd.nist.gov/vuln/detail/CVE-2024-42285"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-4323","cve":"CVE-2024-4323","aliases":[],"title":"Fluent Bit: \"Linguistic Lumberjack\" - memory corruption parsing trace requests in the embedded HTTP server","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Fluent Bit","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"\"Linguistic Lumberjack\" - memory corruption parsing trace requests in the embedded HTTP server -> RCE","attack_vector":"Network (remote)","remediation":"Data-plane: DaemonSet on every GPU node - fleet rollout + disable /api/v1/traces","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-4323"],"status":"curated","published":"2024-05-20"},{"id":"CVE-2024-43864","cve":"CVE-2024-43864","aliases":["net/mlx5e fix CT entry update leaks of modify header context"],"title":"Linux kernel mlx5_core TC connection tracking offload: Updating a connection-tracking entry allocates a replacement","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_core TC connection tracking offload","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Updating a connection-tracking entry allocates a replacement modify-header context; if that allocation fails - which it does once you exceed the firmware's maximum, a state an attacker can drive by opening many connections - the error pointer is stored and later dereferenced on free, panicking the kernel, and the old context is leaked. Remote traffic volume alone is enough to reach it on a node doing hardware conntrack offload.","attack_vector":"Remote, unauthenticated: open enough tracked connections through an mlx5 host doing CT offload to exhaust the firmware's modify-header capacity.","remediation":"Upgrade the host kernel to 6.11 or a stable backport (6.6.45, 6.10.4). Rolling reboot. Interim: disable hardware connection-tracking offload on the mlx5 interfaces or cap conntrack table size, both live config changes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-43864","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2024/CVE-2024-43864.json"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-08-21"},{"id":"CVE-2024-44970","cve":"CVE-2024-44970","aliases":["net/mlx5e SHAMPO invalid WQ linked list unlink"],"title":"Linux kernel mlx5_core RX datapath (SHAMPO): SHAMPO can deliver completion entries with zero consumed strides","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_core RX datapath (SHAMPO)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"SHAMPO can deliver completion entries with zero consumed strides for a work queue entry that was already fully consumed and unlinked, causing a second unlink that corrupts the receive work queue linked list. Remote, unauthenticated corruption of RX ring state on the host NIC - the node's networking dies and the kernel is in an inconsistent state.","attack_vector":"Unauthenticated remote sender of network traffic to an mlx5 host with SHAMPO/HW-GRO active.","remediation":"Upgrade the host kernel to 6.11 or a stable backport (6.1.105, 6.6.46, 6.10.5). Rolling reboot fleet-wide. Same interim mitigation as the other SHAMPO bugs: turn off rx-gro-hw via ethtool, a live config change.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-44970","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2024/CVE-2024-44970.json"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-04"},{"id":"CVE-2024-45410","cve":"CVE-2024-45410","aliases":[],"title":"Traefik: Traefik-added X-Forwarded-* headers can be spoofed by the client and are trusted by the backend","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Traefik","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Traefik-added X-Forwarded-* headers can be spoofed by the client and are trusted by the backend; authentication bypass at the application layer","attack_vector":"Unauthenticated network","remediation":"Rolling Traefik upgrade; explicitly configure trusted forwarded headers","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-45410"],"status":"curated","published":"2024-09-19"},{"id":"CVE-2024-46717","cve":"CVE-2024-46717","aliases":["net/mlx5e SHAMPO incorrect page release"],"title":"Linux kernel mlx5_core RX datapath (SHAMPO / HW-GRO): A remote sender can make the mlx5 receive path release a SHAMPO","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_core RX datapath (SHAMPO / HW-GRO)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A remote sender can make the mlx5 receive path release a SHAMPO header page while it is still in use, giving a double release and DMA into freed memory. This runs in NAPI softirq on every inbound frame, before any socket lookup or credential check - so an unauthenticated peer anywhere that can route packets to the node corrupts host kernel memory on the primary datacenter NIC. Note that NVD rates this 5.5 while the assigning kernel.org CNA rates it 9.8; the CNA score is the one that reflects reachability.","attack_vector":"Unauthenticated remote attacker able to send network traffic to a node running mlx5 with HW-GRO/SHAMPO enabled. No account, no VM, no adjacency needed.","remediation":"Upgrade the host kernel to 6.11 or a stable backport (6.1.109, 6.6.50, 6.10.9). Rolling reboot of every ConnectX/BlueField host. Immediate mitigation without reboot: disable HW-GRO on the mlx5 interfaces (ethtool -K <dev> rx-gro-hw off), which takes the SHAMPO path out of service at some throughput cost.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46717","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2024/CVE-2024-46717.json"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-18"},{"id":"CVE-2024-46946","cve":"CVE-2024-46946","aliases":[],"title":"langchain-experimental: Arbitrary code execution in 0.1.17–0.3.0","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"langchain-experimental","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Arbitrary code execution in 0.1.17–0.3.0","attack_vector":"Untrusted prompt input","remediation":"Upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46946"],"status":"curated","published":"2024-09-19"},{"id":"CVE-2024-47167","cve":"CVE-2024-47167","aliases":[],"title":"Gradio: SSRF from the file-upload/proxy path","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Gradio","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"SSRF from the file-upload/proxy path","attack_vector":"Unauthenticated network","remediation":"Upgrade to 4.44+","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-47167"],"status":"curated","published":"2024-10-10"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-1284","CWE-125"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-47408","cve":"CVE-2024-47408","aliases":[],"title":"Linux kernel SMC-R/SMC-D (CLC proposal parsing, smcd_v2_ext_offset): The SMC server trusted an offset field taken","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel SMC-R/SMC-D (CLC proposal parsing, smcd_v2_ext_offset)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The SMC server trusted an offset field taken straight out of the connecting client's CLC proposal message. Exceeding the maximum turns it into an arbitrary displacement into the parsing buffer - out-of-bounds access chosen by an unauthenticated remote peer. SMC matters here because it is the transparent RDMA acceleration path: LD_PRELOAD smc_run in front of an ordinary TCP application and its sockets silently become RoCE, so this parser sits in front of workloads whose operators do not know they are running an RDMA protocol stack at all.","attack_vector":"Remote, unauthenticated. The CLC proposal is the first SMC message a client sends; parsing happens before anything is established.","remediation":"Kernel update bounds-checking smcd_v2_ext_offset. Immediate mitigation: if SMC is not deliberately used, ensure the smc module is not loaded and that AF_SMC is not reachable - on many distro kernels it autoloads, so check rather than assume.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=48d5a8a304a643613dab376a278f29d3e22f7c34","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2024/CVE-2024-47408.json"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-47575","cve":"CVE-2024-47575","aliases":[],"title":"Fortinet FortiManager: \"FortiJump\" - missing authentication in fgfmd","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Fortinet FortiManager","year":"2024","cvss_score":9.8,"severity":"critical","kev":true,"impact":"\"FortiJump\" - missing authentication in fgfmd -> unauthenticated RCE on the fleet manager","attack_vector":"Network (remote)","remediation":"Control-plane: patch; FortiManager compromise means whole-fabric compromise","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-47575"],"status":"curated","published":"2024-10-23"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-125","CWE-193"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-47695","cve":"CVE-2024-47695","aliases":[],"title":"Linux kernel RTRS client (rtrs-clt init_conns connection-id bound): When connection setup fails partway through, the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel RTRS client (rtrs-clt init_conns connection-id bound)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"When connection setup fails partway through, the cleanup loop starts at cid == con_num rather than con_num - 1 and indexes one past the end of the connection array. RTRS is the RDMA transport under RNBD block devices, so this is the block-storage path of an RDMA cluster; the kernel CNA rates it network-reachable and unauthenticated, because a peer that makes connection establishment fail at the right point drives the out-of-bounds access remotely.","attack_vector":"Remote. A server-side peer that fails connection establishment partway through the multi-connection setup.","remediation":"Kernel update resetting cid to con_num - 1 before the cleanup loop. If RNBD/RTRS is not in use, keep the modules unloaded.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=01b9be936ee8839ab9f83a7e84ee02ac6c8303c4","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2024/CVE-2024-47695.json"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-47943","cve":"CVE-2024-47943","aliases":[],"title":"Rittal IoT Interface and CMC III Processing Unit - firmware upgrade signature check: The admin web interface verifies","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Rittal IoT Interface and CMC III Processing Unit - firmware upgrade signature check","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The admin web interface verifies patch files with an HMAC-style check whose key is a long string hard-coded in the firmware - and the firmware is freely downloadable, so anyone can extract the key and sign their own malicious update. That is a complete bypass of firmware integrity on the device that monitors cabinet temperature, humidity, door state, smoke and access for a rack or a row. An attacker who lands code there gets persistence on facility hardware nobody re-images, the ability to falsify environmental telemetry so a real thermal excursion in a GPU row never raises an alarm, and control of whatever access and door outputs the CMC III drives. The falsified-telemetry angle is the one operators underestimate: your temperature monitoring is the thing that is supposed to tell you a cooling attack is underway, and this lets the attacker turn it into a liar while the racks cook. The IoT Interface variant is a general-purpose gateway, so the same implant is a pivot point onto the facility network.","attack_vector":"Access to the admin web interface to upload the crafted patch - so an authenticated or otherwise-reachable path to the device on the facility VLAN. CMC III processing units and IoT Interfaces are typically on the monitoring network alongside PDU cards and environmental sensors, reachable from the DCIM collector and often from anything else on that segment. Given the same product family's history of backdoor accounts and injection flaws, reaching the admin interface is not a high bar.","remediation":"Rittal firmware update - a small device, quick flash, no cooling impact, so this is one of the more tractable items on this list. Do it across every CMC III processing unit and IoT Interface, and note that the fleet count in a large hall is per-row or per-cabinet, so it is a technician-day of work rather than a single change. Because the flaw defeats firmware integrity, any unit you believe was reachable by an untrusted host should be re-flashed from vendor media rather than trusted after an in-place update. Then isolate: monitoring devices on their own VLAN, admin interfaces reachable only from a jump host, and no route from tenant or corporate networks.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-47943","https://r.sec-consult.com/rittaliot","https://seclists.org/fulldisclosure/2024/Oct/4"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-48063","cve":"CVE-2024-48063","aliases":[],"title":"PyTorch (`torch.distributed` RemoteModule / RPC): Deserialization RCE across the distributed RPC channel","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"PyTorch (`torch.distributed` RemoteModule / RPC)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Deserialization RCE across the distributed RPC channel (vendor-disputed as intended behavior)","attack_vector":"Any host that can reach the RPC port of a multi-node training job — i.e. a co-tenant on the same fabric","remediation":"No patch — vendor position is that `torch.distributed` is a trusted-network protocol. Remediation is network isolation: per-tenant VRF/VLAN on the training fabric, never expose RPC ports across tenants","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-48063"],"status":"curated","fleet":{"ubiquity":"Very common - `torch.distributed` RPC / RemoteModule underpins multi-node training and RL rollout fanout on GPU clusters","remediation_pain":"Config/network change, not a patch - PyTorch disputes it, so remediation is isolating the RPC plane per tenant, which is an architectural change across the fleet","pain_class":"other","why_fleet_wide":"`RemoteModule` deserializes attacker-controlled data, so anyone who can reach a training job's RPC port executes code on every rank - one reachable worker owns the whole distributed job"},"published":"2024-10-29"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-1284","CWE-125"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-49568","cve":"CVE-2024-49568","aliases":[],"title":"Linux kernel SMC-R/SMC-D (CLC proposal parsing, v2_ext_offset / eid_cnt / ism_gid_cnt): The same unvalidated-offset","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel SMC-R/SMC-D (CLC proposal parsing, v2_ext_offset / eid_cnt / ism_gid_cnt)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The same unvalidated-offset pattern, this time across three remote-controlled fields at once - the V2 extension offset plus two counts that determine how many entries the server then walks. An unauthenticated peer sets the offset out of range and the count high, and the server reads far past its buffer. Three fields is worse than one: the counts turn a single bad read into a bounded-by-the-attacker sweep.","attack_vector":"Remote, unauthenticated, in the first message of the SMC handshake.","remediation":"Kernel update validating all three fields before use. Confirm whether SMC is actually in use in your fleet; if not, keep the module unloaded.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=295a92e3df32e72aff0f4bc25c310e349d07ffbf","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2024/CVE-2024-49568.json"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-4985","cve":"CVE-2024-4985","aliases":[],"title":"GitHub Enterprise Server: Forged SAML response with encrypted assertions enabled","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"GitHub Enterprise Server","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Forged SAML response with encrypted assertions enabled -> provision/gain site-admin access","attack_vector":"Network (remote)","remediation":"Control-plane: GHES upgrade immediately","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-4985"],"status":"curated","published":"2024-05-20"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-50106","cve":"CVE-2024-50106","aliases":[],"title":"Linux NFS server (nfsd, laundromat vs free_stateid race): A race between the delegation laundromat and a client-issued","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux NFS server (nfsd, laundromat vs free_stateid race)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A race between the delegation laundromat and a client-issued FREE_STATEID gives a use-after-free in nfsd. A tenant that can time FREE_STATEID against delegation expiry crashes the file server, and a UAF of this shape is a plausible escalation primitive on the storage node.","attack_vector":"Any NFSv4 client that can reach the server and hold a delegation - so any tenant compute node with the export mounted.","remediation":"Update the storage server kernel to one carrying the fix and reboot. Disabling delegations (echo 0 > /proc/fs/nfsd/... or the leasetime knob) narrows the window but is not a fix.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-50106","https://git.kernel.org/pub/scm/linux/security/vulns.git"],"status":"curated"},{"id":"CVE-2024-53138","cve":"CVE-2024-53138","aliases":["net/mlx5e kTLS incorrect page refcounting"],"title":"Linux kernel mlx5_core kTLS TX offload: The kTLS TX path mixes get_page() and page_ref_inc() when acquiring references","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_core kTLS TX offload","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The kTLS TX path mixes get_page() and page_ref_inc() when acquiring references but only put_page() when releasing. With large folios the page is dereferenced too many times, producing a use-after-free. Found in the wild via sendfile() with zero-copy over NFS. Any node terminating TLS in hardware on ConnectX is exposed - storage gateways and object-store frontends in a GPU cloud typically are.","attack_vector":"Remote, unauthenticated, over a TLS connection served by mlx5 hardware kTLS TX offload with large-folio pages in play.","remediation":"Upgrade the host kernel to 6.12 or one of the many stable backports (5.4.287, 5.10.231, 5.15.174, 6.1.119, 6.6.63, 6.11.10) - the wide backport range means most distros already ship a fix. Rolling reboot. Interim: disable kTLS TX offload on mlx5 interfaces (ethtool -K <dev> tls-hw-tx-offload off), a live config change at a CPU cost.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53138","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2024/CVE-2024-53138.json"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-12-04"},{"id":"CVE-2024-53676","cve":"CVE-2024-53676","aliases":[],"title":"HPE Insight Remote Support (directory traversal to RCE): Directory traversal allowing unauthenticated remote code","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HPE Insight Remote Support (directory traversal to RCE)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Directory traversal allowing unauthenticated remote code execution. Public proof-of-concept code exists, and this one has seen real-world exploitation attention - treat it as actively targeted.","attack_vector":"Unauthenticated network access to Insight RS.","remediation":"Patch per HPESBGN04731 without waiting for a maintenance window. If IRS was exposed, hunt for webshells and unexpected files under the application directories before declaring it clean.","references":["https://support.hpe.com/hpesc/public/docDisplay?docId=hpesbgn04731en_us&docLocale=en_US"],"status":"curated"},{"id":"CVE-2024-54085","cve":"CVE-2024-54085","aliases":[],"title":"AMI MegaRAC SPx (Redfish Host Interface): Unauthenticated auth bypass, full BMC takeover, malicious firmware flash.","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx (Redfish Host Interface)","year":"2024","cvss_score":9.8,"severity":"critical","kev":true,"impact":"Unauthenticated auth bypass, full BMC takeover, malicious firmware flash. Persists below the OS and survives reimaging","attack_vector":"Network / Redfish host interface, unauthenticated","remediation":"Out-of-band BMC flash on every node; requires an ODM rebase of the AMI fix (Supermicro / Lenovo / HPE / ASRock each ship their own build, availability lags AMI by months), bricking risk on interrupted flash","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-54085"],"status":"curated","fleet":{"ubiquity":"universal - AMI MegaRAC is the OEM BMC stack shipped under most Supermicro/ASRock Rack/Gigabyte/Quanta GPU servers","remediation_pain":"firmware-flash - BMC firmware image per node, applied out-of-band, and the OEM must first rebase AMI's fix into its own build; realistically a rolling node-drain because a bad flash bricks the board","pain_class":"firmware-flash","why_fleet_wide":"Unauthenticated Redfish auth bypass by spoofing the X-Server-Addr/Host header gives full BMC takeover on every node running the same OEM image; the BMC sits below the hypervisor, so an implant survives OS reimaging and GPU node rebuilds. First BMC CVE ever added to CISA KEV (2025-06-25)."},"published":"2025-03-11"},{"id":"CVE-2024-5452","cve":"CVE-2024-5452","aliases":[],"title":"PyTorch Lightning: RCE via deserialization of untrusted checkpoint","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"PyTorch Lightning","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"RCE via deserialization of untrusted checkpoint","attack_vector":"Customer-supplied `.ckpt` file","remediation":"Tenant-owned library. Provider action is scanning uploaded checkpoints at the object-store boundary","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-5452"],"status":"curated","published":"2024-06-06"},{"id":"CVE-2024-55591","cve":"CVE-2024-55591","aliases":[],"title":"Fortinet FortiOS/FortiProxy: Auth bypass via crafted Node.js websocket requests","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Fortinet FortiOS/FortiProxy","year":"2024","cvss_score":9.8,"severity":"critical","kev":true,"impact":"Auth bypass via crafted Node.js websocket requests -> super-admin privileges","attack_vector":"Network (remote)","remediation":"Control-plane: firmware; review all admin and local users","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-55591"],"status":"curated","published":"2025-01-14"},{"id":"CVE-2024-5660","cve":"CVE-2024-5660","aliases":["TFV-12","Hardware Page Aggregation bypass","HPA erratum"],"title":"Arm Neoverse V1 / V2 / V3 / V3AE / N2 and Cortex-A77/A78/A710/X1-X925 cores with Hardware Page Aggregation enabled","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Arm Neoverse V1 / V2 / V3 / V3AE / N2 and Cortex-A77/A78/A710/X1-X925 cores with Hardware Page Aggregation enabled…","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A guest can defeat Stage-2 translation, meaning the hypervisor's memory isolation for that VM stops holding. On a multi-tenant Arm host this is a tenant-to-hypervisor escape, and from there a read of every other VM resident on the socket. It also breaks Granule Protection Table checks, so if you are relying on Arm CCA / Realm confidentiality to sell 'the operator cannot see your model', that guarantee is void on affected silicon. Neoverse V2 is the Grace core, V1/V2 are Graviton3/Graviton4, N2 ships in several Arm server parts, so this touches most of the Arm side of an AI fleet.","attack_vector":"Unprivileged or kernel-level code inside any guest VM on an affected Arm host, with no special device access. No physical access, no network reachability, nothing but a VM on the box.","remediation":"This is silicon errata, so the fix is firmware discipline, not a code patch: EL3 firmware must set CPUECTLR_EL1[46]=1 to disable hardware page aggregation on every affected core. That means an updated TF-A / BL31 build from the platform OEM, flashed to each node, with a reboot and therefore a job drain. Ampere, NVIDIA and cloud OEMs each ship their own TF-A fork, so you wait on your board vendor, not on Arm. Budget a rolling reboot of the whole Arm fleet. There is a small performance cost to losing HPA on large-page workloads; measure before assuming it is free.","references":["https://developer.arm.com/Arm%20Security%20Center/Arm%20CPU%20Vulnerability%20CVE-2024-5660","https://trustedfirmware-a.readthedocs.io/en/latest/security_advisories/security-advisory-tfv-12.html","https://nvd.nist.gov/vuln/detail/CVE-2024-5660"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-12-10"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-56640","cve":"CVE-2024-56640","aliases":[],"title":"Linux kernel (net/smc): The server-side listen worker frees a connection outside the socket lock, so smc_conn_free()","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/smc)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The server-side listen worker frees a connection outside the socket lock, so smc_conn_free() and the link/link-group refcount drops can run twice for the same connection. The refcount saturates (addition-on-zero / underflow warnings) and the link group and link are released while still in use - a remote-driven use-after-free of the RDMA link state that a connecting peer can provoke by failing device negotiation at the right moment.","attack_vector":"Remote, pre-authentication: the double free happens in smc_listen_work / smc_listen_find_device, the server side of SMC connection setup, so any fabric peer able to reach an SMC-capable listener and abort or fail device negotiation drives it. Requires SMC in use on the listening node; the module autoloads from an unprivileged socket(AF_SMC, ...) call.","remediation":"Boot a kernel carrying the fix commits (takes the socket lock across the listen-path connection teardown). Interim: keep tenant-reachable services off SMC listeners and blacklist the smc module on nodes not using SMC-R.","references":["https://git.kernel.org/stable/c/f502a88fdd415647a1f2dc45fac71b9c522a052b","https://git.kernel.org/stable/c/673d606683ac70bc074ca6676b938bff18635226","https://nvd.nist.gov/vuln/detail/CVE-2024-56640"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-56656","cve":"CVE-2024-56656","aliases":[],"title":"Linux bnxt_en driver (5760X / P7 aggregation ID mask): The bnxt_en driver mishandles the aggregation ID mask on 5760X","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux bnxt_en driver (5760X / P7 aggregation ID mask)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The bnxt_en driver mishandles the aggregation ID mask on 5760X (P7) chips, oopsing the kernel. P7 is the Thor2 generation — the 400G-class Broadcom NIC going into current AI server designs — so this specifically hits new GPU-node builds rather than legacy hardware. A kernel oops on a training node kills the job for every rank in the collective, not just that node.","attack_vector":"Triggered by received traffic patterns hitting the hardware GRO/LRO path. Effectively remote from anything that can send traffic to the node.","remediation":"Kernel or driver package upgrade plus a host reboot. Standard rolling-reboot across the fleet, but coordinate with running jobs — a mid-training reboot is more expensive than the patch. If you build your own kernels, this is a backport-and-rebuild.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-56656"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-12-27"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-56718","cve":"CVE-2024-56718","aliases":[],"title":"Linux kernel (net/smc): A link-down work item can be queued before the link group is freed but run after, so the worker","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/smc)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A link-down work item can be queued before the link group is freed but run after, so the worker operates on a freed link group - the published crash is list corruption (prev->next NULL) inside smc_link_down_work, i.e. a write through a freed pointer from a kworker. A peer that can flap an SMC-R link while connections are closing turns fabric noise into host memory corruption.","attack_vector":"Driven from the RDMA fabric: link-down work is scheduled from link/port events on the RoCE or IB device, and the race is against link-group teardown that tenants drive by closing SMC connections. No credentials are needed on either side - the fabric peer only has to cause a link event, and the local side only has to be running SMC-R (the smc module autoloads on an unprivileged socket(AF_SMC, ...)).","remediation":"Boot a kernel carrying the fix commits (takes a link-group reference across the link-down work). Interim: blacklist the smc module on nodes not using SMC-R, and keep untrusted tenants off the RDMA fabric segment that can generate link events.","references":["https://git.kernel.org/stable/c/bec2f52866d511e94c1c37cd962e4382b1b1a299","https://git.kernel.org/stable/c/841b1824750d3b8d1dc0a96b14db4418b952abbc","https://nvd.nist.gov/vuln/detail/CVE-2024-56718"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-57936","cve":"CVE-2024-57936","aliases":[],"title":"Linux kernel (drivers/infiniband/hw/bnxt_re): The driver advertises support for 13 scatter-gather entries per work","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/hw/bnxt_re)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The driver advertises support for 13 scatter-gather entries per work request while its internal WQE structure only holds 6, so a tenant that posts a legal large-SGE send writes past the end of the WQE and corrupts adjacent kernel memory. Reported upstream as traffic failures and system crashes - this is attacker-shaped heap corruption from an ordinary verbs post_send.","attack_vector":"Any tenant container holding /dev/infiniband/uverbs* on a Broadcom Gen P7 bnxt_re adapter can trigger it by posting a send work request with more than 6 SGEs - the count the stack is told is legal. No fabric peer or privilege is required; the overflow happens on the local post path. Only applies where bnxt_re Gen P7 hardware is deployed.","remediation":"No fixed version is recorded in this entry; boot a stable kernel carrying the max-SGE fix (commits 3de1b50f055d / 9a479088e0c8). Interim: remove /dev/infiniband device nodes from tenant containers on bnxt_re Gen P7 nodes, or drain those nodes of untrusted tenants until patched.","references":["https://git.kernel.org/stable/c/3de1b50f055dc2ca7072a526cdda21f691c22dd9","https://git.kernel.org/stable/c/9a479088e0c8f6140b8c7752b563bc8c6c6dcc8c","https://nvd.nist.gov/vuln/detail/CVE-2024-57936"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-58240","cve":"CVE-2024-58240","aliases":[],"title":"Linux kernel (net/tls): The synchronous decrypt path shared refcounting and completion state with the async path, so a","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/tls)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The synchronous decrypt path shared refcounting and completion state with the async path, so a decrypt that was never actually asynchronous still went through the async wake-up dance. The kernel CNA scores this as a remotely reachable high-impact issue in the record decrypt path; it is also the prerequisite for the async-decrypt use-after-free fixes that follow it.","attack_vector":"The decrypt path is driven by whatever peer is sending records to a kTLS RX socket, so exposure follows every kTLS listener on the node. Practically, patch it as part of the tls async-decrypt cluster (CVE-2024-26582 / -26583 / -26584 / -26800 / CVE-2025-40176) rather than on its own - splitting them leaves the later fixes applied on top of the state this one cleans up.","remediation":"Boot a kernel carrying the linked stable commits, together with the rest of the tls async-decrypt series. Interim: disable async crypto offload for kTLS so the synchronous path is used.","references":["https://git.kernel.org/stable/c/f32ab6cb544d01cdc150de5df389ac685aa507b0","https://git.kernel.org/stable/c/5f7761c94b1db7570cab328c9f940a33284509b7","https://nvd.nist.gov/vuln/detail/CVE-2024-58240"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-6800","cve":"CVE-2024-6800","aliases":[],"title":"GitHub Enterprise Server: XML signature wrapping with publicly exposed federation metadata","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"GitHub Enterprise Server","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"XML signature wrapping with publicly exposed federation metadata -> forge a site-admin session","attack_vector":"Network (remote)","remediation":"Control-plane: GHES upgrade; review the admin list","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-6800"],"status":"curated","published":"2024-08-20"},{"id":"CVE-2024-9053","cve":"CVE-2024-9053","aliases":[],"title":"vLLM (`AsyncEngineRPCServer`): Unsafe deserialization on RPC entrypoints","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (`AsyncEngineRPCServer`)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unsafe deserialization on RPC entrypoints → RCE","attack_vector":"Unauthenticated network to the RPC server port","remediation":"Upgrade past 0.6.0","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-9053"],"status":"curated","published":"2025-03-20"},{"id":"CVE-2024-9070","cve":"CVE-2024-9070","aliases":[],"title":"BentoML (runner server): Deserialization RCE on the internal runner server","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"BentoML (runner server)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Deserialization RCE on the internal runner server","attack_vector":"Network to the runner port — reachable by a co-tenant in a flat cluster network","remediation":"Upgrade past 1.3.4.post1; runner ports must be namespace-local","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-9070"],"status":"curated","published":"2025-03-20"},{"id":"CVE-2024-9486","cve":"CVE-2024-9486","aliases":[],"title":"Kubernetes Image Builder: VM images built with the Proxmox provider ship default credentials","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes Image Builder","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"VM images built with the Proxmox provider ship default credentials; node takeover from the network","attack_vector":"Unauthenticated network reaching an affected node","remediation":"Rebuild and redeploy every affected node image; this is a full fleet reimage, not a package update","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated","published":"2024-10-15"},{"id":"CVE-2025-11200","cve":"CVE-2025-11200","aliases":[],"title":"MLflow (auth): Weak password requirements","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"MLflow (auth)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Weak password requirements → authentication bypass","attack_vector":"Unauthenticated network","remediation":"Upgrade; do not use MLflow's own auth as the tenant boundary","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-11200"],"status":"curated","published":"2025-10-29"},{"id":"CVE-2025-11201","cve":"CVE-2025-11201","aliases":[],"title":"MLflow (model creation): Directory traversal on model creation","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"MLflow (model creation)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Directory traversal on model creation → RCE","attack_vector":"Network to the tracking server","remediation":"Upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-11201"],"status":"curated","published":"2025-10-29"},{"id":"CVE-2025-15379","cve":"CVE-2025-15379","aliases":[],"title":"MLflow (serving container init): Command injection in `_install_model_dependencies`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"MLflow (serving container init)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Command injection in `_install_model_dependencies`","attack_vector":"Customer-supplied model requirements consumed when the provider builds a serving container","remediation":"Provider-owned in managed serving: the tenant's requirements file executes shell in the build","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-15379"],"status":"curated","published":"2026-03-30"},{"id":"CVE-2025-1550","cve":"CVE-2025-1550","aliases":[],"title":"Keras (`Model.load_model`): Arbitrary code execution from a crafted `.keras` archive even with `safe_mode=True`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Keras (`Model.load_model`)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Arbitrary code execution from a crafted `.keras` archive even with `safe_mode=True`","attack_vector":"Customer-supplied Keras model file","remediation":"Upgrade Keras in base images; no host-level patch. The safe-mode flag is not a security boundary","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-1550"],"status":"curated","fleet":{"ubiquity":"Very common - Keras 3.0.0-3.8.0 ships in most TF/Keras training images","remediation_pain":"**Image rebuild** across the fleet to Keras 3.9.0+; no host restart, but every image and every cached model must be re-vetted","pain_class":"other","why_fleet_wide":"`Model.load_model` executes arbitrary Python from a crafted `.keras` config *even with `safe_mode=True`*, so any model artifact pulled from a hub or a customer bucket is RCE on the loading GPU node"},"published":"2025-03-11"},{"id":"CVE-2025-1793","cve":"CVE-2025-1793","aliases":[],"title":"LlamaIndex (vector store integrations): SQL injection across multiple vector store integrations","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LlamaIndex (vector store integrations)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"SQL injection across multiple vector store integrations","attack_vector":"Attacker-influenced query/metadata reaching the store","remediation":"Upgrade past 0.12.21","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-1793"],"status":"curated","published":"2025-06-05"},{"id":"CVE-2025-1945","cve":"CVE-2025-1945","aliases":[],"title":"picklescan (model scanner): Scanner fails to detect malicious pickles when ZIP flag bits are flipped","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"picklescan (model scanner)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Scanner fails to detect malicious pickles when ZIP flag bits are flipped","attack_vector":"Customer-supplied PyTorch archive submitted to a provider's \"we scan your models\" control","remediation":"Direct hit on the provider's own control plane. Upgrade picklescan and stop treating a clean scan as proof of safety","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-1945"],"status":"curated","published":"2025-03-10"},{"id":"CVE-2025-1974","cve":"CVE-2025-1974","aliases":[],"title":"ingress-nginx: \"IngressNightmare\": unauthenticated RCE in the admission controller, reachable from any pod","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"\"IngressNightmare\": unauthenticated RCE in the admission controller, reachable from any pod, giving cluster-wide secret access and full cluster takeover","attack_vector":"Any pod on the cluster network, no credentials needed","remediation":"Emergency controller upgrade, or delete the ValidatingWebhookConfiguration and network-restrict the admission port. Rolling controller upgrade, no GPU drain. The highest-priority item in this table for a multi-tenant neocloud","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated","fleet":{"ubiquity":"Very common - ingress-nginx is the default ingress on most managed and self-run K8s clusters, including the control planes neoclouds put in front of GPU tenants","remediation_pain":"`daemon-restart` (rolling deployment upgrade of the controller) - cheap on the ingress pods themselves, but the *cleanup* is the pain: every secret in every namespace must be assumed stolen and rotated","pain_class":"daemon-restart","why_fleet_wide":"Anything on the pod network - i.e. any tenant workload - can hit the validating admission controller, load a shared library into the controller pod and read all secrets in all namespaces, which is cluster takeover"},"published":"2025-03-25"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-21805","cve":"CVE-2025-21805","aliases":[],"title":"Linux kernel (drivers/infiniband/ulp/rtrs): A remote client corrupts kernel linked lists on the RDMA block-storage","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/ulp/rtrs)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A remote client corrupts kernel linked lists on the RDMA block-storage server. An IB event handler is registered on every connection but never unregistered, so repeated connect/disconnect cycles leave stale handlers linked into device-wide lists - list corruption and stale-pointer execution on the target host, rated critical and network-reachable by the kernel CNA.","attack_vector":"Target-side and pre-authentication: whoever can reach the rtrs/rnbd server's listener drives it, entirely from the connection-establishment path (the CM request handler). Repeat connect/disconnect is the whole exploit. Applies to nodes exporting storage over rtrs/rnbd - if that listener is reachable from tenant networks, treat it as tenant-reachable.","remediation":"No fixed release is published in this record - apply the listed stable fix commits or run a current stable kernel. Interim: firewall the rtrs/rnbd server port to the storage network only, or stop exporting rtrs targets on affected hosts until patched.","references":["https://git.kernel.org/stable/c/5a79cc9bc961fafe90787f86e8f53ba6fad8d63b","https://git.kernel.org/stable/c/1af2c769032b6b334cd2a867d7d8c7cbbc527b2d","https://nvd.nist.gov/vuln/detail/CVE-2025-21805"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-21850","cve":"CVE-2025-21850","aliases":[],"title":"Linux kernel (drivers/nvme/target): The target disables a namespace without waiting for in-flight I/O to drain, so a","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/nvme/target)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The target disables a namespace without waiting for in-flight I/O to drain, so a request can be submitted against an already torn-down queue and dereference freed state - a general protection fault that panics the storage node. Any routine namespace reconfiguration becomes a coin flip on crashing the box while clients are attached.","attack_vector":"The crash needs two halves: an operator or automation disabling/reconfiguring an exported namespace (configfs, host root), and a connected client with I/O in flight at that instant. The client side is fully controlled by whoever is attached to the subsystem, so a tenant peer that keeps a steady stream of I/O open holds the window open permanently and turns every namespace disable into a node panic. Requires nvmet configured and exporting namespaces.","remediation":"No fixed version is listed on this record - boot a kernel carrying the linked stable commits. Interim: quiesce or disconnect clients before disabling a namespace, and treat namespace reconfiguration on unpatched target nodes as a maintenance operation rather than an online one.","references":["https://git.kernel.org/stable/c/cc0607594f6813342b27c752c6fb6f6eb9980cb5","https://git.kernel.org/stable/c/4082326807072b71496501b6a0c55ffe8d5092a5","https://nvd.nist.gov/vuln/detail/CVE-2025-21850"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-21927","cve":"CVE-2025-21927","aliases":["nvme-tcp header digest memory corruption","malicious NVMe/TCP target attacks initiator"],"title":"Linux kernel - NVMe/TCP host (initiator), drivers/nvme/host/tcp.c: Nvme_tcp_recv_pdu() did not validate the PDU header","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel - NVMe/TCP host (initiator), drivers/nvme/host/tcp.c","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Nvme_tcp_recv_pdu() did not validate the PDU header length, so with header digests enabled a target can send a packet declaring an invalid header length (for example 255) and make nvme_tcp_verify_hdgst() access memory outside the allocation and overwrite it with the calculated digest. The attack direction is target-to-initiator: a compromised or rogue storage target corrupts kernel memory on every GPU compute node that mounts from it. One storage compromise becomes fleet-wide kernel compromise, and header digests - a data-integrity feature operators turn on deliberately - are the precondition.","attack_vector":"The attacker controls an NVMe/TCP target the victim connects to, or can spoof/inject into that TCP connection, and returns a PDU with an out-of-range header length. Because NVMe/TCP has no transport authentication by default, an on-path attacker or anyone who can win a race to the discovery address can pose as the target.","remediation":"Host reboot / kernel upgrade on all NVMe/TCP initiator nodes - that is the GPU compute fleet, not just storage, so plan a full rolling drain. Immediate mitigations: disable header digests on affected initiators (nvme connect option, applied on reconnect, no reboot) to remove the precondition, and enable TLS for NVMe/TCP where supported so the target's identity is proven. Also verify that discovery addresses cannot be hijacked on the storage network.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2025/CVE-2025-21927.json","https://nvd.nist.gov/vuln/detail/CVE-2025-21927"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-04-01"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-22088","cve":"CVE-2025-22088","aliases":[],"title":"Linux kernel (drivers/infiniband/hw/erdma): Use-after-free while accepting an inbound RDMA connection. The connection","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/hw/erdma)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Use-after-free while accepting an inbound RDMA connection. The connection endpoint is dropped and then dereferenced again inside the accept path, so a peer opening connections against a listener reaches freed kernel memory. Rated critical and network-reachable by the kernel CNA.","attack_vector":"Pre-authentication and fully remote: any peer that can reach an erdma listener drives the accept path - no host account, no device node, no tenant cooperation. Requires the erdma driver (Alibaba Cloud elastic RDMA), so this matters if any part of the fleet runs on Alibaba Cloud instances with RDMA enabled.","remediation":"No fixed release is published in this record - apply the listed stable fix commits or run a current stable kernel image for those instances. Interim: restrict which sources can reach erdma listeners (security groups / firewall), or unload erdma on hosts that do not use it.","references":["https://git.kernel.org/stable/c/bc1db4d8f1b0dc480d7d745a60a8cc94ce2badd4","https://git.kernel.org/stable/c/667a628ab67d359166799fad89b3c6909599558a","https://nvd.nist.gov/vuln/detail/CVE-2025-22088"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-23310","cve":"CVE-2025-23310","aliases":[],"title":"NVIDIA Triton: Stack buffer overflow","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"NVIDIA Triton","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Stack buffer overflow → code execution","attack_vector":"Unauthenticated network to the inference port","remediation":"Patch; part of the Aug-2025 Triton cluster disclosed by Wiz","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23310"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-121"],"published":"2025-08-06"},{"id":"CVE-2025-23311","cve":"CVE-2025-23311","aliases":[],"title":"NVIDIA Triton: Stack overflow via crafted request","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"NVIDIA Triton","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Stack overflow via crafted request → code execution","attack_vector":"Unauthenticated network","remediation":"Patch","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23311"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-121"],"published":"2025-08-06"},{"id":"CVE-2025-23316","cve":"CVE-2025-23316","aliases":[],"title":"NVIDIA Triton Inference Server (Python backend): Attacker-controlled input in the Python backend","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"NVIDIA Triton Inference Server (Python backend)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Attacker-controlled input in the Python backend → remote code execution","attack_vector":"Unauthenticated network to an exposed Triton HTTP/gRPC port","remediation":"Patch to the fixed Triton release and rebuild NGC-derived images. Provider-published Triton images must be rebuilt, not just re-tagged","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23316"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-78"],"published":"2025-09-17"},{"id":"CVE-2025-25257","cve":"CVE-2025-25257","aliases":[],"title":"Fortinet FortiWeb: Unauthenticated SQL injection","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Fortinet FortiWeb","year":"2025","cvss_score":9.8,"severity":"critical","kev":true,"impact":"Unauthenticated SQL injection -> pre-auth code execution on the WAF","attack_vector":"Network (remote)","remediation":"Control-plane: firmware upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-25257"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2025-07-17"},{"id":"CVE-2025-25291","cve":"CVE-2025-25291","aliases":[],"title":"GitLab (ruby-saml): ReXML/Nokogiri parser differential","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"GitLab (ruby-saml)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"ReXML/Nokogiri parser differential -> SAML authentication bypass and account takeover","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade GitLab/ruby-saml; rotate the IdP signing certificate","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-25291"],"status":"curated","published":"2025-03-12"},{"id":"CVE-2025-27520","cve":"CVE-2025-27520","aliases":[],"title":"BentoML: RCE via insecure deserialization","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"BentoML","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"RCE via insecure deserialization","attack_vector":"Network to the serving endpoint","remediation":"Upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-27520"],"status":"curated","published":"2025-04-04"},{"id":"CVE-2025-32375","cve":"CVE-2025-32375","aliases":[],"title":"BentoML: Insecure deserialization RCE prior to 1.4.8","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"BentoML","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Insecure deserialization RCE prior to 1.4.8","attack_vector":"Network to the serving endpoint","remediation":"Upgrade to 1.4.8+","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-32375"],"status":"curated","published":"2025-04-09"},{"id":"CVE-2025-32434","cve":"CVE-2025-32434","aliases":[],"title":"PyTorch (`torch.load`): RCE via unsafe deserialization even with `weights_only=True`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"PyTorch (`torch.load`)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"RCE via unsafe deserialization even with `weights_only=True`","attack_vector":"Customer-supplied `.pt`/`.pth` checkpoint file loaded by any tenant job or by a provider-run model-import service","remediation":"No host patch. Bump `torch>=2.6.0` in every base image the provider ships; if the tenant pins an old torch in their own image, the provider cannot remediate — advise migration to safetensors. Shared responsibility: provider owns base images, tenant owns pinned envs","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-32434"],"status":"curated","fleet":{"ubiquity":"Universal - PyTorch ≤2.5.1 is in essentially every AI container image in every GPU fleet","remediation_pain":"**Image rebuild across the fleet** (`node-drain` for anything long-running) - the fix is `torch>=2.6.0` inside every tenant and platform image; you cannot hot-patch a library already imported into a running training job","pain_class":"node-drain","why_fleet_wide":"`torch.load(weights_only=True)` - the setting everyone was told was the safe one - is bypassable, so *every* checkpoint-loading path in the fleet is an RCE sink; 130+ public PoCs"},"published":"2025-04-18"},{"id":"CVE-2025-33222","cve":"CVE-2025-33222","aliases":[],"title":"NVIDIA Isaac Launchable: Hard-coded credentials in Isaac Launchable give an unauthenticated network attacker code","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Isaac Launchable","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Hard-coded credentials in Isaac Launchable give an unauthenticated network attacker code execution and privilege escalation - scored 9.8. Hard-coded credentials are unpatchable by configuration: the fix has to be a new build, and any deployment still running the old artifact stays exploitable regardless of what you change around it.","attack_vector":"Network, unauthenticated, no user interaction. The credentials are in the artifact, so anyone with the artifact has them.","remediation":"Update to the fixed Isaac Launchable release per bulletin 5749 and rotate anything the embedded credential could reach. Cost: redeploy. Critically, patching alone is insufficient - assume the credential is public and revoke it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33222","https://github.com/NVIDIA/product-security/tree/main/2025/5749"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-798"],"published":"2025-12-23"},{"id":"CVE-2025-33223","cve":"CVE-2025-33223","aliases":[],"title":"NVIDIA Isaac Launchable: Execution with unnecessary privileges lets an unauthenticated network attacker reach code","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Isaac Launchable","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Execution with unnecessary privileges lets an unauthenticated network attacker reach code execution, privilege escalation, information disclosure and data tampering. Scored 9.8.","attack_vector":"Network, unauthenticated, no user interaction.","remediation":"Update to the fixed release per bulletin 5749 and redeploy. Cost: redeploy only, but review what privileges the workload actually needs - the underlying pattern is over-privileged execution, which recurs.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33223","https://github.com/NVIDIA/product-security/tree/main/2025/5749"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-250"],"published":"2025-12-23"},{"id":"CVE-2025-33224","cve":"CVE-2025-33224","aliases":[],"title":"NVIDIA Isaac Launchable: A second over-privileged execution path with the same unauthenticated network reach and 9.8","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Isaac Launchable","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A second over-privileged execution path with the same unauthenticated network reach and 9.8 score.","attack_vector":"Network, unauthenticated, no user interaction.","remediation":"Update to the fixed release per bulletin 5749 and redeploy. Cost: redeploy.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33224","https://github.com/NVIDIA/product-security/tree/main/2025/5749"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-250"],"published":"2025-12-23"},{"id":"CVE-2025-37089","cve":"CVE-2025-37089","aliases":[],"title":"HPE StoreOnce (command injection RCE): Unauthenticated remote code execution on the backup appliance","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HPE StoreOnce (command injection RCE)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unauthenticated remote code execution on the backup appliance.","attack_vector":"Unauthenticated network access to StoreOnce.","remediation":"Upgrade StoreOnce per HPESBST04847. Ships with the auth-bypass fix in the same update.","references":["https://support.hpe.com/hpesc/public/docDisplay?docId=hpesbst04847en_us&docLocale=en_US"],"status":"curated"},{"id":"CVE-2025-37090","cve":"CVE-2025-37090","aliases":[],"title":"HPE StoreOnce (server-side request forgery): SSRF from the backup appliance, letting an unauthenticated attacker pivot","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HPE StoreOnce (server-side request forgery)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"SSRF from the backup appliance, letting an unauthenticated attacker pivot requests into internal networks the appliance can reach.","attack_vector":"Unauthenticated network access.","remediation":"Upgrade StoreOnce per HPESBST04847.","references":["https://support.hpe.com/hpesc/public/docDisplay?docId=hpesbst04847en_us&docLocale=en_US"],"status":"curated"},{"id":"CVE-2025-37092","cve":"CVE-2025-37092","aliases":[],"title":"HPE StoreOnce (command injection RCE): Second unauthenticated command-injection RCE path on StoreOnce","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HPE StoreOnce (command injection RCE)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Second unauthenticated command-injection RCE path on StoreOnce.","attack_vector":"Unauthenticated network access.","remediation":"Upgrade StoreOnce per HPESBST04847.","references":["https://support.hpe.com/hpesc/public/docDisplay?docId=hpesbst04847en_us&docLocale=en_US"],"status":"curated"},{"id":"CVE-2025-37093","cve":"CVE-2025-37093","aliases":[],"title":"HPE StoreOnce (authentication bypass): Unauthenticated attacker bypasses authentication on StoreOnce entirely, gaining","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HPE StoreOnce (authentication bypass)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unauthenticated attacker bypasses authentication on StoreOnce entirely, gaining full control of the backup appliance. Backup systems hold copies of everything and are a primary ransomware target - this is the bug that makes your recovery path attackable.","attack_vector":"Network access to the StoreOnce management interface. No credentials.","remediation":"Upgrade StoreOnce Software per HPESBST04847 as a priority. Appliance upgrade with a service window. Verify backup immutability/retention-lock settings while you are there - authentication bypass plus mutable backups is the ransomware worst case.","references":["https://support.hpe.com/hpesc/public/docDisplay?docId=hpesbst04847en_us&docLocale=en_US"],"status":"curated"},{"id":"CVE-2025-37095","cve":"CVE-2025-37095","aliases":[],"title":"HPE StoreOnce (directory traversal information disclosure): Unauthenticated directory traversal disclosing files","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HPE StoreOnce (directory traversal information disclosure)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unauthenticated directory traversal disclosing files from the backup appliance.","attack_vector":"Unauthenticated network access.","remediation":"Upgrade StoreOnce per HPESBST04847.","references":["https://support.hpe.com/hpesc/public/docDisplay?docId=hpesbst04847en_us&docLocale=en_US"],"status":"curated"},{"id":"CVE-2025-37096","cve":"CVE-2025-37096","aliases":[],"title":"HPE StoreOnce (command injection RCE): Third unauthenticated command-injection RCE path on StoreOnce","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HPE StoreOnce (command injection RCE)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Third unauthenticated command-injection RCE path on StoreOnce.","attack_vector":"Unauthenticated network access.","remediation":"Upgrade StoreOnce per HPESBST04847.","references":["https://support.hpe.com/hpesc/public/docDisplay?docId=hpesbst04847en_us&docLocale=en_US"],"status":"curated"},{"id":"CVE-2025-37099","cve":"CVE-2025-37099","aliases":[],"title":"HPE Insight Remote Support (remote code execution): Unauthenticated remote code execution on the Insight RS server","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HPE Insight Remote Support (remote code execution)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unauthenticated remote code execution on the Insight RS server. IRS has inbound reach to your whole HPE estate and outbound reach to HPE, so it is a high-value pivot in both directions.","attack_vector":"Unauthenticated network access to Insight RS below v7.15.0.646.","remediation":"Upgrade Insight RS to 7.15.0.646. Application upgrade with restart. IRS rarely needs broad network exposure - firewall it to the devices it actually monitors.","references":["https://support.hpe.com/hpesc/public/docDisplay?docId=hpesbgn04878en_us&docLocale=en_US"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-38211","cve":"CVE-2025-38211","aliases":[],"title":"Linux kernel (drivers/infiniband/core): The iWARP connection manager frees the work objects it is currently executing","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/core)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The iWARP connection manager frees the work objects it is currently executing from. If the last reference to a connection id is dropped inside an event handler, that handler's own work item is freed underneath the workqueue, corrupting kernel workqueue internals rather than just one connection. Rated critical and network-reachable by the kernel CNA.","attack_vector":"Driven by connection events from the fabric - a peer connecting, rejecting or disconnecting at the moment the local application destroys its connection id. No local credentials required; the remote side controls the timing half of the race. Applies to any iWARP-capable provider (irdma, siw, erdma, cxgb4), including tenant-initiated connections through /dev/infiniband/rdma_cm.","remediation":"No fixed release is published in this record - apply the listed stable fix commits or run a current stable kernel. Interim: disable iWARP providers not in use (unload siw / restrict irdma to RoCE mode where supported) and limit which peers can open iWARP connections to the node.","references":["https://git.kernel.org/stable/c/013dcdf6f03bcedbaf1669e3db71c34a197715b2","https://git.kernel.org/stable/c/bf7eff5e3a36c54bbe8aff7fd6dd7c07490b81c5","https://nvd.nist.gov/vuln/detail/CVE-2025-38211"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-38246","cve":"CVE-2025-38246","aliases":[],"title":"Linux bnxt_en driver (XDP redirect list flush): List corruption in the XDP redirect path, found crashing production","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux bnxt_en driver (XDP redirect list flush)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"List corruption in the XDP redirect path, found crashing production systems. Same family as the other bnxt XDP defects: if you run XDP for packet steering in front of inference endpoints, the Broadcom driver's redirect path has repeatedly been the weak spot, and the failure mode is a host crash rather than a graceful drop.","attack_vector":"Traffic through an attached XDP program using redirect on a Broadcom NIC.","remediation":"Kernel/driver upgrade plus host reboot. Companion fix CVE-2025-38439 (wrong DMA unmap length on XDP_REDIRECT) is in the same area — take both. Interim: detach XDP from Broadcom interfaces.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38246","https://nvd.nist.gov/vuln/detail/CVE-2025-38439"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-07-09"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-20","CWE-835"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-38264","cve":"CVE-2025-38264","aliases":[],"title":"Linux kernel NVMe-oF TCP host (nvme-tcp R2T PDU request-list handling): Nvme_tcp_handle_r2t() did not check that the","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel NVMe-oF TCP host (nvme-tcp R2T PDU request-list handling)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Nvme_tcp_handle_r2t() did not check that the request identified by an inbound Ready-to-Transfer PDU was not already on a list, so a malicious target sends a crafted R2T and injects a loop into the initiator's request list. The commit message says it plainly - a malicious R2T PDU. Every GPU node connected to that target is affected, and the corruption is in the block-layer request path, so it sits directly under the filesystem holding checkpoints and datasets.","attack_vector":"Remote, from the target side. A hostile or compromised NVMe/TCP target, or an attacker who can inject into an unencrypted NVMe/TCP session on the storage network.","remediation":"Kernel update on compute nodes validating the request state in nvme_tcp_handle_r2t(). Consider NVMe/TCP over TLS on the storage path so R2T PDUs cannot be injected by an on-path attacker, and treat the storage network as a trust boundary rather than as infrastructure.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=0bf04c874fcb1ae46a863034296e4b33d8fbd66c","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2025/CVE-2025-38264.json"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-38430","cve":"CVE-2025-38430","aliases":[],"title":"Linux NFS server (nfsd, nfsd4_spo_must_allow): nfsd4_spo_must_allow examines NFSv4 compound state without first","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux NFS server (nfsd, nfsd4_spo_must_allow)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"nfsd4_spo_must_allow examines NFSv4 compound state without first checking the request actually is a v4 compound, so a non-compound request drives it into invalid memory. Remote, unauthenticated, and it crashes or corrupts the file server every tenant depends on.","attack_vector":"Any host that can send RPC to the nfsd port. No mount or credential required.","remediation":"Update the storage server kernel to a release with the fix and reboot. Restrict which subnets can reach port 2049 in the interim - it does not eliminate the bug but it reduces who can send the malformed request.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38430","https://git.kernel.org/pub/scm/linux/security/vulns.git"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-38471","cve":"CVE-2025-38471","aliases":[],"title":"Linux kernel (net/tls): The strparser kept a stale reference to an skb that TCP had already coalesced away, and the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/tls)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The strparser kept a stale reference to an skb that TCP had already coalesced away, and the check that all queued skbs share decrypt state then read freed slab memory. A remote peer controls the send pattern that makes TCP coalesce, so this is a network-reachable use-after-free in the kTLS receive path.","attack_vector":"Remote: the peer's segmentation pattern drives TCP's skb compaction while the local kTLS reader waits for a record. Any kTLS RX socket the peer can reach is in scope - tenant workloads, storage clients, control-plane connections. No local privilege or device node needed.","remediation":"Boot a kernel carrying the linked stable commits. Interim: terminate TLS in userspace for connections to untrusted peers.","references":["https://git.kernel.org/stable/c/730fed2ff5e259495712518e18d9f521f61972bb","https://git.kernel.org/stable/c/1f3a429c21e0e43e8b8c55d30701e91411a4df02","https://nvd.nist.gov/vuln/detail/CVE-2025-38471"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-38734","cve":"CVE-2025-38734","aliases":[],"title":"Linux kernel (net/smc): The SMC listen worker keeps touching the SMC socket after smc_listen_out() has handed it off","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/smc)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The SMC listen worker keeps touching the SMC socket after smc_listen_out() has handed it off and it may already be released, giving a use-after-free on the server side of connection setup. An unauthenticated peer that connects and disconnects at the right moment corrupts host kernel memory from a kworker context.","attack_vector":"Remote and pre-authentication: the fault is inside smc_listen_work, the handshake worker for inbound SMC connections, so any fabric or IP peer that can reach an SMC-capable listening socket drives it. The victim only needs the smc module loaded, which happens on the first unprivileged socket(AF_SMC, ...) anywhere on the node.","remediation":"Boot a kernel carrying the fix commits. Interim: do not expose SMC-capable listeners to untrusted tenants or peers, and blacklist the smc module (`install smc /bin/false`) on nodes that do not use SMC-R.","references":["https://git.kernel.org/stable/c/070b4af44c4b6e4c35fb1ca7001a6a88fd2d318f","https://git.kernel.org/stable/c/85545f1525f9fa9bf44fec77ba011024f15da342","https://nvd.nist.gov/vuln/detail/CVE-2025-38734"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-476","CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-39682","cve":"CVE-2025-39682","aliases":[],"title":"Linux kernel (net/tls): A zero-length record already sitting on the rx_list breaks the invariant that zero-copy decrypt","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/tls)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A zero-length record already sitting on the rx_list breaks the invariant that zero-copy decrypt never has to queue an skb. The receive path then tries to queue a record it decrypted straight into the user buffer and has no skb to work with, corrupting the receive state of a connection a remote peer controls.","attack_vector":"Remote: the peer sends a zero-length TLS record and then a record of a different type, which is entirely within its control on any kTLS connection. Applies to any tenant-facing or fabric-facing kTLS RX socket on the node; no local privilege required.","remediation":"Boot a kernel carrying the linked stable commits. Interim: terminate TLS in userspace for connections to untrusted peers.","references":["https://git.kernel.org/stable/c/2902c3ebcca52ca845c03182000e8d71d3a5196f","https://git.kernel.org/stable/c/c09dd3773b5950e9cfb6c9b9a5f6e36d06c62677","https://nvd.nist.gov/vuln/detail/CVE-2025-39682"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-131","CWE-193"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-39758","cve":"CVE-2025-39758","aliases":[],"title":"Linux kernel SoftiWARP transmit path (siw_qp_tx, siw_tcp_sendpages byte accounting): After do_tcp_sendpages() was","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel SoftiWARP transmit path (siw_qp_tx, siw_tcp_sendpages byte accounting)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"After do_tcp_sendpages() was inlined, the sendmsg byte count passed for each page no longer matched the bvec length actually set up, so the transmit path pushed the wrong number of bytes per page. Wrong-length sends on a zero-copy RDMA transmit path mean bytes adjacent to the intended payload go onto the wire, and the stream desynchronises against what the peer expects. The kernel CNA scores it network-reachable with full confidentiality and integrity impact.","attack_vector":"Remote-facing. The corruption occurs on data leaving the node over an established SoftiWARP connection, so a peer that can induce the pathological send pattern observes the extra bytes.","remediation":"Kernel update correcting the byte count in siw_tcp_sendpages(). Unload siw where it is not required.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=42ebc16d9d2563f1a1ce0f05b643ee68d54fabf8","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2025/CVE-2025-39758.json"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787","CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-39946","cve":"CVE-2025-39946","aliases":[],"title":"Linux kernel (net/tls): When the socket buffer is too small to hold a whole record, kTLS parses early and re-parses as","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/tls)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"When the socket buffer is too small to hold a whole record, kTLS parses early and re-parses as more bytes arrive. A bogus record header made the parse fail without aborting the stream, so each retry copied more data into the same skb and eventually overflowed the allocated space - a remote out-of-bounds write in the receive path.","attack_vector":"Purely remote and pre-authentication with respect to the TLS session: a peer that can send to a kTLS RX socket manipulates the receive buffer state (syzbot did it with small out-of-band sends followed by a large normal send) and then supplies an invalid record length. Any tenant-facing or fabric-facing kTLS listener on the node is in scope; no local access needed.","remediation":"Boot a kernel carrying the linked stable commits. Interim: terminate TLS in userspace for peer-facing services, or avoid small SO_RCVBUF settings on kTLS sockets - noting that a hostile peer influences the condition regardless.","references":["https://git.kernel.org/stable/c/b36462146d86b1f22e594fe4dae611dffacfb203","https://git.kernel.org/stable/c/4cefe5be73886f383639fe0850bb72d5b568a7b9","https://nvd.nist.gov/vuln/detail/CVE-2025-39946"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-40176","cve":"CVE-2025-40176","aliases":[],"title":"Linux kernel (net/tls): If the skb clone that pins the input buffer for an async decrypt cannot be allocated, kTLS","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/tls)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"If the skb clone that pins the input buffer for an async decrypt cannot be allocated, kTLS proceeds with the decrypt anyway. The result is a use-after-free on the skb and, worse, the crypto engine writing plaintext into the caller's userspace buffer after recv() has already returned - so decrypted bytes land in whatever that address space has since reused the page.","attack_vector":"Remote and unauthenticated relative to the TLS data path: a peer sending records to any kTLS RX socket drives the async decrypt, and the failing clone allocation is inducible with memory pressure a co-tenant can generate. No device node or privilege needed; applies to tenant traffic and to node storage/control-plane connections using kTLS.","remediation":"Boot a kernel carrying the linked stable commits. Interim: disable async crypto offload for kTLS (avoid cryptd-backed AEAD drivers) or terminate TLS in userspace on exposed nodes.","references":["https://git.kernel.org/stable/c/9f83fd0c179e0f458e824e417f9d5ad53443f685","https://git.kernel.org/stable/c/c61d4368197d65c4809d9271f3b85325a600586a","https://nvd.nist.gov/vuln/detail/CVE-2025-40176"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-40212","cve":"CVE-2025-40212","aliases":[],"title":"Linux NFS server (nfsd, nfsd_set_fh_dentry): A refcount leak in the pseudo-root filehandle path lets a client drive the","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux NFS server (nfsd, nfsd_set_fh_dentry)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A refcount leak in the pseudo-root filehandle path lets a client drive the reference count until state is mishandled, giving remote memory corruption on the server. Reached through ordinary NFSv4 LOOKUP traversal of the exported pseudo-filesystem.","attack_vector":"Any NFSv4 client that can reach the server and walk the export pseudo-root.","remediation":"Update the storage server kernel and reboot. This is in the standard NFSv4 lookup path, so there is no export-level mitigation.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-40212","https://git.kernel.org/pub/scm/linux/security/vulns.git"],"status":"curated"},{"id":"CVE-2025-40350","cve":"CVE-2025-40350","aliases":[],"title":"Linux kernel mlx5_core RX datapath (striding RQ + XDP multi-buffer): The mlx5 driver assumed an XDP program could","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_core RX datapath (striding RQ + XDP multi-buffer)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The mlx5 driver assumed an XDP program could not change the xdp_buff layout. It can, via bpf_xdp_adjust_head/tail - and when it shrinks non-linear data the driver hits a BUG_ON, or builds a malformed skb. Any node running an XDP program on mlx5 (load balancers, DDoS scrubbers, CNI dataplanes like Cilium) can be crashed or memory-corrupted by remote packets. Jumbo-MTU routed fabrics are standard in datacenters, so the attacker does not need to be L2-adjacent.","attack_vector":"Unauthenticated remote sender, provided the target node runs an XDP program on an mlx5 interface with striding RQ.","remediation":"Upgrade the host kernel to 6.18 or a stable backport (6.6.115, 6.12.56, 6.17.6). Rolling reboot. Interim: unload the XDP program from mlx5 interfaces if you can accept the performance/feature loss - a live config change.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-40350","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2025/CVE-2025-40350.json"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-12-16"},{"id":"CVE-2025-41426","cve":"CVE-2025-41426","aliases":[],"title":"Vertiv (stack-based buffer overflow, code execution): A stack overflow gives an attacker code execution on the Vertiv","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Vertiv (stack-based buffer overflow, code execution)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A stack overflow gives an attacker code execution on the Vertiv facility device - persistent control of rack power/environmental infrastructure.","attack_vector":"Network access to the affected device.","remediation":"Apply the Vertiv firmware update per ICSA-25-140-10 and isolate facility devices onto their own network segment.","references":["https://www.cisa.gov/news-events/ics-advisories/icsa-25-140-10","https://www.vertiv.com/en-us/support/security-support-center/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-46412","cve":"CVE-2025-46412","aliases":["CVE-2025-41426","ICSA-25-140-10"],"title":"Vertiv Liebert RDU101 (<=1.9.0.0) and Liebert IS-UNITY (<=8.4.1.0) communication cards: Authentication bypass plus","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Vertiv Liebert RDU101 (<=1.9.0.0) and Liebert IS-UNITY (<=8.4.1.0) communication cards","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Authentication bypass plus a stack-based buffer overflow giving code execution on the card that sits inside Liebert cooling and power equipment across most of the datacenter industry. This is the worst shape a facility bug can take: an unauthenticated attacker on the facility network gets a persistent foothold on a device physically wired into thermal control, and from there can command the attached unit, spoof telemetry upward to the DCIM/BMS so dashboards stay green, and pivot laterally to every other card on the same VLAN. Applied against the cooling serving a GPU hall this is a fleet-wide availability kill and a hardware-damage risk, not an IT incident. Because the card also relays alarms, the same access lets an attacker blind the operator during the event. For a multi-tenant operator it is also a tenant-handoff problem: a card compromised under one tenant's occupancy stays compromised across the next tenant, because nobody re-flashes communication cards between customers.","attack_vector":"Unauthenticated, network-reachable, low complexity - CISA rates it exploitable remotely. The realistic exposure is the facility/BMS VLAN. RDU101 and IS-UNITY cards are also routinely reachable from the DCIM collector and, in far too many sites, from a vendor remote-support VPN concentrator that the mechanical contractor owns. Shodan-class internet exposure of Liebert cards is a recurring finding, so check whether yours are NAT'd out for 'remote monitoring'.","remediation":"Patchable, and this one is worth the window: update RDU101 to v1.9.1.2_0000001 and IS-UNITY to v8.4.3.1_00160. The card firmware update does not require draining the cooling unit, but it does require the card to be reachable and briefly offline, so schedule it as a monitoring outage rather than a cooling outage - budget a per-card touch across the whole fleet, which for a large hall is dozens of cards and a real technician-day cost. Alongside the patch, segment: cards on a dedicated VLAN, no route to tenant or corporate networks, no inbound from the internet, and inventory every Liebert card by firmware version so you can prove the fleet is clean. Leased colo: this is the landlord's card in the landlord's unit - send them the CISA advisory ID and require written confirmation of the firmware level.","references":["https://www.cisa.gov/news-events/ics-advisories/icsa-25-140-10","https://nvd.nist.gov/vuln/detail/CVE-2025-46412","https://nvd.nist.gov/vuln/detail/CVE-2025-41426","https://www.vertiv.com/en-us/support/security-support-center/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-47277","cve":"CVE-2025-47277","aliases":[],"title":"vLLM (`PyNcclPipe` KV transfer): RCE via the KV cache transfer integration","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (`PyNcclPipe` KV transfer)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"RCE via the KV cache transfer integration","attack_vector":"Unauthenticated network to the KV pipe port between prefill/decode nodes","remediation":"Upgrade past 0.8.4. Disaggregated-prefill deployments expose a new unauthenticated plane on the tenant fabric","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-47277"],"status":"curated","published":"2025-05-20"},{"id":"CVE-2025-49655","cve":"CVE-2025-49655","aliases":[],"title":"Keras: Deserialization of untrusted data in 3.11.0–3.11.2","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Keras","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Deserialization of untrusted data in 3.11.0–3.11.2 → malicious model runs arbitrary code","attack_vector":"Customer-supplied model file","remediation":"Upgrade to 3.11.3+","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-49655"],"status":"curated","published":"2025-10-17"},{"id":"CVE-2025-49825","cve":"CVE-2025-49825","aliases":[],"title":"Teleport: Remote authentication bypass in Teleport Community Edition (<=17.5.1)","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Teleport","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Remote authentication bypass in Teleport Community Edition (<=17.5.1)","attack_vector":"Network (remote)","remediation":"Control-plane + data-plane: upgrade the proxy AND every node agent; rotate CAs","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-49825"],"status":"curated","published":"2025-06-17"},{"id":"CVE-2025-53521","cve":"CVE-2025-53521","aliases":["K000156741"],"title":"F5 BIG-IP (APM access policy): Specific malicious traffic against a virtual server with a BIG-IP APM access policy","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"F5 BIG-IP (APM access policy)","year":"2025","cvss_score":9.8,"severity":"critical","kev":true,"impact":"Specific malicious traffic against a virtual server with a BIG-IP APM access policy configured leads to remote code execution. This is part of F5's October 2025 mass-remediation advisory issued after the nation-state breach of F5's internal systems and BIG-IP source code, and it's confirmed in CISA's KEV catalog — treat any internet- or cluster-VPN-facing BIG-IP APM instance as a priority target.","attack_vector":"Remote, against a virtual server that has an APM access policy applied — F5 has not disclosed full exploitation prerequisites, consistent with a critical RCE that's actively tracked as exploited.","remediation":"Software upgrade to the fixed BIG-IP version per F5 K000156741, then reboot/failover. Because this is part of the breach-remediation batch, treat the whole BIG-IP estate as needing review — not just this one CVE — and prioritize any unit that's End of Technical Support, since F5 does not evaluate or patch those.","references":["https://my.f5.com/manage/s/article/K000156741","https://www.cisa.gov/known-exploited-vulnerabilities-catalog?field_cve=CVE-2025-53521"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-10-15"},{"id":"CVE-2025-62863","cve":"CVE-2025-62863","aliases":["AMP-SB-0007","AmpereOne UEFI-MM PCIe driver OOB write"],"title":"Ampere AmpereOne AC03 before 3.5.9.3, AC04 before 4.4.5.2, AmpereOne M before 5.4.5.1","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Ampere AmpereOne AC03 before 3.5.9.3, AC04 before 4.4.5.2, AmpereOne M before 5.4.5.1 - UEFI Management Mode PCIe…","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A malformed SMC from the normal world produces an out-of-bounds write inside the S-EL0 UEFI-MM secure partition. That is code execution in the secure world reached from the OS - the boundary AmpereOne uses to protect runtime firmware services. Once there, the attacker can rewrite firmware state, forge attestation, or persist across a tenant handoff. The bulletin's siblings give a secure-partition-context OOB write (CVE-2025-62864) and an S-EL0 info leak (CVE-2025-62862), so the whole SMC surface should be treated as compromised until patched.","attack_vector":"Host kernel or hypervisor issuing SMC calls on an AmpereOne node. On bare-metal AmpereOne rental this is the tenant. No physical or network access needed.","remediation":"Update to the fixed firmware for your part - AC03 3.5.9.3, AC04 4.4.5.2, AmpereOne M 5.4.5.1 - from the board OEM. Flash + reboot + drain per node; because this is in UEFI-MM, it ships as part of the platform firmware bundle and the ODM must integrate Ampere's release before you can install it. Verify the running version after reboot; AmpereOne firmware version strings are per-SKU and easy to get wrong on a mixed fleet.","references":["https://amperecomputing.com/products/security-bulletins/amp-sb-0007","https://nvd.nist.gov/vuln/detail/CVE-2025-62863","https://amperecomputing.com/products/product-security"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-12-16"},{"id":"CVE-2025-63389","cve":"CVE-2025-63389","aliases":[],"title":"Ollama (API auth): Critical authentication bypass on API endpoints through v0.12.3","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ollama (API auth)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Critical authentication bypass on API endpoints through v0.12.3","attack_vector":"Unauthenticated network to the Ollama API","remediation":"Upgrade; assume any tenant-reachable Ollama instance is fully controllable","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-63389"],"status":"curated","published":"2025-12-18"},{"id":"CVE-2025-6543","cve":"CVE-2025-6543","aliases":["CTX694788"],"title":"Citrix NetScaler ADC / Gateway (configured as VPN Gateway, ICA Proxy, CVPN, RDP Proxy, or AAA virtual server): A memory","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Citrix NetScaler ADC / Gateway (configured as VPN Gateway, ICA Proxy, CVPN, RDP Proxy, or AAA virtual server)","year":"2025","cvss_score":9.8,"severity":"critical","kev":true,"impact":"A memory overflow in the Gateway/AAA code path corrupts control flow, letting an attacker crash the appliance (denial of service) and, per CISA's KEV listing, this has been exploited in the wild — meaning it's already a live threat to any NetScaler used as the VPN entry point into a cluster's management network.","attack_vector":"Remote, targeting a NetScaler configured as a Gateway/AAA virtual server (i.e. the VPN/ICA-proxy entry point) — exact prerequisites are undisclosed by Citrix, consistent with an actively-exploited memory-corruption bug.","remediation":"Software upgrade to the fixed NetScaler ADC/Gateway build per Citrix advisory CTX694788, then reboot. Given confirmed in-the-wild exploitation, patch ahead of the normal cycle — this is the box guarding VPN access into the cluster's OOB/management network, not just a data-plane load balancer.","references":["https://support.citrix.com/support-home/kbsearch/article?articleNumber=CTX694788","https://www.cisa.gov/known-exploited-vulnerabilities-catalog?field_cve=CVE-2025-6543"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-06-25"},{"id":"CVE-2025-66405","cve":"CVE-2025-66405","aliases":[],"title":"Portkey AI Gateway: Gateway resolves the destination baseURL from attacker-controlled precedence","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Portkey AI Gateway","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Gateway resolves the destination baseURL from attacker-controlled precedence → request hijack","attack_vector":"Tenant request to the shared gateway","remediation":"Upgrade to 1.14.0+","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-66405"],"status":"curated","published":"2025-12-01"},{"id":"CVE-2025-67038","cve":"CVE-2025-67038","aliases":["ICSA-26-069-02"],"title":"Lantronix EDS5000 serial-to-Ethernet device server: Root command execution on the device server","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Lantronix EDS5000 serial-to-Ethernet device server","year":"2025","cvss_score":9.8,"severity":"critical","kev":true,"impact":"Root command execution on the device server. The HTTP RPC module writes a log line whenever a login fails, and it builds that log line by shell-concatenating the attempted username — so a failed login with a malicious username is enough to run arbitrary commands as root. This CVE is in CISA's Known Exploited Vulnerabilities catalog, meaning it has been used in real attacks.","attack_vector":"No valid credentials needed — the trigger is a failed authentication attempt, so the attacker just needs network reachability to the device's HTTP interface and controls the username field.","remediation":"Firmware flash is required; there's no config toggle to disable the vulnerable logging path. Given the confirmed exploitation, treat this as urgent — patch every EDS5000 in the fleet ahead of routine maintenance windows, one device at a time (each flash drops the serial sessions it's bridging).","references":["https://www.cisa.gov/news-events/ics-advisories/icsa-26-069-02","https://www.cisa.gov/known-exploited-vulnerabilities-catalog?field_cve=CVE-2025-67038"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2026-03-11"},{"id":"CVE-2025-6802","cve":"CVE-2025-6802","aliases":["CVE-2025-6793","CVE-2025-6794","CVE-2025-6795","CVE-2025-6796","CVE-2025-6797","CVE-2025-6798","CVE-2025-6799","CVE-2025-6800","CVE-2025-6801","CVE-2025-6803","CVE-2025-6804","CVE-2025-6805","CVE-2025-6806","CVE-2025-6807","CVE-2025-8426"],"title":"Marvell QConvergeConsole (QLogic Fibre Channel / FC-NVMe / CNA HBA management web console), 5.5.0.78 and earlier","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Marvell QConvergeConsole (QLogic Fibre Channel / FC-NVMe / CNA HBA management web console), 5.5.0.78 and earlier","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A 2025 batch of sixteen ZDI advisories against the same Tomcat/GWT console, most of them reachable with no authentication at all. getFileFromURL gives unrestricted file upload leading to code execution as the console's service account (SYSTEM or root); saveAsText gives directory-traversal RCE; the rest are arbitrary file write, arbitrary file deletion and arbitrary file read. QConvergeConsole is the box that flashes QLogic HBA firmware and boot-from-SAN parameters across the fleet, so owning it means owning the HBA firmware and the SAN login identity of every server it manages - a persistence layer beneath the tenant OS, and a way to re-point a node's boot LUN. The file-deletion variants alone are a fleet availability event.","attack_vector":"Any host that can reach the QConvergeConsole web listener (typically TCP 8080/8443 on a management server or on individual hosts running the agent). No credentials for the unauthenticated set; the older 2020 cluster's 'authenticated' variants were shown to be reachable by bypassing the console's own auth.","remediation":"Upgrade to the fixed Marvell QConvergeConsole release, but treat that as a stopgap: this codebase has now produced two large unauthenticated-RCE clusters (2020 and 2025) and has no meaningful hardening. The durable fix for a GPU fleet is to remove QConvergeConsole from production hosts entirely and manage QLogic HBAs with the qaucli/qlogic CLI driven from configuration management, then firewall the console port. Software-only change, no HBA flash and no array downtime. Any host whose console was internet- or tenant-reachable should have its HBA firmware reflashed from a known-good vendor image before reuse.","references":["https://www.zerodayinitiative.com/advisories/ZDI-25-464/","https://www.zerodayinitiative.com/advisories/ZDI-25-450/","https://nvd.nist.gov/vuln/detail/CVE-2025-6802"],"status":"curated","published":"2025-07-07"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-68811","cve":"CVE-2025-68811","aliases":[],"title":"Linux NFS-over-RDMA server (svcrdma, svc_rdma_copy_inline_range): The inline copy path adds a page index where it","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux NFS-over-RDMA server (svcrdma, svc_rdma_copy_inline_range)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The inline copy path adds a page index where it should add a byte offset, so memcpy lands outside the current page and writes into unrelated kernel memory on the NFS server. Triggered by ordinary RDMA inline traffic from a client.","attack_vector":"Any NFS/RDMA client that can reach the storage server over the fabric.","remediation":"Update the storage server kernel to one carrying the rc_pageoff fix and reboot. Serve affected exports over TCP as a stopgap.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git","https://nvd.nist.gov/vuln/detail/CVE-2025-68811"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-71068","cve":"CVE-2025-71068","aliases":[],"title":"Linux NFS-over-RDMA server (svcrdma, svc_rdma_copy_inline_range): svc_rdma_copy_inline_range indexes rq_pages with an","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux NFS-over-RDMA server (svcrdma, svc_rdma_copy_inline_range)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"svc_rdma_copy_inline_range indexes rq_pages with an unvalidated rc_curpage, so a crafted inline RDMA message walks the server past the end of its page array. Kernel memory corruption on the storage server, triggered from the RDMA fabric.","attack_vector":"Any NFS/RDMA client on the same fabric as the storage server.","remediation":"Update the storage server kernel and reboot. Interim mitigation is to serve the affected exports over TCP instead of RDMA.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git","https://nvd.nist.gov/vuln/detail/CVE-2025-71068"],"status":"curated"},{"id":"CVE-2025-7775","cve":"CVE-2025-7775","aliases":[],"title":"Citrix NetScaler: Memory overflow","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Citrix NetScaler","year":"2025","cvss_score":9.8,"severity":"critical","kev":true,"impact":"Memory overflow -> pre-authentication remote code execution and/or DoS","attack_vector":"Network (remote)","remediation":"Control-plane: emergency firmware","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-7775"],"status":"curated","published":"2025-08-26"},{"id":"CVE-2026-0300","cve":"CVE-2026-0300","aliases":[],"title":"Palo Alto PAN-OS: Buffer overflow in the User-ID Captive Portal","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Palo Alto PAN-OS","year":"2026","cvss_score":9.8,"severity":"critical","kev":true,"impact":"Buffer overflow in the User-ID Captive Portal -> unauthenticated arbitrary code execution","attack_vector":"Network (remote)","remediation":"Control-plane: emergency patch; disable the Captive Portal if unused","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-0300"],"status":"curated","published":"2026-05-06"},{"id":"CVE-2026-0545","cve":"CVE-2026-0545","aliases":[],"title":"MLflow (jobs API): `/ajax-api/3.0/jobs/*` unauthenticated even with basic-auth enabled","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"MLflow (jobs API)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"`/ajax-api/3.0/jobs/*` unauthenticated even with basic-auth enabled","attack_vector":"Unauthenticated network","remediation":"Upgrade; auth coverage gaps recur across MLflow releases","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-0545"],"status":"curated","published":"2026-04-03"},{"id":"CVE-2026-12481","cve":"CVE-2026-12481","aliases":[],"title":"Keras (Lambda layer): Arbitrary code execution via Lambda-layer deserialization in 3.14.0","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Keras (Lambda layer)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Arbitrary code execution via Lambda-layer deserialization in 3.14.0","attack_vector":"Customer-supplied model containing a Lambda layer","remediation":"No safe fix short of refusing Lambda layers. Provider should reject models declaring Lambda layers at ingest","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-12481"],"status":"curated","published":"2026-07-03"},{"id":"CVE-2026-1281","cve":"CVE-2026-1281","aliases":[],"title":"Ivanti Endpoint Manager Mobile: Code injection","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ivanti Endpoint Manager Mobile","year":"2026","cvss_score":9.8,"severity":"critical","kev":true,"impact":"Code injection -> unauthenticated remote code execution","attack_vector":"Network (remote)","remediation":"Control-plane: patch immediately; the MDM plane reaches operator devices","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-1281"],"status":"curated","published":"2026-01-29"},{"id":"CVE-2026-1340","cve":"CVE-2026-1340","aliases":[],"title":"Ivanti Endpoint Manager Mobile: Code injection","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ivanti Endpoint Manager Mobile","year":"2026","cvss_score":9.8,"severity":"critical","kev":true,"impact":"Code injection -> unauthenticated remote code execution","attack_vector":"Network (remote)","remediation":"Control-plane: patch immediately; assume compromise if internet-facing","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-1340"],"status":"curated","published":"2026-01-29"},{"id":"CVE-2026-15969","cve":"CVE-2026-15969","aliases":[],"title":"SGLang (`/load_lora_adapter_from_tensors`): Unauthenticated RCE bypassing `SafeUnpickler`'s incomplete denylist","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"SGLang (`/load_lora_adapter_from_tensors`)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unauthenticated RCE bypassing `SafeUnpickler`'s incomplete denylist","attack_vector":"Customer-supplied LoRA adapter posted to the serving API","remediation":"Upgrade. LoRA adapters are executable content — a denylist unpickler is not a boundary","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-15969"],"status":"curated","published":"2026-07-30"},{"id":"CVE-2026-21643","cve":"CVE-2026-21643","aliases":[],"title":"Fortinet FortiClient EMS: Unauthenticated SQL injection","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Fortinet FortiClient EMS","year":"2026","cvss_score":9.8,"severity":"critical","kev":true,"impact":"Unauthenticated SQL injection -> code execution via crafted HTTP requests","attack_vector":"Network (remote)","remediation":"Control-plane: patch the endpoint-management server","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-21643"],"status":"curated","published":"2026-02-06"},{"id":"CVE-2026-22778","cve":"CVE-2026-22778","aliases":[],"title":"vLLM (image error echo): Error path returns sensitive content on invalid image input","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (image error echo)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Error path returns sensitive content on invalid image input","attack_vector":"Unauthenticated network to the multimodal endpoint","remediation":"Upgrade past 0.14.1","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-22778"],"status":"curated","published":"2026-02-02"},{"id":"CVE-2026-23112","cve":"CVE-2026-23112","aliases":["CVE-2026-52989","CVE-2022-50717"],"title":"Linux kernel nvmet-tcp - PDU iovec construction and H2C Transfer Tag handling: nvmet_tcp_build_pdu_iovec() walks past","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel nvmet-tcp - PDU iovec construction and H2C Transfer Tag handling","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"nvmet_tcp_build_pdu_iovec() walks past the command's scatterlist when a PDU length or offset exceeds sg_cnt and then builds a bvec from garbage sg->length/offset values, and in the related defect the error is not propagated so the socket receive loop reads network data into an uninitialised iterator. The older sibling is the same shape: the Transfer Tag in an H2C_DATA PDU is used as an array index with no bounds check. All three give an initiator on the storage network control over where the target kernel copies received data - memory corruption in the process that owns every tenant's namespace mappings, and a straightforward crash of the entire target if you only want the outage.","attack_vector":"Any host that can complete an NVMe/TCP connection to the target on TCP 4420 and send a malformed PDU. A provisioned namespace is not required for the connection-level PDU handling; on a fabric with no in-band auth, any host on the storage VLAN qualifies.","remediation":"Kernel update on target nodes, reboot, arrays offline for the duration unless you can fail initiators to a peer target first. There is no runtime workaround - the parsing happens before any policy check. Between now and the maintenance window, the storage VLAN ACL is the control: only known initiator IPs reach 4420. Note the 2022 Transfer Tag bounds check applies to long-lived kernels many operators are still running under distro LTS, so check your actual kernel rather than assuming a recent distro release covers it.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git","https://nvd.nist.gov/vuln/detail/CVE-2026-23112","https://nvd.nist.gov/vuln/detail/CVE-2022-50717"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-02-13"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-23240","cve":"CVE-2026-23240","aliases":[],"title":"Linux kernel (net/tls): Closing a kTLS socket cancelled the transmit work item, but the write-space callback could","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/tls)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Closing a kTLS socket cancelled the transmit work item, but the write-space callback could re-schedule it immediately afterwards from softirq. The worker then runs against a freed TLS context - a use-after-free on socket teardown that a remote peer helps time, since the re-schedule comes from the delayed-ACK path. Kernel heap corruption reachable by any tenant that uses kTLS, with the timing supplied by the far end of the connection.","attack_vector":"Any local process that attaches kTLS to a socket (setsockopt TLS_TX - no special privilege, available to every tenant container) and closes it while data is still in flight. The re-schedule window is opened by the peer's ACK behaviour, so a cooperating or hostile remote endpoint on the fabric can widen it deliberately by controlling when it acknowledges. No device node, no CAP_NET_ADMIN.","remediation":"Boot a kernel carrying the fix commits below (no fixed stable version published in the record). Interim control: none that preserves kTLS - the only real lever is blacklisting the tls ULP module so tenants fall back to userspace TLS until the fleet is rebooted.","references":["https://git.kernel.org/stable/c/a5de36d6cee74a92c1a21b260bc507e64bc451de","https://git.kernel.org/stable/c/854cd32bc74fe573353095e90958490e4e4d641b","https://nvd.nist.gov/vuln/detail/CVE-2026-23240"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-23450","cve":"CVE-2026-23450","aliases":[],"title":"Linux kernel (net/smc): An inbound SYN handled in softirq reads the smc_sock out of the listening TCP socket's","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/smc)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"An inbound SYN handled in softirq reads the smc_sock out of the listening TCP socket's sk_user_data while a concurrent close is clearing it and freeing that smc_sock, so the packet path dereferences NULL or a dangling pointer. A remote host that keeps sending connection attempts to an SMC listener that is closing gets an unauthenticated use-after-free in the network receive path - the strongest primitive in this shard.","attack_vector":"Remote, pre-authentication, packet-driven: the race is in smc_tcp_syn_recv_sock, called from tcp_v4_rcv / tcp_check_req and from the SYN-cookie path, so nothing more than SYN traffic to an SMC-capable listening port is required. No tenant privilege and no RDMA access is needed on the attacker side; the victim only needs an SMC listener, and the smc module autoloads from any unprivileged socket(AF_SMC, ...).","remediation":"Update to 5.15.203 or later on that branch, or any kernel carrying the fix commits. Interim: stop exposing SMC listeners to untrusted networks and tenants, and blacklist the smc module (`install smc /bin/false`) on nodes not using SMC.","references":["https://git.kernel.org/stable/c/f315277856caeafcd996c2611afc085ca2d53275","https://git.kernel.org/stable/c/f00fc26c8a06442b225a350fe000c0a11483e6a3","https://nvd.nist.gov/vuln/detail/CVE-2026-23450"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-24178","cve":"CVE-2026-24178","aliases":[],"title":"NVIDIA FLARE SDK: Unauthenticated remote code execution","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA FLARE SDK","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unauthenticated remote code execution","attack_vector":"Network-adjacent unauthenticated","remediation":"Emergency upgrade of FLARE; redeploy federated-learning services","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24178","https://github.com/NVIDIA/product-security/tree/main/2026/5819"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-639"],"published":"2026-04-28"},{"id":"CVE-2026-24207","cve":"CVE-2026-24207","aliases":[],"title":"Triton Inference Server: Missing authentication","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Missing authentication -> full unauthorized control of the server","attack_vector":"Network-adjacent unauthenticated","remediation":"Emergency Triton upgrade + redeploy all serving images; gate the endpoint behind authn immediately","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24207","https://github.com/NVIDIA/product-security/tree/main/2026/5828"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-288"],"published":"2026-05-20"},{"id":"CVE-2026-24254","cve":"CVE-2026-24254","aliases":[],"title":"NVIDIA Dynamo: Unauthenticated remote code execution","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Dynamo","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unauthenticated remote code execution","attack_vector":"Network-adjacent unauthenticated","remediation":"Emergency Dynamo upgrade; redeploy; gate endpoints behind authn and network policy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24254","https://github.com/NVIDIA/product-security/tree/main/2026/5842"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-288"],"published":"2026-08-04"},{"id":"CVE-2026-24270","cve":"CVE-2026-24270","aliases":[],"title":"NVIDIA AIStore: Missing authentication on API endpoints","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA AIStore","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Missing authentication on API endpoints -> full object-store access","attack_vector":"Network-adjacent unauthenticated inside the cluster","remediation":"Emergency upgrade of AIStore; enforce authn + network policy; audit stored tenant data","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24270","https://github.com/NVIDIA/product-security/tree/main/2026/5849"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-290"],"published":"2026-07-01"},{"id":"CVE-2026-24858","cve":"CVE-2026-24858","aliases":[],"title":"Fortinet (FortiOS/FortiManager/FortiProxy): Auth bypass via alternate path using a FortiCloud account and a registered","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Fortinet (FortiOS/FortiManager/FortiProxy)","year":"2026","cvss_score":9.8,"severity":"critical","kev":true,"impact":"Auth bypass via alternate path using a FortiCloud account and a registered device","attack_vector":"Network (remote)","remediation":"Control-plane: fleet-wide firmware; audit FortiCloud account linkage","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24858"],"status":"curated","published":"2026-01-27"},{"id":"CVE-2026-25089","cve":"CVE-2026-25089","aliases":[],"title":"Fortinet FortiSandbox: OS command injection","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Fortinet FortiSandbox","year":"2026","cvss_score":9.8,"severity":"critical","kev":true,"impact":"OS command injection -> unauthenticated command execution via crafted requests","attack_vector":"Network (remote)","remediation":"Control-plane: emergency firmware upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-25089"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2026-06-09"},{"id":"CVE-2026-26190","cve":"CVE-2026-26190","aliases":[],"title":"Milvus (port 9091): Management port 9091 exposed by default enabling compromise","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Milvus (port 9091)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Management port 9091 exposed by default enabling compromise","attack_vector":"Unauthenticated network — default config","remediation":"Upgrade past 2.5.27/2.6.10 and firewall 9091. Insecure default, so an unmodified deployment is exposed","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-26190"],"status":"curated","published":"2026-02-13"},{"id":"CVE-2026-3055","cve":"CVE-2026-3055","aliases":[],"title":"Citrix NetScaler ADC/Gateway: Insufficient input validation as SAML IdP","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Citrix NetScaler ADC/Gateway","year":"2026","cvss_score":9.8,"severity":"critical","kev":true,"impact":"Insufficient input validation as SAML IdP -> out-of-bounds read / memory disclosure","attack_vector":"Network (remote)","remediation":"Control-plane: patch; rotate SAML signing material and kill sessions","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-3055"],"status":"curated","published":"2026-03-23"},{"id":"CVE-2026-3059","cve":"CVE-2026-3059","aliases":[],"title":"SGLang (multimodal ZMQ broker): Unauthenticated RCE via `pickle.loads()` on the ZMQ broker","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"SGLang (multimodal ZMQ broker)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unauthenticated RCE via `pickle.loads()` on the ZMQ broker","attack_vector":"Unauthenticated network from any host that can reach the broker socket","remediation":"Upgrade. Providers running SGLang as a managed endpoint own this; tenants running their own own the patch but the provider owns fabric isolation","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-3059"],"status":"curated","fleet":{"ubiquity":"Very common - affects the disaggregated/multimodal serving paths that large GPU deployments specifically use","remediation_pain":"`daemon-restart` plus network re-segmentation - the ZMQ broker binds 0.0.0.0 with zero auth, so patching alone is insufficient without isolating the internal serving plane","pain_class":"daemon-restart","why_fleet_wide":"`pickle.loads()` runs immediately on any payload received by an unauthenticated all-interfaces ZMQ broker, so anything with pod-network reach owns every disaggregated serving node - the multi-node serving fabric is the blast radius"},"published":"2026-03-12"},{"id":"CVE-2026-3060","cve":"CVE-2026-3060","aliases":[],"title":"SGLang (encoder parallel disaggregation): Unauthenticated RCE via `pickle.loads()` in the disaggregation module","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"SGLang (encoder parallel disaggregation)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unauthenticated RCE via `pickle.loads()` in the disaggregation module","attack_vector":"Unauthenticated network on the intra-cluster fabric","remediation":"Upgrade; disaggregated serving multiplies unauthenticated internal planes","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-3060"],"status":"curated","published":"2026-03-12"},{"id":"CVE-2026-30623","cve":"CVE-2026-30623","aliases":[],"title":"LiteLLM (MCP server creation): RCE via MCP server registration","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LiteLLM (MCP server creation)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"RCE via MCP server registration","attack_vector":"Network user able to add an MCP server","remediation":"Upgrade; MCP registration is a code-execution primitive","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-30623"],"status":"curated","published":"2026-07-15"},{"id":"CVE-2026-31228","cve":"CVE-2026-31228","aliases":[],"title":"Kubeflow (ART component): RCE in the robustness evaluation function","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Kubeflow (ART component)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"RCE in the robustness evaluation function","attack_vector":"Customer-supplied evaluation config","remediation":"Upgrade ART","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-31228"],"status":"curated","published":"2026-05-12"},{"id":"CVE-2026-31229","cve":"CVE-2026-31229","aliases":[],"title":"Kubeflow (Adversarial Robustness Toolbox component): Insecure deserialization in the Kubeflow model-loading component","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Kubeflow (Adversarial Robustness Toolbox component)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Insecure deserialization in the Kubeflow model-loading component","attack_vector":"Customer-supplied model evaluated by a robustness pipeline","remediation":"Upgrade ART past 1.20.1","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-31229"],"status":"curated","published":"2026-05-12"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94","CWE-88"],"fleet":{"pain_class":"hot-patch"},"id":"CVE-2026-31230","cve":"CVE-2026-31230","aliases":[],"title":"Adversarial Robustness Toolbox (Kubeflow component, robustness_evaluation_fgsm_pytorch.py): The ART Kubeflow evaluation","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Adversarial Robustness Toolbox (Kubeflow component, robustness_evaluation_fgsm_pytorch.py)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The ART Kubeflow evaluation script passes the --clip_values and --input_shape arguments through eval(), so whoever controls those strings executes arbitrary Python inside the pipeline pod. On a GPU cluster that pod usually holds the pipeline service account and mounted dataset volumes, which is the payoff. This is the third of a family whose two siblings, CVE-2026-31228 and CVE-2026-31229, are already in the database.","attack_vector":"Anyone who can influence the pipeline configuration or the automated scripts that invoke the ART robustness evaluation step - so a tenant with pipeline-authoring rights, or an upstream component that builds the argument list from user-supplied data.","remediation":"Upgrade the adversarial-robustness-toolbox package past 1.20.1 wherever the Kubeflow ART component is deployed. Note this is a pip package pulled into pipeline images, so patching Kubeflow itself does not fix it - you have to rebuild the component images.","references":["https://github.com/Trusted-AI/adversarial-robustness-toolbox","https://nvd.nist.gov/vuln/detail/CVE-2026-31230"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-31402","cve":"CVE-2026-31402","aliases":[],"title":"Linux NFS server (nfsd, NFSv4.0 LOCK replay cache): A denied NFSv4.0 LOCK whose conflicting owner string is large","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux NFS server (nfsd, NFSv4.0 LOCK replay cache)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A denied NFSv4.0 LOCK whose conflicting owner string is large writes up to 944 bytes past a fixed 112-byte replay buffer, corrupting kernel slab memory on the NFS server. Two cooperating clients trigger it with no credentials, so a tenant can crash or potentially take over the file server that all jobs mount.","attack_vector":"Two NFSv4.0 clients that can reach the server: one takes a lock with a long owner string, the other requests a conflicting lock. Explicitly unauthenticated per the upstream analysis, so any compute node with the export reachable is enough.","remediation":"Update the storage server kernel to one carrying the nfsd replay-cache bounds fix and reboot it (fail over the export first if you run HA). Where no fixed kernel is available yet, force clients to NFSv4.1+ with vers=4.1 or higher, which does not use this replay cache.","references":["https://git.kernel.org/stable/c/f9fcb4441f6c02bb20c2eb340101e27dfe23607c","https://nvd.nist.gov/vuln/detail/CVE-2026-31402"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-415"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-31533","cve":"CVE-2026-31533","aliases":[],"title":"Linux kernel (net/tls): When the crypto engine backlogs a kTLS encrypt request, both the async completion callback and","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/tls)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"When the crypto engine backlogs a kTLS encrypt request, both the async completion callback and the synchronous error path clean up the same record. The double decrement corrupts the pending-operation sentinel, so the socket stops waiting for outstanding async encryptions entirely - a later send then frees the TLS record while the crypto callback is still holding it, and the callback fires on freed memory. Kernel heap use-after-free reachable by any process using kTLS, which on these nodes is the storage and control plane.","attack_vector":"Any local process that owns a socket and calls setsockopt(TLS_TX) - no privilege beyond owning the socket, so every tenant container has it - can drive this by pushing enough data to backlog the crypto engine. Loading the fleet's crypto queue is easy for a co-tenant, and the CNA scores it network-reachable because a peer that drives sustained TLS traffic contributes to the same backlog condition. No device node and no CAP_NET_ADMIN required.","remediation":"Update to 5.15.203 / 6.1.169 / 6.6.135 / 6.8 or later. Interim control: there is no clean one - kTLS is attachable by any socket owner. If a fleet cannot be rebooted promptly, disabling the kTLS ULP (blacklist tls) removes the reachable path at the cost of falling back to userspace TLS.","references":["https://git.kernel.org/stable/c/414fc5e5a5aff776c150f1b86770e0a25a35df3a","https://git.kernel.org/stable/c/02f3ecadb23558bbe068e6504118f1b712d4ece0","https://nvd.nist.gov/vuln/detail/CVE-2026-31533"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-34159","cve":"CVE-2026-34159","aliases":[],"title":"llama.cpp (RPC `deserialize_tensor`): RPC backend skips all bounds validation","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"llama.cpp (RPC `deserialize_tensor`)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"RPC backend skips all bounds validation","attack_vector":"Unauthenticated network to the RPC port","remediation":"Rebuild past b8492. Same unauthenticated-RPC class as 2024 — the design has not changed","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-34159"],"status":"curated","published":"2026-04-01"},{"id":"CVE-2026-35616","cve":"CVE-2026-35616","aliases":[],"title":"Fortinet FortiClient EMS: Improper access control","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Fortinet FortiClient EMS","year":"2026","cvss_score":9.8,"severity":"critical","kev":true,"impact":"Improper access control -> unauthenticated code execution via crafted requests","attack_vector":"Network (remote)","remediation":"Control-plane: patch FortiClientEMS 7.4.5-7.4.6","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-35616"],"status":"curated","published":"2026-04-04"},{"id":"CVE-2026-39808","cve":"CVE-2026-39808","aliases":[],"title":"Fortinet FortiSandbox: OS command injection via crafted HTTP requests","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Fortinet FortiSandbox","year":"2026","cvss_score":9.8,"severity":"critical","kev":true,"impact":"OS command injection via crafted HTTP requests -> unauthenticated code execution","attack_vector":"Network (remote)","remediation":"Control-plane: emergency firmware; same window as CVE-2026-25089","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-39808"],"status":"curated","published":"2026-04-14"},{"id":"CVE-2026-42208","cve":"CVE-2026-42208","aliases":[],"title":"LiteLLM proxy: SQL injection in a database query path","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LiteLLM proxy","year":"2026","cvss_score":9.8,"severity":"critical","kev":true,"impact":"SQL injection in a database query path","attack_vector":"Network to the LiteLLM proxy","remediation":"**[KEV]** Patch to 1.83.7+ immediately — known exploited. Any AI gateway a neocloud runs for tenants is in scope","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-42208"],"status":"curated","published":"2026-05-08"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-362","CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-43198","cve":"CVE-2026-43198","aliases":[],"title":"Linux kernel (net/ipv4): A child socket created from an inbound handshake is inserted into the TCP hash table before","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/ipv4)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A child socket created from an inbound handshake is inserted into the TCP hash table before its IPv6 private-data pointer is fixed up, so other CPUs can find and use a socket that still points at the listener's ipv6_pinfo. A remote peer that opens connections to any listener - including the SMC/MPTCP paths that share this code - drives other CPUs into type-confused socket state.","attack_vector":"Remote and pre-authentication: the window is inside tcp_v4_syn_recv_sock / tcp_v6_syn_recv_sock during the three-way handshake, so ordinary connection traffic to any TCP listener on the node is enough, with no credentials. This is core TCP rather than an optional module - it surfaced in this seam because net/smc/af_smc.c is one of the callers - so every node is exposed regardless of whether SMC or RDMA is in use.","remediation":"Boot a kernel carrying the fix commits (moves the IPv6 child fixup into tcp_v6_mapped_child_init, before ehash insertion). There is no meaningful interim control - the path is core TCP accept handling; patch and reboot.","references":["https://git.kernel.org/stable/c/fe89b2f05b854847784f91127319172945c1fadd","https://git.kernel.org/stable/c/858d2a4f67ff69e645a43487ef7ea7f28f06deae","https://nvd.nist.gov/vuln/detail/CVE-2026-43198"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-43465","cve":"CVE-2026-43465","aliases":["net/mlx5e RX XDP multi-buf frag counting for striding RQ"],"title":"Linux kernel mlx5_core RX datapath (striding RQ, page_pool): A regression introduced by the fix for CVE-2025-40350","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_core RX datapath (striding RQ, page_pool)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A regression introduced by the fix for CVE-2025-40350: dropped XDP fragments stopped being counted driver-side, so page_pool reference counts go negative and mlx5 releases 64 fragments against a refcount of 63. Remote packets corrupt page-pool refcounting in the host kernel. Worth flagging to operators as a pattern - patching the earlier RX bug without moving to a current stable re-exposes you.","attack_vector":"Unauthenticated remote sender to a node running XDP multi-buffer on mlx5 striding RQ, on a kernel that carries the CVE-2025-40350 fix but not this one.","remediation":"Upgrade the host kernel to 7.0 or a stable backport (6.18.19, 6.19.9). Rolling reboot. Do not stop at the CVE-2025-40350 fix level - verify your running kernel carries both.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43465","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-43465.json"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-05-08"},{"id":"CVE-2026-44484","cve":"CVE-2026-44484","aliases":[],"title":"PyTorch Lightning: Reintroduced unsafe deserialization in 2.6.2","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"PyTorch Lightning","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Reintroduced unsafe deserialization in 2.6.2","attack_vector":"Customer-supplied checkpoint","remediation":"Tenant-owned; provider blocks pickle-format checkpoints at ingest","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-44484"],"status":"curated","published":"2026-05-14"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-45898","cve":"CVE-2026-45898","aliases":[],"title":"Linux kernel (drivers/infiniband/core): The iWARP connection manager returns work items to a free list while the same","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/core)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The iWARP connection manager returns work items to a free list while the same items are still queued on the workqueue, so they get re-initialised under the workqueue's feet. The result is workqueue list corruption and a kernel BUG - a hard panic that takes the whole shared node down and every tenant's job with it. Corrupting kernel list pointers is also a classic escalation primitive, and the CNA rates it network-reachable with no privileges.","attack_vector":"Driven by ordinary connection churn on the fabric: any tenant using rdma_cm, or a remote peer opening and closing iWARP connections quickly, hits it. The upstream reproducer is plain ucmatose stress on an Intel E830 in iWARP mode. Requires iw_cm in use - irdma in iWARP mode, cxgb4, or soft-iWARP siw. Not reachable on pure RoCE/IB paths.","remediation":"Update to a stable kernel carrying 38c5b49fffa1 (or eb715133e0ae / a6b9e793e74e) and reboot. Interim: run irdma in RoCE rather than iWARP mode where the fabric allows it, and unload/blacklist siw on nodes that do not need soft-iWARP.","references":["https://git.kernel.org/stable/c/38c5b49fffa1b760959af74f11806eeb3ef4706d","https://git.kernel.org/stable/c/eb715133e0ae12514bba4d2d5ce1dee774476056","https://nvd.nist.gov/vuln/detail/CVE-2026-45898"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-46135","cve":"CVE-2026-46135","aliases":["CVE-2026-64534","CVE-2026-64535"],"title":"Linux kernel nvmet-tcp - ICReq/teardown race and data-digest error paths: Three lifecycle bugs an initiator can drive","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel nvmet-tcp - ICReq/teardown race and data-digest error paths","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Three lifecycle bugs an initiator can drive on purpose. Send an ICReq and immediately close the socket and the target's queue state gets flipped back to LIVE after teardown already started, defeating the DISCONNECTING guard and leading to use-after-free. Negotiate data digests and deliberately mismatch one on a non-final H2C_DATA PDU and the target either underflows a refcount into a permanent workqueue deadlock or leaves a half-uninitialised command that teardown uninits a second time. The workqueue deadlock variant is the one that hurts operationally: the target stops serving I/O for every tenant and only a reboot clears it, so a single initiator - including a tenant VM with a legitimately provisioned namespace - can take the whole storage node down on demand and repeat it after every restart.","attack_vector":"Any host that can open an NVMe/TCP connection to the target. For the digest variants the attacker needs a connection that negotiates data digests, which any initiator can request; no valid namespace or credential is required for the ICReq race.","remediation":"Kernel update on the target nodes and a reboot. Disabling data digests removes two of the three but not the ICReq race, and costs you the end-to-end integrity check, so it is a poor trade. Because these are cheap remote denial-of-service primitives rather than deep exploitation, prioritise them by blast radius: any target serving more than one tenant should be patched before targets serving one. Rate-limiting or per-initiator connection caps on the storage VLAN reduce the repeat rate but do not fix it.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git","https://nvd.nist.gov/vuln/detail/CVE-2026-46135","https://nvd.nist.gov/vuln/detail/CVE-2026-64535"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-05-28"},{"id":"CVE-2026-46325","cve":"CVE-2026-46325","aliases":["RDMA/rxe iova-to-va conversion error","Soft-RoCE MR page size mismatch"],"title":"Linux kernel - RDMA/rxe memory region translation, drivers/infiniband/sw/rxe/rxe_mr.c: Rxe mishandles memory regions","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel - RDMA/rxe memory region translation, drivers/infiniband/sw/rxe/rxe_mr.c","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Rxe mishandles memory regions whose page size differs from the system PAGE_SIZE - it steps the page list by mr->page_size while each stored entry actually represents PAGE_SIZE of memory. The result is that an IO virtual address resolves to the wrong physical page, so a legitimate-looking remote access lands on memory outside the intended region. This is the most dangerous shape a memory-registration bug can take: the rkey check passes, the access is authorised, and the data returned or overwritten belongs to something else entirely. It matters concretely on ARM64 GPU hosts (64K pages) and anywhere hugepages are used for RDMA buffers - which is standard practice for training-job memory. A reported kernel panic is the visible symptom; silent cross-region reads and writes are the security consequence.","attack_vector":"A remote peer issues ordinary RDMA reads or writes against a memory region registered with a page size differing from the host PAGE_SIZE. No malformed packets needed - the mistranslation happens in the victim's own code path. Reachability is whatever the rxe endpoint's reachability is; combined with the rkey weaknesses described in the ReDMArk entry, an attacker who guesses into a region gets misdirected access on top of unauthorised access.","remediation":"Host reboot / kernel upgrade. Interim: unload and blacklist rdma_rxe if Soft-RoCE is not in deliberate use (config change, no downtime). If rxe is required, avoid registering regions whose page size differs from PAGE_SIZE until patched - in practice that means not backing RDMA buffers with hugepages on affected kernels, which costs performance but is a same-day application/config change. Prioritise ARM64 GPU hosts and any node with 64K pages.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2026/CVE-2026-46325.json","https://nvd.nist.gov/vuln/detail/CVE-2026-46325"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-06-09"},{"id":"CVE-2026-47627","cve":"CVE-2026-47627","aliases":[],"title":"NVIDIA Triton Inference Server: A path traversal scored 9.8 (network, no privileges, full","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A path traversal scored 9.8 (network, no privileges, full confidentiality/integrity/availability impact). NVIDIA's summary understates it as denial of service while the vector says complete compromise; treat the vector as authoritative and patch this first among the Triton set. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5865. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47627","https://github.com/NVIDIA/product-security/tree/main/2026/5865"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-22"],"published":"2026-08-18"},{"id":"CVE-2026-49468","cve":"CVE-2026-49468","aliases":[],"title":"LiteLLM proxy: Host-header parsing flaw in the proxy","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LiteLLM proxy","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Host-header parsing flaw in the proxy","attack_vector":"Unauthenticated network","remediation":"Upgrade to 1.84.0+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-49468"],"status":"curated","published":"2026-06-22"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-306","CWE-78"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2026-49980","cve":"CVE-2026-49980","aliases":[],"title":"rclone (rcd remote control server): An unauthenticated request to the rclone remote-control server instantiates a","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"rclone (rcd remote control server)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"An unauthenticated request to the rclone remote-control server instantiates a backend defined inline in the request, and some backend options run commands. That is unauthenticated arbitrary command execution as whatever user rclone runs as - typically the data-mover service account holding credentials to every storage system it touches.","attack_vector":"Anyone who can reach the rclone rcd HTTP port. Data movers are commonly run on a shared node with the port bound broadly, so a tenant on the same network is enough.","remediation":"Upgrade rclone to the fixed release and restart every rcd/serve instance. Treat any exposed instance as compromised and rotate all remote credentials in its config. Bind rcd to loopback, require --rc-user/--rc-pass, and never expose it on a tenant-reachable interface.","references":["https://github.com/rclone/rclone/security/advisories/GHSA-qw24-gh76-8rvv","https://nvd.nist.gov/vuln/detail/CVE-2026-49980"],"status":"curated"},{"id":"CVE-2026-51080","cve":"CVE-2026-51080","aliases":[],"title":"Proxmox VE (libpve-storage-perl XXE): XML external entity injection in the Proxmox storage library, reachable","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Proxmox VE (libpve-storage-perl XXE)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"XML external entity injection in the Proxmox storage library, reachable unauthenticated, leading to full compromise. Proxmox is increasingly used as the hypervisor for smaller GPU clouds where vSphere licensing does not pencil out.","attack_vector":"Unauthenticated network access to the affected Proxmox service.","remediation":"Upgrade libpve-storage-perl past v9.1.1 / v8.3.7 via the Proxmox enterprise or no-subscription repo. Package update plus service restart; no VM downtime required.","references":["https://forum.proxmox.com/threads/proxmox-virtual-environment-security-advisories.149331/post-849970"],"status":"curated","fleet":{"pain_class":"daemon-restart"}},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-390","CWE-908"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-52989","cve":"CVE-2026-52989","aliases":[],"title":"Linux kernel NVMe-oF TCP target (nvmet-tcp, unpropagated PDU iovec build errors): Nvmet_tcp_build_pdu_iovec() detects","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel NVMe-oF TCP target (nvmet-tcp, unpropagated PDU iovec build errors)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Nvmet_tcp_build_pdu_iovec() detects an out-of-bounds PDU length or offset, flags a fatal error - and returns void. The caller never learns, sets the queue into receive-data state anyway, and the socket loop then reads attacker-supplied network bytes into an uninitialised iov_iter. A remote initiator that sends a malformed PDU length therefore steers kernel writes through an iterator whose contents are whatever was on the stack. The kernel CNA rates it network, unauthenticated, full CIA, and the reasoning is visible in the fix.","attack_vector":"Remote, unauthenticated, by sending a PDU with an out-of-range length or data offset to the nvmet-tcp listener.","remediation":"Kernel update making the iovec builder return an error and the callers honour it. Same network containment applies: the storage target's listener must be unreachable from tenant networks.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=046fa5c72d15cd8e2d592e275697ea399d8f76b0","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-52989.json"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-53176","cve":"CVE-2026-53176","aliases":["IB/isert short login PDU","iSER target pre-authentication crash"],"title":"Linux kernel - iSER (iSCSI Extensions for RDMA) target, drivers/infiniband/ulp/isert/ib_isert.c: The iSER target","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel - iSER (iSCSI Extensions for RDMA) target, drivers/infiniband/ulp/isert/ib_isert.c","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The iSER target accepted login PDUs shorter than ISER_HEADERS_LEN and parsed them anyway, giving out-of-bounds access in the kernel before any authentication has occurred. Login is by definition the pre-auth surface, so anyone who can reach the iSER target port gets remote kernel compromise on the storage node with no credentials at all. A storage node in a GPU cluster typically mounts and serves many tenants' volumes, so compromising it is equivalent to compromising every dataset it fronts.","attack_vector":"Connect to the iSER target and send a truncated login PDU. Fully pre-authentication, no valid initiator identity required, reachable from anywhere on the storage fabric. If the storage fabric is not separated from the tenant fabric - a common shortcut - this is reachable from tenant workloads directly.","remediation":"Host reboot / kernel upgrade on iSER target nodes, treated as urgent given it is pre-auth and scored 9.8. Immediate config controls while you schedule the reboot: restrict the iSER/iSCSI target port to known initiator addresses at the switch and host firewall, and place storage targets on a fabric partition tenants cannot reach. If iSER is unused, unload ib_isert and disable the target configuration.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2026/CVE-2026-53176.json","https://nvd.nist.gov/vuln/detail/CVE-2026-53176"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-06-25"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-415"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-53355","cve":"CVE-2026-53355","aliases":[],"title":"Linux kernel (net/rds): If RDS/IB queue-pair setup fails after the send ring is allocated but before the receive ring","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/rds)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"If RDS/IB queue-pair setup fails after the send ring is allocated but before the receive ring is, the unwind frees the send ring without clearing the pointer. The teardown logic uses NULL to decide what it still owns, so a later shutdown pass frees the same ring again - a double free / use-after-free of RDMA send state reachable by whoever can make connection setup fail partway.","attack_vector":"Driven from the fabric: RDS/IB connection setup runs from the RDMA CM event handler when a peer connects, and the failure window is a partial setup (resource exhaustion, device limits, a peer that aborts mid-setup), followed by the normal shutdown path. A tenant can also drive repeated RDS connection attempts locally - socket(AF_RDS, ...) is unprivileged and autoloads rds/rds_rdma via the net-pf-21 alias. Requires the RDS RDMA transport in use over an IB/RoCE device.","remediation":"Boot a kernel carrying the fix commits (clears i_sends after vfree in the setup unwind). Interim: blacklist rds and rds_rdma, or deny socket family 21 to tenants; RDS over IB is rarely a deliberate dependency on an AI cluster.","references":["https://git.kernel.org/stable/c/66cccec111421a10efdc2c74499d15b93e7acae5","https://git.kernel.org/stable/c/29d940026dce39e3018dab6f67c9427249321270","https://nvd.nist.gov/vuln/detail/CVE-2026-53355"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-53363","cve":"CVE-2026-53363","aliases":[],"title":"Linux kernel (net/xfrm): IPTFS fragment consumption loses the shared-page marker, so ESP concludes the payload pages","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"IPTFS fragment consumption loses the shared-page marker, so ESP concludes the payload pages are privately owned and decrypts in place over pages that are still referenced elsewhere - typically read-only page-cache pages. Inbound IPsec traffic therefore silently rewrites memory belonging to other work on the node: file cache corruption for any tenant on that host, and kernel panics once the damaged pages are used.","attack_vector":"Driven by inbound ESP traffic on an IPTFS-mode SA, so any peer that can put packets on the SA reaches it - a node-to-node encryption peer inside the cluster, a compromised node, or the far end of a tenant overlay tunnel. Conditional on IP-TFS mode being in use (CONFIG_XFRM_IPTFS / xfrm_iptfs loaded and an SA configured with mode iptfs). No tenant device node is required; a tenant on the far side of the tunnel or a compromised peer node is enough.","remediation":"Boot a kernel carrying the fix commits below (no fixed stable version is published in the record - match by commit with your distro kernel). Interim control: stop using IPTFS mode for node-to-node and tenant-overlay SAs and fall back to plain ESP tunnel mode until the fleet is patched.","references":["https://git.kernel.org/stable/c/dd66f7f6e360ee82cd905517726f8e9091265de5","https://git.kernel.org/stable/c/c885d111ed9f5a0a1f3cc4e87a50db6518abaa6c","https://nvd.nist.gov/vuln/detail/CVE-2026-53363"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-53398","cve":"CVE-2026-53398","aliases":[],"title":"Linux NFS server (nfsd, SECINFO_NO_NAME decode): A truncated SECINFO_NO_NAME operation leaves sin_exp uninitialized and","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux NFS server (nfsd, SECINFO_NO_NAME decode)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A truncated SECINFO_NO_NAME operation leaves sin_exp uninitialized and the error path then acts on it, corrupting kernel memory on the file server. A malformed compound from an unauthenticated client is enough.","attack_vector":"Any host that can send NFSv4 traffic to the server's RPC port.","remediation":"Update the storage server kernel to one with the decode cleanup fix and reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53398","https://git.kernel.org/pub/scm/linux/security/vulns.git"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2026-53805","cve":"CVE-2026-53805","aliases":["GHSA-478c-rj3v-9229"],"title":"NVIDIA GEN3C (Spatial Intelligence Lab) inference API server: The inference API server runs Python pickle.loads()","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GEN3C (Spatial Intelligence Lab) inference API server","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The inference API server runs Python pickle.loads() directly on raw HTTP request bodies at /request-inference and /seed-model, with no authentication in front of either. Anyone who can reach the port executes arbitrary code as the inference process - which is the process that holds the GPU, the model weights on disk, and whatever cloud credentials the pod was handed. There is no escalation step and no credential to steal first; the attacker is inside the serving process on the first request. On a shared cluster this also becomes a lateral-movement path, because any neighbouring pod that can route to the service gets the same unauthenticated shell.","attack_vector":"Network reach to the GEN3C inference API port. No account, no credentials, no user interaction, low complexity. In practice that means every pod on the cluster network, plus the public internet wherever the service was fronted by a LoadBalancer or a permissive NodePort. Nothing about the endpoint signals that it is dangerous, so it tends to be left open on the assumption that it is internal.","remediation":"There is no NVIDIA PSIRT bulletin and no fixed release tag for this - GEN3C is a research repository, so the fix exists only as upstream commits (pull requests 62 and 63, commit db2ffe12). Rebuild your image from a revision that includes that commit and redeploy the service. Because there is no version to pin, treat this as a supply-chain check rather than a patch: pin the commit hash in your build and re-verify it. Immediately and regardless of patching, put the inference port behind an authenticated proxy and a network policy - do not leave it reachable from general cluster traffic.","references":["https://www.vulncheck.com/advisories/nvidia-sil-gen3c-unauthenticated-rce-via-pickle-deserialization-in-inference-api","https://github.com/advisories/GHSA-478c-rj3v-9229","https://github.com/nv-tlabs/GEN3C/commit/db2ffe12ced12ddafcec5e0422ee46ce8520746b","https://nvd.nist.gov/vuln/detail/CVE-2026-53805"],"status":"curated"},{"id":"CVE-2026-5760","cve":"CVE-2026-5760","aliases":[],"title":"SGLang (`/v1/rerank`): RCE via a malicious `tokenizer.chat_template` rendered as Jinja2","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"SGLang (`/v1/rerank`)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"RCE via a malicious `tokenizer.chat_template` rendered as Jinja2","attack_vector":"Customer-supplied model file — the chat template inside the model repo is the payload","remediation":"Upgrade. Jinja chat templates are code; scanning the weights does not cover the tokenizer config","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-5760"],"status":"curated","published":"2026-04-20"},{"id":"CVE-2026-59309","cve":"CVE-2026-59309","aliases":[],"title":"VMware vCenter (VMware Directory Service authentication bypass): An unauthenticated attacker with network access","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"VMware vCenter (VMware Directory Service authentication bypass)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"An unauthenticated attacker with network access to vCenter bypasses authentication entirely and gains access to the system. vCenter owns every VM and host in the cluster, so this is total virtualization-estate compromise. Public reporting describes active exploitation at scale across many countries.","attack_vector":"Network access to vCenter. No credentials required.","remediation":"Apply the Broadcom fix immediately - this and its sibling directory-traversal bug are being exploited in the wild. vCenter appliance patch plus restart. Given the exploitation reports, do not treat patching alone as sufficient: hunt for new/modified SSO accounts, unexpected scheduled tasks and altered vpxd logs before calling it clean.","references":["https://support.broadcom.com/web/ecx/support-content-notification/-/external/content/SecurityAdvisories/0/38017"],"status":"curated"},{"id":"CVE-2026-59310","cve":"CVE-2026-59310","aliases":[],"title":"VMware vCenter (Syslog server directory traversal to RCE): Directory traversal in the vCenter syslog server letting","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"VMware vCenter (Syslog server directory traversal to RCE)","year":"2026","cvss_score":9.8,"severity":"critical","kev":true,"impact":"Directory traversal in the vCenter syslog server letting an unauthenticated network attacker execute arbitrary code on vCenter. Chained with the authentication bypass in the same advisory, this is the full remote-takeover pair - and it is being exploited in the wild.","attack_vector":"Network access to vCenter. Unauthenticated.","remediation":"Apply the Broadcom fix immediately. vCenter patch and restart. Assume compromise on any vCenter that was network-exposed and unpatched during the exploitation window; rotate vCenter and ESXi credentials afterwards.","references":["https://support.broadcom.com/web/ecx/support-content-notification/-/external/content/SecurityAdvisories/0/38017"],"status":"curated"},{"id":"CVE-2026-63077","cve":"CVE-2026-63077","aliases":[],"title":"JetBrains TeamCity: Deserialization in the agent polling protocol","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"JetBrains TeamCity","year":"2026","cvss_score":9.8,"severity":"critical","kev":true,"impact":"Deserialization in the agent polling protocol -> unauthenticated remote code execution on the CI server","attack_vector":"Network (remote)","remediation":"Control-plane: URGENT upgrade to 2026.1.3/2025.11.7; rotate all build secrets and signing keys","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63077"],"status":"curated","published":"2026-07-27"},{"id":"CVE-2026-63886","cve":"CVE-2026-63886","aliases":["LIO iSCSI CHAP_R base64 overflow","iscsi_target_auth pre-auth heap overflow"],"title":"Linux kernel - LIO iSCSI target CHAP authentication, drivers/target/iscsi/iscsi_target_auth.c","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel - LIO iSCSI target CHAP authentication, drivers/target/iscsi/iscsi_target_auth.c","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Chap_server_compute_hash() allocated the client digest buffer at the hash's digest size, then passed a base64 CHAP_R response to the decoder without checking whether the input could produce more output than that. Up to 127 base64 characters decode to 95 bytes - a 63-byte overflow for SHA-256, 79 bytes for MD5 - and the existing length check fired only after the write had already happened. This is a heap overflow reached during CHAP authentication, meaning pre-authentication by definition: anyone who can open an iSCSI session to the target gets a controlled kernel heap write on the storage node. The HEX branch of the same switch already validated its length, so the bug was a straightforward omission on the base64 path.","attack_vector":"Open an iSCSI session to the LIO target and send a CHAP_R value with the '0b' base64 prefix and a long payload. No valid credentials required - the overflow happens while the target is still working out whether your credentials are valid. Reachable from any host that can reach the iSCSI port.","remediation":"Host reboot / kernel upgrade on all LIO iSCSI target nodes - urgent, pre-auth, 9.8. Interim controls: restrict the iSCSI target port to known initiator addresses at the firewall and switch, and if CHAP is not required, note that disabling it does not help since the parser is reached during login negotiation. Move iSCSI targets off tenant-reachable networks. Where the iSCSI target is legacy, decommissioning it is the cleanest fix.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2026/CVE-2026-63886.json","https://nvd.nist.gov/vuln/detail/CVE-2026-63886"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-07-19"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-64033","cve":"CVE-2026-64033","aliases":[],"title":"Linux kernel (drivers/infiniband/ulp/rtrs): On the RTRS server, a failure while publishing a new session's sysfs","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/ulp/rtrs)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"On the RTRS server, a failure while publishing a new session's sysfs entries drops the last reference to the session object and then keeps using it - a use-after-free in the storage target's own address space, driven by a peer establishing a session. The CNA scores it network-reachable with no privileges and full compromise.","attack_vector":"Any host that can reach the RTRS server on the storage fabric - i.e. a tenant node or a compromised client of an RNBD/RTRS block export - triggers this by opening sessions and driving the setup failure path. Requires the rtrs_server module loaded, so it only affects nodes acting as RTRS/RNBD targets.","remediation":"Update to 5.15.209 or later, or a stable kernel carrying 01e42aabaf76 / 548f3956e53a, and reboot. Interim: stop exporting RTRS/RNBD targets from the affected nodes, or restrict which fabric peers can reach the RTRS listener until patched.","references":["https://git.kernel.org/stable/c/01e42aabaf7632beb4bf235c7238b96c746d4144","https://git.kernel.org/stable/c/548f3956e53a7f7bde912d8129010b8986d5e602","https://nvd.nist.gov/vuln/detail/CVE-2026-64033"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787","CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-64046","cve":"CVE-2026-64046","aliases":[],"title":"Linux kernel (net/tls): A particular kTLS ring state builds a scatterlist whose chain link points directly at another","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/tls)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A particular kTLS ring state builds a scatterlist whose chain link points directly at another chain link. The scatterlist iterator does not resolve chained links recursively, so this is illegal input to the crypto API - the AEAD walk runs off into whatever the second link descriptor happens to be. The practical outcome is the crypto engine reading and writing memory outside the record's buffers during TLS encryption, on the same code path that carries storage and control traffic.","attack_vector":"Any unprivileged local process that enables kTLS on its socket with setsockopt(TLS_TX) and drives the send ring into the wrapped state (ring end at zero with a non-zero start). Every tenant container can do this - kTLS attachment needs no capability. TLS 1.3 records make the condition easier to hit because they consume the reserved wrap slot for content-type chaining.","remediation":"Boot a kernel carrying the fix commits below - the record's version field (5.5) is the introducing release, not a fix. Interim control: blacklist the tls ULP module so tenants cannot attach kTLS, at the cost of userspace TLS performance.","references":["https://git.kernel.org/stable/c/49a5faaa471ddcd37b6893970c9916eb836e7c31","https://git.kernel.org/stable/c/91359966e247c0244c66d50bbb8e74aefa4321c3","https://nvd.nist.gov/vuln/detail/CVE-2026-64046"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-193","CWE-787"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-64047","cve":"CVE-2026-64047","aliases":[],"title":"Linux kernel (net/tls): When the kTLS transmit scatterlist ring wraps, the chain link that stitches the tail back to","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/tls)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"When the kTLS transmit scatterlist ring wraps, the chain link that stitches the tail back to the head is planted one entry short of the true last slot. The crypto layer is then handed a scatter list that describes the wrong memory, so a TLS record is encrypted from - and written into - buffers the socket does not own. That is attacker-influenceable memory corruption plus plaintext confusion on the encrypted control and storage plane: bytes from an unrelated buffer can end up inside a record on the wire.","attack_vector":"Any local process that owns a socket and enables kTLS via setsockopt(TLS_TX), which every tenant container can do without privilege. The bug needs the send ring to wrap, which happens naturally with sustained scatter-gather sends; a tenant can arrange it deliberately with a chosen write pattern. The CNA scores it network-reachable because a peer's flow-control behaviour shapes when the ring wraps. No device node required.","remediation":"Boot a kernel carrying the fix commits below - the version field in the record (5.5) marks where the flaw was introduced, not a fixed release, so match by commit against your vendor kernel. Interim control: blacklist the tls ULP so tenant workloads use userspace TLS until nodes are rebooted.","references":["https://git.kernel.org/stable/c/73963a375885d5ccb7def39fd0b4f542e0f343dd","https://git.kernel.org/stable/c/47110c3a9ac247b688657337f5981efcfcb240dc","https://nvd.nist.gov/vuln/detail/CVE-2026-64047"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-64102","cve":"CVE-2026-64102","aliases":["RDMA/siw MPA FPDU length underflow"],"title":"Linux kernel - RDMA/siw (soft-iWARP) MPA framing, drivers/infiniband/sw/siw/siw_qp_rx.c: The siw receive path decodes","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel - RDMA/siw (soft-iWARP) MPA framing, drivers/infiniband/sw/siw/siw_qp_rx.c","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The siw receive path decodes the MPA framed-PDU length from the wire and feeds it into signed arithmetic against the per-fragment remainder without rejecting an underflow first. A remote peer that sends a crafted FPDU length drives the subsequent length math negative, which downstream is used to size receive operations - the classic path to out-of-bounds kernel access from a network packet. Same reachability profile as the Read Response bug in the same file: remote, unauthenticated relative to the connection, over plain TCP, giving full compromise of the victim kernel.","attack_vector":"Any peer on an established siw connection sends an MPA FPDU whose declared length underflows the remaining fragment length. Because siw rides ordinary TCP, this is reachable from anything the node peers with, including across routed networks if the RDMA port is not firewalled. No local access and no credentials on the victim.","remediation":"Host reboot / kernel upgrade. Same practical advice as the sibling siw bug: unload and blacklist the siw module wherever soft-iWARP is not deliberately in use - a config change with zero downtime that eliminates both findings at once. Where siw is required, upgrade the kernel and reboot on a rolling drain; also firewall the iWARP TCP ports to known peers as an interim control.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2026/CVE-2026-64102.json","https://nvd.nist.gov/vuln/detail/CVE-2026-64102"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-07-19"},{"id":"CVE-2026-64122","cve":"CVE-2026-64122","aliases":[],"title":"Linux kernel mlx5_core TX timeout devlink health reporter: The TX timeout recovery handler accesses the netdev pointer","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_core TX timeout devlink health reporter","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The TX timeout recovery handler accesses the netdev pointer after the channel and its send queues have already been torn down and freed - a use-after-free in the code path that is supposed to recover the NIC from a stall. So the failure mode is: the NIC wedges under load, the recovery machinery fires, and instead of recovering it corrupts host kernel memory.","attack_vector":"Remote and unauthenticated in effect - an attacker who can drive enough traffic to induce a TX timeout on the mlx5 interface reaches the recovery path. No credentials on the host.","remediation":"Upgrade the host kernel to a build carrying the fix (mainline 7.x and current stable series). Rolling reboot of the fleet. There is no useful config mitigation - you cannot safely turn off TX timeout recovery.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-64122","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-64122.json"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-07-19"},{"id":"CVE-2026-64268","cve":"CVE-2026-64268","aliases":["RDMA/siw Read Response overrun","soft-iWARP RREAD sink overflow"],"title":"Linux kernel - RDMA/siw (soft-iWARP), drivers/infiniband/sw/siw/siw_qp_rx.c: Siw places inbound Read Response segments","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel - RDMA/siw (soft-iWARP), drivers/infiniband/sw/siw/siw_qp_rx.c","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Siw places inbound Read Response segments at the sink buffer without checking the running total against the sink length on continuation segments - the sink is validated only on the first fragment and the cumulative length only on the last. A connected peer that answers an outstanding RDMA READ with segments that never set the DDP Last flag, carrying more payload than was requested, walks the write pointer past the validated buffer and writes attacker-controlled data out of bounds in the kernel. siw runs iWARP over ordinary routable TCP, so the attacker is simply the far end of an established connection and needs no local privilege on the victim. Remote kernel memory corruption with full confidentiality, integrity and availability loss - the worst-scored item in this slice.","attack_vector":"The attacker is the remote peer of an siw connection, reachable over normal TCP - which means anything the victim connects out to, including a storage target or a peer node in a job, can attack back. They respond to any RDMA READ the victim issues with a stream of Read Response segments with DDP Last clear and total length exceeding the RREAD length. No authentication is involved; the connection itself is the only prerequisite. Note siw is a software iWARP provider commonly loaded for testing, for nodes without an RNIC, and by blktests/nvme-over-fabrics setups - it is often present on GPU nodes that do not intend to use it.","remediation":"Host reboot / kernel upgrade to a version carrying the fix (backported across stable trees) - no firmware, no switch work. Immediate zero-downtime mitigation: if you are not deliberately using soft-iWARP, blacklist and unload the module (rmmod siw; add to modprobe blacklist), which removes the attack surface entirely and costs nothing on the vast majority of GPU nodes. Audit with lsmod across the fleet - siw is frequently loaded by dependency rather than by intent. If siw is in use, treat the kernel upgrade as urgent and drain nodes in waves.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2026/CVE-2026-64268.json","https://nvd.nist.gov/vuln/detail/CVE-2026-64268"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-07-25"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-64534","cve":"CVE-2026-64534","aliases":[],"title":"Linux kernel (drivers/nvme/target): A client connected to your NVMe-oF TCP target can drive a reference-count underflow","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/nvme/target)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A client connected to your NVMe-oF TCP target can drive a reference-count underflow on a target-side request object, giving a use-after-free on the shared storage node and, past that, a permanently wedged workqueue. Once the workqueue is stuck the target stops serving every subsystem it exports, so one tenant's malformed traffic becomes a storage outage for everyone homed on that node.","attack_vector":"Reachable by any peer on the IP/storage fabric that can open a TCP connection to the nvmet-tcp listener - the trigger is a command that fails nvmet_req_init (an unsupported opcode is enough) followed by a deliberately wrong data digest, both fully under the sender's control and both landing before any meaningful authorization on the command. Requires the node to be running nvmet with a tcp port enabled and data digest negotiated; no tenant device node is needed, only network reach to the target port.","remediation":"No fixed version is listed on this record beyond the 5.10.261 / 5.12 stable branches - boot a kernel carrying the linked stable commits. Interim: restrict the nvmet-tcp listener to a management/storage VLAN that tenant workloads cannot address, or stop exporting the nvmet subsystem until the node is patched.","references":["https://git.kernel.org/stable/c/22ec7a9fe9153d2737ee9b2fa6d2e43a1491decf","https://git.kernel.org/stable/c/ba35b1c674ca3841c0dfadd698f2c1b3ec542d4e","https://nvd.nist.gov/vuln/detail/CVE-2026-64534"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-64535","cve":"CVE-2026-64535","aliases":[],"title":"Linux kernel NVMe-oF TCP target (nvmet-tcp, data-digest mismatch handling): With data digests enabled, a digest","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel NVMe-oF TCP target (nvmet-tcp, data-digest mismatch handling)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"With data digests enabled, a digest mismatch on a non-final H2C_DATA PDU makes the error handler call nvmet_req_uninit() - dropping the submission-queue percpu reference - without marking the command completed. Queue teardown then walks the command list, still sees the command as needing data, and touches it again: use-after-free on the storage target that serves the cluster's datasets and checkpoints. The attacker only needs to be able to open an NVMe/TCP connection and send a bad digest, so this is unauthenticated remote memory corruption in the process that has every tenant's namespaces attached.","attack_vector":"Remote and unauthenticated. Any host that can reach the nvmet-tcp listener - and on most clusters the storage VLAN is reachable from compute nodes, which means from tenant containers with host networking or from a compromised co-tenant job.","remediation":"Kernel update on the storage target nodes making the digest-error path mark the command completed. Interim controls that actually work: turn off data digests on the affected subsystems (removes the trigger), and restrict the nvmet listener to the storage network with an explicit host-NQN allow list rather than accepting any connecting initiator.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=088ee46c18d99baef453afd74181dd40ade044ad","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-64535.json"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-64541","cve":"CVE-2026-64541","aliases":[],"title":"Linux kernel SMC-R connection data control (smc_cdc_rx_handler socket lifetime): The CDC receive handler looks the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel SMC-R connection data control (smc_cdc_rx_handler socket lifetime)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The CDC receive handler looks the connection up by token under the link group's lock, drops the lock, and only then dereferences the connection and its socket - holding no reference across the gap. A concurrent close on the local side frees the socket in that window, and the remote peer decides when the CDC message arrives, so a peer on the RDMA fabric times its send against a closing connection to get a use-after-free. The token is carried in the wire message, meaning the peer also chooses which connection to aim at.","attack_vector":"Remote over the SMC-R RDMA link. Requires an established link group with the target, which any peer that completes an SMC handshake has.","remediation":"Kernel update pinning the socket across the lock drop. Keep SMC off tenant-reachable paths if it is not deliberately in use.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=1951bffbc6493ec34cff3956b29d4bc6606904a6","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-64541.json"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-64566","cve":"CVE-2026-64566","aliases":[],"title":"Linux kernel (net/xfrm): The same ownership-marker bug as CVE-2026-53363, in the other IPTFS frag-transfer helper.","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The same ownership-marker bug as CVE-2026-53363, in the other IPTFS frag-transfer helper. Inner packets built from a page-pool-backed receive path look privately owned to ESP, so a nested transport-mode SA decrypts in place and writes over pages the outer IPTFS skb still references. That is attacker-triggered kernel memory corruption from the wire, with a panic as the visible outcome and silent data corruption as the quiet one.","attack_vector":"Inbound ESP/IPTFS traffic. Any peer holding the SA - a peer node on the cluster fabric, or a tenant endpoint terminating an overlay tunnel - drives it by sending traffic that lands on a page-pool receive path (i.e. a normal modern NIC driver). Conditional on IPTFS mode plus a nested transport-mode SA. No local access and no privilege on the target node is required.","remediation":"Boot a kernel carrying the fix commits below (no fixed stable version is published in the record). Both this and CVE-2026-53363 must land - they are two separate call sites. Interim control: drop IPTFS mode and nested transport-mode SAs from the encryption design until patched.","references":["https://git.kernel.org/stable/c/d8aaf06b29f5a0b6186cf68d21c7d63678ee3891","https://git.kernel.org/stable/c/ffd64e0717efd83fbf3396ab4e5ac6d795dac4d0","https://nvd.nist.gov/vuln/detail/CVE-2026-64566"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-65791","cve":"CVE-2026-65791","aliases":["CVE-2026-65679","CVE-2026-65796","CVE-2026-65681"],"title":"Windows iSCSI Target Service (Windows Server 2012 through Windows Server 2025 / Windows 10 1607+): Three heap-based","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Windows iSCSI Target Service (Windows Server 2012 through Windows Server 2025 / Windows 10 1607+)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Three heap-based buffer overflows allowing an unauthorized attacker to execute code over the network against the Windows iSCSI Target Service, plus a null-dereference denial of service in the same August 2026 batch. Unauthenticated network RCE against the service that owns every exported virtual disk on the box - the attacker gets SYSTEM on the storage server and, with it, read/write to every tenant VHD it serves. Relevant to GPU operators running Windows-based storage nodes or Hyper-V clusters backing GPU VMs, and to the long tail of Server 2012-era boxes still exporting iSCSI for management infrastructure.","attack_vector":"Any host that can reach TCP 3260 on the Windows storage server. No credentials.","remediation":"Apply the August 2026 Windows cumulative update on every server running the iSCSI Target Service role and reboot - which drops all iSCSI sessions and any VM or host booting from those LUNs, so drain first. If the role is enabled but unused (a common leftover on general-purpose Windows servers), remove the role instead of patching it; that permanently deletes the exposure. Restrict 3260 to known initiator addresses with Windows Firewall regardless of patch state. Note Server 2012 is out of mainstream support - confirm your ESU covers this batch or plan the migration.","references":["https://msrc.microsoft.com/update-guide/vulnerability/CVE-2026-65791","https://nvd.nist.gov/vuln/detail/CVE-2026-65791"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-11"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","fleet":{"pain_class":"node-drain"},"id":"CVE-2026-68160","cve":"CVE-2026-68160","aliases":[],"title":"CephFS kernel client (ceph.ko, ceph_handle_caps): The kernel trusts snap_trace_len straight off the wire, so a","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"CephFS kernel client (ceph.ko, ceph_handle_caps)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The kernel trusts snap_trace_len straight off the wire, so a malicious or compromised MDS returns 0xFFFFFFFF and the client reads far past the message buffer before any authentication state is checked. Result is a kernel out-of-bounds read on every compute node mounting CephFS - crash at minimum, memory disclosure at worst.","attack_vector":"Anything that can speak MDS protocol to a CephFS kernel client: a compromised or spoofed MDS, or an attacker with a foothold on the storage fabric who can inject cap messages toward compute nodes.","remediation":"Patch the kernel on all CephFS client nodes and reboot them. In the meantime run msgr2 secure mode so cap messages cannot be injected by a non-cluster party, and keep the Ceph public network unreachable from tenant workloads.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68160","https://git.kernel.org/pub/scm/linux/security/vulns.git"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-787"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-68426","cve":"CVE-2026-68426","aliases":[],"title":"Linux kernel (net/xfrm): When IPsec crypto offload takes a GSO segment asynchronously, the segment is unlinked from the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"When IPsec crypto offload takes a GSO segment asynchronously, the segment is unlinked from the list but the list head's tail pointer still points at it. The generic transmit path then writes through that pointer into memory the crypto engine already owns and may have freed - a use-after-free write in the normal egress path of every encrypted flow on the node. This is corruption on the send side of fabric encryption, so it damages whatever the slab hands out next, across tenants.","attack_vector":"Triggered by ordinary large (GSO) traffic leaving a node through an IPsec-offloaded device whenever the async crypto path claims the final segment. Conditional on IPsec offload being enabled (hardware or async software crypto engine). Any tenant that can generate bulk TCP egress over the encrypted fabric drives the condition; nothing privileged is required and no attacker-supplied packet content is needed, which also makes it fire on its own under load.","remediation":"Boot a kernel carrying the fix commits below (no fixed stable version published). Interim control: disable IPsec crypto offload on the NIC (turn off the esp-hw-offload feature with ethtool, or drop XFRM_OFFLOAD from SA configuration) so segments are not stolen asynchronously.","references":["https://git.kernel.org/stable/c/33e1b0d25ca0d2818c635ff80e6aa0d295e08a98","https://git.kernel.org/stable/c/bbca7cc3b2b4b10afbfee99b81d9ee78f5423046","https://nvd.nist.gov/vuln/detail/CVE-2026-68426"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-131"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-72041","cve":"CVE-2026-72041","aliases":[],"title":"Linux kernel (net/xfrm): The ESP-in-TCP send path mis-tracked scatter-gather message offsets and socket memory charges","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The ESP-in-TCP send path mis-tracked scatter-gather message offsets and socket memory charges whenever a send completed only partially, rather than using the helper that keeps the message consistent at every step. Wrong offsets on a partially-sent message are the same defect family as CVE-2026-52935 in this file, and the CNA scores this one as remotely reachable with no privileges. Read it as: the transmit state describing where encrypted bytes live can disagree with reality, with socket accounting drifting alongside it.","attack_vector":"Any process that attaches the espintcp ULP to a TCP socket (no capability required beyond owning the socket) and sends data that the peer only partially consumes - a slow or deliberately stalling peer on the fabric is enough to force the partial-send path repeatedly. Conditional on CONFIG_INET_ESPINTCP. Be aware the upstream commit message is unusually terse ('fixes some bugs in skmsg accounting') and does not spell out the memory-safety consequence; the 9.8 network/no-privilege score is the CNA's, not derived from the message.","remediation":"Boot a kernel carrying the fix commits below (no fixed stable version published); take it together with CVE-2026-52935, which touches the same partial-send state. Interim control: disable CONFIG_INET_ESPINTCP where IPsec-over-TCP is not required.","references":["https://git.kernel.org/stable/c/54d73f18f8919735f4d04d6f43374f75756c0180","https://git.kernel.org/stable/c/14c0b42c8a2cd9b5361bbff45b52f69c62c6a286","https://nvd.nist.gov/vuln/detail/CVE-2026-72041"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-72129","cve":"CVE-2026-72129","aliases":["nvmet-rdma inline data nonzero offset","NVMe-oF RDMA inline scatterlist underflow"],"title":"Linux kernel - NVMe-oF RDMA target, drivers/nvme/target/rdma.c: Nvmet_rdma_use_inline_sg() accepted any host-controlled","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel - NVMe-oF RDMA target, drivers/nvme/target/rdma.c","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Nvmet_rdma_use_inline_sg() accepted any host-controlled inline data offset that satisfied off + len <= inline_data_size, but the mapping still assumed the data started in the first inline page. When a port is configured with inline_data_size larger than PAGE_SIZE - a legitimate, performance-motivated setting, allowed up to 16K - an offset past PAGE_SIZE makes the length calculation underflow to roughly 4 GiB, and the block backend reads far past the intended page. There is also a double-free path where a persistent inline page is mistaken for an allocated scatterlist and freed. Remote, host-controlled, on the NVMe-oF-over-RDMA target that fronts tenant storage: this is the disaggregated-storage isolation break in its most direct form.","attack_vector":"A connected NVMe-oF initiator submits a command with an inline data offset in the range between PAGE_SIZE and the port's configured inline_data_size. Requires the target port to be configured with inline_data_size > PAGE_SIZE, which operators do deliberately to cut latency on small writes - so the hardened, performance-tuned configuration is the vulnerable one.","remediation":"Host reboot / kernel upgrade on NVMe-oF RDMA targets. Immediate config-only mitigation with a measurable but acceptable cost: set the port's inline_data_size to PAGE_SIZE or less (nvmet configfs, applied on port re-enable, brief reconnect for attached initiators) - this removes the precondition entirely without a reboot. Then schedule the kernel upgrade and restore the larger inline size afterwards if the latency mattered.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2026/CVE-2026-72129.json","https://nvd.nist.gov/vuln/detail/CVE-2026-72129"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-08-15"},{"id":"CVE-2026-72130","cve":"CVE-2026-72130","aliases":["nvmet-auth short AUTH_RECEIVE heap write","DH-HMAC-CHAP SUCCESS1 overflow"],"title":"Linux kernel - NVMe-oF target DH-HMAC-CHAP authentication, drivers/nvme/target/fabrics-cmd-auth.c: The sibling of the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel - NVMe-oF target DH-HMAC-CHAP authentication, drivers/nvme/target/fabrics-cmd-auth.c","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The sibling of the previous finding and worse - this one writes. nvmet_execute_auth_receive() trusted the AUTH_RECEIVE allocation length after checking only that it was non-zero, so a remote initiator supplying a one-byte allocation length reaches the fixed-size response builders with an undersized buffer and triggers a 16-byte heap out-of-bounds write. The code merely warned about the short length and then formatted the response anyway. A controlled heap write from a remote peer on a storage node is a direct path to kernel code execution and thus to every tenant volume the target serves.","attack_vector":"A remote NVMe-oF initiator with access to an auth-enabled target sends an AUTH_RECEIVE with a one-byte allocation length while the exchange is in the SUCCESS1 or FAILURE1 state. Reachable only when in-band DH-HMAC-CHAP is configured - so again, this is a risk you take on by following the standard hardening advice on an unpatched kernel.","remediation":"Host reboot / kernel upgrade on NVMe-oF targets. Patch before enabling DH-HMAC-CHAP; if auth is already on and the kernel is unpatched, treat the upgrade as urgent and restrict target reachability to known initiators at the firewall in the meantime. Roll targets in waves with multipath initiators so tenants see no I/O interruption.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2026/CVE-2026-72130.json","https://nvd.nist.gov/vuln/detail/CVE-2026-72130"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-08-15"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-415"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-72137","cve":"CVE-2026-72137","aliases":[],"title":"Linux kernel (net/xfrm): NAT-keepalive frees the keepalive skb whenever the IPv4/IPv6 send helper returns an error","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"NAT-keepalive frees the keepalive skb whenever the IPv4/IPv6 send helper returns an error, including after the output path has already taken ownership and consumed it. Any send failure on a keepalive turns into a double free - kernel heap corruption on a node doing NAT-traversal IPsec.","attack_vector":"Remote-influenced rather than remote-executed: the keepalive timer runs on its own, and whether the send fails after handoff depends on network conditions an off-path party can shape (route loss, neighbour failure, MTU/ICMP responses). Conditional on NAT-T keepalives being configured on SAs, which is standard when IPsec crosses NAT or cloud gateways. No local privilege or device node needed.","remediation":"Boot a kernel carrying the linked stable commits. Interim: disable NAT keepalives on SAs that do not need them (no NAT in the path), which removes the code path entirely.","references":["https://git.kernel.org/stable/c/d0a4dc7efa825bce60a8da8f7d43c864a159abde","https://git.kernel.org/stable/c/5b0c4c916f202b8fd13d12afb6af62b385622f81","https://nvd.nist.gov/vuln/detail/CVE-2026-72137"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-72217","cve":"CVE-2026-72217","aliases":[],"title":"Linux SUNRPC (xdr_buf_to_bvec, nfsd write path): xdr_buf_to_bvec stores a bio_vec before checking the slot is in range","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux SUNRPC (xdr_buf_to_bvec, nfsd write path)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"xdr_buf_to_bvec stores a bio_vec before checking the slot is in range, so a client-controlled RPC payload size drives an out-of-bounds write into adjacent kernel slab memory on the NFS server. The overflowing values come straight from the client, which makes this a remote kernel memory corruption on the shared file server.","attack_vector":"Any NFS client that can send writes to the server - i.e. any tenant compute node with the export mounted.","remediation":"Update the NFS server kernel to one with the bound-check-before-store fix in SUNRPC and reboot. This is in the write path, so there is no useful config workaround short of making the export read-only.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-72217","https://git.kernel.org/pub/scm/linux/security/vulns.git"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-362","CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-72451","cve":"CVE-2026-72451","aliases":[],"title":"Linux kernel (net/xfrm): The IPsec input path validates a security association before taking the state lock, so a state","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The IPsec input path validates a security association before taking the state lock, so a state that is killed in the window between the check and the insert still gets published into the per-CPU input cache. Later packets then resolve against a dead SA - traffic keeps being accepted and decrypted under credentials the control plane believes it has revoked, and the cached dead state is a use-after-free on the decrypt hot path. Revoking a tenant's SA stops meaning the tenant's traffic stops being accepted.","attack_vector":"The insertion happens on the packet input path, so it is driven by anyone who can send ESP traffic that resolves an SA on the node: a peer on the RDMA/IP fabric, another node, or a tenant terminating an overlay tunnel. The race partner is an ordinary SA deletion - exactly what an IKE daemon does on rekey or teardown, so the window opens on its own during normal rekey churn and can be widened by a peer forcing rekeys. No local access to the victim node is required.","remediation":"Update to 6.12.97 or later on the 6.12 stable series, or a vendor kernel carrying the fix commits below. Interim control: none that is real - slowing SA rekey/teardown churn narrows the window but does not close it. Treat this as a mandatory reboot on any node running IPsec on the fabric.","references":["https://git.kernel.org/stable/c/6dab4dec9a49121d079981ac913569f232c06b06","https://git.kernel.org/stable/c/a1a3360a0c44b8b5c134db2a9d0667b61c9cf523","https://nvd.nist.gov/vuln/detail/CVE-2026-72451"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-72463","cve":"CVE-2026-72463","aliases":[],"title":"Linux kernel (net/xfrm): Async ESP resumption holds a reference on the original skb","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Async ESP resumption holds a reference on the original skb->dev, but the receive callback swaps skb->dev to the tunnel device without taking one. The resumption then drops a reference it never took - use-after-free on the tunnel netdevice plus a permanent refcount leak on the original device, so the node both corrupts memory and can never unregister the interface.","attack_vector":"Remote and packet-driven: inbound ESP that goes through async crypto (hardware or async software AEAD) and lands on a vti/xfrm tunnel device. Any peer on the fabric sending ESP to the node reaches it - no local access, no device node. Conditional on tunnel devices with an rcv_cb (vti/vti6/xfrmi) and async crypto being in use, which is the normal shape for offloaded IPsec on smart NICs.","remediation":"Update to 6.19.x / 6.20 or later, or a kernel carrying the linked stable commits. Interim: disable IPsec crypto offload (ip xfrm state ... without offload, or ethtool -K <dev> esp-hw-offload off) so the synchronous path is used, and avoid vti-style tunnel devices on affected nodes.","references":["https://git.kernel.org/stable/c/63a30015199912bd5055bead8001b1ae68a67cdb","https://git.kernel.org/stable/c/8045c0df98d4f14c54e5cb875f1c9c0ce89fe4ff","https://nvd.nist.gov/vuln/detail/CVE-2026-72463"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-362","CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-72494","cve":"CVE-2026-72494","aliases":[],"title":"Linux kernel (drivers/infiniband/hw/irdma): The driver signalled completion of control-plane requests through an","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/hw/irdma)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The driver signalled completion of control-plane requests through an unordered flag, with no memory barriers. A waiter can therefore observe the request as finished and release the request context while the completing side is still writing to it, or miss the wakeup entirely. The kernel CNA rates this as network-reachable with no privileges and full confidentiality, integrity and availability loss.","attack_vector":"Reached through irdma control-plane requests, which are driven by tenant verbs activity (QP/MR/CQ operations from /dev/infiniband/uverbs*) and by connection events arriving from peers on the fabric. Only nodes with Intel E810/E830 irdma are affected.","remediation":"No fixed version is listed in the record - take the stable kernel carrying bde37aed0724 (or d9c8c45e6d2f) and reboot. Interim: irdma nodes carrying untrusted tenants should have /dev/infiniband/* removed from those containers until patched.","references":["https://git.kernel.org/stable/c/bde37aed0724c0139dea177f3aae8d989b6babb1","https://git.kernel.org/stable/c/d9c8c45e6d2f438a3c8e643ae78b59454fa0fadd","https://nvd.nist.gov/vuln/detail/CVE-2026-72494"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-7301","cve":"CVE-2026-7301","aliases":[],"title":"SGLang (scheduler ROUTER socket): ROUTER socket binds `0.0.0.0` by default and `pickle.loads()` incoming messages","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"SGLang (scheduler ROUTER socket)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"ROUTER socket binds `0.0.0.0` by default and `pickle.loads()` incoming messages","attack_vector":"Unauthenticated network — default configuration is exploitable","remediation":"Upgrade and override the bind address. The insecure default means an unmodified deployment is exposed","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-7301"],"status":"curated","published":"2026-05-18"},{"id":"CVE-2026-7304","cve":"CVE-2026-7304","aliases":[],"title":"SGLang (custom logit processor): `dill.loads` on user objects when `--enable-custom-logit-processor` is set","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"SGLang (custom logit processor)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"`dill.loads` on user objects when `--enable-custom-logit-processor` is set → unauthenticated RCE","attack_vector":"Unauthenticated network to the serving port","remediation":"Upgrade; never enable custom logit processors on a tenant-facing endpoint","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-7304"],"status":"curated","published":"2026-05-18"},{"id":"CVE-2026-74269","cve":"CVE-2026-74269","aliases":[],"title":"Linux bnxt_en driver (XDP head-grow underflow): Head underflow when an XDP program grows the packet head on a Broadcom","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux bnxt_en driver (XDP head-grow underflow)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Head underflow when an XDP program grows the packet head on a Broadcom NIC, crashing the host. Found by the kernel's own XDP test suite, which is a reminder that the XDP fast paths on vendor NIC drivers get materially less real-world coverage than the normal receive path — and AI-serving front-ends are one of the few places they run at scale.","attack_vector":"An attached XDP program that grows the packet head, processing received traffic.","remediation":"Kernel/driver upgrade plus host reboot. Interim: detach head-growing XDP programs from Broadcom interfaces.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-74269"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-15"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-74345","cve":"CVE-2026-74345","aliases":[],"title":"Linux kernel SoftiWARP connection manager (siw_cm, endpoint/socket disassociation): A malformed MPA request during","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel SoftiWARP connection manager (siw_cm, endpoint/socket disassociation)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A malformed MPA request during iWARP connection setup causes the new endpoint to be torn down, and siw_socket_disassoc() drops the last reference and frees the endpoint while the caller then clears the now-dangling socket pointer. KASAN caught the use-after-free in the connection-manager work handler. The whole sequence happens during connection establishment, so it is reachable before any application-level authentication - a remote peer that can reach the siw listener gets a kernel use-after-free by sending a deliberately broken handshake.","attack_vector":"Remote, pre-authentication. Anyone who can complete a TCP connection to the SoftiWARP listening port on the node - which on a flat cluster network is every other tenant's workload.","remediation":"Kernel update moving the socket-pointer clear inside siw_socket_disassoc(). Because siw is a loadable software provider, unloading or blacklisting the siw module on nodes that do not need SoftiWARP eliminates the listener immediately, no reboot required - do that first, patch on the next maintenance window.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=b28d513393f81e2de00f82970487a9d001557e4e","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-74345.json"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-193","CWE-125"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-74361","cve":"CVE-2026-74361","aliases":[],"title":"Linux kernel (drivers/nvme/host): An off-by-one in the Flexible Data Placement index check accepts a placement index","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/nvme/host)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"An off-by-one in the Flexible Data Placement index check accepts a placement index one past the end of the configuration array, so a value that came from userspace is used to read out of bounds in the I/O submission path. The tenant supplying the index gets an out-of-bounds read of adjacent kernel memory and, at minimum, a node-destabilising fault.","attack_vector":"Reachable by an unprivileged process on the node that can issue I/O to an FDP-capable NVMe namespace and set the write placement hint - no /dev/nvme passthrough or CAP_SYS_ADMIN needed, since the placement index rides on ordinary writes. Conditional on the namespace having FDP enabled, which is increasingly common on modern datacenter SSDs used for tenant scratch space.","remediation":"No fixed version is listed on this record - boot a kernel carrying the linked stable commits. Interim: disable FDP on namespaces exposed to tenants, or do not hand tenants direct block devices on FDP-enabled drives until the node is patched.","references":["https://git.kernel.org/stable/c/5e406928404d67a8da8aa3ae21732e1ea1a04118","https://git.kernel.org/stable/c/5d0e7d2af884b91329235abb16652ae4eead8079","https://nvd.nist.gov/vuln/detail/CVE-2026-74361"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-74384","cve":"CVE-2026-74384","aliases":[],"title":"Linux kernel (drivers/nvme/host): The multipath current-path array is sized by the count of possible NUMA nodes but","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/nvme/host)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The multipath current-path array is sized by the count of possible NUMA nodes but indexed by the actual NUMA node ID. On any machine whose node IDs are sparse rather than 0..N-1, ordinary I/O writes eight bytes past the end of the slab allocation - silent heap corruption of whatever object sits next, driven by the normal data path rather than by an error case.","attack_vector":"No attacker action is required beyond running I/O: any tenant's reads and writes through an NVMe multipath device index the array by the NUMA node of the CPU they land on. Reachability is a hardware/topology condition, not a permission one - it needs a system with non-contiguous NUMA node IDs (the report is from POWER9 with nodes 0, 8, 252-255; check `ls /sys/devices/system/node/` on non-x86 or accelerator-heavy nodes before assuming you are clear). Dense-node x86 hosts are not affected.","remediation":"No fixed version is listed on this record - boot a kernel carrying the linked stable commits. Interim: audit node IDs on every architecture in the fleet and prioritise patching hosts with sparse NUMA topology; there is no configuration knob that avoids the bad index.","references":["https://git.kernel.org/stable/c/7e7b167e65610dfa7564d449474f4b477f9d4c1c","https://git.kernel.org/stable/c/316b5f1168264844aa125959de1d6da2b1905795","https://nvd.nist.gov/vuln/detail/CVE-2026-74384"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-74394","cve":"CVE-2026-74394","aliases":["RDMA/srpt immediate data integer overflow","SRP target"],"title":"Linux kernel - SRP target (srpt), drivers/infiniband/ulp/srpt/ib_srpt.c: An integer overflow in the immediate-data","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel - SRP target (srpt), drivers/infiniband/ulp/srpt/ib_srpt.c","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"An integer overflow in the immediate-data length check on the SRP target lets a remote initiator bypass the bound and reach kernel memory it should not. The srpt target serves block storage over InfiniBand to compute clients, so a single tenant that can connect as an initiator gets kernel-level compromise of the storage node - and through it, access to every other tenant's volumes exported from the same target.","attack_vector":"A remote SRP initiator submits a command whose immediate data length is chosen so the length arithmetic wraps, defeating the check. Any host allowed to connect to the target can do it; on fabrics without per-tenant partitioning that is any host on the subnet.","remediation":"Host reboot / kernel upgrade on SRP target nodes. Interim: enforce InfiniBand partitioning so only authorised initiators can reach the target port (SM config change, effective on the next sweep, no reload), and restrict target ACLs to known initiator GUIDs. Where SRP has been superseded by NVMe-oF, decommission the srpt target and unload the module - the cleanest fix.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2026/CVE-2026-74394.json","https://nvd.nist.gov/vuln/detail/CVE-2026-74394"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-08-15"},{"id":"CVE-2026-74474","cve":"CVE-2026-74474","aliases":[],"title":"Linux VXLAN driver (transmit-path header pulls): `vxlan_xmit()`, `arp_reduce()` and `vxlan_mdb_entry_skb_get()`","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux VXLAN driver (transmit-path header pulls)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"`vxlan_xmit()`, `arp_reduce()` and `vxlan_mdb_entry_skb_get()` validate header availability with `pskb_may_pull()`, but on the transmit path `skb->data` points at the MAC header, so the offset accounting is wrong by the Ethernet header length and the driver reads past the validated region. This is the Linux kernel VXLAN data path — the software VTEP used by Linux-based switch OSes (SONiC, Cumulus), by container overlay networks, and by any host doing VXLAN encapsulation itself. Memory corruption in the encapsulation path of a multi-tenant overlay is as close to the centre of the tenant-isolation boundary as this database gets.","attack_vector":"Traffic traversing the VXLAN transmit path on an affected host or switch — reachable from inside a tenant overlay, since tenants generate the frames being encapsulated.","remediation":"Kernel upgrade plus host reboot; on Linux-based switch NOSes it arrives as a NOS image upgrade plus switch reload, so stage across MLAG/ECMP pairs. No config workaround short of not using VXLAN. Part of a 2026 cluster of VXLAN data-path fixes — CVE-2026-74475, CVE-2026-74473, CVE-2026-74406, CVE-2026-63993 — so take the whole batch in one upgrade rather than chasing them individually.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-74474","https://nvd.nist.gov/vuln/detail/CVE-2026-74473","https://nvd.nist.gov/vuln/detail/CVE-2026-63993"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-08-15"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-74493","cve":"CVE-2026-74493","aliases":[],"title":"Linux kernel (net/smc): Link-group termination drops conns_lock after finding a connection but before taking a socket","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/smc)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Link-group termination drops conns_lock after finding a connection but before taking a socket reference, so a concurrent close can free the socket the termination worker is about to write to. KASAN confirms a slab-use-after-free write from smc_lgr_terminate_work - a fabric event that overlaps a tenant closing its connection turns into host memory corruption.","attack_vector":"Reachable whenever link-group termination overlaps connection close. Termination is driven from the fabric side (link down, device event, peer-initiated teardown) while the close is an unprivileged tenant operation, so neither half needs privilege. The write lands in a kworker, so the blast radius is the node, not the tenant. Requires SMC-R in use; the module autoloads from an unprivileged socket(AF_SMC, ...).","remediation":"Boot a kernel carrying the fix commits (takes the socket reference while conns_lock still protects the tree entry). Interim: blacklist the smc module on nodes not running SMC-R, and keep untrusted tenants off the fabric segment that can drive link-group termination.","references":["https://git.kernel.org/stable/c/ff44f2df57fb5560bdc75eb977867643e764a262","https://git.kernel.org/stable/c/bea8dc14de2d56aca749d368563e6888217710a5","https://nvd.nist.gov/vuln/detail/CVE-2026-74493"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-74556","cve":"CVE-2026-74556","aliases":["libiscsi_tcp SCSI Response data segment overflow","open-iscsi initiator conn->data overflow"],"title":"Linux kernel - iSCSI TCP initiator, drivers/scsi/libiscsi_tcp.c: The iSCSI initiator receives PDU data segments into a","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel - iSCSI TCP initiator, drivers/scsi/libiscsi_tcp.c","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The iSCSI initiator receives PDU data segments into a fixed 8192-byte conn->data buffer. The login, text, reject and async-event paths all check the DataSegmentLength against that size, but the SCSI Command Response path copies sense and response data in without the same check. The only remaining bound is the initiator's advertised MaxRecvDataSegmentLength, which open-iscsi defaults to 262144 - so a target returning a SCSI Response with a data segment between 8193 and 262144 bytes overflows the buffer by up to 250 KB. Target-to-initiator again: one compromised or rogue iSCSI target achieves kernel memory corruption on every compute node that mounts from it.","attack_vector":"The attacker controls an iSCSI target the victim connects to, or can inject into the TCP session, and returns a SCSI Command Response with an oversized data segment. Because open-iscsi's default MaxRecvDataSegmentLength is far above the buffer size, the vulnerable window is wide in stock configurations rather than requiring unusual tuning.","remediation":"Host reboot / kernel upgrade on all iSCSI initiator nodes - the compute fleet, so plan a rolling drain. Immediate config mitigation that closes the window without a reboot: set MaxRecvDataSegmentLength to 8192 or below in /etc/iscsi/iscsid.conf and reconnect sessions, which caps the receive length at the buffer size. Also verify targets are authenticated (mutual CHAP) so a rogue target cannot attract sessions.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2026/CVE-2026-74556.json","https://nvd.nist.gov/vuln/detail/CVE-2026-74556"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-08-15"},{"fleet":{"pain_class":"firmware-flash"},"id":"NCVD-2014-001-supermicro-ipmi-bmc-firmware-wpc","cve":null,"aliases":["PSBlock","Supermicro port 49152 password disclosure"],"title":"Supermicro IPMI BMC firmware (WPCM450 / X8-X9 generation): An unauthenticated HTTP GET for /PSBlock on port 49152","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro IPMI BMC firmware (WPCM450 / X8-X9 generation)","year":"2014","cvss_score":9.8,"severity":"critical","kev":false,"impact":"An unauthenticated HTTP GET for /PSBlock on port 49152 returns the BMC's user database with passwords in cleartext. No exploit and no authentication are involved - it is a file read. Those credentials are usually reused across the whole fleet and often match the IPMI accounts on every other chassis the operator bought at the same time, so one exposed node yields the management plane for all of them: power control, virtual media boot of an attacker image, and KVM to the console. No CVE was ever assigned to this, so it appears in no NVD-derived feed and no scanner keyed on CVE identifiers.","attack_vector":"Anyone who can reach TCP/49152 on the BMC. At disclosure roughly 32,000 hosts were reachable directly from the internet; inside a datacenter, any tenant with a route to the management VLAN qualifies.","remediation":"Update BMC firmware to the fixed generation from Supermicro, then rotate every IPMI credential - the old ones must be assumed public, and a firmware update does not invalidate them. Verify port 49152 is closed after the update rather than trusting the release notes. Put the management network behind its own VLAN with no tenant route, and confirm that from a tenant workload rather than from a jump host. Nodes too old to receive fixed firmware should be treated as permanently exposed and segmented accordingly.","references":["https://nmap.org/nsedoc/scripts/supermicro-ipmi-conf.html","https://svn.nmap.org/nmap/scripts/supermicro-ipmi-conf.nse"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-23"],"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2023-010-mlflow-mlflow-server-mlflow-ui-m","cve":null,"aliases":["GHSA-83fm-w79m-64r5"],"title":"MLflow (mlflow server / mlflow ui, Model Registry): REMOTE FILE ACCESS on the host running the tracking and registry","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"MLflow (mlflow server / mlflow ui, Model Registry)","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"REMOTE FILE ACCESS on the host running the tracking and registry server. Anyone who can query an mlflow server or mlflow ui instance below 2.3.1 can traverse to files outside the intended artifact path. In an AI cluster the tracking server is a high-value target precisely because it is boring infrastructure: it typically holds or can reach object-store credentials, database connection strings, and the pointers to every team's model artifacts, and it is very often deployed on an internal network with no auth in front of it on the theory that it is 'just experiment tracking'. Reading files off that host converts into credentials for the artifact store, and from there into the ability to plant models that other pipelines will load and execute. Affects only deployments that actually run the mlflow server or mlflow ui commands; managed offerings that do not invoke them are out of scope.","attack_vector":"Network, unauthenticated in the common deployment. Anyone able to send requests to the tracking or registry server is in scope — which is everyone, unless the operator has separately put a VPC boundary, IP allowlist or auth middleware in front of it, since MLflow ships none.","remediation":"Upgrade MLflow to 2.3.1 or later and restart the tracking/registry servers. Independently of the patch, put an actual authentication and authorization layer in front of MLflow and restrict inbound access by network policy or IP allowlist — the vendor's own guidance is that these servers are not designed to be exposed. Rotate any credentials reachable from the server host if it has been broadly reachable.","references":["https://github.com/mlflow/mlflow/security/advisories/GHSA-83fm-w79m-64r5"],"status":"curated"},{"id":"CVE-2018-3679","cve":"CVE-2018-3679","aliases":[],"title":"Intel Data Center Manager SDK (reference UI): The DCM SDK's reference UI allows an unauthenticated remote attacker","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel Data Center Manager SDK (reference UI)","year":"2018","cvss_score":9.6,"severity":"critical","kev":false,"impact":"The DCM SDK's reference UI allows an unauthenticated remote attacker to execute code with administrator privileges. Reference UIs get shipped into production more often than vendors expect - if you built anything on the DCM SDK, check whether the sample UI went with it.","attack_vector":"Unauthenticated remote attacker able to reach the reference UI.","remediation":"Upgrade the DCM SDK past 5.0 and remove the reference UI from any production deployment. Application-level change; no host reboot or firmware.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-3679","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00143.html"],"status":"curated","published":"2018-09-12"},{"id":"CVE-2019-15897","cve":"CVE-2019-15897","aliases":[],"title":"BeeGFS (beegfs-ctl / metadata server): Authentication bypass by talking directly to a BeeGFS metadata server. BeeGFS is","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"BeeGFS (beegfs-ctl / metadata server)","year":"2019","cvss_score":9.6,"severity":"critical","kev":false,"impact":"Authentication bypass by talking directly to a BeeGFS metadata server. BeeGFS is a common choice for AI-training scratch storage because it is fast and easy to stand up, and its threat model assumes the storage network is private. Anyone who reaches the metadata server can act against the filesystem's metadata — which means other tenants' namespaces on a shared BeeGFS deployment.","attack_vector":"Network access to a BeeGFS metadata server. The advisory notes such servers are typically not exposed externally — but inside a GPU cluster, 'not externally exposed' still means reachable by every tenant workload on the storage VLAN.","remediation":"Upgrade BeeGFS past 7.1.3 and enable connection authentication (the shared-secret `connAuthFile`), then restart the metadata, storage and client services — a coordinated restart that stalls I/O, so drain jobs first. Enabling connAuth is the load-bearing step; the version upgrade alone does not help if authentication stays off.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-15897"],"status":"curated","fleet":{"pain_class":"node-drain"},"tags":["tenant-isolation"],"published":"2019-12-05"},{"id":"CVE-2021-21538","cve":"CVE-2021-21538","aliases":["DSA-2021-082"],"title":"Dell iDRAC9 (Virtual Console / authentication): An attacker with no credentials lands directly inside the server's","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC9 (Virtual Console / authentication)","year":"2021","cvss_score":9.6,"severity":"critical","kev":false,"impact":"An attacker with no credentials lands directly inside the server's Virtual Console - the same screen-and-keyboard the operator uses. From there they see whatever the tenant is running, drive the BIOS/boot menu, and pair it with Virtual Media to boot a node off an image they supply. On a bare-metal cloud that is a silent cross-tenant handoff failure: the previous or a neighbouring tenant's session is visible and controllable without ever touching the production network. Only iDRAC9 firmware in the 4.40.00.00-4.40.09.99 band is affected, so this is a narrow-window regression that is easy to miss in a mixed-vintage fleet.","attack_vector":"Anything routable to the iDRAC address on the out-of-band management VLAN, no account and no prior foothold. If the OOB network is flat across racks, a single compromised jump host or a mis-scoped VPN split-tunnel reaches every node in the affected firmware band.","remediation":"Flash iDRAC9 to 4.40.10.00 or later. This is per-node but fully out-of-band (iDRAC web UI, racadm, Redfish SimpleUpdate, or OME) and does NOT require a host reboot - the iDRAC resets itself and you lose OOB reachability for roughly two to five minutes while the host keeps running its jobs. No drain needed. Immediate config-only mitigation while you roll: disable Virtual Console and Virtual Media under iDRAC Settings, and ACL the management VLAN down to the jump hosts.","references":["https://www.dell.com/support/kbdoc/000186420","https://nvd.nist.gov/vuln/detail/CVE-2021-21538"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2021-07-29"},{"id":"CVE-2021-21596","cve":"CVE-2021-21596","aliases":[],"title":"Dell OpenManage Enterprise (remote code execution): Remote code execution on the OpenManage Enterprise console","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Dell OpenManage Enterprise (remote code execution)","year":"2021","cvss_score":9.6,"severity":"critical","kev":false,"impact":"Remote code execution on the OpenManage Enterprise console. OME is the fleet-wide control plane that already holds credentials for, and can push firmware to, every iDRAC it manages - so compromising it is not compromising one node, it is compromising the mechanism that drives all of them. An attacker in OME can trigger firmware deployment, mount Virtual Media, and power-cycle at fleet scale from a single box. Affects OME versions before 3.6.2 and the corresponding OME-Modular builds.","attack_vector":"An attacker with access to the immediate subnet the OME appliance sits on. That is usually the management network segment, so the practical question is who else lives on the same VLAN as your OME appliance - jump hosts, monitoring, DCIM, and often a broader IT segment than anyone intends.","remediation":"Upgrade the OME appliance to 3.6.2 or later. This is a single appliance upgrade, not a per-node campaign, so it is cheap in rollout terms - no node reboots, no job drain, only the OME console's own downtime. The structural fix is network placement: put OME on its own segment with an explicit allowlist rather than sharing the general management VLAN, and treat it as tier-0 infrastructure because it holds BMC credentials for the whole fleet.","references":["https://www.dell.com/support/kbdoc/000189673","https://nvd.nist.gov/vuln/detail/CVE-2021-21596"],"status":"curated","published":"2021-08-09"},{"id":"CVE-2022-24422","cve":"CVE-2022-24422","aliases":["DSA-2022-068"],"title":"Dell iDRAC9 (VNC server): Unauthenticated access to the iDRAC VNC console","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC9 (VNC server)","year":"2022","cvss_score":9.6,"severity":"critical","kev":false,"impact":"Unauthenticated access to the iDRAC VNC console. Same practical outcome as an unauthenticated KVM: the attacker watches and drives the tenant's console, can reach the boot menu and firmware setup, and can chain into Virtual Media to boot an attacker image below the hypervisor. Affects iDRAC9 5.00.00.00 up to 5.10.10.00, so it hits the 15G PowerEdge generation that a lot of first-wave GPU fleets standardised on.","attack_vector":"Anything routable to the iDRAC VNC port on the management VLAN, unauthenticated - but only where the iDRAC VNC server has actually been enabled. It is off by default, so the real exposure is the subset of nodes where someone turned it on for remote hands.","remediation":"Flash iDRAC9 to 5.10.10.00 or later - out-of-band, per-node, no host reboot and no job drain (iDRAC self-resets, brief OOB blackout only). Because VNC is opt-in, the fastest mitigation is config-only and costs nothing: audit which nodes have the iDRAC VNC server enabled and turn it off. That closes the hole fleet-wide in minutes while the firmware campaign runs on its own schedule.","references":["https://www.dell.com/support/kbdoc/en-us/000199267/dsa-2022-068-dell-idrac9-security-update-for-an-improper-authentication-vulnerability","https://nvd.nist.gov/vuln/detail/CVE-2022-24422"],"status":"curated","published":"2022-05-26"},{"id":"CVE-2022-41924","cve":"CVE-2022-41924","aliases":[],"title":"Tailscale (Windows client): Local API bound to a TCP socket","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Tailscale (Windows client)","year":"2022","cvss_score":9.6,"severity":"critical","kev":false,"impact":"Local API bound to a TCP socket -> a malicious website reconfigures tailscaled and achieves RCE","attack_vector":"Network (remote)","remediation":"Control-plane: force-update all operator clients","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-41924"],"status":"curated","published":"2022-11-23"},{"id":"CVE-2023-3043","cve":"CVE-2023-3043","aliases":["AMI-SA-2023010"],"title":"AMI MegaRAC SPx 12 / SPx 13 (BMC network service): The twin of CVE-2023-37293: a stack smash in the BMC's","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx 12 / SPx 13 (BMC network service)","year":"2023","cvss_score":9.6,"severity":"critical","kev":false,"impact":"The twin of CVE-2023-37293: a stack smash in the BMC's network-facing path that hands an unauthenticated attacker execution on the management controller. Once there the attacker can hold power control over the node indefinitely, mount virtual media to boot an attacker image, read whatever the BMC can see, and write persistent firmware. The persistence is the point - this survives OS reinstall, disk replacement and node rebuild, so a fleet operator who rotates a suspect node back into the pool reintroduces the implant.","attack_vector":"Adjacent network, unauthenticated, low complexity. Reachable from anything sharing the BMC's broadcast domain. No tenant credentials, no BMC credentials, no clicking required.","remediation":"Firmware flash to SPx_12.7 / SPx_13.6 - out-of-band, per node, gated on your ODM publishing a rebased image. Budget for a fleet-wide BMC flash campaign, not a patch window: BMC images are not managed by your OS config management and each SKU needs its own tested build. Config-only interim mitigation is VLAN isolation of the BMC plane plus removing any tenant-network route to it; disabling IPMI-over-LAN does not close this one because the exposure is the BMC's own network stack.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2023010.pdf","https://nvd.nist.gov/vuln/detail/CVE-2023-3043"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-01-09"},{"id":"CVE-2023-37293","cve":"CVE-2023-37293","aliases":["AMI-SA-2023010"],"title":"AMI MegaRAC SPx 12 / SPx 13 (BMC network service): Unauthenticated code execution inside the BMC, reached","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx 12 / SPx 13 (BMC network service)","year":"2023","cvss_score":9.6,"severity":"critical","kev":false,"impact":"Unauthenticated code execution inside the BMC, reached with a crafted packet and nothing else. The attacker lands below the hypervisor on a processor that stays powered whether or not the node is booted, and from there owns the node's power state, its console, its virtual media and its BIOS flash. On a GPU fleet this is the classic rack-wide implant: reimaging the tenant OS, replacing the NVMe, or rebuilding the node from the provisioning system does not remove it. AMI rates the scope as changed, meaning the BMC compromise is expected to reach beyond the BMC into the host.","attack_vector":"Any host on the same L2 segment as the BMC management interface - no credentials, no user interaction, low attack complexity. In practice that means any compromised node, switch, PDU, out-of-band jump box or IPMI-speaking tool that shares the management VLAN. If BMCs are flat-VLANed across a hall (a common shortcut in leased colo and neocloud builds), one foothold reaches every BMC in the hall.","remediation":"BMC firmware flash to MegaRAC SPx_12.7 / SPx_13.6 or later - out-of-band, per node, with real bricking risk if power or the transfer is interrupted mid-write. You cannot take AMI's build directly: your ODM (Supermicro, Quanta, Gigabyte, Wiwynn, ASRock Rack, Inventec) must rebase and requalify, and that lag has historically run six months to over a year, longer on white-box SKUs whose vendor has stopped shipping BMC images. The host generally stays up through the flash but KVM/vMedia sessions drop and the BMC reboots. Until the image exists, the only real control is network: put BMCs on an isolated management VLAN with no route from tenant, storage or corporate networks, and reach them only through a bastion.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2023010.pdf","https://nvd.nist.gov/vuln/detail/CVE-2023-37293","https://www.ami.com/security-advisories/"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-01-09"},{"id":"CVE-2024-24593","cve":"CVE-2024-24593","aliases":[],"title":"ClearML API server: CSRF against the API server","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"ClearML API server","year":"2024","cvss_score":9.6,"severity":"critical","kev":false,"impact":"CSRF against the API server","attack_vector":"Logged-in operator visiting a malicious page","remediation":"Upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-24593"],"status":"curated","published":"2024-02-06"},{"id":"CVE-2024-27892","cve":"CVE-2024-27892","aliases":[],"title":"Arista EOS (OpenConfig gNMI Set authorization): A gNMI Set request that authorization should have rejected is executed","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (OpenConfig gNMI Set authorization)","year":"2024","cvss_score":9.6,"severity":"critical","kev":false,"impact":"A gNMI Set request that authorization should have rejected is executed, so a caller writes switch configuration it has no right to write. Model-driven management is how large fabrics are actually operated now, and gNMI is the write path — if its authorization does not hold, your RBAC on the fabric does not exist. Pairs with CVE-2024-27890 (same defect, separate advisory) and CVE-2025-1260 on the gNOI side.","attack_vector":"A client able to reach the gNMI endpoint with credentials whose authorization should have been insufficient. Requires OpenConfig to be configured.","remediation":"EOS upgrade plus reload. Interim: restrict gNMI/gNOI reachability with a control-plane ACL to only the automation hosts that legitimately write config — live config change and worth doing permanently.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-27892","https://nvd.nist.gov/vuln/detail/CVE-2024-27890"],"status":"curated","published":"2026-06-04"},{"id":"CVE-2024-34359","cve":"CVE-2024-34359","aliases":[],"title":"llama-cpp-python: RCE via Jinja2 template in a GGUF model's metadata (`Llama` class)","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"llama-cpp-python","year":"2024","cvss_score":9.6,"severity":"critical","kev":false,"impact":"RCE via Jinja2 template in a GGUF model's metadata (`Llama` class)","attack_vector":"Customer-supplied GGUF model file — the chat template inside it is the payload","remediation":"Upgrade; GGUF chat templates are executable and are not covered by pickle scanning","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-34359"],"status":"curated","published":"2024-05-14"},{"id":"CVE-2024-35225","cve":"CVE-2024-35225","aliases":[],"title":"Jupyter Server Proxy: Unauthenticated web access to a user's proxied processes","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Jupyter Server Proxy","year":"2024","cvss_score":9.6,"severity":"critical","kev":false,"impact":"Unauthenticated web access to a user's proxied processes","attack_vector":"Network attacker reaching the hub","remediation":"Upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-35225"],"status":"curated","published":"2024-06-11"},{"cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:N/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-125"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-58006","cve":"CVE-2024-58006","aliases":[],"title":"Linux kernel (drivers/pci/controller/dwc): A PCIe BAR window can end up larger than the memory actually backing it, so","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/pci/controller/dwc)","year":"2024","cvss_score":9.6,"severity":"critical","kev":false,"impact":"A PCIe BAR window can end up larger than the memory actually backing it, so accesses past the real allocation sail through the inbound address translation unit untranslated. The device on the other end of the link gets to read and write host memory that was never meant to be exposed through that BAR - a straight boundary break across the PCIe link, which is why the kernel CNA scored it 9.6 with a changed scope.","attack_vector":"Only applies to a machine running Linux in PCIe ENDPOINT mode on a DesignWare controller - i.e. the box is the device, not the host. The peer that exploits it is the PCIe host on the other side of the link, pre-authentication, by re-driving BAR setup so a second pci_epc_set_bar() shrinks the window without a matching clear_bar(). This is not the usual configuration for a GPU-cluster compute node; audit it if you run DWC-based endpoint cards, smartNICs or DPUs that terminate a host link, and treat the connected host as the attacker.","remediation":"Update to a kernel carrying the fix (no fixed_in published by the CNA - track the stable commits below into your 6.1.y / 6.6.y / 6.12.y branch). Interim: do not allow an endpoint function driver to re-invoke set_bar() on an already-configured BAR, and unbind endpoint function drivers on any DWC endpoint you are not actively using.","references":["https://git.kernel.org/stable/c/b5cacfd067060c75088363ed3e19779078be2755","https://git.kernel.org/stable/c/3229c15d6267de8e704b4085df8a82a5af2d63eb","https://nvd.nist.gov/vuln/detail/CVE-2024-58006"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-6385","cve":"CVE-2024-6385","aliases":[],"title":"GitLab: Attacker can trigger a CI pipeline as another user","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"GitLab","year":"2024","cvss_score":9.6,"severity":"critical","kev":false,"impact":"Attacker can trigger a CI pipeline as another user -> jobs run with that user's token","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; rotate CI job and runner registration tokens","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-6385"],"status":"curated","published":"2024-07-11"},{"id":"CVE-2026-50540","cve":"CVE-2026-50540","aliases":[],"title":"Kata Containers: kata-runtime host code execution via an untrusted input path","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kata Containers","year":"2026","cvss_score":9.6,"severity":"critical","kev":false,"impact":"kata-runtime host code execution via an untrusted input path","attack_vector":"Any tenant workload / malicious image under Kata","remediation":"Emergency Kata upgrade; node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-50540"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2026-08-07"},{"id":"CVE-2026-8037","cve":"CVE-2026-8037","aliases":[],"title":"Progress LoadMaster (ADC): OS command injection in the API","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Progress LoadMaster (ADC)","year":"2026","cvss_score":9.6,"severity":"critical","kev":true,"impact":"OS command injection in the API -> unauthenticated remote code execution on the load balancer","attack_vector":"Adjacent network","remediation":"Control-plane: patch; the ADC terminates tenant-facing TLS","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-8037"],"status":"curated","published":"2026-06-04"},{"id":"CVE-2025-26385","cve":"CVE-2025-26385","aliases":["ICSA-26-027-04"],"title":"Johnson Controls Metasys Application and Data Server (ADS) deployed with SQL Express: Command injection on the Metasys","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Johnson Controls Metasys Application and Data Server (ADS) deployed with SQL Express","year":"2025","cvss_score":9.5,"severity":"critical","kev":false,"impact":"Command injection on the Metasys ADS that yields remote SQL execution. The ADS is the site's building-automation database and supervisory server: it holds the point database, the schedules, the trends and the operator accounts for every NAE/SNE/SNC engine driving air handling in the building. Getting arbitrary SQL - and, through it, the usual SQL-Server-to-OS escalation paths - means an attacker owns the system that both commands and reports on cooling. They can rewrite schedules and setpoints so the change persists, and rewrite the trend history so the postmortem shows nothing anomalous. For a hall of 40 kW+ racks that is fleet-wide availability risk with an intact-looking dashboard, which is the worst combination for an operator trying to diagnose why a training run just died.","attack_vector":"Network access to the Metasys ADS. Metasys servers are Windows machines that facilities teams routinely domain-join and expose to the corporate network for the Site Management Portal, so the realistic path is a corporate-network foothold rather than direct internet exposure - though internet-published SMP instances do exist. Anyone who compromises a facilities workstation is one hop away.","remediation":"Vendor software update to the ADS per the Johnson Controls advisory - a server-side patch and a normal Windows change window, no controller firmware and no cooling downtime, so schedule it. Additionally: get the ADS off the corporate network segment entirely, restrict SQL Express to localhost, run the Metasys service account with least privilege, and require a jump host for SMP access. Leased colo: the Metasys ADS is the landlord's building server and typically serves all tenants - you cannot patch it, so make its version and network placement a contractual disclosure and demand notification when JCI publishes an advisory.","references":["https://www.cisa.gov/news-events/ics-advisories/icsa-26-027-04","https://nvd.nist.gov/vuln/detail/CVE-2025-26385","https://www.johnsoncontrols.com/trust-center/cybersecurity/security-advisories"],"status":"curated"},{"id":"CVE-2026-44946","cve":"CVE-2026-44946","aliases":[],"title":"Rancher: SAML assertion replay: the ACS handler does not enforce one-time use, so a captured assertion logs","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Rancher","year":"2026","cvss_score":9.5,"severity":"critical","kev":false,"impact":"SAML assertion replay: the ACS handler does not enforce one-time use, so a captured assertion logs an attacker in","attack_vector":"Unauthenticated network in a MITM or log-access position","remediation":"Emergency Rancher upgrade; invalidate all SAML sessions","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-44946"],"status":"curated","published":"2026-06-30"},{"id":"CVE-2023-25131","cve":"CVE-2023-25131","aliases":["JVN#95119483"],"title":"CyberPower PowerPanel Business Local/Remote/Management v4.8.6 and earlier (Windows and Linux): A default password","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"CyberPower PowerPanel Business Local/Remote/Management v4.8.6 and earlier (Windows and Linux)","year":"2023","cvss_score":9.4,"severity":"critical","kev":false,"impact":"A default password that ships enabled. This is the most boring vulnerability in the facility layer and probably the most exploited class of them in practice - power management software gets installed once by whoever racked the gear and never revisited.","attack_vector":"Unauthenticated remote access to the PowerPanel Business interface using published default credentials.","remediation":"Upgrade past v4.8.6 and change the credential. Then go and check every other piece of power-management software you inherited with a build: default credentials on facility gear are the single highest-yield internal audit an operator can run, and it costs a day.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25131"],"status":"curated","published":"2023-04-24"},{"id":"CVE-2023-3128","cve":"CVE-2023-3128","aliases":[],"title":"Grafana: Azure AD accounts validated on the mutable, non-unique email claim","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Grafana","year":"2023","cvss_score":9.4,"severity":"critical","kev":false,"impact":"Azure AD accounts validated on the mutable, non-unique email claim -> account takeover / auth bypass","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; pin the OAuth tenant and allowed groups","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-3128"],"status":"curated","published":"2023-06-22"},{"id":"CVE-2023-4966","cve":"CVE-2023-4966","aliases":[],"title":"Citrix NetScaler ADC/Gateway: \"CitrixBleed\" - memory overread leaking valid session tokens","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Citrix NetScaler ADC/Gateway","year":"2023","cvss_score":9.4,"severity":"critical","kev":true,"impact":"\"CitrixBleed\" - memory overread leaking valid session tokens -> MFA bypass","attack_vector":"Network (remote)","remediation":"Control-plane: patch AND terminate every existing session; tokens survive patching","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-4966"],"status":"curated","published":"2023-10-10"},{"id":"CVE-2024-0964","cve":"CVE-2024-0964","aliases":[],"title":"Gradio: Remotely triggerable local file include via a JSON value in an API request","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Gradio","year":"2024","cvss_score":9.4,"severity":"critical","kev":false,"impact":"Remotely triggerable local file include via a JSON value in an API request","attack_vector":"Unauthenticated network to the demo","remediation":"Upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0964"],"status":"curated","published":"2024-02-05"},{"cvss_vector":"CVSS:4.0/AV:N/AC:L/AT:N/PR:N/UI:P/VC:H/VI:H/VA:H/SC:H/SI:H/SA:H","cwe":["CWE-94","CWE-352"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2025-62593","cve":"CVE-2025-62593","aliases":["GHSA-q279-jhrf-cc6v","ShadowRay 2.0"],"title":"Ray (dashboard job submission API, browser-origin guard): Ray's only defense against browser-driven job submission was","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ray (dashboard job submission API, browser-origin guard)","year":"2025","cvss_score":9.4,"severity":"critical","kev":true,"impact":"Ray's only defense against browser-driven job submission was a User-Agent check, which Firefox and Safari let an attacker route around via DNS rebinding. A developer merely visiting a malicious web page hands the attacker remote code execution on their Ray cluster - including laptop-local dev clusters that were never meant to be reachable. CISA added this to the KEV catalog on 2026-08-17, so treat it as actively exploited.","attack_vector":"A remote web page loaded in the victim's Firefox or Safari, which then rebinds DNS to reach a Ray dashboard on localhost or on any network the victim's browser can route to. No credentials on the Ray side.","remediation":"Upgrade Ray to 2.52.0 or later, then enable token authentication (RAY_AUTH_MODE=token) since the 2.52.0 fix alone does not authenticate the endpoints. Restart head and workers. Given the KEV listing, also audit job history and node processes on any cluster that was reachable from a developer workstation.","references":["https://github.com/ray-project/ray/security/advisories/GHSA-q279-jhrf-cc6v","https://nvd.nist.gov/vuln/detail/CVE-2025-62593"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-4404","cve":"CVE-2026-4404","aliases":[],"title":"Harbor: Hard-coded default credentials give web UI access to the whole registry","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Harbor","year":"2026","cvss_score":9.4,"severity":"critical","kev":false,"impact":"Hard-coded default credentials give web UI access to the whole registry","attack_vector":"Unauthenticated network","remediation":"Upgrade Harbor immediately and change the default password; treat every hosted image as potentially tampered with and re-verify signatures","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-4404"],"status":"curated","published":"2026-03-23"},{"id":"CVE-2026-53488","cve":"CVE-2026-53488","aliases":[],"title":"containerd: CRI plugin propagates unvalidated image LABEL values into container config","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2026","cvss_score":9.4,"severity":"critical","kev":false,"impact":"CRI plugin propagates unvalidated image LABEL values into container config; highest-severity containerd issue to date","attack_vector":"Malicious image","remediation":"Emergency rolling containerd upgrade with node drain; gate tenant image sources","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53488"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2026-07-01"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:L","cwe":["CWE-787","CWE-668"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-63830","cve":"CVE-2026-63830","aliases":[],"title":"Linux kernel (net/core, net/tls): The bitmap that marks sk_msg scatterlist entries as externally owned was not carried","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/core, net/tls)","year":"2026","cvss_score":9.4,"severity":"critical","kev":false,"impact":"The bitmap that marks sk_msg scatterlist entries as externally owned was not carried along when entries are shifted, split, transferred or compacted - including the partial tail entry produced when kTLS splits an open record. An entry backed by a splice-supplied page-cache page can therefore arrive in a new slot with its ownership bit clear, at which point a sockmap BPF verdict is handed the page as writable program data. Writes then land in the shared page cache: file contents belonging to other workloads on the node get modified through what should be a read-only zero-copy reference.","attack_vector":"Requires a sk_msg/sockmap BPF program attached on the node - a service mesh or CNI datapath (Cilium and similar), not something a tenant can attach itself - and a socket doing splice or sendfile-style zero-copy, with kTLS in the mix for the open-record-split variant. The corrupted target is the page cache, which is shared across every tenant on the host, so the blast radius crosses the tenant boundary even though the trigger sits in the platform's own datapath. Say plainly: the tenant is the victim here more often than the attacker, unless tenants can influence what the mesh's BPF program writes.","remediation":"Boot a kernel carrying the fix commits below (no fixed stable version published). Interim control: disable sk_msg/sockmap acceleration in the service mesh or CNI (Cilium's socket-level load balancing and sockops redirection) on nodes carrying shared file-backed data until patched.","references":["https://git.kernel.org/stable/c/f126eed589eec6f201405abbc398844042ef6d57","https://git.kernel.org/stable/c/31a110642b5fb5e61940cbcfb503445ac4f28017","https://nvd.nist.gov/vuln/detail/CVE-2026-63830"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:4.0/AV:N/AC:L/AT:N/PR:L/UI:N/VC:H/VI:H/VA:H/SC:H/SI:H/SA:H","cwe":["CWE-426"],"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2026-048-cloudnativepg-instance-manager-p","cve":null,"aliases":["GHSA-x8c2-3p4r-v9r6","CVE-2026-55769 (reserved)"],"title":"CloudNativePG instance manager (PostgreSQL connection search_path): The owner of any managed database — a role","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"CloudNativePG instance manager (PostgreSQL connection search_path)","year":"2026","cvss_score":9.4,"severity":"critical","kev":false,"impact":"The owner of any managed database — a role CloudNativePG creates by default at bootstrap, so typically the tenant themselves — escalates to PostgreSQL superuser across every database in the cluster and then to OS command execution inside the instance pod. The instance manager opens superuser connections without pinning search_path in the startup packet, so resolution falls back through ALTER ROLE and ALTER DATABASE, both of which the database owner controls. The tenant plants overloads of built-in operators such as = and > in the public schema and re-points search_path at them. On the next reconcile, routine introspection queries run as the cluster postgres role and, although the relation is schema-qualified as pg_catalog.pg_extension, the operators in the same query are not — so the planted function bodies execute with superuser rights. From there: COPY ... FROM PROGRAM for shell inside the pod, then the pod ServiceAccount token, with the remaining blast radius set by that ServiceAccount's RBAC and the surrounding cloud workload identity. This is the CVE-2018-1058 pattern, and unlike the sibling password advisory it needs no unusual configuration.","attack_vector":"Network, low privileges: a role holding DATABASE OWNER on any CNPG-managed database. That role is created by default at cluster bootstrap and is normally the one handed to the application or tenant. Exploitation completes on the operator's own next reconcile, so no user interaction is needed.","remediation":"Upgrade CloudNativePG to 1.28.4, 1.29.2 or 1.30.0 and roll the operator and instance pods. Where you cannot upgrade yet, pin a fixed search_path on operator connections, or make sure no untrusted party holds DATABASE OWNER on a managed database. After patching, inspect the public schema of managed databases for planted operator or function overloads, since the fix does not remove what is already there, and rotate the instance pod ServiceAccount token if you find any.","references":["https://github.com/cloudnative-pg/cloudnative-pg/security/advisories/GHSA-x8c2-3p4r-v9r6","https://github.com/cloudnative-pg/cloudnative-pg/pull/10774"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2021-37678","cve":"CVE-2021-37678","aliases":[],"title":"TensorFlow / Keras: Arbitrary code execution via unsafe YAML deserialization of model config","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"TensorFlow / Keras","year":"2021","cvss_score":9.3,"severity":"critical","kev":false,"impact":"Arbitrary code execution via unsafe YAML deserialization of model config","attack_vector":"Customer-supplied model YAML","remediation":"Historical but still present in pinned tenant envs; patch base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-37678"],"status":"curated","published":"2021-08-12"},{"id":"CVE-2023-1177","cve":"CVE-2023-1177","aliases":[],"title":"MLflow (tracking server): Path traversal (`\\..\\filename`)","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"MLflow (tracking server)","year":"2023","cvss_score":9.3,"severity":"critical","kev":false,"impact":"Path traversal (`\\..\\filename`) → arbitrary file read","attack_vector":"Unauthenticated network to the MLflow tracking server","remediation":"Upgrade to 2.2.1+; MLflow has no auth by default","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-1177"],"status":"curated","fleet":{"ubiquity":"Common - MLflow is the default experiment/model registry alongside GPU training clusters","remediation_pain":"`daemon-restart` - upgrade the tracking server; the pain is credential rotation, since leaked SSH keys and cloud creds must be assumed compromised","pain_class":"daemon-restart","why_fleet_wide":"Unauthenticated LFI via the Model Versions API reads any file the server can read - SSH keys, cloud credentials - and one tracking server usually fronts every team's models and artifact buckets"},"published":"2023-03-24"},{"id":"CVE-2023-24509","cve":"CVE-2023-24509","aliases":[],"title":"Arista EOS (redundant supervisor, RPR/SSO): On modular chassis with dual supervisors running RPR or SSO redundancy","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (redundant supervisor, RPR/SSO)","year":"2023","cvss_score":9.3,"severity":"critical","kev":false,"impact":"On modular chassis with dual supervisors running RPR or SSO redundancy, an existing unprivileged user can log in to the standby supervisor as root. The standby has full access to the chassis state and becomes the active supervisor on failover, so this is a straight unprivileged-to-root path on your largest, most central switches — typically the spines.","attack_vector":"Authenticated but unprivileged user with network access to the standby supervisor.","remediation":"EOS upgrade. Requires a reload of the supervisors — do the standby first, fail over, then the former active, which keeps the chassis forwarding throughout. Until patched, restrict management reachability to the standby supervisor's address.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-24509"],"status":"curated","published":"2023-04-13"},{"cvss_vector":"CVSS:4.0/AV:N/AC:L/AT:N/PR:N/UI:N/VC:H/VI:H/VA:H/SC:N/SI:N/SA:N","cwe":["CWE-269"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-4976","cve":"CVE-2023-4976","aliases":[],"title":"Pure Storage FlashBlade management interface authentication: An attacker authenticates to the FlashBlade management","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Pure Storage FlashBlade management interface authentication","year":"2023","cvss_score":9.3,"severity":"critical","kev":false,"impact":"An attacker authenticates to the FlashBlade management interface as a local account through a path that was never meant to accept logins, and lands with privileged access to the array. Every filesystem and object bucket the array serves is then reachable.","attack_vector":"Network reach to the FlashBlade management interface. No valid operator credential is needed - the unintended authentication method is the way in.","remediation":"Upgrade Purity//FB to the fixed release from Pure's security bulletin. Before and after, restrict the management interface to an administrative network, and review array audit logs for logins that did not come from a known admin.","references":["https://www.purestorage.com/security","https://nvd.nist.gov/vuln/detail/CVE-2023-4976"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.0/AV:N/AC:L/PR:N/UI:N/S:C/C:H/I:L/A:N","cwe":["CWE-862","CWE-598"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2023-6020","cve":"CVE-2023-6020","aliases":["GHSA-6cxr-8q3m-jwrr","ShadowRay family"],"title":"Ray (dashboard /static/ file handler): Path traversal under the dashboard's /static/ route lets an unauthenticated","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ray (dashboard /static/ file handler)","year":"2023","cvss_score":9.3,"severity":"critical","kev":false,"impact":"Path traversal under the dashboard's /static/ route lets an unauthenticated caller read any file the Ray process can open. On a GPU head node that typically means cloud instance credentials, kubeconfigs, SSH keys and other tenants' checkpoint paths - enough to pivot into the rest of the cluster. Part of the same 2023 Ray CVE cluster as the ShadowRay job-submission bug.","attack_vector":"Anyone who can issue HTTP requests to the Ray dashboard port (8265). No authentication required.","remediation":"Upgrade Ray to 2.8.1 or later and restart the head node. Because the dashboard has no authentication by design in these versions, also restrict port 8265 with a NetworkPolicy or firewall and rotate any credential files that lived on the head node while it was exposed.","references":["https://github.com/advisories/GHSA-6cxr-8q3m-jwrr","https://nvd.nist.gov/vuln/detail/CVE-2023-6020"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-22252","cve":"CVE-2024-22252","aliases":[],"title":"VMware ESXi / Workstation / Fusion: Use-after-free in the XHCI USB controller","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"VMware ESXi / Workstation / Fusion","year":"2024","cvss_score":9.3,"severity":"critical","kev":false,"impact":"Use-after-free in the XHCI USB controller - guest-to-host code execution as the VMX process","attack_vector":"Tenant VM guest (local admin inside the VM)","remediation":"ESXi patch + host reboot with evacuation. Interim mitigation: remove USB controllers from tenant VMs, which is cheap and usually harmless for GPU workloads","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-22252"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-03-05"},{"id":"CVE-2024-22253","cve":"CVE-2024-22253","aliases":[],"title":"VMware ESXi / Workstation / Fusion: Use-after-free in the UHCI USB controller","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"VMware ESXi / Workstation / Fusion","year":"2024","cvss_score":9.3,"severity":"critical","kev":false,"impact":"Use-after-free in the UHCI USB controller - guest-to-host code execution as the VMX process","attack_vector":"Tenant VM guest","remediation":"ESXi patch + host reboot with evacuation; same USB-controller removal mitigation","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-22253"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-03-05"},{"id":"CVE-2024-3573","cve":"CVE-2024-3573","aliases":[],"title":"MLflow (LFI via URI parsing): Local file inclusion — read arbitrary files","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"MLflow (LFI via URI parsing)","year":"2024","cvss_score":9.3,"severity":"critical","kev":false,"impact":"Local file inclusion — read arbitrary files","attack_vector":"Unauthenticated network to the tracking server","remediation":"Upgrade; one of ~15 traversal variants, each a bypass of the last","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-3573"],"status":"curated","published":"2024-04-16"},{"cvss_vector":"CVSS:4.0/AV:N/AC:L/AT:N/PR:N/UI:N/VC:H/VI:H/VA:N/SC:N/SI:N/SA:N","cwe":["CWE-269"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2024-55949","cve":"CVE-2024-55949","aliases":[],"title":"MinIO (admin IAM import API): The IAM import API can be driven to grant an attacker administrative policy, converting a","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"MinIO (admin IAM import API)","year":"2024","cvss_score":9.3,"severity":"critical","kev":false,"impact":"The IAM import API can be driven to grant an attacker administrative policy, converting a low-privileged or unauthenticated position into full control of users, policies and every bucket in the deployment. That is total collapse of the tenancy model on the object store.","attack_vector":"Reachable against the MinIO admin API endpoint over the network.","remediation":"Upgrade to RELEASE.2024-12-18T13-15-44Z or later and restart all nodes. Then dump the IAM configuration and diff it against your intended state, remove any policy attachments you did not create, and rotate root and admin credentials.","references":["https://github.com/minio/minio/security/advisories/GHSA-cwq8-g58r-32hg","https://nvd.nist.gov/vuln/detail/CVE-2024-55949"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-22224","cve":"CVE-2025-22224","aliases":[],"title":"VMware ESXi / Workstation: TOCTOU race leading to an out-of-bounds write in VMX - full VM escape to host code execution","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"VMware ESXi / Workstation","year":"2025","cvss_score":9.3,"severity":"critical","kev":true,"impact":"TOCTOU race leading to an out-of-bounds write in VMX - full VM escape to host code execution; exploited in the wild as a zero-day [KEV]","attack_vector":"Tenant VM guest (local admin inside the VM)","remediation":"ESXi patch + host reboot with vMotion evacuation. GPU-passthrough VMs cannot vMotion, so this is a tenant-visible drain. Top-priority escape of the 2025 set","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-22224"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-03-04"},{"id":"CVE-2025-33187","cve":"CVE-2025-33187","aliases":[],"title":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware: An attacker with privileged access reaches","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware","year":"2025","cvss_score":9.3,"severity":"critical","kev":false,"impact":"An attacker with privileged access reaches SoC-protected areas through SROOT, ending in code execution below the OS. At 9.3 this is the most severe entry in the GB10 firmware set. These sit in the GB10 root-of-trust chain, so a successful exploit undermines the platform's own attestation and secure-boot story rather than just the OS above it.","attack_vector":"Local access to the DGX Spark. Several of the set need no privileges at all; the rest need host root. This is a desk-side developer box, so physical and local access assumptions are much weaker than for a racked DGX.","remediation":"Apply the DGX Spark firmware update from bulletin 5720. Cost: flash plus reboot, low drain cost given the form factor, but not live-patchable and root-of-trust firmware cannot be rolled back once applied.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33187","https://github.com/NVIDIA/product-security/tree/main/2025/5720"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-269"],"fleet":{"pain_class":"firmware-flash"},"published":"2025-11-25"},{"id":"CVE-2025-41236","cve":"CVE-2025-41236","aliases":[],"title":"VMware ESXi / Workstation / Fusion: Integer overflow in the VMXNET3 virtual NIC","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"VMware ESXi / Workstation / Fusion","year":"2025","cvss_score":9.3,"severity":"critical","kev":false,"impact":"Integer overflow in the VMXNET3 virtual NIC - VM escape to host code execution (Pwn2Own Berlin 2025)","attack_vector":"Tenant VM guest","remediation":"ESXi patch + host reboot with evacuation. VMXNET3 is the default adapter, so there is no practical configuration mitigation","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-41236"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-07-15"},{"id":"CVE-2025-41237","cve":"CVE-2025-41237","aliases":[],"title":"VMware ESXi / Workstation / Fusion: Integer underflow in VMCI leading to an out-of-bounds write","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"VMware ESXi / Workstation / Fusion","year":"2025","cvss_score":9.3,"severity":"critical","kev":false,"impact":"Integer underflow in VMCI leading to an out-of-bounds write - VM escape (Pwn2Own Berlin 2025)","attack_vector":"Tenant VM guest","remediation":"ESXi patch + host reboot with evacuation","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-41237"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-07-15"},{"id":"CVE-2025-41238","cve":"CVE-2025-41238","aliases":[],"title":"VMware ESXi / Workstation / Fusion: Heap overflow in the PVSCSI controller","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"VMware ESXi / Workstation / Fusion","year":"2025","cvss_score":9.3,"severity":"critical","kev":false,"impact":"Heap overflow in the PVSCSI controller - out-of-bounds write and VM escape (Pwn2Own Berlin 2025)","attack_vector":"Tenant VM guest","remediation":"ESXi patch + host reboot with evacuation","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-41238"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-07-15"},{"id":"CVE-2025-53696","cve":"CVE-2025-53696","aliases":["CVE-2025-53695"],"title":"Software House iSTAR Ultra firmware verification and web application (tested through 6.9.2): The controller verifies","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Software House iSTAR Ultra firmware verification and web application (tested through 6.9.2)","year":"2025","cvss_score":9.3,"severity":"critical","kev":false,"impact":"The controller verifies its firmware at boot, but the verification skips portions of the image - so an attacker who can push firmware gets persistent, boot-surviving code on the door controller that the device's own integrity check will happily bless. The companion issue is an authenticated OS command injection in the web application that escalates to root, which is a plausible way to reach the firmware write in the first place. This is the persistence problem in its purest form: you can re-image the panel, but if the attacker's implant is in the unverified region and you restore from a backup taken after compromise, it comes back. For a GPU operator the practical meaning is that door control - the boundary protecting drives, console ports and the OOB switch inside the cage - can be silently and durably owned, and no amount of log review will show it because the implant controls the logs. This is also a tenant-handoff failure: a panel compromised during one customer's tenancy stays compromised for the next one, since nobody re-flashes access-control hardware between tenants.","attack_vector":"The command-injection path requires an authenticated session to the iSTAR web application, so realistically credential theft, a default or shared integrator account, or chaining from one of the unauthenticated iSTAR issues. The firmware verification gap is then exercised by an attacker who already has that access. All of it lives on the physical-security VLAN.","remediation":"Move to firmware later than 6.9.2 per Johnson Controls' guidance and confirm with the vendor that the verification gap is closed in the version you land on - the disclosure notes later firmware may also be affected, so a version number alone is not evidence. Because the flaw defeats integrity verification, patching is not sufficient for a panel you believe was reachable while vulnerable: replace it or have the vendor perform a full low-level reflash, and rebuild its configuration from a known-good source rather than a device backup. Operationally, rotate every credential on the access-control system, audit the account list for integrator and service accounts, and require MFA on any path to the iSTAR web application.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-53696","https://nvd.nist.gov/vuln/detail/CVE-2025-53695","https://www.johnsoncontrols.com/trust-center/cybersecurity/security-advisories"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-64513","cve":"CVE-2025-64513","aliases":[],"title":"Milvus: Unauthenticated attacker exploits the server directly","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Milvus","year":"2025","cvss_score":9.3,"severity":"critical","kev":false,"impact":"Unauthenticated attacker exploits the server directly","attack_vector":"Unauthenticated network to Milvus","remediation":"Upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-64513"],"status":"curated","published":"2025-11-10"},{"id":"CVE-2026-20794","cve":"CVE-2026-20794","aliases":[],"title":"Intel Data Center GPU driver for VMware ESXi (buffer overflow): A buffer overflow in the Intel datacenter graphics","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel Data Center GPU driver for VMware ESXi (buffer overflow)","year":"2026","cvss_score":9.3,"severity":"critical","kev":false,"impact":"A buffer overflow in the Intel datacenter graphics driver running in the ESXi device-driver ring lets a privileged local actor escalate and execute code. This is the GPU driver in the hypervisor - directly relevant to anyone running Intel Data Center GPU Flex/Max under ESXi.","attack_vector":"Local privileged access on the ESXi host.","remediation":"Update the Intel Data Center Graphics Driver for ESXi to 2.0.2 or later. Driver VIB update plus host reboot, so it lands as part of a rolling host maintenance pass.","references":["https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01402.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-22696","cve":"CVE-2026-22696","aliases":["dcap-qvl QE Identity verification bypass","DCAP quote verification flaw"],"title":"Phala dcap-qvl - the Rust/npm/Python DCAP quote verification library used to verify Intel SGX and TDX attestation","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Phala dcap-qvl - the Rust/npm/Python DCAP quote verification library used to verify Intel SGX and TDX attestation…","year":"2026","cvss_score":9.3,"severity":"critical","kev":false,"impact":"The library fetches Quoting Enclave identity collateral from the provisioning service but never verifies its signature against the certificate chain, and does not enforce MRSIGNER, ISVPRODID or ISVSVN policy on the QE report. An attacker forges QE Identity data to whitelist a Quoting Enclave that is not Intel's, then signs arbitrary quotes that the verifier accepts as genuine. That is total forgery of SGX and TDX remote attestation: a machine can claim to be running an attested confidential workload while running anything at all. For an operator or a customer relying on attestation to prove that model weights only decrypt inside a genuine TDX trust domain, the proof is worthless. This is the highest-scored item in this set and it is a pure software bug in the verifier, not a CPU flaw.","attack_vector":"Whoever controls the machine claiming to be attested, plus the ability to serve or influence the collateral the verifier fetches. No CPU vulnerability, no physical access, no privileged position on the verifier - the verifier simply accepts a forged identity. Anyone running a confidential-compute service whose verification path uses this library is exposed.","remediation":"Pure software update, no firmware and no reboot: upgrade to dcap-qvl 0.3.9 or later (npm @phala/dcap-qvl-node / -web 0.3.4+). Fast and cheap to deploy, which is the good news. The important operator action is inventory: DCAP quote verification is usually buried inside a confidential-computing framework or a TEE-attestation SaaS rather than being a dependency anyone declared deliberately, so audit which library your attestation path actually calls. Any attestation accepted by a vulnerable version before the upgrade should be treated as unverified and re-attested - the fix does not retroactively invalidate quotes you already trusted.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-22696","https://github.com/Phala-Network/dcap-qvl/security/advisories/GHSA-796p-j2gh-9m2q"],"status":"curated","published":"2026-01-26"},{"id":"CVE-2026-24834","cve":"CVE-2026-24834","aliases":[],"title":"Kata Containers: Kata with Cloud Hypervisor allows a user to break the VM isolation boundary","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kata Containers","year":"2026","cvss_score":9.3,"severity":"critical","kev":false,"impact":"Kata with Cloud Hypervisor allows a user to break the VM isolation boundary","attack_vector":"Any tenant workload running under Kata","remediation":"Emergency Kata upgrade; drain and recreate Kata pods. The whole point of Kata is this boundary, so treat as critical","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24834"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2026-02-19"},{"id":"CVE-2026-48797","cve":"CVE-2026-48797","aliases":[],"title":"Backpropagate (single-GPU LLM fine-tuning library) - Reflex web UI: The optional web UI exposes a training control","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Backpropagate (single-GPU LLM fine-tuning library) - Reflex web UI","year":"2026","cvss_score":9.3,"severity":"critical","kev":false,"impact":"The optional web UI exposes a training control surface without authentication, scored 9.3. On a GPU box this means an unauthenticated caller drives training jobs on your hardware - the practical outcomes are GPU theft for someone else's workload, poisoning of a fine-tune in progress, and read access to whatever the training process can see.","attack_vector":"Network, unauthenticated. Anyone who can reach the UI port. Fine-tuning UIs get bound to 0.0.0.0 during experimentation far more often than anyone admits.","remediation":"Upgrade past 1.1.1, or disable the Reflex web UI entirely if you do not use it. Cost: restart the service. The standing control is a network policy that stops research workloads binding tooling ports to anything but localhost.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-48797"],"status":"curated","fleet":{"pain_class":"daemon-restart"},"published":"2026-06-17"},{"id":"CVE-2026-53475","cve":"CVE-2026-53475","aliases":[],"title":"Assisted Migration Agent (hardcoded insecure TLS to vCenter): The agent hardcodes insecure TLS when talking to vCenter","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Assisted Migration Agent (hardcoded insecure TLS to vCenter)","year":"2026","cvss_score":9.3,"severity":"critical","kev":false,"impact":"The agent hardcodes insecure TLS when talking to vCenter, so a machine-in-the-middle harvests vCenter administrator credentials in transit - unauthorized admin access to the virtualization estate.","attack_vector":"Network position between the migration agent and vCenter.","remediation":"Update the assisted-migration-agent to a build including the upstream fix, and rotate the vCenter admin credentials the agent used. Retire the agent when the migration completes rather than leaving it deployed.","references":["https://access.redhat.com/security/cve/CVE-2026-53475"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-787","CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-63938","cve":"CVE-2026-63938","aliases":[],"title":"Linux kernel (arch/x86/kvm/svm): Page State Change requests from a confidential guest were validated against the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (arch/x86/kvm/svm)","year":"2026","cvss_score":9.3,"severity":"critical","kev":false,"impact":"Page State Change requests from a confidential guest were validated against the maximum possible scratch-area size rather than the size the guest's own pointer actually leaves available, so a guest can make the host walk PSC entries beyond the buffer. Result is guest-directed out-of-bounds access in host kernel memory from inside an encrypted VM.","attack_vector":"Issued by the guest itself: a SEV-SNP guest places its scratch pointer at a non-zero offset in the GHCB shared buffer and submits a PSC request whose entry indices run past the effective end. No host privilege, no VMM involvement. Conditional on the node running SEV-SNP guests on kvm_amd.","remediation":"Boot a kernel with the referenced stable commits applied; the record carries no fixed release string, so verify by commit. Until then, keep SEV-SNP tenants off the affected nodes or downgrade those workloads to non-confidential VMs.","references":["https://git.kernel.org/stable/c/5198f70c09a5f6e9e5f5a0a2c6b388f24294b176","https://git.kernel.org/stable/c/75c8d1d7291268b479794fba5808971dc2f5eaf3","https://nvd.nist.gov/vuln/detail/CVE-2026-63938"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-787","CWE-131"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-63939","cve":"CVE-2026-63939","aliases":[],"title":"Linux kernel (arch/x86/kvm/svm): KVM computed the usable size of the guest-provided GHCB scratch area wrongly, so a","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (arch/x86/kvm/svm)","year":"2026","cvss_score":9.3,"severity":"critical","kev":false,"impact":"KVM computed the usable size of the guest-provided GHCB scratch area wrongly, so a guest that points its scratch pointer partway into the shared buffer gets KVM to read and write past the end of it. That is a guest-controlled host-kernel buffer overflow out of an encrypted VM - the escape primitive the whole SEV design is meant to prevent.","attack_vector":"A malicious SEV-ES / SEV-SNP guest sets an in-GHCB scratch pointer at an offset inside the shared buffer and then issues an operation (notably a Page State Change request) whose payload runs past the real end of the area. Reachable purely from guest ring 0 on any node running AMD confidential VMs; no host access needed.","remediation":"Update to a kernel containing the referenced stable commits (no fixed version string was published with the record - match the commit hash). Interim: drain SEV-ES/SNP tenants off unpatched nodes, or disable SEV-ES/SNP guest types in the scheduler until patched.","references":["https://git.kernel.org/stable/c/6ca9400d36005ffdca25f80186bea781c7e1dc4c","https://git.kernel.org/stable/c/9f0a9e780f02c02d025a190f1885e1d1d73b87bd","https://nvd.nist.gov/vuln/detail/CVE-2026-63939"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-191","CWE-787"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-63940","cve":"CVE-2026-63940","aliases":[],"title":"Linux kernel (arch/x86/kvm/svm): A confidential guest can hand KVM a port-I/O request with length or count zero","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (arch/x86/kvm/svm)","year":"2026","cvss_score":9.3,"severity":"critical","kev":false,"impact":"A confidential guest can hand KVM a port-I/O request with length or count zero, underflowing the size arithmetic that sizes the GHCB scratch area and pushing host-kernel accesses past the end of that buffer. The SEV-ES/SNP boundary that is supposed to keep an encrypted tenant away from host memory is what fails here.","attack_vector":"Driven entirely from inside a running SEV-ES / SEV-SNP guest: the guest issues a VMGEXIT port-I/O exit through its own GHCB with len/count of 0. Needs no host account and no VMM cooperation. Only applies to nodes actually running AMD SEV-ES/SNP guests under kvm_amd; a container tenant with no VM of its own cannot reach it.","remediation":"Boot a kernel carrying the referenced stable commits - the kernel CNA published no fixed release string for this record, so match on the commit in your distro's changelog. Interim control: stop scheduling SEV-ES/SNP guests on unpatched nodes, or run those tenants as ordinary (non-confidential) VMs until the fix lands.","references":["https://git.kernel.org/stable/c/3b6035bc6bff20e89752ce4358bc4c9a9d5883f2","https://git.kernel.org/stable/c/2254972d4d69e279ba4e87bf0968eb08ad0d3c92","https://nvd.nist.gov/vuln/detail/CVE-2026-63940"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-72329","cve":"CVE-2026-72329","aliases":[],"title":"Linux liquidio driver (Marvell/Cavium, cached VF pci_dev lookup table): The LiquidIO PF caches VF `pci_dev` pointers","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux liquidio driver (Marvell/Cavium, cached VF pci_dev lookup table)","year":"2026","cvss_score":9.3,"severity":"critical","kev":false,"impact":"The LiquidIO PF caches VF `pci_dev` pointers without taking a reference, so the cached pointers can dangle and are then dereferenced when handling a VF function-level-reset request. A VF triggering an FLR — something a tenant does simply by resetting their own device — drives a use-after-free in the host kernel. FLR is the operation an operator relies on to *clean up* between tenants, so the mechanism intended to enforce the handoff boundary is the one that breaks it.","attack_vector":"A tenant holding a LiquidIO VF issuing a function-level reset, or any path that triggers `OCTEON_VF_FLR_REQUEST` handling on the PF.","remediation":"Kernel upgrade plus host reboot on nodes with Marvell LiquidIO adapters. Rolling drain. Note that LiquidIO is end-of-life hardware still present in older inference fleets — if you are running it, weigh replacing the adapters against maintaining kernel patches for a driver that is no longer actively developed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-72329"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-08-15"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-362","CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-72495","cve":"CVE-2026-72495","aliases":[],"title":"Linux kernel (drivers/infiniband/hw/bnxt_re): A user context could request the write-combine doorbell page repeatedly","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/hw/bnxt_re)","year":"2026","cvss_score":9.3,"severity":"critical","kev":false,"impact":"A user context could request the write-combine doorbell page repeatedly and concurrently, although the driver supports exactly one, with no lock protecting the state. Racing requests corrupt the doorbell page index mapping and strand the allocated index when the mmap insert fails - a tenant influencing which doorbell page it or another context ends up mapped to. The CNA scored it scope-changed and rated it 9.3.","attack_vector":"Two threads inside one tenant container holding /dev/infiniband/uverbs* on a Broadcom bnxt_re NIC issuing the WC-page allocation request simultaneously. No fabric peer, no host root.","remediation":"No fixed version is listed in the record - take the stable kernel carrying 478c4d24193f (or da406b8b49c1 / 441baa790434) and reboot. Interim: drop /dev/infiniband/* from untrusted containers on bnxt_re nodes.","references":["https://git.kernel.org/stable/c/478c4d24193fe3e6aa2accd4874ae43000e4a217","https://git.kernel.org/stable/c/da406b8b49c1dfe661a497483940d7ee781430db","https://nvd.nist.gov/vuln/detail/CVE-2026-72495"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-74439","cve":"CVE-2026-74439","aliases":[],"title":"Linux kernel (drivers/iommu/intel): The VT-d scalable-mode context entry is zeroed while its Present bit is still set","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/intel)","year":"2026","cvss_score":9.3,"severity":"critical","kev":false,"impact":"The VT-d scalable-mode context entry is zeroed while its Present bit is still set, and the PASID directory pages are freed before the IOMMU is told to stop using them. The hardware can keep walking a live-looking entry into memory the kernel has already reallocated, so a tenant's device DMAs into arbitrary reused host memory - a direct tenant-to-host DMA escape, not just a crash.","attack_vector":"Reached on the normal teardown path whenever a passthrough device with a scalable-mode PASID table is released - a tenant closing its /dev/vfio/* device fd, or a VM exiting, is enough to run device_pasid_table_teardown(). Requires VT-d scalable mode (the default on modern Intel platforms with PASID/SVA); no host root.","remediation":"Update to a stable kernel carrying commits e9e83bcf / 58871810. The CNA's fixed-version data for this record is not usable as a version target, so verify the commits are in your distro kernel. Interim: disable VT-d scalable mode / PASID (intel_iommu=sm_off) on nodes that hand devices to tenants, accepting the loss of SVA.","references":["https://git.kernel.org/stable/c/e9e83bcfe37dc719182500dd823c03ab57d934f0","https://git.kernel.org/stable/c/588718101e8449605f1c7e858fecb7cfa701cdab","https://nvd.nist.gov/vuln/detail/CVE-2026-74439"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-74517","cve":"CVE-2026-74517","aliases":[],"title":"Linux kernel (arch/x86/kvm): The I/O APIC's delayed EOI work was cancelled only after vCPUs were freed, so the work","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (arch/x86/kvm)","year":"2026","cvss_score":9.3,"severity":"critical","kev":false,"impact":"The I/O APIC's delayed EOI work was cancelled only after vCPUs were freed, so the work item can fire against destroyed vCPUs and deliver an interrupt through freed memory. KASAN shows a slab use-after-free read in the APIC fast-delivery path - a host kernel UAF that a local process can schedule at will, which is a privilege-escalation primitive on a shared node.","attack_vector":"Any process that can open /dev/kvm: create a VM with an in-kernel I/O APIC, assert a line so the delayed EOI work is queued, then destroy the VM and let the worker run against the freed vCPUs. No guest cooperation needed beyond running the VM. In a cluster this matters wherever tenants get /dev/kvm directly or via nested virtualization.","remediation":"Update to a kernel with the fix (the record lists 6.13 as a fixed release; verify against the referenced commits for your stable branch). Interim: keep /dev/kvm out of tenant containers and disable nested virt for tenant VMs.","references":["https://git.kernel.org/stable/c/ed56a6b58222f9c1f4115a0bd2788dd6ed6022e2","https://git.kernel.org/stable/c/9910e835580fef3bef53b70241dd00c4bffad693","https://nvd.nist.gov/vuln/detail/CVE-2026-74517"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-125","CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-74573","cve":"CVE-2026-74573","aliases":[],"title":"Linux kernel (drivers/iommu/arm/arm-smmu-v3): On Arm hosts a virtual device is mapped to only the first of its Stream","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/arm/arm-smmu-v3)","year":"2026","cvss_score":9.3,"severity":"critical","kev":false,"impact":"On Arm hosts a virtual device is mapped to only the first of its Stream IDs, so a guest's invalidation requests never reach the ATC and IOTLB entries belonging to its other streams. Stale device-side translations survive an unmap and the device keeps DMAing into pages the guest already released - an open DMA window into recycled memory. A device with no streams at all makes the kernel read a zero-size pointer out of bounds.","attack_vector":"A tenant VMM using iommufd vDEVICE on an Arm SMMUv3 host (Grace-class AI nodes) creates a vDEVICE for a passthrough device that presents more than one Stream ID, then relies on guest-driven invalidation. Guest-driven and conditional on Arm SMMUv3 with nested translation; not reachable on x86 hosts.","remediation":"Update to a stable kernel carrying commits 3808bab5 / 0acbc621. Interim on Grace/Arm nodes: do not enable nested translation (iommufd vDEVICE) for tenant guests, and restrict passthrough to single-Stream-ID devices until patched.","references":["https://git.kernel.org/stable/c/3808bab5d95ae79e333e11f6a73d178e084c645d","https://git.kernel.org/stable/c/0acbc621341aca4eb94d9c2f43e1ab273ff088f0","https://nvd.nist.gov/vuln/detail/CVE-2026-74573"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:4.0/AV:N/AC:L/AT:P/PR:N/UI:N/VC:H/VI:H/VA:H/SC:N/SI:N/SA:N","cwe":["CWE-287"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2026-33322","cve":"CVE-2026-33322","aliases":[],"title":"MinIO (OIDC authentication): JWT algorithm confusion in the OIDC login path lets an attacker present a token the server","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"MinIO (OIDC authentication)","year":"2026","cvss_score":9.2,"severity":"critical","kev":false,"impact":"JWT algorithm confusion in the OIDC login path lets an attacker present a token the server validates under the wrong algorithm, authenticating as any identity the IdP could issue - including an administrator. Full takeover of the object store and every tenant's buckets in it.","attack_vector":"Anyone who can reach the MinIO console/STS login endpoint on a deployment configured with OIDC.","remediation":"Upgrade to the release in GHSA-5cx5-wh4m-82fh and restart all MinIO nodes. Rotate any long-lived service accounts and STS credentials issued before the upgrade, and check the audit log for logins that do not match a real IdP session.","references":["https://github.com/minio/minio/security/advisories/GHSA-5cx5-wh4m-82fh","https://nvd.nist.gov/vuln/detail/CVE-2026-33322"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:4.0/AV:N/AC:L/AT:P/PR:N/UI:N/VC:H/VI:H/VA:H/SC:N/SI:N/SA:N","cwe":["CWE-306","CWE-15"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2026-41176","cve":"CVE-2026-41176","aliases":[],"title":"rclone (rc API, options/set): options/set is exposed pre-authentication and can rewrite the running instance's auth","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"rclone (rc API, options/set)","year":"2026","cvss_score":9.2,"severity":"critical","kev":false,"impact":"options/set is exposed pre-authentication and can rewrite the running instance's auth settings, so an attacker first turns the remaining protections off and then drives the rest of the rc API - up to command execution. One unauthenticated call converts a partly-protected data mover into a fully open one.","attack_vector":"Any host with network reach to the rclone rc endpoint.","remediation":"Upgrade rclone and restart every rc/serve process. Rotate remote credentials for exposed instances. Bind the rc listener to loopback and put it behind an authenticating proxy rather than relying on rclone's own flags alone.","references":["https://github.com/rclone/rclone/security/advisories/GHSA-25qr-6mpr-f7qx","https://nvd.nist.gov/vuln/detail/CVE-2026-41176"],"status":"curated"},{"cvss_vector":"CVSS:4.0/AV:N/AC:L/AT:P/PR:N/UI:N/VC:H/VI:H/VA:H/SC:N/SI:N/SA:N","cwe":["CWE-78","CWE-306","CWE-94"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2026-41179","cve":"CVE-2026-41179","aliases":[],"title":"rclone (rc API, operations/fsinfo): operations/fsinfo is reachable without authentication and accepts an","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"rclone (rc API, operations/fsinfo)","year":"2026","cvss_score":9.2,"severity":"critical","kev":false,"impact":"operations/fsinfo is reachable without authentication and accepts an attacker-defined backend, so a caller can point it at a WebDAV remote whose bearer_token_command runs a shell command. Same practical outcome as the rcd RCE: code execution on the data-mover host and access to every credential in its config.","attack_vector":"Any host that can reach the rclone rc HTTP endpoint.","remediation":"Upgrade rclone and restart all rc/serve instances. Rotate the credentials in the rclone config for anything the host could reach. Enforce authentication on the rc endpoint and keep it off shared networks.","references":["https://github.com/rclone/rclone/security/advisories/GHSA-jfwf-28xr-xw6q","https://nvd.nist.gov/vuln/detail/CVE-2026-41179"],"status":"curated"},{"id":"CVE-2026-47243","cve":"CVE-2026-47243","aliases":[],"title":"Kata Containers: runtime-rs standalone virtio-fs path is vulnerable to a guest-to-host escape","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kata Containers","year":"2026","cvss_score":9.2,"severity":"critical","kev":false,"impact":"runtime-rs standalone virtio-fs path is vulnerable to a guest-to-host escape","attack_vector":"Any tenant VM under Kata","remediation":"Emergency Kata upgrade; node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47243"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2026-08-07"},{"id":"CVE-2026-49445","cve":"CVE-2026-49445","aliases":[],"title":"Cilium: With L7 enabled, the embedded Envoy exposes a world-accessible admin.sock on the cluster","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2026","cvss_score":9.2,"severity":"critical","kev":false,"impact":"With L7 enabled, the embedded Envoy exposes a world-accessible admin.sock on the cluster; full dataplane control","attack_vector":"Any pod on the cluster network","remediation":"Emergency rolling Cilium upgrade; no GPU drain but expect brief dataplane blips per node","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-49445"],"status":"curated","published":"2026-07-15"},{"id":"CVE-2017-16727","cve":"CVE-2017-16727","aliases":["ICSA-17-355-01"],"title":"Moxa NPort W2150A / W2250A wireless device server: The device ships with an empty default password, so anyone who can","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Moxa NPort W2150A / W2250A wireless device server","year":"2017","cvss_score":9.1,"severity":"critical","kev":false,"impact":"The device ships with an empty default password, so anyone who can reach it on the network can log in as an unauthorized user with no credential at all and take over the serial device server.","attack_vector":"No authentication required — just network reachability to a device still running the default (blank) password.","remediation":"Firmware upgrade past 1.11 plus setting a real administrator password on every device — the firmware fix stops shipping the box in an unauthenticated state, but existing deployed units also need someone to actually set a password during the upgrade. Track this as a fleet-wide credential-rotation project, not just a flash.","references":["http://www.securityfocus.com/bid/102254","https://ics-cert.us-cert.gov/advisories/ICSA-17-355-01"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2017-12-22"},{"id":"CVE-2018-6440","cve":"CVE-2018-6440","aliases":[],"title":"Brocade Fabric OS (proxy service information disclosure): Unauthenticated remote attackers can obtain sensitive","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Brocade Fabric OS (proxy service information disclosure)","year":"2018","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Unauthenticated remote attackers can obtain sensitive information from the Fabric OS proxy service. Pre-auth information disclosure on a SAN switch typically yields fabric topology and configuration — which is the reconnaissance an attacker needs to know which zone to attack to reach a specific tenant's storage.","attack_vector":"Unauthenticated, remote to the FOS proxy service.","remediation":"Fabric OS upgrade plus reboot per fabric. Immediate control is management-network isolation for all FC switch management interfaces.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-6440"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2018-12-03"},{"id":"CVE-2019-16261","cve":"CVE-2019-16261","aliases":[],"title":"Tripp Lite PDUMH15AT / SU750XL PDU: The PDU accepts unauthenticated POST requests to its /Forms/ endpoints, which can","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Tripp Lite PDUMH15AT / SU750XL PDU","year":"2019","cvss_score":9.1,"severity":"critical","kev":false,"impact":"The PDU accepts unauthenticated POST requests to its /Forms/ endpoints, which can be used to change the manager or admin password, or directly shut off power to an outlet. Anyone who can reach the PDU's web interface can cut power to whatever rack that outlet feeds, with no login at all.","attack_vector":"Fully remote and unauthenticated — a crafted POST request to the /Forms/ directory is all that's needed to flip an outlet or take over the admin account.","remediation":"Software/firmware upgrade — Tripp Lite (now Eaton) shipped a fixed release after this was reported; confirm every deployed unit is past 12.04.0053 (PDUMH15AT) / 12.04.0052 (SU750XL). Flash each PDU; this directly powers a rack, so schedule around a window where a brief power-monitoring interruption is acceptable, and audit for units still on vulnerable firmware since these are frequently forgotten 'fringe' infrastructure.","references":["https://blog.korelogic.com/blog/2019/08/19/unpatched_fringe_infrastructure_bits"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2019-09-12"},{"id":"CVE-2019-4169","cve":"CVE-2019-4169","aliases":["IBM X-Force 158702"],"title":"IBM OpenPower firmware OP910/OP920 - OpenBMC IPMI credential handling: The original default BMC password kept working","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"IBM OpenPower firmware OP910/OP920 - OpenBMC IPMI credential handling","year":"2019","cvss_score":9.1,"severity":"critical","kev":false,"impact":"The original default BMC password kept working over IPMI after an operator changed it. Every hardening runbook says 'change the default BMC password' and on these firmware levels doing so accomplished nothing for the IPMI path - the fleet stayed openable with a credential printed in the vendor documentation. This is the cleanest example in the cluster of why BMC posture cannot be assessed from configuration intent: you have to test that the old credential is actually dead.","attack_vector":"Network access to the BMC's IPMI interface with the publicly documented default credential. No prior foothold required.","remediation":"Fixed in later OpenPower firmware; delivery is a per-node system firmware update with a maintenance window. The transferable lesson is a test, not a patch: after any BMC credential rotation, actively attempt an IPMI and Redfish login with the old and default credentials and alert if either succeeds. Make that a recurring fleet check, not a one-off - it catches this class of bug on any vendor.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-4169","https://exchange.xforce.ibmcloud.com/vulnerabilities/158702"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2019-08-26"},{"id":"CVE-2020-4926","cve":"CVE-2020-4926","aliases":[],"title":"IBM Spectrum Scale 5.1 core / IBM Elastic Storage System 6.1: Unauthorized access to user data, or injection of","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"IBM Spectrum Scale 5.1 core / IBM Elastic Storage System 6.1","year":"2020","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Unauthorized access to user data, or injection of arbitrary data into the communication between cluster nodes. This is the core GPFS daemon protocol, not an add-on layer — an attacker positioned on the storage cluster network can read other tenants' data or write data that nodes accept as legitimate. For a shared training filesystem, data injection is also a training-data poisoning vector.","attack_vector":"An attacker with access to the inter-node communication path of the Spectrum Scale cluster — the back-end storage network.","remediation":"Upgrade Spectrum Scale / ESS to a fixed level. This is a coordinated cluster upgrade; GPFS supports rolling node upgrades but the version-compatibility window means planning, and a full-cluster restart is sometimes unavoidable. Enable and verify GPFS cluster-level authentication and encryption in transit, which is a config change and the thing that actually removes the exposure.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-4926"],"status":"curated","tags":["tenant-isolation"],"published":"2022-05-24"},{"id":"CVE-2021-1577","cve":"CVE-2021-1577","aliases":[],"title":"Cisco APIC / Cloud APIC (API endpoint): Unauthenticated arbitrary file read and write on the APIC","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Cisco APIC / Cloud APIC (API endpoint)","year":"2021","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Unauthenticated arbitrary file read and write on the APIC — the ACI fabric controller. File write on the controller is effectively fabric takeover: you can plant credentials, alter policy state, and reprogram forwarding for every tenant on the pod.","attack_vector":"Unauthenticated, remote to the APIC's API endpoint. No credentials.","remediation":"APIC software upgrade across the controller cluster (rolling, one APIC at a time, fabric keeps forwarding). Restrict APIC API reachability to a dedicated management network as a durable control — a config change you should make regardless of this CVE.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1577"],"status":"curated","published":"2021-08-25"},{"id":"CVE-2021-22794","cve":"CVE-2021-22794","aliases":["SEVD-2022-095-01"],"title":"Schneider Electric StruxureWare Data Center Expert (DCE) v7.8.1 and prior: Path traversal to remote code execution","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Schneider Electric StruxureWare Data Center Expert (DCE) v7.8.1 and prior","year":"2021","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Path traversal to remote code execution on the DCIM appliance. DCE is the aggregation point for a site's entire power and cooling estate - it holds working credentials for every UPS, PDU, NMC, CRAC and sensor it polls. Compromise here is not one device; it is the keys to the whole facility layer, and from there PHYSICAL control of power and cooling.","attack_vector":"Network access to the DCE appliance. DCE typically sits on the facility management network with broad reachability by design, since it must poll everything.","remediation":"Upgrade DCE to v7.9.0 or later. Appliance upgrade with a service restart - hours, not a power maintenance window. Afterwards, rotate every device credential DCE stored, because they were all reachable. Long term, DCE should be the most tightly segmented host you run: inbound access from a jump host only, no internet egress.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-22794"],"status":"curated","fleet":{"pain_class":"daemon-restart"},"published":"2022-04-13"},{"id":"CVE-2021-22795","cve":"CVE-2021-22795","aliases":["SEVD-2022-095-01"],"title":"Schneider Electric StruxureWare Data Center Expert (DCE) v7.8.1 and prior: OS command injection over the network","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Schneider Electric StruxureWare Data Center Expert (DCE) v7.8.1 and prior","year":"2021","cvss_score":9.1,"severity":"critical","kev":false,"impact":"OS command injection over the network on the DCIM appliance - same blast radius as the path traversal above, reached by a different route. An attacker running commands on DCE inherits its trust relationship with every piece of power and cooling gear at the site.","attack_vector":"Remote, over the network to the DCE appliance.","remediation":"Upgrade to DCE v7.9.0 or later and rotate all stored device credentials. If the appliance was reachable from an untrusted network, rebuild it - DCE keeps polling credentials in a form an attacker with shell can read.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-22795"],"status":"curated","published":"2022-04-13"},{"id":"CVE-2021-26731","cve":"CVE-2021-26731","aliases":[],"title":"Lanner IAC-AST2500A BMC firmware: An authenticated BMC user escalates to root code execution on the controller","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lanner IAC-AST2500A BMC firmware","year":"2021","cvss_score":9.1,"severity":"critical","kev":false,"impact":"An authenticated BMC user escalates to root code execution on the controller. It is included alongside its unauthenticated siblings because it closes a different door: an operator who mitigates the unauthenticated bugs by putting the BMC behind a bastion still has this one live for anyone holding a valid BMC credential, including monitoring accounts and any credential shared across the fleet. The outcome is the same - out-of-band power, console, virtual media and firmware-level persistence. Command injection and stack buffer overflows in the modifyUserb_func handler of spx_restservice, reachable after authentication through the user-modification path.","attack_vector":"An authenticated attacker reaching the BMC REST service. Given how commonly BMC credentials are shared across a whitebox fleet, one leaked password is fleet-wide reach.","remediation":"Firmware flash from Lanner or the board integrator, with the same sourcing difficulty as the rest of this cluster. Config-only actions that help now: rotate BMC credentials to per-node unique values so a single leak is not fleet-wide, remove any BMC accounts issued to tenants or third parties, and keep the BMC REST service off any network segment reachable from tenant workloads. Note that this is one of at least nine published CVEs against this one BMC module's REST service, which is itself the signal - the codebase was not audited before shipping.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26731","https://www.nozominetworks.com/labs/vulnerability-advisories/cve-2021-26731/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-28506","cve":"CVE-2021-28506","aliases":[],"title":"Arista EOS (gNOI): gNOI APIs bypass authentication, allowing an unauthenticated factory reset of the switch","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (gNOI)","year":"2021","cvss_score":9.1,"severity":"critical","kev":false,"impact":"gNOI APIs bypass authentication, allowing an unauthenticated factory reset of the switch — instant fabric-wide outage primitive","attack_vector":"Network","remediation":"EOS upgrade; interim mitigation is a service ACL restricting gNOI, though CVE-2021-28507 shows those ACLs were themselves bypassable","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-28506"],"status":"curated","published":"2022-01-14"},{"id":"CVE-2021-47348","cve":"CVE-2021-47348","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Memory is handed to a consumer without being initialised","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2021","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Memory is handed to a consumer without being initialised or cleared in the amdgpu display core (DC/DM). Whatever the previous owner left behind is readable - and on a GPU node the previous owner is very often a different tenant's job. This is the classic residual-data leak between workloads sharing a card: model weights, activations, keys or tokens from the prior tenant can surface in a fresh allocation. Upstream fix: drm/amd/display: Avoid HDCP over-read and corruption","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-47348","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-21"},{"id":"CVE-2022-0670","cve":"CVE-2022-0670","aliases":[],"title":"Ceph Manager (volumes plugin): Owner of one CephFS share can read/write any share or the entire file system","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ceph Manager (volumes plugin)","year":"2022","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Owner of one CephFS share can read/write any share or the entire file system - cross-tenant","attack_vector":"Network (remote)","remediation":"Data-plane: ceph-mgr upgrade across the cluster; audit CephFS share access","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-0670"],"status":"curated","published":"2022-07-25"},{"id":"CVE-2022-0715","cve":"CVE-2022-0715","aliases":["TLStorm"],"title":"APC Smart-UPS SMT/SMC/SMX/SCL/SMTL series - firmware update signing: Firmware images are signed with a key that leaked","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"APC Smart-UPS SMT/SMC/SMX/SCL/SMTL series - firmware update signing","year":"2022","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Firmware images are signed with a key that leaked, so an attacker can flash arbitrary firmware onto the UPS and have it accepted as genuine. This is the worst of the three TLStorm bugs for an operator: the implant lives in the power device, survives every reimage of every server behind it, and is invisible to anything you run on the compute plane. An attacker who owns the UPS owns a kill switch for the racks it feeds, on a timer of their choosing.","attack_vector":"Network access to the UPS - either directly on the management network or via the intercepted cloud channel. Also reachable by anyone who can push a firmware image through the vendor's own update path.","remediation":"Flash to fixed firmware, which changes the signing scheme. Because the compromise persists in firmware, patching alone does not prove cleanliness on a unit you believe was targeted - the honest answer there is re-flash from a known-good image and verify the reported firmware version out of band. Treat UPS firmware as part of your supply chain: version-inventory it, and refuse units that cannot report a verifiable firmware version.","references":["https://www.se.com/ww/en/download/document/SEVD-2022-067-02/","https://www.armis.com/research/tlstorm/"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["physical-impact"],"published":"2022-03-09"},{"id":"CVE-2022-23131","cve":"CVE-2022-23131","aliases":[],"title":"Zabbix: Unverified user login in session data (SAML SSO enabled)","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Zabbix","year":"2022","cvss_score":9.1,"severity":"critical","kev":true,"impact":"Unverified user login in session data (SAML SSO enabled) -> unauthenticated admin takeover","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; rotate the Zabbix session secret","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-23131"],"status":"curated","published":"2022-01-13"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:N","cwe":["CWE-918"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2022-24856","cve":"CVE-2022-24856","aliases":["GHSA-www6-hf2v-v9m9"],"title":"FlyteConsole (cors_proxy endpoint): FlyteConsole's cors_proxy forwards attacker-chosen URLs, so anyone who reaches the","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"FlyteConsole (cors_proxy endpoint)","year":"2022","cvss_score":9.1,"severity":"critical","kev":false,"impact":"FlyteConsole's cors_proxy forwards attacker-chosen URLs, so anyone who reaches the console can make the server fetch internal endpoints on their behalf - most usefully the cloud instance metadata service. That yields the node's IAM role credentials, which on a GPU cluster generally unlock the object stores holding every tenant's datasets and checkpoints. Headers may be passed along to the attacker-chosen destination too.","attack_vector":"Any user who can reach FlyteConsole. Exploitable without authentication when the console is exposed to the internet.","remediation":"Upgrade FlyteConsole to 0.52.0 or later, which removes cors_proxy entirely, and redeploy. As an interim measure take FlyteConsole off the internet. If the console was ever publicly reachable, rotate the node/pod IAM credentials it could have surfaced.","references":["https://github.com/flyteorg/flyteconsole/security/advisories/GHSA-www6-hf2v-v9m9","https://nvd.nist.gov/vuln/detail/CVE-2022-24856"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2022-31247","cve":"CVE-2022-31247","aliases":[],"title":"Rancher: Anyone who can create role template bindings escalates privileges cluster-wide","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Rancher","year":"2022","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Anyone who can create role template bindings escalates privileges cluster-wide","attack_vector":"Cluster user with cluster-owner or project-owner rights","remediation":"Upgrade Rancher; audit RoleTemplateBindings","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31247"],"status":"curated","published":"2022-09-07"},{"id":"CVE-2023-23947","cve":"CVE-2023-23947","aliases":[],"title":"Argo CD: Improper authorization lets a user modify resources outside their permitted projects","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2023","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Improper authorization lets a user modify resources outside their permitted projects","attack_vector":"Any authenticated Argo CD user","remediation":"Rolling Argo CD upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-23947"],"status":"curated","published":"2023-02-16"},{"id":"CVE-2023-25132","cve":"CVE-2023-25132","aliases":["JVN#95119483"],"title":"CyberPower PowerPanel Business - default.cmd file upload: Unrestricted upload of a dangerous file type into default.cmd","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"CyberPower PowerPanel Business - default.cmd file upload","year":"2023","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Unrestricted upload of a dangerous file type into default.cmd - the script PowerPanel runs on a power event. Same nasty shape as the PowerChute issue: the attacker's code runs at the moment the UPS signals, across every host the software controls, with elevated privilege. The trigger is a power event, which an attacker with UPS access can also cause.","attack_vector":"An attacker who can write to the PowerPanel Business installation - reachable via the default-credential issue in the same advisory.","remediation":"Upgrade past v4.8.6. Independently: shutdown scripts invoked by power-management software are privileged code and should be under change control with restricted write permissions, on every host, regardless of vendor.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25132"],"status":"curated","published":"2023-04-24"},{"id":"CVE-2023-25725","cve":"CVE-2023-25725","aliases":[],"title":"HAProxy (before 2.7.3): HAProxy's HTTP/1 header parser accepts empty header field names, which can be used to make","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HAProxy (before 2.7.3)","year":"2023","cvss_score":9.1,"severity":"critical","kev":false,"impact":"HAProxy's HTTP/1 header parser accepts empty header field names, which can be used to make legitimate headers silently disappear after parsing. An attacker can use this to smuggle requests past access-control rules that were supposed to inspect those headers — bypassing ACLs meant to keep unauthorized traffic away from backend inference/storage services.","attack_vector":"Remote — an attacker sends a crafted HTTP/1 request with an empty header field name to a HAProxy instance doing header-based ACL enforcement.","remediation":"Software upgrade to HAProxy 2.7.3 or later, then reload/restart the process. HAProxy typically runs as a software component rather than an appliance, so this is a package upgrade + service restart across whichever hosts run it in front of the cluster; a graceful reload avoids dropping in-flight connections if your HAProxy version supports it.","references":["https://git.haproxy.org/?p=haproxy-2.7.git%3Ba=commit%3Bh=a0e561ad7f29ed50c473f5a9da664267b60d1112","https://www.haproxy.org/"],"status":"curated","fleet":{"pain_class":"daemon-restart"},"published":"2023-02-14"},{"id":"CVE-2023-28863","cve":"CVE-2023-28863","aliases":[],"title":"AMI MegaRAC SPx12/SPx13: Insufficient verification of data authenticity — firmware image signature can be subverted","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx12/SPx13","year":"2023","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Insufficient verification of data authenticity — firmware image signature can be subverted; enables persistent implant","attack_vector":"Local/network firmware update path","remediation":"BMC flash; also requires operational control so that only signed, vendor-verified images reach the update endpoint","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-28863"],"status":"curated","published":"2023-04-18"},{"id":"CVE-2023-31127","cve":"CVE-2023-31127","aliases":["libspdm session establishment bypass"],"title":"DMTF libspdm - SPDM session establishment (reference implementation used in GPU/device attestation): A device","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"DMTF libspdm - SPDM session establishment (reference implementation used in GPU/device attestation)","year":"2023","cvss_score":9.1,"severity":"critical","kev":false,"impact":"A device supporting both DHE and PSK sessions with mutual authentication can be driven into establishing a session with KEY_EXCHANGE plus PSK_FINISH, bypassing mutual authentication entirely. Scored 9.1 with a changed scope. This matters far beyond libspdm itself: SPDM is the protocol underneath device attestation and the encrypted host-to-device channel in confidential GPU computing and in PCIe/CXL IDE, and libspdm is the reference code many silicon and firmware vendors derived from. If your accelerator's attestation stack forked libspdm before 2.3.1, ask the vendor directly.","attack_vector":"An attacker on an adjacent path with low privileges - in the device context, something able to speak SPDM to the responder, which means the platform firmware, a DPU, or a compromised device on the link.","remediation":"Update to libspdm 2.3.1 or later. Cost for an operator is indirect and slow: you do not deploy libspdm, your GPU/DPU/NIC firmware vendor does, so the action is a supply-chain question to each vendor and then a firmware update on their schedule. Firmware flash means node drain. Track it as a vendor-management item rather than a patch you can apply.","references":["https://github.com/DMTF/libspdm/security/advisories/GHSA-qw76-4v8p-xq9f","https://nvd.nist.gov/vuln/detail/CVE-2023-31127"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2023-05-08"},{"id":"CVE-2023-3267","cve":"CVE-2023-3267","aliases":["ZDI-23-1149"],"title":"CyberPower PowerPanel Enterprise DCIM - remote backup location username field: OS command injection through","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"CyberPower PowerPanel Enterprise DCIM - remote backup location username field","year":"2023","cvss_score":9.1,"severity":"critical","kev":false,"impact":"OS command injection through the remote-backup username field, executing as NT AUTHORITY\\SYSTEM. Chained behind either authentication bypass above, this is unauthenticated SYSTEM on the host that manages your UPS estate.","attack_vector":"Authenticated user adding a remote backup location - trivially reachable given the two auth bypasses in the same advisory batch.","remediation":"Upgrade PowerPanel Enterprise. If the host was reachable from an untrusted network, rebuild it rather than patch - SYSTEM-level RCE on a Windows DCIM host means credential theft from the whole box.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-3267"],"status":"curated","published":"2023-08-14"},{"id":"CVE-2023-34329","cve":"CVE-2023-34329","aliases":[],"title":"AMI MegaRAC SPx12 (BMC&C): Auth bypass by spoofing the HTTP header","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx12 (BMC&C)","year":"2023","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Auth bypass by spoofing the HTTP header; combined with CVE-2023-34330 gives unauthenticated RCE as root on the BMC. Below-OS persistence","attack_vector":"Network, BMC web/Redfish","remediation":"BMC firmware update per node via ODM build; combined-chain risk means treat as critical even where the BMC is on a management VLAN","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-34329"],"status":"curated","fleet":{"ubiquity":"universal - same MegaRAC OEM footprint","remediation_pain":"firmware-flash - full BMC image update per node, out-of-band, cannot run while a tenant job holds the host","pain_class":"firmware-flash","why_fleet_wide":"Chaining the HTTP-header auth spoof with the code-injection primitive yields pre-OS code execution on the BMC; identical firmware across a homogeneous GPU fleet means one exploit works on every node."},"published":"2023-07-18"},{"id":"CVE-2023-3961","cve":"CVE-2023-3961","aliases":[],"title":"Samba: Path traversal in client pipe names","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Samba","year":"2023","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Path traversal in client pipe names -> connect to Unix sockets outside the private directory","attack_vector":"Network (remote)","remediation":"Data-plane: smbd upgrade on file-serving nodes","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-3961"],"status":"curated","published":"2023-11-03"},{"id":"CVE-2023-48023","cve":"CVE-2023-48023","aliases":[],"title":"Ray (`/log_proxy`): SSRF from the dashboard","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ray (`/log_proxy`)","year":"2023","cvss_score":9.1,"severity":"critical","kev":false,"impact":"SSRF from the dashboard","attack_vector":"Unauthenticated network to the dashboard","remediation":"No vendor patch (same disputed position); isolate","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-48023"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"published":"2023-11-28"},{"id":"CVE-2023-51786","cve":"CVE-2023-51786","aliases":[],"title":"Lustre (incorrect access control, 2.13.x-2.15.x before 2.15.4): Incorrect access control in Lustre lets an attacker","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Lustre (incorrect access control, 2.13.x-2.15.x before 2.15.4)","year":"2023","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Incorrect access control in Lustre lets an attacker escalate privileges and obtain sensitive information. This is the modern-release equivalent of the 2019 family and it lands on the versions actually deployed in current AI-training clusters — 2.15.x is the long-term-support line most sites run today. On a shared training filesystem, 'obtain sensitive information' means another tenant's datasets, checkpoints and model weights.","attack_vector":"An attacker with Lustre client access on versions 2.13.x, 2.14.x, or 2.15.x before 2.15.4.","remediation":"Upgrade to Lustre 2.15.4 or later — client and server packages, with a coordinated restart of the storage cluster. Enable Lustre nodemap with `admin`/`trusted` set to off for tenant clients and squash root, so a compromised client cannot act as root against the filesystem; that is a config change you can make ahead of the upgrade and it is the durable control.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-51786"],"status":"curated","tags":["tenant-isolation"],"published":"2024-03-07"},{"id":"CVE-2024-0003","cve":"CVE-2024-0003","aliases":[],"title":"Pure Storage FlashArray Purity (remote administrative account creation): An attacker uses a remote administrative","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Pure Storage FlashArray Purity (remote administrative account creation)","year":"2024","cvss_score":9.1,"severity":"critical","kev":false,"impact":"An attacker uses a remote administrative service to create a privileged account on the array - persistent backdoor access to the storage system.","attack_vector":"Remote network access with high privilege to the administrative service.","remediation":"Apply the Purity update, then enumerate array accounts and remove any you did not create. Account review matters more than the version bump here.","references":["https://purestorage.com/security"],"status":"curated"},{"id":"CVE-2024-0004","cve":"CVE-2024-0004","aliases":[],"title":"Pure Storage FlashArray Purity (array admin command execution): A user holding the array admin role executes arbitrary","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Pure Storage FlashArray Purity (array admin command execution)","year":"2024","cvss_score":9.1,"severity":"critical","kev":false,"impact":"A user holding the array admin role executes arbitrary commands remotely and escalates beyond the intended management boundary onto the array's underlying OS.","attack_vector":"Authenticated array admin.","remediation":"Apply the Purity update from Pure's security page.","references":["https://purestorage.com/security"],"status":"curated"},{"id":"CVE-2024-0005","cve":"CVE-2024-0005","aliases":[],"title":"Pure Storage FlashArray / FlashBlade Purity (SNMP configuration command injection): A crafted SNMP configuration yields","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Pure Storage FlashArray / FlashBlade Purity (SNMP configuration command injection)","year":"2024","cvss_score":9.1,"severity":"critical","kev":false,"impact":"A crafted SNMP configuration yields arbitrary remote command execution on the array. Affects FlashBlade too, which is the platform commonly used for AI training data lakes.","attack_vector":"Authenticated high-privilege user able to set SNMP configuration.","remediation":"Apply the Purity update, and audit existing SNMP configuration on arrays for injected content - the payload persists in configuration across the upgrade.","references":["https://purestorage.com/security"],"status":"curated"},{"id":"CVE-2024-12378","cve":"CVE-2024-12378","aliases":["Arista Security Advisory 0116"],"title":"Arista EOS (secure VXLAN / Tunnelsec agent): After the Tunnelsec agent restarts, traffic that should be encrypted","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (secure VXLAN / Tunnelsec agent)","year":"2024","cvss_score":9.1,"severity":"critical","kev":false,"impact":"After the Tunnelsec agent restarts, traffic that should be encrypted inside secure VXLAN tunnels goes out in cleartext. You configured tunnel encryption between sites or between pods, the agent bounced, and now every tenant's overlay traffic is readable by anyone with a tap on the underlay — with no alarm, because the tunnels still look up. This is the worst kind of fabric bug: the control plane reports healthy while the confidentiality guarantee is silently gone.","attack_vector":"Passive attacker with access to the underlay path the VXLAN tunnels traverse — a dark-fibre tap, a transit provider, a compromised intermediate switch, or another tenant with underlay visibility. Requires the Tunnelsec agent to have restarted, which happens on upgrade, crash, or config change.","remediation":"EOS upgrade plus a switch reload on every VTEP running secure VXLAN. Until then, monitor Tunnelsec agent restarts and treat any restart as a confidentiality incident for traffic since that moment. There is no config workaround that keeps encryption on across a restart.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-12378"],"status":"curated","tags":["tenant-isolation"],"published":"2025-05-08"},{"id":"CVE-2024-21887","cve":"CVE-2024-21887","aliases":[],"title":"Ivanti Connect Secure: Command injection in web components","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ivanti Connect Secure","year":"2024","cvss_score":9.1,"severity":"critical","kev":true,"impact":"Command injection in web components; chained with CVE-2023-46805 for unauthenticated root RCE","attack_vector":"Network (remote)","remediation":"Control-plane: rebuild from a factory image; rotate every VPN secret and certificate","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21887"],"status":"curated","published":"2024-01-12"},{"id":"CVE-2024-22120","cve":"CVE-2024-22120","aliases":[],"title":"Zabbix: Unsanitized clientip in the audit log","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Zabbix","year":"2024","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Unsanitized clientip in the audit log -> time-based blind SQL injection around script execution","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; rotate stored host and IPMI credentials","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-22120"],"status":"curated","published":"2024-05-17"},{"id":"CVE-2024-32752","cve":"CVE-2024-32752","aliases":["CVE-2017-17704","ICSA-24-158-04"],"title":"Software House iSTAR door controllers (firmware before 6.6.B) and the IP-ACM Ethernet Door Module link: The iSTAR","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Software House iSTAR door controllers (firmware before 6.6.B) and the IP-ACM Ethernet Door Module link","year":"2024","cvss_score":9.1,"severity":"critical","kev":false,"impact":"The iSTAR controllers do not support authenticated communications with the iSTAR Configuration Utility, and the older IP-ACM door-module link used a fixed AES key and a fixed IV restarted on every message - which leaks enough structure that door-unlock commands can be replayed or forged outright. Both bugs land in the same place: the wire between the controller and its door hardware or configuration tool is not trustworthy, so an attacker on that network can open doors without ever authenticating to anything. Standing in the hall or the cage, that person reaches drives, console ports, the out-of-band management switch and the physical console of machines running customer workloads. For a multi-tenant GPU operator this collapses the cage boundary that the whole bare-metal isolation story rests on, and it does so without generating a failed-authentication event anywhere, because there is no authentication to fail.","attack_vector":"A host on the physical-security network that can see traffic between the iSTAR controller and the IP-ACM modules or the configuration utility. That means passive capture and replay, or active injection, from the security VLAN - typically shared with CCTV and the integrator's tooling. Physical access to structured cabling in a back-of-house space is an equally valid vector and is not covered by most tenants' security policies because the cabling is the landlord's.","remediation":"Update iSTAR controller firmware to 6.6.B or later, which introduces authenticated ICU communications, and retire the affected IP-ACM generation in favour of the IP-ACM v2 hardware that supports proper encryption. That is firmware plus a hardware refresh of the door modules - a capital project with a security integrator and per-door downtime, not a patch. Until it is done: physically secure every cable run and panel between the controller and the door modules, put the physical-security VLAN behind a firewall with an allow-list, and enable 802.1X or port security on the switch ports serving door hardware so a rogue device cannot join the segment. Treat any hall where door-module cabling runs through space you do not control as a hall with a broken cage boundary.","references":["https://www.cisa.gov/news-events/ics-advisories/icsa-24-158-04","https://nvd.nist.gov/vuln/detail/CVE-2024-32752","https://nvd.nist.gov/vuln/detail/CVE-2017-17704","https://systemoverlord.com/2017/12/18/cve-2017-17704-broken-cryptography-in-istar-ultra-ip-acm-by-software-house.html"],"status":"curated"},{"id":"CVE-2024-3411","cve":"CVE-2024-3411","aliases":["VU#163057"],"title":"The IPMI 2.0 authenticated-session mechanism as specified and as implemented across multiple vendors: An attacker","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"The IPMI 2.0 authenticated-session mechanism as specified and as implemented across multiple vendors","year":"2024","cvss_score":9.1,"severity":"critical","kev":false,"impact":"An attacker hijacks an established IPMI session by spoofing packets with a predicted session ID, bypassing authentication entirely. What they inherit is the privilege of whoever's session they stole - typically an operator or automation account with power control, boot-device selection, sensor access and often virtual media. This is a spec-level weakness rather than one vendor's bug, so it is one of the few entries here that an operator should assume applies to their whole heterogeneous fleet, not just the boards from one manufacturer. Session IDs are predictable and BMC random number generation is weak, so the values that are supposed to make a session unforgeable are guessable. CERT/CC tracks it as VU#163057 and the affected list spans vendor BMC implementations built to the Intel IPMI specification.","attack_vector":"Network reachability to the BMC's IPMI-over-LAN port (UDP 623) with the ability to observe or infer session state. Unauthenticated in effect, since the whole point is that authentication is bypassed. Anything on the out-of-band management VLAN qualifies, as does anything that can reach a BMC exposed by a routing mistake.","remediation":"Firmware updates exist from individual vendors, but there is no single fix because the weakness is in the protocol's assumptions - so the durable remediation is to stop using IPMI-over-LAN. Disable IPMI-over-LAN on the BMC and drive management through Redfish over TLS instead; that is a config-only change on most modern BMCs and is the single highest-value action in this entire database for an operator who still has UDP 623 open. Where legacy tooling forces IPMI, restrict UDP 623 to an explicit allowlist of management hosts at the switch, and treat the management VLAN as a network where session hijacking is assumed possible.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-3411","https://kb.cert.org/vuls/id/163057"],"status":"curated"},{"id":"CVE-2024-37287","cve":"CVE-2024-37287","aliases":[],"title":"Kibana: Prototype pollution via ML/Alerting connectors + write access to internal ML indices","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Kibana","year":"2024","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Prototype pollution via ML/Alerting connectors + write access to internal ML indices -> arbitrary code execution","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade the observability UI tier; restrict ML roles","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-37287"],"status":"curated","published":"2024-08-13"},{"id":"CVE-2024-3829","cve":"CVE-2024-3829","aliases":[],"title":"Qdrant (snapshot recovery): Arbitrary file read and write during snapshot recovery","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Qdrant (snapshot recovery)","year":"2024","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Arbitrary file read and write during snapshot recovery","attack_vector":"Customer-supplied snapshot file","remediation":"Upgrade; snapshots must be treated as untrusted archives","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-3829"],"status":"curated","published":"2024-06-03"},{"id":"CVE-2024-45763","cve":"CVE-2024-45763","aliases":["DSA-2024-449"],"title":"Dell Enterprise SONiC (OS command injection): OS command injection giving arbitrary command execution on the switch's","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell Enterprise SONiC (OS command injection)","year":"2024","cvss_score":9.1,"severity":"critical","kev":false,"impact":"OS command injection giving arbitrary command execution on the switch's underlying Linux. Chained behind the authentication-bypass in the same advisory batch, an unauthenticated attacker goes from the network to root on the switch in two steps. On SONiC the 'switch' is a fairly normal Linux box with the ASIC SDK attached, so root there means arbitrary forwarding-table manipulation for every tenant on the device.","attack_vector":"Remote attacker holding high-privilege access — but see CVE-2024-45764, which supplies that access without credentials.","remediation":"Same fix as the rest of DSA-2024-449: NOS image upgrade and reboot on every Dell Enterprise SONiC switch running 4.1.x or 4.2.x. Assume any device that was reachable pre-patch may hold persistence and consider a clean re-image rather than an in-place upgrade.","references":["https://www.dell.com/support/kbdoc/en-us/000245655/dsa-2024-449-security-update-for-dell-enterprise-sonic-distribution-vulnerabilities","https://nvd.nist.gov/vuln/detail/CVE-2024-45763"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-11-08"},{"id":"CVE-2024-45765","cve":"CVE-2024-45765","aliases":["DSA-2024-449"],"title":"Dell Enterprise SONiC (privilege boundary in CLI): High-privilege OS commands can be run by users holding less","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell Enterprise SONiC (privilege boundary in CLI)","year":"2024","cvss_score":9.1,"severity":"critical","kev":false,"impact":"High-privilege OS commands can be run by users holding less privileged roles. Whatever role separation you built for your NOC — read-only, operator, admin — does not hold on the switch. Relevant to any operator who gives tenants or contractors scoped switch access.","attack_vector":"Authenticated user with a low-privilege SONiC role.","remediation":"NOS image upgrade and reboot per switch. Until then, do not issue low-privilege SONiC accounts to anyone you would not give admin, because the distinction is not enforced.","references":["https://www.dell.com/support/kbdoc/en-us/000245655/dsa-2024-449-security-update-for-dell-enterprise-sonic-distribution-vulnerabilities","https://nvd.nist.gov/vuln/detail/CVE-2024-45765"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-11-08"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:H","cwe":["CWE-1284","CWE-125"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-49571","cve":"CVE-2024-49571","aliases":[],"title":"Linux kernel SMC-R/SMC-D (CLC proposal parsing, iparea_offset / ipv6_prefixes_cnt): Third instance of the same class in","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel SMC-R/SMC-D (CLC proposal parsing, iparea_offset / ipv6_prefixes_cnt)","year":"2024","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Third instance of the same class in the same handshake parser - the IP-area offset and the IPv6 prefix count are taken from the remote client without bounds checks, giving an unauthenticated peer an out-of-bounds read on the server. That three separate patches were needed for one message format is the real finding: the SMC CLC parser was written assuming a cooperative peer, and an operator should treat the whole surface as untrusted rather than patching field by field.","attack_vector":"Remote, unauthenticated, first message of the SMC handshake.","remediation":"Kernel update validating iparea_offset and ipv6_prefixes_cnt. Given the pattern, the durable control is not exposing AF_SMC on tenant-reachable interfaces at all unless SMC acceleration is a deliberate design choice.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=47ce46349672a7e0c361bfe39ed0b22e824ef4fb","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2024/CVE-2024-49571.json"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:H","cwe":["CWE-190"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-53146","cve":"CVE-2024-53146","aliases":[],"title":"Linux NFS server (nfsd, NFSv4 COMPOUND tag decode): An NFSv4 COMPOUND tag length near U32_MAX overflows the length+4","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux NFS server (nfsd, NFSv4 COMPOUND tag decode)","year":"2024","cvss_score":9.1,"severity":"critical","kev":false,"impact":"An NFSv4 COMPOUND tag length near U32_MAX overflows the length+4 arithmetic, producing a wrong allocation and out-of-bounds access on the server. A single crafted COMPOUND from an unauthenticated client crashes the shared file server.","attack_vector":"Any host that can send an NFSv4 COMPOUND to the server's RPC port.","remediation":"Update the storage server kernel and reboot. No config workaround - COMPOUND is the base NFSv4 transport unit.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53146","https://git.kernel.org/pub/scm/linux/security/vulns.git"],"status":"curated"},{"id":"CVE-2024-7776","cve":"CVE-2024-7776","aliases":[],"title":"ONNX (`download_model`): Arbitrary file overwrite","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"ONNX (`download_model`)","year":"2024","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Arbitrary file overwrite","attack_vector":"Customer-supplied model reference","remediation":"Upgrade past 1.16.1","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-7776"],"status":"curated","published":"2025-03-20"},{"id":"CVE-2024-9487","cve":"CVE-2024-9487","aliases":[],"title":"GitHub Enterprise Server: Improper signature verification","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"GitHub Enterprise Server","year":"2024","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Improper signature verification -> SAML SSO bypass, unauthorized user provisioning and instance access","attack_vector":"Network (remote)","remediation":"Control-plane: GHES upgrade; audit newly created users","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-9487"],"status":"curated","published":"2024-10-10"},{"id":"CVE-2025-0108","cve":"CVE-2025-0108","aliases":[],"title":"Palo Alto PAN-OS: Management web interface auth bypass invoking PHP scripts","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Palo Alto PAN-OS","year":"2025","cvss_score":9.1,"severity":"critical","kev":true,"impact":"Management web interface auth bypass invoking PHP scripts; actively exploited","attack_vector":"Network (remote)","remediation":"Control-plane: patch; management plane on an out-of-band network only","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0108"],"status":"curated","published":"2025-02-12"},{"id":"CVE-2025-10263","cve":"CVE-2025-10263","aliases":["TFV-17","AMP-SB-0008","XSA-493","TLBI+DSB completes too early"],"title":"Arm Neoverse N1 / N2 / V1 / V2 / V3 / V3AE, Cortex-A76/A77/A78/A710, Cortex-X1-X925, C1-Ultra/Premium","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Arm Neoverse N1 / N2 / V1 / V2 / V3 / V3AE, Cortex-A76/A77/A78/A710, Cortex-X1-X925, C1-Ultra/Premium; Trusted…","year":"2025","cvss_score":9.1,"severity":"critical","kev":false,"impact":"A store issued on one core can land after another core has already invalidated the translation and completed its TLBI+DSB sequence. The write then commits through a stale translation into memory owned by a higher exception level - a guest writing into hypervisor memory, or a hypervisor writing into EL3/secure memory. In practice that is a write primitive into page tables and other memory-management structures you assumed were fenced off. Neoverse N1 is Ampere Altra and Graviton2; V1/V2 are Graviton3/Graviton4 and Grace. This is close to a worst case for a shared Arm host.","attack_vector":"Code running in a guest or in the host kernel on an affected Arm core, racing another core's TLB maintenance. Purely local, no device or network access needed, but it is a race, so exploitation needs control of scheduling on at least two cores - trivially satisfied by any tenant with more than one vCPU.","remediation":"Requires coordinated firmware and OS updates, and you need both halves. TF-A platforms must build with WORKAROUND_CVE_2025_10263 enabled, which makes the affected TLBI+DSB sequences execute twice; the hypervisor and guest kernels need the matching arm64 errata workaround (Xen shipped XSA-493, Red Hat shipped a long errata chain). Flash + reboot + job drain on every node, and the OEM has to ship the BL31 build first. Ampere published AMP-SB-0008 for Altra and Altra Max. Expect a measurable cost on TLB-shootdown-heavy workloads because the barrier sequence is now doubled.","references":["https://developer.arm.com/documentation/112137","https://trustedfirmware-a.readthedocs.io/en/latest/security_advisories/security-advisory-tfv-17.html","http://xenbits.xen.org/xsa/advisory-493.html","https://amperecomputing.com/products/product-security"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-06-09"},{"id":"CVE-2025-1260","cve":"CVE-2025-1260","aliases":["Arista Security Advisory 21098"],"title":"Arista EOS (OpenConfig gNOI authorization): The gNOI equivalent of the gNMI authorization bypass: operations","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (OpenConfig gNOI authorization)","year":"2025","cvss_score":9.1,"severity":"critical","kev":false,"impact":"The gNOI equivalent of the gNMI authorization bypass: operations that should have been rejected run anyway. gNOI covers reboot, image install, certificate rotation and factory reset — so an under-privileged caller can reload switches or push images, not merely edit config.","attack_vector":"A client reaching the gNOI endpoint on a switch with OpenConfig configured, holding credentials that should not authorize the operation.","remediation":"EOS upgrade plus reload. Immediately restrict gNOI endpoint reachability by ACL and re-issue any certificates that could have been rotated by an unauthorized caller.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-1260"],"status":"curated","published":"2025-03-04"},{"id":"CVE-2025-15031","cve":"CVE-2025-15031","aliases":[],"title":"MLflow (pyfunc tar extraction): Arbitrary file write from crafted tar entries","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"MLflow (pyfunc tar extraction)","year":"2025","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Arbitrary file write from crafted tar entries","attack_vector":"Customer-supplied model archive","remediation":"Upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-15031"],"status":"curated","published":"2026-03-18"},{"id":"CVE-2025-23317","cve":"CVE-2025-23317","aliases":[],"title":"NVIDIA Triton (HTTP server): Attacker can start a reverse shell from the HTTP server","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"NVIDIA Triton (HTTP server)","year":"2025","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Attacker can start a reverse shell from the HTTP server","attack_vector":"Unauthenticated network to the HTTP port","remediation":"Patch. Chains with CVE-2025-23319/23320 into full takeover of the inference host","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23317"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:H/A:H","cwe":["CWE-122"],"published":"2025-08-06"},{"id":"CVE-2025-45378","cve":"CVE-2025-45378","aliases":[],"title":"Dell CloudLink (restricted shell breakout): A privileged user breaks out of the restricted shell into a full command","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Dell CloudLink (restricted shell breakout)","year":"2025","cvss_score":9.1,"severity":"critical","kev":false,"impact":"A privileged user breaks out of the restricted shell into a full command shell on the CloudLink server and escalates. Reachable over the network when SSH is enabled with web credentials.","attack_vector":"SSH to the CloudLink appliance with a known privileged password.","remediation":"Upgrade CloudLink past 8.1.2 (fixed in 8.2). Disable SSH on the appliance if you do not operationally need it - that removes the network path entirely.","references":["https://www.dell.com/support/kbdoc/en-us/000384363/dsa-2025-374-security-update-for-dell-cloudlink-multiple-security-vulnerabilities"],"status":"curated"},{"id":"CVE-2025-46364","cve":"CVE-2025-46364","aliases":[],"title":"Dell CloudLink (CLI escape): A privileged user with a known password escapes the CLI and takes control of the CloudLink","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Dell CloudLink (CLI escape)","year":"2025","cvss_score":9.1,"severity":"critical","kev":false,"impact":"A privileged user with a known password escapes the CLI and takes control of the CloudLink server. CloudLink is the key manager for encrypted volumes - control of it is control of the data-at-rest keys for everything it protects.","attack_vector":"Network access plus a known privileged credential.","remediation":"Upgrade CloudLink to 8.1.1 or later. Appliance upgrade with a service window. Because this is a KMS, also plan key rotation if you cannot rule out prior access - patching does not undo a key compromise.","references":["https://www.dell.com/support/kbdoc/en-us/000384363/dsa-2025-374-security-update-for-dell-cloudlink-multiple-security-vulnerabilities"],"status":"curated"},{"id":"CVE-2025-54576","cve":"CVE-2025-54576","aliases":[],"title":"OAuth2-Proxy: skip_auth_routes route matching flaw","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"OAuth2-Proxy","year":"2025","cvss_score":9.1,"severity":"critical","kev":false,"impact":"skip_auth_routes route matching flaw -> authentication bypass on protected paths","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade the ingress auth sidecar; re-audit every skip rule","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-54576"],"status":"curated","published":"2025-07-30"},{"id":"CVE-2025-6000","cve":"CVE-2025-6000","aliases":[],"title":"HashiCorp Vault: Root-namespace operator with write on sys/audit gains code execution on the Vault host","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HashiCorp Vault","year":"2025","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Root-namespace operator with write on sys/audit gains code execution on the Vault host via the plugin directory","attack_vector":"Network (remote)","remediation":"Control-plane: URGENT - Vault holds tenant and BMC creds; upgrade + rotate root token","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-6000"],"status":"curated","published":"2025-08-01"},{"id":"CVE-2025-67039","cve":"CVE-2025-67039","aliases":["ICSA-26-069-02"],"title":"Lantronix EDS3000PS serial-to-Ethernet device server: Full bypass of the management-page login","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Lantronix EDS3000PS serial-to-Ethernet device server","year":"2025","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Full bypass of the management-page login. Appending a specific suffix to a management URL, combined with a crafted Authorization header, gets an attacker straight into admin functionality with no valid credentials — including whatever serial console sessions the device is bridging.","attack_vector":"Purely network-based, no credentials required — the attacker just needs to reach the device's web management port and knows the URL/header trick published in the advisory.","remediation":"Firmware flash to the fixed release; this is an auth-check logic bug, not something you can compensate for with a password change. Roll out per device; each flash briefly interrupts the serial bridging that device provides.","references":["https://www.cisa.gov/news-events/ics-advisories/icsa-26-069-02"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2026-03-11"},{"id":"CVE-2026-0257","cve":"CVE-2026-0257","aliases":[],"title":"Palo Alto PAN-OS: GlobalProtect portal/gateway auth bypass","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Palo Alto PAN-OS","year":"2026","cvss_score":9.1,"severity":"critical","kev":true,"impact":"GlobalProtect portal/gateway auth bypass -> establish an unauthorized VPN connection","attack_vector":"Network (remote)","remediation":"Control-plane: patch; review VPN session and tunnel logs for rogue connections","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-0257"],"status":"curated","published":"2026-05-13"},{"id":"CVE-2026-14890","cve":"CVE-2026-14890","aliases":[],"title":"SGLang (expert-parallel backup ZMQ PULL): Unauthenticated, unvalidated deserialization on a routable interface","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"SGLang (expert-parallel backup ZMQ PULL)","year":"2026","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Unauthenticated, unvalidated deserialization on a routable interface","attack_vector":"Co-tenant on the cluster fabric","remediation":"Upgrade; the EP backup plane binds routable by default","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-14890"],"status":"curated","published":"2026-07-16"},{"id":"CVE-2026-25199","cve":"CVE-2026-25199","aliases":[],"title":"Apache CloudStack Proxmox extension (cross-tenant instance access): The extension keys CloudStack instances to Proxmox","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Apache CloudStack Proxmox extension (cross-tenant instance access)","year":"2026","cvss_score":9.1,"severity":"critical","kev":false,"impact":"The extension keys CloudStack instances to Proxmox VMs using a user-editable setting (proxmox_vmid), so a tenant edits that field and gains access to another tenant's instance. A straightforward cross-tenant break in a multi-tenant IaaS control plane.","attack_vector":"Any authenticated CloudStack tenant able to set instance settings.","remediation":"Upgrade Apache CloudStack past 4.22.0.0 or disable the Proxmox extension. Until patched, restrict who can edit instance settings - the vulnerable field is user-writable by design.","references":["https://lists.apache.org/thread/n8mt5b7wkpysstb8w7rr9f02kc5cq2xm"],"status":"curated"},{"id":"CVE-2026-35030","cve":"CVE-2026-35030","aliases":[],"title":"LiteLLM (JWT auth): Auth bypass when `enable_jwt_auth` is set","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LiteLLM (JWT auth)","year":"2026","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Auth bypass when `enable_jwt_auth` is set","attack_vector":"Unauthenticated network","remediation":"Upgrade to 1.83.0+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-35030"],"status":"curated","published":"2026-04-06"},{"id":"CVE-2026-41475","cve":"CVE-2026-41475","aliases":["CVE-2023-51773","CVE-2026-26264","CVE-2025-66624","CVE-2026-41502","CVE-2026-41503","CVE-2026-21878","CVE-2018-10238"],"title":"BACnet Stack open-source C library (bacnet-stack) embedded in third-party controllers and gateways: A run","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"BACnet Stack open-source C library (bacnet-stack) embedded in third-party controllers and gateways","year":"2026","cvss_score":9.1,"severity":"critical","kev":false,"impact":"A run of out-of-bounds reads and length underflows in the decoders for WritePropertyMultiple, ReadPropertyMultiple, WriteProperty and NPDU handling, all reachable by an unauthenticated attacker sending a truncated or malformed request. The reason this matters more than the individual CVSS scores suggest is supply chain: bacnet-stack is the reference C implementation that a long tail of controller, gateway and sensor vendors embed in their firmware without ever telling the customer. Your CRAH controller, your BACnet router, your rear-door heat exchanger's BACnet interface and your environmental gateway may all be running the same library, and none of them appear in a search for 'bacnet-stack'. One crafted packet can therefore crash a heterogeneous set of devices simultaneously - a fleet-wide loss of thermal control in a hall where 40-140 kW racks have minutes of margin. Earlier issues in the same library reached memory corruption rather than just reads, so treat the class as potentially more than DoS on older embedded builds.","attack_vector":"Unauthenticated BACnet/IP on the facility network. No credentials, no interaction, and in several of these cases the trigger is a single truncated request. Broadcast-reachable services widen this further. The practical problem is that you cannot enumerate affected devices from the outside - you need each vendor to disclose whether they embed the library and at what version, and most will not answer quickly.","remediation":"You cannot patch this yourself in the general case. The library fix is upstream (bacnet-stack 1.4.3 / 1.5.0 and later), but the code is compiled into vendor firmware, so remediation means each device vendor rebuilding and shipping firmware, then a per-device flash by the controls contractor. Expect that most embedded devices in your hall will never get a fixed build. That makes segmentation the actual answer: BACnet on an isolated VLAN, no untrusted hosts on it, no BBMD bridging to anything else. In parallel, use this as a procurement lever - make an SBOM for the BACnet stack a requirement in new controller purchases, because right now operators have no way to answer 'am I affected' and that is the real finding.","references":["https://github.com/bacnet-stack/bacnet-stack/security/advisories/GHSA-cvv4-v3g6-4jmv","https://nvd.nist.gov/vuln/detail/CVE-2026-41475","https://nvd.nist.gov/vuln/detail/CVE-2023-51773","https://github.com/bacnet-stack/bacnet-stack/security/advisories/GHSA-phjh-v45p-gmjj"],"status":"curated"},{"id":"CVE-2026-46043","cve":"CVE-2026-46043","aliases":["RDMA/rxe payload_size underflow","Soft-RoCE BTH pad validation"],"title":"Linux kernel - RDMA/rxe (Soft-RoCE) receive path, drivers/infiniband/sw/rxe/rxe_recv.c: Rxe_rcv() checked only that an","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel - RDMA/rxe (Soft-RoCE) receive path, drivers/infiniband/sw/rxe/rxe_recv.c","year":"2026","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Rxe_rcv() checked only that an inbound packet was at least header_size bytes, but payload_size() then subtracts the attacker-controlled BTH pad field and the ICRC size from the packet length. A short packet, or one carrying a forged non-zero pad, makes that subtraction underflow and hands a bogus length to everything downstream in the receive path. Since Soft-RoCE rides UDP/4791, a single unauthenticated datagram from anywhere that can reach the node crashes or corrupts the kernel - no connection, no handshake, no credentials. On a shared fabric one packet from one tenant takes down a GPU node and every job on it.","attack_vector":"Send a crafted UDP datagram to port 4791 on any host with rdma_rxe loaded. The BTH pad field is set by the attacker, so even a packet long enough to pass the header check can drive the payload length negative. Entirely pre-authentication and reachable from any source the network permits, including across routed segments if 4791 is not filtered.","remediation":"Host reboot / kernel upgrade. Faster and cheaper: unload and blacklist rdma_rxe on every node that is not deliberately running Soft-RoCE (config change, zero downtime) - this closes the entire rxe family of remote packet bugs. If rxe is genuinely required, firewall UDP/4791 to known peers immediately as a stopgap, then upgrade the kernel on a rolling drain. Note the follow-on CVE-2026-46133 shows the first attempt at this fix was incomplete, so verify you are on a kernel carrying both.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2026/CVE-2026-46043.json","https://nvd.nist.gov/vuln/detail/CVE-2026-46043"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["fabric-dos"],"published":"2026-05-27"},{"id":"CVE-2026-53186","cve":"CVE-2026-53186","aliases":["RDMA/srp SRP_RSP sense buffer overrun","SCSI RDMA Protocol initiator"],"title":"Linux kernel - SRP (SCSI RDMA Protocol) initiator, drivers/infiniband/ulp/srp/ib_srp.c: The SRP initiator copied the","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel - SRP (SCSI RDMA Protocol) initiator, drivers/infiniband/ulp/srp/ib_srp.c","year":"2026","cvss_score":9.1,"severity":"critical","kev":false,"impact":"The SRP initiator copied the sense data out of an SRP_RSP without bounding the copy by the length actually received, so a malicious or compromised SRP target can overrun the initiator's buffer and read host kernel memory or crash the node. This inverts the usual threat direction - here the storage array attacks its clients. In a GPU cluster where an array or a software SRP target serves many compute nodes, one compromised target reaches every node that mounts from it, which is a fleet-wide blast radius from a single storage compromise.","attack_vector":"The attacker controls or has compromised an SRP target the victim connects to, and returns an SRP_RSP whose declared sense length exceeds what was received. No credentials on the victim are needed beyond it being a normal client of the target. Also reachable by an attacker who can inject into the SRP connection using the RDMA packet-injection primitives.","remediation":"Host reboot / kernel upgrade on all SRP initiator nodes. Interim: verify SRP targets are on a management-isolated storage fabric with per-tenant partitioning so an untrusted party cannot stand up a rogue target and attract connections; that is a switch/SM config change. If SRP is legacy in your environment and NVMe-oF has replaced it, unload ib_srp and remove the initiator configuration entirely - a config change that eliminates the surface.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2026/CVE-2026-53186.json","https://nvd.nist.gov/vuln/detail/CVE-2026-53186"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-06-25"},{"id":"CVE-2026-64269","cve":"CVE-2026-64269","aliases":["RDMA/rtrs-srv chunk overrun","RTRS server rdma_write_sg unbounded length"],"title":"Linux kernel - RDMA/rtrs server (RDMA Transport, used by RNBD block storage), drivers/infiniband/ulp/rtrs/rtrs-srv.c","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel - RDMA/rtrs server (RDMA Transport, used by RNBD block storage), drivers/infiniband/ulp/rtrs/rtrs-srv.c","year":"2026","cvss_score":9.1,"severity":"critical","kev":false,"impact":"When the RTRS server answers a READ it builds the RDMA WRITE source scatter/gather entry with a length taken straight off the wire descriptor, which the remote peer filled itself, and the source lkey used is the PD-wide local_dma_lkey rather than a key bound to that chunk's mapping - so the verbs layer never constrains the transfer to the chunk size. A peer advertising a length larger than max_chunk_size makes the NIC read past the chunk's mapped region and ship the result back over the fabric. With no IOMMU or in passthrough mode - a common configuration on high-performance storage nodes chasing latency - that returns adjacent host memory to the attacker. This is exactly the disaggregated-storage tenant-isolation break operators worry about: one client reads memory belonging to the server and, by extension, to other clients.","attack_vector":"The attacker is an RTRS client connected to the server - the normal trust position for a storage tenant. They set desc[0].len in the read descriptor larger than the negotiated chunk size; before the fix only a zero length was rejected. With a translating IOMMU the over-range access faults and drops the connection (denial of service instead of disclosure), so IOMMU configuration decides whether this reads memory or just breaks the session.","remediation":"Host reboot / kernel upgrade on RTRS/RNBD server nodes. Interim config changes that materially reduce impact: enable the IOMMU in translating (not passthrough) mode on storage servers, which converts disclosure into a connection abort - this is a kernel command-line change requiring a reboot anyway, so fold it into the same maintenance window and expect a small latency cost. Restrict RTRS server ports to authenticated client subnets. If RNBD/RTRS is not in use, ensure the modules are not loaded.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2026/CVE-2026-64269.json","https://nvd.nist.gov/vuln/detail/CVE-2026-64269"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-07-25"},{"id":"CVE-2026-64319","cve":"CVE-2026-64319","aliases":["nvmet-auth DHCHAP_REPLY bounds","NVMe-oF in-band auth heap overread"],"title":"Linux kernel - NVMe-oF target DH-HMAC-CHAP authentication, drivers/nvme/target/fabrics-cmd-auth.c: Nvmet_auth_reply()","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel - NVMe-oF target DH-HMAC-CHAP authentication, drivers/nvme/target/fabrics-cmd-auth.c","year":"2026","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Nvmet_auth_reply() reads the variable-length array using attacker-supplied hash-length and DH-value-length fields without checking they fit the allocated transfer length. A malicious initiator sends a DHCHAP_REPLY with a small transfer length but large hl/dhvlen and drives out-of-bounds heap reads of up to 526 bytes past the buffer, with the out-of-bounds pointer handed straight to the crypto layer. The bitter irony for operators: this is in the authentication code you were told to turn on to fix NQN spoofing, and it is exploitable pre-authentication - so enabling in-band auth opens this surface rather than closing it. Still worth enabling auth, but only on a patched kernel.","attack_vector":"Any peer that can reach an auth-enabled NVMe-oF target sends a crafted DHCHAP_REPLY during the authentication exchange. By definition this happens before authentication completes, so no credentials are required. Discovered by an automated vulnerability-discovery engine, which suggests more of this class is coming in the same file.","remediation":"Host reboot / kernel upgrade on NVMe-oF targets before or at the same time as enabling DH-HMAC-CHAP. If you have already enabled in-band auth on an unpatched kernel, prioritise the upgrade; disabling auth is not a good trade because it reopens NQN spoofing. Network-level control in the interim: restrict target ports to known initiator addresses. Sequence the rollout as patch first, then enable auth.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2026/CVE-2026-64319.json","https://nvd.nist.gov/vuln/detail/CVE-2026-64319"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-07-25"},{"id":"CVE-2026-64320","cve":"CVE-2026-64320","aliases":["nvmet pre-auth OOB heap read","NVMe-oF Discovery Get Log Page memory disclosure"],"title":"Linux kernel - NVMe-oF target discovery controller, drivers/nvme/target/discovery.c: The discovery controller validated","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel - NVMe-oF target discovery controller, drivers/nvme/target/discovery.c","year":"2026","cvss_score":9.1,"severity":"critical","kev":false,"impact":"The discovery controller validated only the dword alignment of the host-supplied 64-bit Log Page Offset before adding it to a small heap buffer and memcpy'ing data straight back over the fabric. The discovery subsystem accepts every Host NQN and explicitly skips DH-HMAC-CHAP, so this is reachable pre-authentication by any TCP, RDMA, or FC peer that can see the target. The kernel CNA's own writeup reports an empirical run on a default nvmet-tcp target leaking 81 canonical kernel pointers in a single Get Log Page response; pointing the offset at unmapped memory panics the target instead. For a neocloud running disaggregated NVMe, one unauthenticated request from any tenant reads the storage node's kernel heap - and repeated requests crash the node serving everyone.","attack_vector":"Connect to the discovery controller (nvme discover, or a raw fabrics command) with any Host NQN and issue Get Log Page with an offset at or beyond the allocated discovery log length. No authentication, no valid identity, no prior connection. Works over NVMe/TCP on port 4420, over NVMe/RDMA, and over Fibre Channel. Purely a read primitive plus a crash primitive, but the leaked kernel pointers defeat KASLR and set up heavier exploitation.","remediation":"Host reboot / kernel upgrade on every nvmet target - urgent, given pre-auth reachability and a 9.1 score. Immediate stopgaps while you schedule it: firewall NVMe/TCP 4420 and the RDMA discovery path to known initiator addresses, and move the discovery controller off any tenant-reachable interface onto a management network (nvmet configfs change, runtime, no reboot, but initiators need reconfiguring). Enabling DH-HMAC-CHAP does not help here because the discovery subsystem bypasses it by design - only the kernel fix or network isolation closes it.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2026/CVE-2026-64320.json","https://nvd.nist.gov/vuln/detail/CVE-2026-64320"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-07-25"},{"id":"CVE-2026-7302","cve":"CVE-2026-7302","aliases":[],"title":"SGLang (multimodal runtime): Unauthenticated path traversal","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"SGLang (multimodal runtime)","year":"2026","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Unauthenticated path traversal → arbitrary file write as the server process","attack_vector":"Unauthenticated network to the serving port","remediation":"Upgrade; run serving processes non-root with read-only root filesystems","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-7302"],"status":"curated","published":"2026-05-18"},{"id":"CVE-2026-7482","cve":"CVE-2026-7482","aliases":[],"title":"Ollama (GGUF model loader): Heap out-of-bounds read from an attacker-supplied GGUF via `/api/create`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ollama (GGUF model loader)","year":"2026","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Heap out-of-bounds read from an attacker-supplied GGUF via `/api/create`","attack_vector":"Customer-supplied model file to an exposed API","remediation":"Upgrade to 0.17.1+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-7482"],"status":"curated","published":"2026-05-04"},{"id":"CVE-2018-8930","cve":"CVE-2018-8930","aliases":["MASTERKEY-1","MASTERKEY-2","MASTERKEY-3"],"title":"AMD EPYC / Ryzen - Hardware Validated Boot enforcement: Hardware Validated Boot is not properly enforced, so an","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD EPYC / Ryzen - Hardware Validated Boot enforcement","year":"2018","cvss_score":9,"severity":"critical","kev":false,"impact":"Hardware Validated Boot is not properly enforced, so an attacker who can reflash the BIOS can install firmware the platform will accept and execute despite it not being legitimately signed. That is a persistent, below-the-OS implant on a server: it survives reimaging, disk replacement and tenant handoff, and nothing running in the OS can see it. On a bare-metal GPU cloud where nodes are recycled between customers, this is the classic 'previous tenant left something behind' scenario.","attack_vector":"Local with BIOS reflash capability - so root plus SPI write access, a compromised BMC, or physical/supply-chain access. Not remote.","remediation":"Fixed in AMD reference firmware (AGESA / PSP / SEV firmware) and delivered only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo and the ODMs each rebuild and requalify AMD's AGESA drop before shipping. **Expect one to six months of OEM lag**, and on end-of-support platforms expect nothing. Applying it is a drain plus full power cycle, not a driver reload. Verify by reading back the PSP/SMU firmware version afterwards rather than trusting the BIOS version string. The durable control on a bare-metal fleet is not the patch but the process: measure firmware between tenants, enable platform SPI write protection, and treat any node whose firmware measurement changed as suspect rather than reprovisioning it blindly.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-8930","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2018-03-22"},{"id":"CVE-2018-8931","cve":"CVE-2018-8931","aliases":["RYZENFALL-1"],"title":"AMD Secure Processor (Ryzen / Ryzen Pro / Ryzen Mobile): Insufficient access control on the Secure Processor lets code","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor (Ryzen / Ryzen Pro / Ryzen Mobile)","year":"2018","cvss_score":9,"severity":"critical","kev":false,"impact":"Insufficient access control on the Secure Processor lets code already running with OS administrator rights reach into the AMD Secure Processor and execute there. The ASP sits below the hypervisor and below Secure Boot, so once an attacker is inside it, everything the platform's security rests on - memory encryption keys, fTPM state, boot measurements - is theirs. On a shared host this is the end of any isolation guarantee you were making to tenants, and the compromise survives an OS reinstall.","attack_vector":"Local, requires OS administrator/root plus the ability to flash or load a signed driver. Not remotely reachable. Realistically this is a post-exploitation depth charge: an attacker who already owns the host uses it to get persistence you cannot wash out.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. AMD's fix was a PSP firmware update carried in AGESA. Note this batch (the CTS-Labs disclosures) targeted client Ryzen silicon rather than EPYC; verify against your actual server SKU before spending a maintenance window on it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-8931","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2018-03-22"},{"id":"CVE-2018-8932","cve":"CVE-2018-8932","aliases":["RYZENFALL-2","RYZENFALL-3","RYZENFALL-4"],"title":"AMD Secure Processor (Ryzen / Ryzen Pro): The same class of Secure Processor access-control failure as RYZENFALL-1","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor (Ryzen / Ryzen Pro)","year":"2018","cvss_score":9,"severity":"critical","kev":false,"impact":"The same class of Secure Processor access-control failure as RYZENFALL-1, covering three further variants. An administrator-level attacker writes into ASP-protected memory and gains execution in the secure coprocessor, defeating the hardware root of trust the rest of the platform is anchored to.","attack_vector":"Local, administrator-privileged. Requires the attacker to already control the OS on the node.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. Client-silicon focused; confirm applicability to your EPYC server SKUs before scheduling.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-8932","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2018-03-22"},{"id":"CVE-2018-8933","cve":"CVE-2018-8933","aliases":["FALLOUT-1","FALLOUT-2","FALLOUT-3"],"title":"AMD EPYC Server - protected memory region access control: Insufficient access control over protected memory regions on","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD EPYC Server - protected memory region access control","year":"2018","cvss_score":9,"severity":"critical","kev":false,"impact":"Insufficient access control over protected memory regions on EPYC server parts lets privileged code read and write memory that the platform reserves for security purposes - including regions used by SMM and the secure processor. An attacker uses it to get persistence and to reach data the hardware was supposed to fence off from the OS entirely.","attack_vector":"Local, requires administrator privilege on the host. EPYC server silicon specifically, which is what makes this batch relevant to a datacenter rather than a desktop.","remediation":"Fixed in AMD reference firmware (AGESA / PSP / SEV firmware) and delivered only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo and the ODMs each rebuild and requalify AMD's AGESA drop before shipping. **Expect one to six months of OEM lag**, and on end-of-support platforms expect nothing. Applying it is a drain plus full power cycle, not a driver reload. Verify by reading back the PSP/SMU firmware version afterwards rather than trusting the BIOS version string.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-8933","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2018-03-22"},{"id":"CVE-2018-8934","cve":"CVE-2018-8934","aliases":["CHIMERA-FW"],"title":"Promontory chipset firmware (AMD Ryzen / Ryzen Pro platforms): A backdoor in the Promontory chipset firmware. The","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Promontory chipset firmware (AMD Ryzen / Ryzen Pro platforms)","year":"2018","cvss_score":9,"severity":"critical","kev":false,"impact":"A backdoor in the Promontory chipset firmware. The chipset sits on the DMA path for USB, SATA and PCIe, so code running there can read and write host memory independently of the CPU and outside the reach of anything the OS enforces. Included as the canonical example of the risk class rather than as an EPYC issue: third-party chipset silicon in your server has its own firmware, its own DMA capability, and usually no attestation story at all.","attack_vector":"Local, requires the ability to load chipset firmware. Ryzen/Ryzen Pro client platforms rather than EPYC servers.","remediation":"Fixed by a chipset firmware update from the OEM, delivered in a BIOS package - drain plus power cycle. Verify applicability before spending a window: this is client-platform silicon and almost certainly not in your EPYC server fleet. The transferable lesson for a datacenter operator is to ask which non-AMD firmware images your server actually loads at boot and who signs them.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-8934","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2018-03-22"},{"id":"CVE-2018-8936","cve":"CVE-2018-8936","aliases":["CHIMERA-adjacent PSP escalation"],"title":"AMD EPYC / Ryzen - Platform Security Processor privilege escalation: A direct privilege escalation into the Platform","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD EPYC / Ryzen - Platform Security Processor privilege escalation","year":"2018","cvss_score":9,"severity":"critical","kev":false,"impact":"A direct privilege escalation into the Platform Security Processor on EPYC server parts. The PSP holds the platform's root of trust, fTPM state and SEV key material, so an attacker who escalates into it owns the security posture of the whole node beneath the hypervisor. Nothing the OS or the hypervisor can do detects or contains it.","attack_vector":"Local, administrator privilege.","remediation":"Fixed in AMD reference firmware (AGESA / PSP / SEV firmware) and delivered only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo and the ODMs each rebuild and requalify AMD's AGESA drop before shipping. **Expect one to six months of OEM lag**, and on end-of-support platforms expect nothing. Applying it is a drain plus full power cycle, not a driver reload. Verify by reading back the PSP/SMU firmware version afterwards rather than trusting the BIOS version string.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-8936","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2018-03-22"},{"id":"CVE-2020-15206","cve":"CVE-2020-15206","aliases":[],"title":"TensorFlow (SavedModel protobuf): Mutating a SavedModel protobuf crashes or corrupts the serving process","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"TensorFlow (SavedModel protobuf)","year":"2020","cvss_score":9,"severity":"critical","kev":false,"impact":"Mutating a SavedModel protobuf crashes or corrupts the serving process","attack_vector":"Customer-supplied SavedModel served by a shared serving tier","remediation":"Patch; isolate per-model serving processes","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-15206"],"status":"curated","published":"2020-09-25"},{"id":"CVE-2021-42114","cve":"CVE-2021-42114","aliases":["Blacksmith"],"title":"PC-DDR4 / LPDDR4X DRAM - Target Row Refresh mitigation: Non-uniform Rowhammer patterns triggered bit flips on every one","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"PC-DDR4 / LPDDR4X DRAM - Target Row Refresh mitigation","year":"2021","cvss_score":9,"severity":"critical","kev":false,"impact":"Non-uniform Rowhammer patterns triggered bit flips on every one of the 40 DDR4 modules the researchers tested, including modules whose TRR implementation had resisted TRRespass. Scored 9.0 with a changed scope. Same operator consequence as TRRespass: an integrity attack on host memory reachable from tenant code.","attack_vector":"Local code on the node able to generate the access pattern. Scored AV:Network by the reporters because remote code paths (JavaScript, network stacks) can drive memory access, but the realistic datacenter path is a tenant workload.","remediation":"No vendor patch. Use ECC and alert on correctable-error rate rather than only on uncorrectable errors; enable increased refresh rate or RFM in BIOS if your platform exposes it, accepting a small memory-bandwidth cost. Cost: BIOS change means drain and reboot. Effectively UNPATCHABLE.","references":["https://comsec.ethz.ch/research/dram/blacksmith/","https://comsec.ethz.ch/wp-content/files/blacksmith_sp22.pdf","https://nvd.nist.gov/vuln/detail/CVE-2021-42114"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"published":"2021-11-16"},{"id":"CVE-2022-22805","cve":"CVE-2022-22805","aliases":["TLStorm"],"title":"APC Smart-UPS SmartConnect family (SMT, SMC, SMTL, SCL, SMX series) - cloud-connected UPS firmware: A heap overflow in","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"APC Smart-UPS SmartConnect family (SMT, SMC, SMTL, SCL, SMX series) - cloud-connected UPS firmware","year":"2022","cvss_score":9,"severity":"critical","kev":false,"impact":"A heap overflow in TLS packet reassembly gives an attacker code execution on the UPS's own controller - the device that decides whether your racks get power. From there an attacker can cut output, refuse to transfer to battery during a utility event, or (as Armis demonstrated on the bench) drive the unit until it physically burns. This is not a monitoring card compromise; it is the power train. A single UPS covering a GPU row kills every training job in that row with no checkpoint.","attack_vector":"Unauthenticated. The UPS initiates an outbound TLS connection to Schneider's cloud service, so an attacker who can intercept or MITM that connection - or who is simply on the same network segment as the UPS management port - reaches the vulnerable parser. No credentials, no prior foothold on the compute network.","remediation":"Firmware flash on every affected UPS, pushed through the Schneider update tool or the cloud service. Cost is real: each unit must be updated individually and some models require the load to be transferred or the unit taken to bypass first, so this is a scheduled electrical maintenance window per unit, not a fleet-wide push. If you cannot patch promptly, block the UPS's outbound path to the SmartConnect cloud and put the management port on an isolated VLAN with no route to the internet.","references":["https://www.se.com/ww/en/download/document/SEVD-2022-067-02/","https://www.armis.com/research/tlstorm/"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["physical-impact"],"published":"2022-03-09"},{"id":"CVE-2022-22806","cve":"CVE-2022-22806","aliases":["TLStorm"],"title":"APC Smart-UPS SmartConnect family (SMT, SMC, SMTL, SCL, SMX series) - TLS state machine: A TLS authentication bypass by","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"APC Smart-UPS SmartConnect family (SMT, SMC, SMTL, SCL, SMX series) - TLS state machine","year":"2022","cvss_score":9,"severity":"critical","kev":false,"impact":"A TLS authentication bypass by capture-replay: a malformed handshake puts the UPS into an unauthenticated-but-connected state, letting an attacker talk to it as if it were the trusted cloud service. Combined with the unsigned-firmware issue below, this is the full chain from network reachability to controlling whether a GPU hall stays energised.","attack_vector":"Unauthenticated, network-adjacent to the UPS management interface or positioned on the path of its outbound cloud connection.","remediation":"Same firmware campaign as CVE-2022-22805 - a per-unit flash with an electrical maintenance window for models that need a bypass transfer. Interim mitigation is network isolation of the UPS management plane and blocking its egress. There is no configuration toggle that removes the vulnerable code path.","references":["https://www.se.com/ww/en/download/document/SEVD-2022-067-02/","https://www.armis.com/research/tlstorm/"],"status":"curated","tags":["physical-impact"],"published":"2022-03-09"},{"id":"CVE-2022-31035","cve":"CVE-2022-31035","aliases":[],"title":"Argo CD: Stored XSS via a `javascript:` link executes in an admin's browser","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2022","cvss_score":9,"severity":"critical","kev":false,"impact":"Stored XSS via a `javascript:` link executes in an admin's browser","attack_vector":"Any user who can create an Application","remediation":"Rolling Argo CD upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31035"],"status":"curated","published":"2022-06-27"},{"id":"CVE-2023-22482","cve":"CVE-2023-22482","aliases":[],"title":"Argo CD: Improper authorization causes the API to accept tokens it should reject","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2023","cvss_score":9,"severity":"critical","kev":false,"impact":"Improper authorization causes the API to accept tokens it should reject","attack_vector":"Any authenticated Argo CD user","remediation":"Rolling Argo CD upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-22482"],"status":"curated","published":"2023-01-26"},{"id":"CVE-2023-4299","cve":"CVE-2023-4299","aliases":["ICSA-23-243-04"],"title":"Digi RealPort protocol (Digi console/terminal servers): RealPort is the protocol Digi console servers use to expose","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Digi RealPort protocol (Digi console/terminal servers)","year":"2023","cvss_score":9,"severity":"critical","kev":false,"impact":"RealPort is the protocol Digi console servers use to expose their serial ports as virtual COM ports over the network. Its authentication can be replayed — an attacker who captures a legitimate auth exchange (e.g. via a network tap or ARP spoof on the management VLAN) can replay it to open a session on connected serial equipment without knowing the real password.","attack_vector":"Requires network visibility into a RealPort authentication exchange (passive capture is enough) and the ability to send the replayed traffic to the target console server — no credential cracking needed.","remediation":"Firmware/software upgrade on the RealPort driver and the console server firmware to a version with replay-resistant authentication; also a network-segmentation fix — RealPort traffic should never traverse a network segment an untrusted party can sniff. Rollout is a firmware flash per device plus a driver update on every host connecting to RealPort ports.","references":["https://www.cisa.gov/news-events/ics-advisories/icsa-23-243-04","https://www.digi.com/getattachment/resources/security/alerts/realport-cves/Dragos-Disclosure-Statement.pdf"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-08-31"},{"id":"CVE-2024-0087","cve":"CVE-2024-0087","aliases":[],"title":"Triton Inference Server: RCE via path traversal on the model-load API","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2024","cvss_score":9,"severity":"critical","kev":false,"impact":"RCE via path traversal on the model-load API","attack_vector":"Unauthenticated/low-priv client of the inference endpoint","remediation":"Upgrade Triton (24.09+); rebuild and redeploy all serving images; restrict model-control API","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0087","https://github.com/NVIDIA/product-security/tree/main/2024/5535"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:H/UI:N/S:C/C:H/I:L/A:H","cwe":["CWE-73"],"published":"2024-05-14"},{"id":"CVE-2024-0095","cve":"CVE-2024-0095","aliases":[],"title":"Triton Inference Server: Improper logging of security events (audit gap)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2024","cvss_score":9,"severity":"critical","kev":false,"impact":"Improper logging of security events (audit gap)","attack_vector":"Client of the inference endpoint","remediation":"Upgrade Triton; redeploy serving images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0095","https://github.com/NVIDIA/product-security/tree/main/2024/5546"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:L/A:N","cwe":["CWE-117"],"published":"2024-06-13"},{"id":"CVE-2024-0132","cve":"CVE-2024-0132","aliases":[],"title":"Container Toolkit: Container escape to host filesystem via TOCTOU in the **default** configuration","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Container Toolkit","year":"2024","cvss_score":9,"severity":"critical","kev":false,"impact":"Container escape to host filesystem via TOCTOU in the **default** configuration","attack_vector":"Any tenant that can run an arbitrary container image on a GPU node","remediation":"Bump nvidia-container-toolkit to 1.16.2+ and restart the container runtime on every GPU node; upgrade GPU Operator to 24.6.2+; evict and re-admit tenant workloads","references":["https://github.com/NVIDIA/product-security/tree/main/2024/5582","https://nvd.nist.gov/vuln/detail/CVE-2024-0132"],"status":"curated","fleet":{"ubiquity":"Universal - all versions <= 1.16.1, i.e. the entire installed base at disclosure","remediation_pain":"`daemon-restart` to 1.16.2 / GPU Operator 24.6.2; hosts that ran untrusted images need `node-drain` + rebuild because the host-filesystem write has already happened","pain_class":"node-drain","why_fleet_wide":"TOCTOU in the mount path lets a crafted container image mount the host filesystem and execute code as root - the canonical \"one CVE, every GPU node in the fleet\" event"},"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:R/S:C/C:H/I:H/A:H","cwe":["CWE-367"],"published":"2024-09-26"},{"id":"CVE-2024-28175","cve":"CVE-2024-28175","aliases":[],"title":"Argo CD: Improper URL protocol filtering in link annotations enables client-side attacks against admins","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2024","cvss_score":9,"severity":"critical","kev":false,"impact":"Improper URL protocol filtering in link annotations enables client-side attacks against admins","attack_vector":"Any user who can annotate an Application","remediation":"Rolling Argo CD upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-28175"],"status":"curated","published":"2024-03-13"},{"id":"CVE-2024-28179","cve":"CVE-2024-28179","aliases":[],"title":"Jupyter Server Proxy: Authentication weakness in proxied-process access","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Jupyter Server Proxy","year":"2024","cvss_score":9,"severity":"critical","kev":false,"impact":"Authentication weakness in proxied-process access","attack_vector":"Network user of a JupyterHub deployment","remediation":"Upgrade; commonly deployed on managed GPU notebook platforms","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-28179"],"status":"curated","published":"2024-03-20"},{"id":"CVE-2024-31989","cve":"CVE-2024-31989","aliases":[],"title":"Argo CD: An unprivileged pod in any namespace can reach the unauthenticated Argo CD Redis on 6379 and poison","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2024","cvss_score":9,"severity":"critical","kev":false,"impact":"An unprivileged pod in any namespace can reach the unauthenticated Argo CD Redis on 6379 and poison the cache, which becomes arbitrary deployment","attack_vector":"Any pod on the cluster network","remediation":"Rolling Argo CD upgrade; enable Redis auth and a NetworkPolicy around it. Critical in a multi-tenant neocloud where any tenant pod is on that network","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-31989"],"status":"curated","published":"2024-05-21"},{"id":"CVE-2024-45764","cve":"CVE-2024-45764","aliases":["DSA-2024-449"],"title":"Dell Enterprise SONiC (authentication): A critical step in authentication is missing, so an unauthenticated remote","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell Enterprise SONiC (authentication)","year":"2024","cvss_score":9,"severity":"critical","kev":false,"impact":"A critical step in authentication is missing, so an unauthenticated remote attacker bypasses the protection mechanism and gets into the switch. SONiC is increasingly the NOS of choice for cost-driven GPU buildouts precisely because it is open and cheap; this is the reminder that the open NOS ecosystem has the same class of front-door bugs as the incumbents, with a shorter advisory history to check against.","attack_vector":"Unauthenticated, remote — reachability to the switch's management services is the only requirement.","remediation":"Upgrade Dell Enterprise SONiC past 4.1.x/4.2.x to a fixed release, which means a NOS image install and switch reboot per device. In a SONiC fabric that is a full image swap, not a patch — budget a maintenance window per leaf and stage it across MLAG pairs. Restrict management-interface reachability in the meantime.","references":["https://www.dell.com/support/kbdoc/en-us/000245655/dsa-2024-449-security-update-for-dell-enterprise-sonic-distribution-vulnerabilities","https://nvd.nist.gov/vuln/detail/CVE-2024-45764"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-11-08"},{"id":"CVE-2025-0282","cve":"CVE-2025-0282","aliases":[],"title":"Ivanti Connect Secure: Stack-based buffer overflow","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ivanti Connect Secure","year":"2025","cvss_score":9,"severity":"critical","kev":true,"impact":"Stack-based buffer overflow -> unauthenticated remote code execution","attack_vector":"Network (remote)","remediation":"Control-plane: emergency patch plus factory reset","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0282"],"status":"curated","published":"2025-01-08"},{"id":"CVE-2025-22457","cve":"CVE-2025-22457","aliases":[],"title":"Ivanti Connect Secure/ZTA: Stack-based buffer overflow","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ivanti Connect Secure/ZTA","year":"2025","cvss_score":9,"severity":"critical","kev":true,"impact":"Stack-based buffer overflow -> unauthenticated RCE, exploited by a China-nexus actor","attack_vector":"Network (remote)","remediation":"Control-plane: patch; assume compromise","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-22457"],"status":"curated","published":"2025-04-03"},{"id":"CVE-2025-23266","cve":"CVE-2025-23266","aliases":["NVIDIAScape"],"title":"Container Toolkit: Container escape to host root via malicious image (LD_PRELOAD in OCI hook)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Container Toolkit","year":"2025","cvss_score":9,"severity":"critical","kev":false,"impact":"Container escape to host root via malicious image (LD_PRELOAD in OCI hook)","attack_vector":"Any tenant that can run an arbitrary container image on a GPU node","remediation":"Emergency: bump nvidia-container-toolkit to 1.17.8+, restart container runtime on every GPU node, upgrade GPU Operator Helm chart; evict and re-admit all tenant workloads; audit for prior exploitation","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23266","https://github.com/NVIDIA/product-security/tree/main/2025/5659"],"status":"curated","fleet":{"ubiquity":"Very common - the standard way to ship driver + toolkit + DCGM on every K8s-based neocloud","remediation_pain":"`node-drain` in practice: the Operator's driver and toolkit DaemonSets restart per node, and a driver-container reload requires evicting every GPU pod","pain_class":"node-drain","why_fleet_wide":"One Helm version bump has to roll across every GPU node pool; until it finishes, every tenant pod on an un-rolled node still holds the escape primitive"},"cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-426"],"published":"2025-07-17"},{"id":"CVE-2025-23350","cve":"CVE-2025-23350","aliases":[],"title":"BlueField (GA firmware): RCE on the DPU via firmware buffer overflow","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"BlueField (GA firmware)","year":"2025","cvss_score":9,"severity":"critical","kev":false,"impact":"RCE on the DPU via firmware buffer overflow","attack_vector":"Privileged network attacker on the DPU management path","remediation":"Flash BlueField firmware out-of-band; DPU reset drops tenant networking, schedule node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23350","https://github.com/NVIDIA/product-security/tree/main/2026/5699"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-drain"},"published":"2026-07-01"},{"id":"CVE-2025-23351","cve":"CVE-2025-23351","aliases":[],"title":"BlueField (GA firmware): RCE on the DPU via firmware buffer overflow","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"BlueField (GA firmware)","year":"2025","cvss_score":9,"severity":"critical","kev":false,"impact":"RCE on the DPU via firmware buffer overflow","attack_vector":"Privileged network attacker","remediation":"Flash BlueField firmware out-of-band; node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23351","https://github.com/NVIDIA/product-security/tree/main/2026/5699"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-drain"},"published":"2026-07-01"},{"id":"CVE-2025-29783","cve":"CVE-2025-29783","aliases":[],"title":"vLLM (Mooncake): Unsafe deserialization over ZMQ/TCP bound to all interfaces","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (Mooncake)","year":"2025","cvss_score":9,"severity":"critical","kev":false,"impact":"Unsafe deserialization over ZMQ/TCP bound to all interfaces","attack_vector":"Unauthenticated network, co-tenant reachable","remediation":"Upgrade; the KV-transfer plane is unauthenticated by design in these versions","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-29783"],"status":"curated","published":"2025-03-19"},{"id":"CVE-2025-33210","cve":"CVE-2025-33210","aliases":[],"title":"NVIDIA Isaac Lab: A deserialization flaw reaches code execution","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Isaac Lab","year":"2025","cvss_score":9,"severity":"critical","kev":false,"impact":"A deserialization flaw reaches code execution; scored 9.0 network with a changed scope. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5733 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33210","https://github.com/NVIDIA/product-security/tree/main/2025/5733"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:R/S:C/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2025-12-16"},{"id":"CVE-2025-33244","cve":"CVE-2025-33244","aliases":[],"title":"Apex: Remote RCE via unsafe pickle deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Apex","year":"2025","cvss_score":9,"severity":"critical","kev":false,"impact":"Remote RCE via unsafe pickle deserialization","attack_vector":"Network attacker / malicious checkpoint","remediation":"Bump Apex in training images; rebuild and redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33244","https://github.com/NVIDIA/product-security/tree/main/2026/5782"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2026-03-24"},{"id":"CVE-2026-65094","cve":"CVE-2026-65094","aliases":[],"title":"NVIDIA BlueField - VIRTIO-Net emulation: A VM user sends a crafted message to the BlueField VIRTIO-Net device and gets","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA BlueField - VIRTIO-Net emulation","year":"2026","cvss_score":9,"severity":"critical","kev":false,"impact":"A VM user sends a crafted message to the BlueField VIRTIO-Net device and gets a write-what-where primitive, reaching code execution in the VIRTIO-Net context on the DPU. The DPU is the component you offloaded tenant network isolation onto - a tenant VM reaching code execution inside it inverts the trust model of the whole design. Scored 9.0 with a changed scope.","attack_vector":"A user inside a tenant VM talking to the emulated virtio-net device its own hypervisor exposed. No host or DPU credentials needed. This is the guest-to-DPU boundary.","remediation":"Update the BlueField VIRTIO-Net firmware/software per bulletin 5815 across GA, LTS23, LTS24 and LTS25 branches as applicable. Cost: a DPU firmware update takes the DPU's dataplane down, which means the host loses network - treat it as a full node drain, not a live update. Sequence carefully: a half-updated DPU fleet has inconsistent offload behaviour.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-65094","https://github.com/NVIDIA/product-security/tree/main/2026/5815"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-123"],"fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-345","CWE-367"],"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2026-055-crossplane-package-manager-cosig","cve":null,"aliases":["GHSA-wfqx-gjrf-g28r"],"title":"Crossplane package manager (cosign signature verification via ImageConfig): SUPPLY CHAIN, TIME-OF-CHECK TO TIME-OF-USE","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Crossplane package manager (cosign signature verification via ImageConfig)","year":"2026","cvss_score":9,"severity":"critical","kev":false,"impact":"SUPPLY CHAIN, TIME-OF-CHECK TO TIME-OF-USE: Crossplane verifies a package's signature and then installs a different package. When a package is referenced by tag rather than digest, the package manager resolves that tag separately for the verification step and for the pull step, so a malicious registry serves a correctly signed image to the verifier and an unsigned one to the installer. Signature verification reports success and unsigned attacker code lands in the cluster with whatever privileges the Crossplane package holds — and Crossplane packages are control-plane extensions that reconcile cloud and infrastructure resources, so that is typically broad. The failure mode is the worst kind for an operator: the control is enabled, the dashboard is green, and it is providing no protection.","attack_vector":"Network, unauthenticated from the attacker's side: requires a malicious or compromised OCI registry able to vary what it serves per request. Only affects users who enable signature verification, install by tag rather than digest, and pull from registries they do not control.","remediation":"Install packages by image digest rather than tag — this defeats the race entirely and is the vendor's stated mitigation as well as general best practice. Upgrade to Crossplane 2.3.3 or 2.2.3, where the tag is resolved once and the resulting digest is used for both verification and fetch. Note the maintainers are not backporting to 1.20, so 1.20 clusters must rely on digest pinning permanently.","references":["https://github.com/crossplane/crossplane/security/advisories/GHSA-wfqx-gjrf-g28r"],"status":"curated"},{"id":"CVE-2024-0105","cve":"CVE-2024-0105","aliases":[],"title":"ConnectX / BlueField firmware: Improper certificate validation","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"ConnectX / BlueField firmware","year":"2024","cvss_score":8.9,"severity":"high","kev":false,"impact":"Improper certificate validation -> malicious firmware/image accepted","attack_vector":"Network-adjacent attacker in the firmware update path","remediation":"Flash NIC/DPU firmware; verify update chain; node reboot with tenant eviction","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0105","https://github.com/NVIDIA/product-security/tree/main/2024/5562"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:L/UI:N/S:C/C:L/I:H/A:H","cwe":["CWE-274"],"fleet":{"pain_class":"node-reboot"},"published":"2024-11-01"},{"id":"CVE-2025-12060","cve":"CVE-2025-12060","aliases":[],"title":"Keras (`utils.get_file`, tar extract): Path traversal on tar extraction","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Keras (`utils.get_file`, tar extract)","year":"2025","cvss_score":8.9,"severity":"high","kev":false,"impact":"Path traversal on tar extraction → arbitrary file write","attack_vector":"Customer-supplied dataset/model URL fetched with `extract=True`","remediation":"Upgrade; a training job with write access to shared mounts can escape into other tenants' paths if mounts are shared","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-12060"],"status":"curated","published":"2025-10-30"},{"id":"CVE-2025-53630","cve":"CVE-2025-53630","aliases":[],"title":"llama.cpp (`gguf_init_from_file_impl`): Integer overflow in GGUF init","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"llama.cpp (`gguf_init_from_file_impl`)","year":"2025","cvss_score":8.9,"severity":"high","kev":false,"impact":"Integer overflow in GGUF init","attack_vector":"Customer-supplied GGUF file","remediation":"Rebuild","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-53630"],"status":"curated","published":"2025-07-10"},{"cvss_vector":"CVSS:4.0/AV:N/AC:L/AT:P/PR:L/UI:N/VC:H/VI:H/VA:N/SC:H/SI:H/SA:H","cwe":["CWE-863"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2026-31892","cve":"CVE-2026-31892","aliases":["GHSA-3wf5-g532-rcrr"],"title":"Argo Workflows (controller, podSpecPatch in Strict/Secure template reference mode): A podSpecPatch on the submitted","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo Workflows (controller, podSpecPatch in Strict/Secure template reference mode)","year":"2026","cvss_score":8.9,"severity":"high","kev":false,"impact":"A podSpecPatch on the submitted Workflow takes precedence over the referenced WorkflowTemplate and is written straight into the pod spec with no validation. Every hardening an admin encoded in the template - restricted service account, dropped capabilities, resource limits - is overridable by the tenant submitting the job. On a shared GPU cluster this is the difference between a tenant running in their own sandbox and running as the platform service account.","attack_vector":"Any principal allowed to create Workflows in a watched namespace. templateReferencing Strict, which the docs present as the mechanism that confines users to approved templates, does not stop it.","remediation":"Upgrade the controller to 3.7.11 or 4.0.2 and restart. Note the fix was incomplete twice over (see CVE-2026-42296 and CVE-2026-54526), so land 3.7.15 / 4.0.6 rather than stopping at the first patched release, and add an admission policy that strips podSpecPatch from tenant Workflows.","references":["https://github.com/argoproj/argo-workflows/security/advisories/GHSA-3wf5-g532-rcrr","https://nvd.nist.gov/vuln/detail/CVE-2026-31892"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:4.0/AV:N/AC:L/AT:P/PR:L/UI:N/VC:H/VI:H/VA:N/SC:H/SI:H/SA:H","cwe":["CWE-284"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2026-54526","cve":"CVE-2026-54526","aliases":["GHSA-48p8-g2fx-3wwm"],"title":"Argo Workflows (controller, ArtifactGC.PodSpecPatch / template reference allow-list): The allow-list that is supposed","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo Workflows (controller, ArtifactGC.PodSpecPatch / template reference allow-list)","year":"2026","cvss_score":8.9,"severity":"high","kev":false,"impact":"The allow-list that is supposed to force users onto admin-approved WorkflowTemplates only inspects top-level WorkflowSpec fields, so a podSpecPatch nested under ArtifactGC slips through untouched and is applied to the pod. A tenant who can submit a Workflow overrides the security settings the platform team pinned in the template - service account, security context, mounts - and runs with privileges they were never granted. Third bypass in the same allow-list, after CVE-2026-31892 and CVE-2026-42296.","attack_vector":"Any user or service account with permission to create Workflow objects in a namespace the controller watches. Works even when the controller runs with templateReferencing set to Strict or Secure.","remediation":"Upgrade the workflow controller to 3.7.15 or 4.0.6 and restart it. Do not treat templateReferencing Strict as a boundary on its own - back it with a Kubernetes admission policy (Kyverno, Gatekeeper or ValidatingAdmissionPolicy) that rejects podSpecPatch, hostNetwork and serviceAccountName overrides on tenant-submitted Workflows.","references":["https://github.com/argoproj/argo-workflows/security/advisories/GHSA-48p8-g2fx-3wwm","https://nvd.nist.gov/vuln/detail/CVE-2026-54526"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:L","cwe":["CWE-327"],"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2025-016-ceph-cephx-authentication-protoc","cve":null,"aliases":["GHSA-7q3q-3975-qw3q","CVE-2025-30156 (reserved)"],"title":"Ceph CephX (authentication protocol): A tenant holding one low-privilege CephX client key ends up with cluster-wide","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ceph CephX (authentication protocol)","year":"2025","cvss_score":8.9,"severity":"high","kev":false,"impact":"A tenant holding one low-privilege CephX client key ends up with cluster-wide daemon credentials. CephX wraps tickets in AES-128-CBC with no MAC and a hard-coded IV, so ciphertext is malleable and identical plaintexts are visibly identical. An attacker asks the monitor for tickets on entity names they choose, using the monitor as an encryption oracle, then splices the returned blocks into forged credentials for Manager, MDS and OSD identities. A second, cheaper path found by CLYSO needs only a single bit-flip in a service ticket to set the allow_all field true. Either way the tenant stops being a tenant and becomes the storage fabric: every other customer's RBD volumes, CephFS trees and RGW buckets on the shared cluster are readable and writable, and OSD-level access lets them tamper with data underneath other tenants' checkpoints and datasets. This is the Kerberos 4 PERILS flaw reappearing in the storage substrate that most GPU clouds run their training data on.","attack_vector":"Adjacent network: the attacker must be able to speak to the Ceph monitors on the cluster/messenger network and must already hold one valid CephX key of any privilege level (a normal tenant client key qualifies). The oracle path additionally wants the ability to observe CephX ciphertext on the wire. No user interaction, no admin caps.","remediation":"Upgrade to Ceph 20.2.4 or 19.2.6 and roll every daemon (mon, mgr, osd, mds, rgw) so the hardened handler is actually in use. Because forged credentials may already exist and are indistinguishable from real ones, treat this as a credential-compromise event: rotate CephX keys after the upgrade rather than only patching. Keep the cluster/messenger network off any tenant-reachable VLAN in the meantime.","references":["https://github.com/ceph/ceph/security/advisories/GHSA-7q3q-3975-qw3q","https://docs.ceph.com/en/latest/security/CVE-2025-30156/"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-264","CWE-782"],"fleet":{"pain_class":"node-drain"},"id":"CVE-2012-0946","cve":"CVE-2012-0946","aliases":[],"title":"NVIDIA UNIX (Linux/FreeBSD/Solaris) GPU driver before 295.40 - /dev/nvidia* device node: The GPU-side twin of","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA UNIX (Linux/FreeBSD/Solaris) GPU driver before 295.40 - /dev/nvidia* device node","year":"2012","cvss_score":8.8,"severity":"high","kev":false,"impact":"The GPU-side twin of CVE-2014-8159. The NVIDIA UNIX driver let anyone holding read/write permission on the GPU device node reach arbitrary system memory locations through it - and /dev/nvidia* is world-accessible by design, because unprivileged users have to run CUDA. On a shared GPU node that means any tenant with a GPU handle, which is every tenant, gets a host-memory read/write primitive. Structurally important for the database because it is the earliest clear demonstration that the GPU device node is a privilege boundary the OS does not enforce: the kernel's memory protections do not apply to what the GPU is asked to touch.","attack_vector":"Local, unprivileged - open /dev/nvidia0. Any process permitted to run CUDA on the node, including inside a container with the GPU mapped in.","remediation":"Upgrade to NVIDIA 295.40 or later. In practice a GPU driver upgrade means unloading nvidia.ko, which requires every process holding a GPU to exit - so the node has to be drained of tenant workloads first, and on a training cluster that means waiting for or killing long-running jobs. This is the operational reason GPU driver versions drift badly in production fleets: the upgrade is cheap technically and expensive in lost job-hours, so it gets deferred.","references":["https://www.nvidia.com/en-us/security/","https://www.openwall.com/lists/oss-security/2012/08/01/6","https://nvd.nist.gov/vuln/detail/CVE-2012-0946"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-190","CWE-119"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2014-8159","cve":"CVE-2014-8159","aliases":[],"title":"Linux kernel InfiniBand uverbs (ib_uverbs / ib_umem_get, drivers/infiniband/core/umem.c): The canonical RDMA isolation","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel InfiniBand uverbs (ib_uverbs / ib_umem_get, drivers/infiniband/core/umem.c)","year":"2014","cvss_score":8.8,"severity":"high","kev":false,"impact":"The canonical RDMA isolation break. Any process that can open /dev/infiniband/uverbsN registers a memory region whose start+length overflows the page-aligned end computation in ib_umem_get(). The HCA then holds a valid rkey/lkey for physical memory the caller never owned, and the tenant reads and writes it at line rate via ordinary RDMA verbs - other tenants' pages, the page cache, and kernel text all included. Because the DMA is issued by the HCA rather than the CPU, it bypasses page tables entirely: no MMU check, no KASLR, no SMAP/SMEP. On a shared GPU node where the uverbs device is passed into tenant containers so jobs can use NCCL/UCX, this is a full read/write primitive over host memory from inside an unprivileged container.","attack_vector":"Local, unprivileged - requires only read/write access to an /dev/infiniband/uverbsN character device. In practice every RDMA-enabled tenant container has exactly that, because the device node must be exposed for NCCL, UCX, MPI or GPUDirect to work. No fabric access and no root needed.","remediation":"Patch the kernel to include commit 8494057ab5e40df590ef6ef7d66324d3ae33356b (IB/uverbs: prevent integer overflow in ib_umem_get address arithmetic) plus the follow-up 66578b0b2f69659f00b6169e6fe7377c4b100d18 that re-permits legitimate registrations starting at 0x0; RHEL 6 kernel-2.6.32-504.12.2 and later carry it. Not live-patchable in practice - the change is in the memory-region registration path, so the fleet needs a rolling kernel upgrade and reboot, one node drained at a time. Where a reboot cannot be scheduled, the only real mitigation is to stop exposing /dev/infiniband/uverbs* to untrusted workloads, which means turning off RDMA for those tenants entirely.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=8494057ab5e40df590ef6ef7d66324d3ae33356b","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=66578b0b2f69659f00b6169e6fe7377c4b100d18","https://access.redhat.com/security/cve/CVE-2014-8159","https://bugzilla.redhat.com/show_bug.cgi?id=1181166"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-119"],"fleet":{"pain_class":"node-drain"},"id":"CVE-2016-3710","cve":"CVE-2016-3710","aliases":["XSA-179"],"title":"QEMU VGA device model (hw/display/vga.c) - banked access to video memory: 'Dark Portal' - the guest sets the VGA bank","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"QEMU VGA device model (hw/display/vga.c) - banked access to video memory","year":"2016","cvss_score":8.8,"severity":"high","kev":false,"impact":"'Dark Portal' - the guest sets the VGA bank register, then changes the access mode so the bounds check and the actual access disagree, and writes past the video memory window into the QEMU process. That yields arbitrary code execution on the host with the device model's privileges; without a stub domain on Xen, that is dom0. Relevant to GPU infrastructure because the emulated display device is present in essentially every VM regardless of whether the tenant has a real GPU - operators who assume 'we do not give tenants graphics' still ship an emulated VGA adapter to every guest, and it is one of the oldest and least-reviewed device models in QEMU.","attack_vector":"Guest OS user or administrator - the tenant - writing to standard VGA registers and the video memory window. No special device assignment required.","remediation":"Update QEMU (fixed in the 2.6 cycle) or apply the XSA-179 patches, then restart every guest so it picks up the new device model - QEMU cannot be patched under a running VM. Live migration to a patched host is the way to do this without tenant-visible downtime, which is precisely the option you do not have for guests holding passed-through GPUs. Removing the emulated VGA adapter from guest configurations where it is not needed eliminates the surface outright and is the cheaper move for a fleet of headless compute VMs.","references":["https://xenbits.xen.org/xsa/advisory-179.html","https://www.openwall.com/lists/oss-security/2016/05/09/2","https://access.redhat.com/security/cve/CVE-2016-3710"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.0/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-74","CWE-78"],"fleet":{"pain_class":"firmware-flash"},"id":"CVE-2016-5685","cve":"CVE-2016-5685","aliases":[],"title":"Dell iDRAC7 / iDRAC8 firmware before 2.40.40.40 - racadm CLI string injection: A string injection escapes the","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC7 / iDRAC8 firmware before 2.40.40.40 - racadm CLI string injection","year":"2016","cvss_score":8.8,"severity":"high","kev":false,"impact":"A string injection escapes the restricted racadm command interface and drops the caller into a real Bash shell on the service processor. Any BMC account - including the low-privilege, read-only kind an operator hands to a customer or an NOC contractor for power-cycling their own node - becomes arbitrary code execution on the management controller. From a shell on the iDRAC an attacker persists across host reinstalls, tampers with the host's firmware update path, and watches or drives the host console. This is the pre-2018 example of why 'we only gave them limited BMC access' is not a boundary: the restricted CLI was the boundary, and it is one quoting bug deep.","attack_vector":"Network, post-auth. Any valid iDRAC credential, at any privilege level, over SSH or the racadm interface.","remediation":"Flash iDRAC7/iDRAC8 firmware to 2.40.40.40 or later - an out-of-band update that does not require host downtime, roughly 10-15 minutes per node with a brief loss of management access, so it can be run against live nodes if you accept that window. The policy change worth making alongside it: stop issuing per-tenant or per-vendor BMC accounts at all. Broker power and console operations through your own control plane so tenants never hold a credential on the service processor, which removes this entire class rather than this one instance of it.","references":["https://downloads.dell.com/solutions/dell-management-solution-resources/Dell%20iDRAC%20team%20Response%20to%20CVE-2016-5685%20root%20shell.pdf","https://www.cve.org/CVERecord?id=CVE-2016-5685","https://nvd.nist.gov/vuln/detail/CVE-2016-5685"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.0/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-284"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2016-6258","cve":"CVE-2016-6258","aliases":["XSA-182"],"title":"Xen x86 PV pagetable update fast paths (arch/x86/mm.c): A 32-bit PV guest administrator gains full host privileges by","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen x86 PV pagetable update fast paths (arch/x86/mm.c)","year":"2016","cvss_score":8.8,"severity":"high","kev":false,"impact":"A 32-bit PV guest administrator gains full host privileges by abusing the fast paths for pagetable entry updates - the hypervisor lets the guest make its own pagetables writable and from there it owns the machine. This is the highest-stature generic Xen escape of the era; it is the bug that forced Qubes OS to abandon PV isolation architecturally. For a multi-tenant operator the point is that it needs no device assignment and no exotic hardware: any tenant renting a PV instance escapes to the host and to every co-tenant on it.","attack_vector":"Guest administrator in a paravirtualised (PV) guest. HVM and PVH guests are not affected.","remediation":"Apply the XSA-182 patches and reboot the hypervisor across the fleet, evacuating tenants node by node. The durable answer is architectural rather than a patch: stop offering PV guests and move tenants to HVM/PVH, which is what the Xen ecosystem did after this. If you are still running PV instances in 2026 because a legacy customer needs them, that is a business decision to price and time-box, not a technical constraint.","references":["https://xenbits.xen.org/xsa/advisory-182.html","https://www.qubes-os.org/news/2016/07/26/xsa-182/","https://access.redhat.com/security/cve/CVE-2016-6258"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2018-0395","cve":"CVE-2018-0395","aliases":[],"title":"Cisco NX-OS / FXOS (LLDP parser): A malformed LLDP frame reloads the switch. LLDP is enabled by default on essentially","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco NX-OS / FXOS (LLDP parser)","year":"2018","cvss_score":8.8,"severity":"high","kev":false,"impact":"A malformed LLDP frame reloads the switch. LLDP is enabled by default on essentially every fabric port and is unauthenticated by design, so any host on any port can drop its leaf. In a training cluster a single leaf reload takes out a whole rack of GPUs mid-job.","attack_vector":"Unauthenticated, adjacent — a single crafted frame from a directly connected host.","remediation":"NX-OS upgrade plus reload. Interim: disable LLDP on tenant-facing ports. Note that if you use LLDP for cabling verification (common in GPU builds, to prove the rail-optimized topology is wired right) disabling it costs you that check, so most operators patch rather than disable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-0395"],"status":"curated","tags":["fabric-dos"],"published":"2018-10-17"},{"id":"CVE-2018-1000400","cve":"CVE-2018-1000400","aliases":[],"title":"CRI-O: Ambient-capability mishandling runs containers with elevated privileges","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"CRI-O","year":"2018","cvss_score":8.8,"severity":"high","kev":false,"impact":"Ambient-capability mishandling runs containers with elevated privileges","attack_vector":"Any tenant workload","remediation":"Upgrade CRI-O; drain node","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-1000400"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2018-05-18"},{"cvss_vector":"CVSS:3.0/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-121"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2018-10907","cve":"CVE-2018-10907","aliases":[],"title":"GlusterFS (brick, server-rpc-fops.c): Multiple stack buffer overflows from fixed-size alloca() allocations in the brick","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"GlusterFS (brick, server-rpc-fops.c)","year":"2018","cvss_score":8.8,"severity":"high","kev":false,"impact":"Multiple stack buffer overflows from fixed-size alloca() allocations in the brick RPC handlers. An authenticated client overflows the brick's stack and gets code execution on the storage server, which owns the data of every tenant on that node.","attack_vector":"Any authenticated gluster client that can mount a volume and send RPCs to a brick.","remediation":"Upgrade glusterfs server to the fixed release and restart the brick processes. Confirm the build has stack protector and FORTIFY_SOURCE enabled, and keep brick ports reachable only from known client subnets.","references":["https://access.redhat.com/security/cve/CVE-2018-10907","https://nvd.nist.gov/vuln/detail/CVE-2018-10907"],"status":"curated"},{"cvss_vector":"CVSS:3.0/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-59"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2018-10928","cve":"CVE-2018-10928","aliases":[],"title":"GlusterFS (brick, gfs3_symlink_req): Symlink creation is not confined to the volume, so a client plants a link pointing","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"GlusterFS (brick, gfs3_symlink_req)","year":"2018","cvss_score":8.8,"severity":"high","kev":false,"impact":"Symlink creation is not confined to the volume, so a client plants a link pointing at a path on the server outside the gluster tree and then reads or writes through it. That reaches other volumes and the server's own filesystem from inside one tenant's mount.","attack_vector":"Any authenticated gluster client with a mounted volume.","remediation":"Upgrade glusterfs server and restart the bricks, including the CVE-2018-14651 follow-up patch. Verify no residual symlinks pointing outside the brick roots survive the upgrade.","references":["https://access.redhat.com/security/cve/CVE-2018-10928","https://nvd.nist.gov/vuln/detail/CVE-2018-10928"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2018-1244","cve":"CVE-2018-1244","aliases":[],"title":"Dell iDRAC7 / iDRAC8 / iDRAC9 (SNMP agent): Command injection in the iDRAC SNMP agent gives an attacker who already","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC7 / iDRAC8 / iDRAC9 (SNMP agent)","year":"2018","cvss_score":8.8,"severity":"high","kev":false,"impact":"Command injection in the iDRAC SNMP agent gives an attacker who already holds an iDRAC account with configuration rights arbitrary command execution as root on the BMC itself. That is the full BMC prize: out-of-band power control, Virtual Media boot, KVM, and an implant that lives on the service processor and survives every host reimage and OS reinstall. The privilege jump matters - it converts a routine monitoring or configuration credential into persistent control of the node beneath the hypervisor.","attack_vector":"An authenticated iDRAC account holding the Configure iDRAC privilege - the kind of account handed to monitoring tooling, an integrator, or a datacenter-remote-hands team, not a full administrator. Reachability is the management VLAN.","remediation":"Flash to iDRAC7/8 2.60.60.60 or iDRAC9 3.21.21.21 or later - out-of-band, per-node, no host reboot and no job drain. Config-only mitigations that reduce blast radius immediately: disable the iDRAC SNMP agent where you are not actually scraping it, and audit which service accounts hold Configure iDRAC rather than read-only. Original Dell TechCenter advisory URL is dead; NVD carries the version data.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-1244","https://www.dell.com/support/kbdoc/en-us/000131409/dsa-2018-004-idrac-vulnerabilities"],"status":"curated","published":"2018-07-02"},{"id":"CVE-2018-15774","cve":"CVE-2018-15774","aliases":[],"title":"Dell iDRAC9 (Redfish): Redfish interface permission-check flaw enabling privilege escalation to admin","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC9 (Redfish)","year":"2018","cvss_score":8.8,"severity":"high","kev":false,"impact":"Redfish interface permission-check flaw enabling privilege escalation to admin","attack_vector":"Network / Redfish, authenticated low-privilege","remediation":"iDRAC firmware update; illustrates that Redfish RBAC bugs are as common as the IPMI ones they replaced","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-15774"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2018-12-13"},{"id":"CVE-2018-3628","cve":"CVE-2018-3628","aliases":[],"title":"Intel AMT (HTTP handler) in Intel CSME firmware: A buffer overflow in AMT's HTTP handler allows arbitrary code","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel AMT (HTTP handler) in Intel CSME firmware","year":"2018","cvss_score":8.8,"severity":"high","kev":false,"impact":"A buffer overflow in AMT's HTTP handler allows arbitrary code execution with AMT privileges. AMT's HTTP handler is what serves the out-of-band management interface, so exploitation gives the attacker the same below-the-OS control AMT itself has.","attack_vector":"Reachable through the AMT network interface on provisioned machines.","remediation":"Fixed in Intel CSME/SPS firmware, which reaches you as an OEM BIOS or firmware package - not as a microcode or OS update. That means: wait for your server vendor to ship it, drain the node, flash, and reboot. OEM availability is the long pole and routinely lags the Intel advisory by one or more quarters on server platforms. Track it per platform SKU, because vendors ship these unevenly across their own product lines. Immediate compensating control is to unprovision AMT where you do not use it and firewall the AMT ports where you do.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-3628","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00112.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2018-07-10"},{"cvss_vector":"CVSS:3.0/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-732"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2018-5490","cve":"CVE-2018-5490","aliases":[],"title":"NetApp Clustered Data ONTAP export policy enforcement (SMBv2/SMBv3): Export policy rules marked read-only are not","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"NetApp Clustered Data ONTAP export policy enforcement (SMBv2/SMBv3)","year":"2018","cvss_score":8.8,"severity":"high","kev":false,"impact":"Export policy rules marked read-only are not actually enforced, so an authenticated SMB client writes to volumes the operator believed were read-only. Shared read-only dataset shares become writable, which means one tenant can poison training data everyone else consumes.","attack_vector":"An authenticated SMBv2 or SMBv3 client of a Clustered Data ONTAP 8.3 release-candidate system. The attacker needs no more than the read access they were legitimately granted.","remediation":"Move off the affected 8.3 RC builds to a fixed ONTAP release. Afterwards, verify write behaviour empirically against each read-only export rather than trusting the policy display, and check the volumes for unexpected modifications.","references":["https://security.netapp.com/advisory/ntap-20150324-0001/","https://nvd.nist.gov/vuln/detail/CVE-2018-5490"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2018-6247","cve":"CVE-2018-6247","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): Any local account on a Windows GPU host can call the driver's escape","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2018","cvss_score":8.8,"severity":"high","kev":false,"impact":"Any local account on a Windows GPU host can call the driver's escape path and hit a NULL dereference. Best case the box bugchecks and you lose every job on it; NVIDIA also rates privilege escalation as possible, which on a kernel driver means SYSTEM.","attack_vector":"Any local user or service account on the host, including anything running inside a Windows container that has the GPU mapped in.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-6247"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2018-04-02"},{"id":"CVE-2018-6248","cve":"CVE-2018-6248","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): Out-of-bounds kernel read/write from an unprivileged escape call","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2018","cvss_score":8.8,"severity":"high","kev":false,"impact":"Out-of-bounds kernel read/write from an unprivileged escape call - a length value is trusted that should not be. This is the shape that turns into a full SYSTEM escalation with enough effort, and a bugcheck with none.","attack_vector":"Any local user on the host with access to the GPU device, including low-privilege service accounts.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-6248"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2018-04-02"},{"id":"CVE-2018-6249","cve":"CVE-2018-6249","aliases":[],"title":"NVIDIA GPU Display Driver (Windows nvlddmkm.sys + Linux nvidia.ko): NULL dereference in the kernel-mode layer reachable","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver (Windows nvlddmkm.sys + Linux nvidia.ko)","year":"2018","cvss_score":8.8,"severity":"high","kev":false,"impact":"NULL dereference in the kernel-mode layer reachable from unprivileged code on Windows, Linux, FreeBSD and Solaris hosts. Loss of the node at minimum; NVIDIA leaves escalation on the table.","attack_vector":"Any local user with a handle on the GPU device node - which on Linux means anyone in the container that got /dev/nvidia*.","remediation":"Install the fixed GPU Display Driver branch on both Windows and Linux nodes. The kernel component (nvlddmkm.sys / nvidia.ko) cannot be hot-swapped under load, so this is a node drain and reboot per host; restart the container runtime afterwards so mounted driver libraries match the kernel module. No VBIOS or BMC flash.","references":["https://usn.ubuntu.com/3662-1/","https://nvd.nist.gov/vuln/detail/CVE-2018-6249"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2018-04-02"},{"id":"CVE-2018-6250","cve":"CVE-2018-6250","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): Same class as the other early-2018 escape bugs: an unprivileged","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2018","cvss_score":8.8,"severity":"high","kev":false,"impact":"Same class as the other early-2018 escape bugs: an unprivileged caller crashes the kernel driver, with escalation to SYSTEM not ruled out by the vendor.","attack_vector":"Any local user on the Windows host.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-6250"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2018-04-02"},{"id":"CVE-2018-6442","cve":"CVE-2018-6442","aliases":[],"title":"Brocade Fabric OS Webtools (firmware update section): A remote authenticated attacker can abuse the Webtools","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Brocade Fabric OS Webtools (firmware update section)","year":"2018","cvss_score":8.8,"severity":"high","kev":false,"impact":"A remote authenticated attacker can abuse the Webtools firmware-update path. Firmware update on a SAN switch is the persistence mechanism — an attacker who can drive it installs an image that survives every subsequent remediation. Companion issue CVE-2018-6436 does the same through the `firmwaredownload` CLI command for a local attacker.","attack_vector":"Authenticated remote user with Webtools access on FOS before 8.2.1 / 8.1.2f / 8.0.2f / 7.4.2d.","remediation":"Fabric OS upgrade plus reboot. Disable Webtools if your operations are CLI/REST-driven — a live config change. Verify installed firmware digests against Broadcom's published values after any suspicious period, because a patched switch running a tampered image is still compromised.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-6442","https://nvd.nist.gov/vuln/detail/CVE-2018-6436"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2018-11-08"},{"id":"CVE-2018-7807","cve":"CVE-2018-7807","aliases":[],"title":"Schneider Electric Data Center Expert 7.5.0 and earlier - zip upload: A crafted zip uploaded through the DCE UI can","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Schneider Electric Data Center Expert 7.5.0 and earlier - zip upload","year":"2018","cvss_score":8.8,"severity":"high","kev":false,"impact":"A crafted zip uploaded through the DCE UI can path-traverse out of its intended directory and write arbitrary files on the appliance. The exploit path is social as much as technical: an authenticated operator is tricked into uploading a file that looks like a normal DCIM import.","attack_vector":"An authenticated DCE user performing what looks like a routine upload of a supplied file.","remediation":"Upgrade past 7.5.0. Operationally, the lesson generalises to every DCIM: treat imported bundles, device definition packs and 'helpful' vendor-supplied files as untrusted input.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-7807"],"status":"curated","published":"2018-11-30"},{"id":"CVE-2018-9281","cve":"CVE-2018-9281","aliases":[],"title":"Eaton UPS 9PX 8000 SP administration panel: CSRF on the change-password function plus reflected XSS: an attacker forces","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Eaton UPS 9PX 8000 SP administration panel","year":"2018","cvss_score":8.8,"severity":"high","kev":false,"impact":"CSRF on the change-password function plus reflected XSS: an attacker forces a silent password change on the UPS admin account and takes over the device. Once they hold the UPS admin account they control shutdown behaviour and output for whatever the unit feeds - a PHYSICAL outcome from a web bug.","attack_vector":"Requires a logged-in UPS administrator to load an attacker-controlled page.","remediation":"Firmware update on the UPS network card where available. On units this age, availability of a fix is not guaranteed - if none exists, the mitigation is a dedicated management VLAN, no browser access to UPS interfaces from general-purpose workstations, and disabling the web interface if the unit can be managed another way.","references":["https://www.bishopfox.com/news/2018/10/eaton-ups-9px-8000-sp-multiple-vulnerabilities/"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2018-10-24"},{"id":"CVE-2019-0140","cve":"CVE-2019-0140","aliases":["INTEL-SA-00255"],"title":"Intel Ethernet 700 Series Controller firmware (X710/XL710/XXV710): Buffer overflow in the adapter firmware of Intel's","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Ethernet 700 Series Controller firmware (X710/XL710/XXV710)","year":"2019","cvss_score":8.8,"severity":"high","kev":false,"impact":"Buffer overflow in the adapter firmware of Intel's 700-series NICs allowing an *unauthenticated* user to escalate privilege. This is the rare NIC-firmware bug where the published impact is escalation rather than denial of service, and it needs no host account — the attack surface is the network the card is plugged into. A compromised NIC sits below the OS, persists across reinstall, and on many server designs carries the NC-SI sideband to the BMC, so the blast radius extends past the host it lives in.","attack_vector":"Unauthenticated attacker able to reach the adapter over the network. X710/XL710/XXV710 are the standard 10/25/40GbE management and storage NICs in the server generations that host GPUs.","remediation":"Flash 700-series NVM firmware to 7.0 or later via Intel's NVM Update Utility or the OEM firmware bundle. Requires a **cold power cycle** to activate, so it is a per-node drain across the fleet. Audit installed NVM versions first (`ethtool -i`) — 700-series cards are old enough that many fleets have never updated them.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-0140"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2019-11-14"},{"id":"CVE-2019-0169","cve":"CVE-2019-0169","aliases":[],"title":"Intel CSME / TXE: A heap overflow in a CSME subsystem reachable by an unauthenticated attacker for privilege escalation","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel CSME / TXE","year":"2019","cvss_score":8.8,"severity":"high","kev":false,"impact":"A heap overflow in a CSME subsystem reachable by an unauthenticated attacker for privilege escalation. Same shape and same consequence as the other network-reachable CSME overflows: code execution in the engine that sits under the OS.","attack_vector":"Unauthenticated attacker with access to the affected interface.","remediation":"Fixed in Intel CSME/SPS firmware, which reaches you as an OEM BIOS or firmware package - not as a microcode or OS update. That means: wait for your server vendor to ship it, drain the node, flash, and reboot. OEM availability is the long pole and routinely lags the Intel advisory by one or more quarters on server platforms. Track it per platform SKU, because vendors ship these unevenly across their own product lines.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-0169","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00241.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2019-12-18"},{"id":"CVE-2019-13071","cve":"CVE-2019-13071","aliases":[],"title":"CyberPower PowerPanel Business Edition 3.4.0 Agent/Center: Cross-site request forgery across all forms in the web","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"CyberPower PowerPanel Business Edition 3.4.0 Agent/Center","year":"2019","cvss_score":8.8,"severity":"high","kev":false,"impact":"Cross-site request forgery across all forms in the web application. An authenticated facility user visiting an attacker-controlled page silently submits changes to the UPS management configuration - including shutdown behaviour. Old and unglamorous, but this software persists in production far longer than anything on the compute side.","attack_vector":"Requires an authenticated PowerPanel user to visit a page the attacker controls.","remediation":"Upgrade past 3.4.0. If you are still running 3.x, the more useful action is a full inventory of power-management software versions - CyberPower's 2023-2025 CVE history means anything this old is carrying many more unlisted problems.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-13071"],"status":"curated","published":"2019-07-10"},{"id":"CVE-2019-1901","cve":"CVE-2019-1901","aliases":[],"title":"Cisco Nexus 9000 ACI Mode (LLDP subsystem): A buffer overflow in the LLDP subsystem of Nexus 9000 switches in ACI mode","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Cisco Nexus 9000 ACI Mode (LLDP subsystem)","year":"2019","cvss_score":8.8,"severity":"high","kev":false,"impact":"A buffer overflow in the LLDP subsystem of Nexus 9000 switches in ACI mode gives an adjacent unauthenticated attacker denial of service or arbitrary code execution with root privileges on the switch. In ACI, LLDP is not optional — it is how the fabric discovers and validates its own topology — so you cannot simply turn it off the way you can on a standalone NX-OS leaf.","attack_vector":"Unauthenticated, adjacent — a crafted LLDP frame from a device on a leaf port.","remediation":"ACI software upgrade across the fabric (APIC plus switches), staged. There is no good config workaround because ACI depends on LLDP; the compensating control is strict physical and port-admission control on leaf front-panel ports. Related memory-leak issue in the same subsystem: CVE-2023-20089.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-1901","https://nvd.nist.gov/vuln/detail/CVE-2023-20089"],"status":"curated","published":"2019-07-31"},{"id":"CVE-2019-19023","cve":"CVE-2019-19023","aliases":[],"title":"Harbor: Privilege escalation in the Harbor registry","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Harbor","year":"2019","cvss_score":8.8,"severity":"high","kev":false,"impact":"Privilege escalation in the Harbor registry","attack_vector":"Authenticated registry user","remediation":"Upgrade Harbor","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-19023"],"status":"curated","published":"2020-03-20"},{"id":"CVE-2019-19642","cve":"CVE-2019-19642","aliases":[],"title":"Supermicro BMC virtual media subsystem on X8STi-F with IPMI firmware 2.06: The researcher's own description","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC virtual media subsystem on X8STi-F with IPMI firmware 2.06","year":"2019","cvss_score":8.8,"severity":"high","kev":false,"impact":"The researcher's own description is the operator-relevant one: a persistent backdoor on the BMC. Virtual media is the single most dangerous BMC feature to lose control of, because it is the mechanism by which an attacker attaches their own boot image to a node and reboots into it. Combined with command injection in the same subsystem, the attacker gets both code on the controller and the ability to boot the host into whatever they want - the complete out-of-band takeover, surviving any host-side remediation. Shell metacharacters in the ShareHost and ShareName fields of the /rpc/setvmdrive.asp handler reach a command interpreter on the controller.","attack_vector":"An authenticated attacker who can send HTTP requests to the BMC's IP address. Any valid BMC credential plus a route to the management network is sufficient.","remediation":"This is legacy X8-generation hardware and firmware updates for it are effectively unavailable, so treat this as a case where flashing is not a real option. The controls that work are structural: disable virtual media on the BMC where the platform allows it, isolate these BMCs onto a management VLAN with no route from tenant or general corporate networks, and plan the hardware out. If X8-era Supermicro boards are still carrying production workloads, that is a fleet-lifecycle decision, not a patching decision.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-19642","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2019/19xxx/CVE-2019-19642.json"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-78"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2019-4715","cve":"CVE-2019-4715","aliases":[],"title":"IBM Spectrum Scale management GUI: Any authenticated GUI user - including a low-privilege monitoring account - runs","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"IBM Spectrum Scale management GUI","year":"2019","cvss_score":8.8,"severity":"high","kev":false,"impact":"Any authenticated GUI user - including a low-privilege monitoring account - runs commands on the management node. That is full control of the storage management plane: filesets, quotas, exports and audit configuration for every tenant on the cluster.","attack_vector":"HTTP access to the Storage Scale GUI with any valid login. Typically the GUI is on the management network, so a foothold anywhere with management-network reach plus one weak GUI credential is sufficient.","remediation":"Upgrade the GUI to the fixed 4.2/5.0 level named in IBM's bulletin. Separately, take the GUI off any network a tenant workload can route to, and cut back read-only GUI accounts that no longer need to exist.","references":["https://www.ibm.com/support/pages/node/1118913","https://nvd.nist.gov/vuln/detail/CVE-2019-4715"],"status":"curated"},{"id":"CVE-2020-10676","cve":"CVE-2020-10676","aliases":[],"title":"Rancher: Incorrectly applied authorization check lets a namespace be moved into a different project","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Rancher","year":"2020","cvss_score":8.8,"severity":"high","kev":false,"impact":"Incorrectly applied authorization check lets a namespace be moved into a different project; cross-tenant boundary break","attack_vector":"Cluster user with namespace access","remediation":"Upgrade Rancher","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-10676"],"status":"curated","published":"2023-12-12"},{"id":"CVE-2020-11485","cve":"CVE-2020-11485","aliases":[],"title":"NVIDIA DGX BMC (AMI firmware): CSRF in the BMC web application","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVIDIA DGX BMC (AMI firmware)","year":"2020","cvss_score":8.8,"severity":"high","kev":false,"impact":"CSRF in the BMC web application. An operator with a BMC session open in a browser can be made to execute BMC actions by loading an attacker's page - disclosure or code execution against out-of-band management without the attacker ever needing network reach to the BMC themselves. DGX-1 before BMC 3.38.30.","attack_vector":"Any attacker who can get a logged-in BMC administrator to visit a web page. The attacker needs no route to the management network at all - the admin's browser is the route.","remediation":"Flash the DGX BMC firmware from NVIDIA's DGX firmware update container (DGX-1 to 3.38.30 or later, DGX-2 to 1.06.06 or later; DGX A100 per the bulletin's table). A BMC flash does not require the host OS to reboot but drops out-of-band management for several minutes and NVIDIA recommends a host power cycle afterwards, so treat it as a per-node maintenance window. Rotate every BMC and IPMI credential after the flash - flashing does not invalidate secrets an attacker already pulled. Keep BMCs on an isolated management VLAN with no route from tenant or job networks.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-11485"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2020-10-29"},{"id":"CVE-2020-11953","cve":"CVE-2020-11953","aliases":[],"title":"Rittal PDU-3C002DEC rack PDU firmware (through 5.15.40): Arbitrary code execution on the rack PDU","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Rittal PDU-3C002DEC rack PDU firmware (through 5.15.40)","year":"2020","cvss_score":8.8,"severity":"high","kev":false,"impact":"Arbitrary code execution on the rack PDU. Once code runs on the PDU, an attacker has a persistent presence on the OOB network that survives every host reimage in the rack, and direct control of outlet state - so this is both a persistence problem and a PHYSICAL availability problem.","attack_vector":"Network access to the PDU management interface.","remediation":"Firmware flash per PDU. Because code execution means possible implantation, a unit you believe was targeted should be re-flashed from vendor image and its stored credentials rotated, not merely updated.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-11953"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2020-07-14"},{"id":"CVE-2020-12138","cve":"CVE-2020-12138","aliases":[],"title":"AMD ATI atillk64.sys - physical memory mapping driver: The AMD ATI atillk64.sys driver exposes routines that map","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD ATI atillk64.sys - physical memory mapping driver","year":"2020","cvss_score":8.8,"severity":"high","kev":false,"impact":"The AMD ATI atillk64.sys driver exposes routines that map physical memory into a caller's virtual address space, and it lets low-privileged users call them. That is arbitrary physical memory read and write handed to any local user - complete bypass of kernel memory protection with no memory-corruption exploit required, because the driver simply offers the capability. Drivers like this are a favourite BYOVD (bring-your-own-vulnerable-driver) primitive precisely because they are signed and they work as designed.","attack_vector":"Local, low-privileged user with the driver loaded. Windows driver; note that an attacker can also *bring* this driver to a host that never shipped it, which is why it matters even if you do not deploy AMD's Windows tooling.","remediation":"Remove or update the driver. On Windows fleets, add atillk64.sys to your vulnerable-driver blocklist (Microsoft's WDAC blocklist covers this class) rather than relying on it not being installed - the BYOVD path means an attacker supplies the driver themselves. Not applicable to Linux ROCm nodes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-12138"],"status":"curated","tags":["tenant-isolation"],"published":"2020-04-27"},{"id":"CVE-2020-12347","cve":"CVE-2020-12347","aliases":[],"title":"Intel Data Center Manager Console: Improper input validation in the DCM Console lets an authenticated user escalate","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel Data Center Manager Console","year":"2020","cvss_score":8.8,"severity":"high","kev":false,"impact":"Improper input validation in the DCM Console lets an authenticated user escalate privilege over the network. Same management-plane exposure profile as the later DCM issues.","attack_vector":"Authenticated user with network access to the DCM console.","remediation":"Upgrade the Intel Data Center Manager software. This is a management-plane application, so the update is an application upgrade and service restart - no node drain, no firmware, no reboot of managed hosts. The real work is deciding what DCM is allowed to reach: it holds credentials for platform power and telemetry across the fleet, so its network exposure matters more than its version. Upgrade to 3.6.2 or later.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-12347","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00430"],"status":"curated","fleet":{"pain_class":"daemon-restart"},"published":"2020-11-12"},{"id":"CVE-2020-14156","cve":"CVE-2020-14156","aliases":[],"title":"OpenBMC phosphor-host-ipmid (user_channel/passwd_mgr.cpp, /etc/ipmi-pass): The file holding IPMI account passwords","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"OpenBMC phosphor-host-ipmid (user_channel/passwd_mgr.cpp, /etc/ipmi-pass)","year":"2020","cvss_score":8.8,"severity":"high","kev":false,"impact":"The file holding IPMI account passwords is written with permissions that let unprivileged BMC processes read it. Because IPMI passwords are stored in a form the daemon can recover (they have to be, for RMCP+ key derivation), reading this file yields usable BMC administrator credentials rather than hashes to crack. Any minor foothold on the BMC promotes straight to BMC admin, and if the fleet reuses BMC credentials across nodes - which most do - one node's compromise becomes the whole fleet's.","attack_vector":"Requires some code execution on the BMC as any local user. That bar is met by any of the unauthenticated network-daemon bugs in this cluster. Not directly reachable from the host or the network.","remediation":"Fixed upstream in phosphor-host-ipmid in April 2020; on your nodes it means a BMC firmware flash, per node, out-of-band, ODM-gated. The compensating control matters more than the patch: stop reusing BMC credentials across the fleet, rotate them per node, and prefer Redfish local accounts or LDAP over IPMI accounts so the ipmi-pass file has nothing valuable in it. If you disable IPMI over LAN for CVE-2021-39296, that also drains most of the value out of this file.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-14156","https://github.com/openbmc/phosphor-host-ipmid/commit/b265455a2518ece7c004b43c144199ec980fc620","https://github.com/openbmc/openbmc/issues/3670"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2020-06-15"},{"id":"CVE-2020-15046","cve":"CVE-2020-15046","aliases":[],"title":"Supermicro BMC web UI user management (cgi/config_user.cgi, X10DRH-iT): An attacker who gets a logged-in BMC","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC web UI user management (cgi/config_user.cgi, X10DRH-iT)","year":"2020","cvss_score":8.8,"severity":"high","kev":false,"impact":"An attacker who gets a logged-in BMC administrator to load a page silently gains a permanent administrator account on that BMC. The account persists after the victim's session ends, which converts a transient phishing-grade interaction into standing out-of-band control of the node - power, console, virtual media, and a launch point for the firmware-level bugs elsewhere in this list. In fleets where one operator's browser session spans many BMCs, a single page load can seed accounts across dozens of controllers. BIOS 2.0a and IPMI firmware 03.40 - no CSRF protection on the call that creates administrator accounts.","attack_vector":"No network position on the management VLAN needed by the attacker directly - instead they need a BMC administrator with an active session to visit attacker-controlled content. That is an unusually low bar in datacenter operations, where staff routinely have BMC tabs open alongside general browsing.","remediation":"Firmware flash to BMC 03.88 or later and BIOS 3.2 or later on X10DRH-iT. Beyond the flash, the durable control is operational rather than technical: BMC administration should happen from a dedicated management workstation or bastion that does not browse the general internet, and BMC accounts should be audited on a schedule so that an injected administrator gets caught. Audit existing BMC user lists now - this bug leaves evidence, in the form of accounts nobody created.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-15046","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2020/15xxx/CVE-2020-15046.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2020-17389","cve":"CVE-2020-17389","aliases":["ZDI QConvergeConsole"],"title":"Marvell QConvergeConsole (QLogic adapter management): Remote code execution on QConvergeConsole, the management","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Marvell QConvergeConsole (QLogic adapter management)","year":"2020","cvss_score":8.8,"severity":"high","kev":false,"impact":"Remote code execution on QConvergeConsole, the management application for QLogic/Marvell FastLinQ and Fibre Channel adapters. QConvergeConsole is the tool that flashes adapter firmware and configures boot-from-SAN across a fleet, so code execution there is a route to pushing adapter firmware to every server it manages. One of a cluster of near-identical ZDI-reported issues (CVE-2020-17387, CVE-2020-17388, CVE-2020-15642 through CVE-2020-15645) in the same version.","attack_vector":"Remote attacker with authentication to the QConvergeConsole service — the advisory notes the existing authentication mechanism can be bypassed.","remediation":"Upgrade QConvergeConsole past 5.5.0.64. Application upgrade on the management host, no server or switch impact. Better: do not leave a fleet-wide adapter-management console running continuously — stand it up for firmware campaigns and shut it down afterwards, which is a process change that removes a permanently exposed high-value target.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-17389","https://nvd.nist.gov/vuln/detail/CVE-2020-15645"],"status":"curated","published":"2020-08-25"},{"cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-294"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2020-25660","cve":"CVE-2020-25660","aliases":[],"title":"Ceph CephX authentication protocol: CephX does not correctly bind client identity, so an attacker who can capture","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ceph CephX authentication protocol","year":"2020","cvss_score":8.8,"severity":"high","kev":false,"impact":"CephX does not correctly bind client identity, so an attacker who can capture cluster traffic can replay an authentication exchange and act as that client. On a flat storage fabric this lets one tenant impersonate another tenant's OSD/MDS session and reach their data.","attack_vector":"An attacker on the Ceph public or cluster network able to observe and re-send traffic - an adjacent compute node on the same storage VLAN is sufficient.","remediation":"Upgrade to Ceph 14.2.14 / 15.2.6 or later and restart mons, OSDs and MDSes. Enable msgr2 secure mode (ms_cluster_mode=secure, ms_service_mode=secure) so cluster traffic is encrypted and authenticated end to end, and put tenant traffic on a separate L2 domain from the storage fabric.","references":["https://access.redhat.com/security/cve/CVE-2020-25660","https://nvd.nist.gov/vuln/detail/CVE-2020-25660"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2020-5208","cve":"CVE-2020-5208","aliases":[],"title":"ipmitool (IPMI LAN response parsing): Reverses the usual direction of BMC risk: here the management station","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ipmitool (IPMI LAN response parsing)","year":"2020","cvss_score":8.8,"severity":"high","kev":false,"impact":"Reverses the usual direction of BMC risk: here the management station is the victim. ipmitool does not validate data returned by the remote BMC, so a malicious or already-compromised BMC overflows buffers in the tool and executes code on the machine running it. Because operators run ipmitool from automation hosts, in loops, across the whole fleet, and usually as root, one compromised BMC escalates into ownership of the box that holds credentials for every other BMC. That is the fastest path from a single node to the entire management plane.","attack_vector":"Any BMC that your tooling talks to. A single compromised or spoofed BMC on the management network is enough - and BMC-to-management-host is a trust direction almost nobody models.","remediation":"Package update to ipmitool 1.8.19 or later on every management/automation host - no firmware, no reboot, just the package and restarting any long-running collector. Cheap fix, high leverage. Also worth reviewing whether your fleet automation really needs to run ipmitool as root, and whether the credentials it holds are scoped per rack rather than fleet-wide.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5208","https://github.com/ipmitool/ipmitool/security/advisories/GHSA-g659-9qxw-p7cp"],"status":"curated","published":"2020-02-05"},{"id":"CVE-2020-7526","cve":"CVE-2020-7526","aliases":["SEVD-2020-192-01"],"title":"APC PowerChute Business Edition (v9.0.x and earlier): PowerChute runs the shutdown script that fires when a UPS reports","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"APC PowerChute Business Edition (v9.0.x and earlier)","year":"2020","cvss_score":8.8,"severity":"high","kev":false,"impact":"PowerChute runs the shutdown script that fires when a UPS reports a power event. Improper input validation means an attacker can get arbitrary code executed at exactly that moment - during a shutdown, with elevated privilege, on every host running the agent. It is a rare shape of bug: the trigger is a power event you cannot prevent, and the payload runs fleet-wide simultaneously.","attack_vector":"Requires the ability to influence the shutdown script content or the event that invokes it - which in practice means access to the PowerChute management server or the UPS that signals it.","remediation":"Upgrade PowerChute. Agent upgrade across every host that runs it, so this is a fleet-wide package push - schedulable, but it touches every node. Separately, treat shutdown scripts as privileged code: version them, restrict who can edit them, and do not let the UPS management network write to them.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-7526"],"status":"curated","tags":["physical-impact"],"published":"2020-08-31"},{"id":"CVE-2020-7569","cve":"CVE-2020-7569","aliases":["CVE-2020-7572","CVE-2020-7573","CVE-2020-28210"],"title":"Schneider Electric EcoStruxure Building Operation WebReports / WebStation V1.9-V3.1: Authenticated file upload","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Schneider Electric EcoStruxure Building Operation WebReports / WebStation V1.9-V3.1","year":"2020","cvss_score":8.8,"severity":"high","kev":false,"impact":"Authenticated file upload with dangerous type on the WebReports component gives code execution on the EcoStruxure Building Operation server, with an XXE and a broken access-control issue alongside it that help an attacker get there and read server-side files on the way. EBO is Schneider's BAS supervisory platform and it commonly integrates the cooling plant, metering and sometimes access control for the site. Code execution on the EBO server means write authority over the Automation Servers (AS/AS-P) beneath it, which are the devices actually commanding air handling. So the physical consequence is the same as any BAS-server compromise: setpoints and fan commands under attacker control, alarms suppressible, and a hall of high-density GPU racks minutes from thermal shutdown. Worth flagging that Schneider is also the vendor for a lot of the power side in the same buildings, so a compromised EBO server frequently sits inside the same trust boundary as the electrical monitoring.","attack_vector":"Requires an authenticated EBO session for the upload path, so the realistic chain is credential theft or a weak/default operator account followed by upload. The reflected/stored XSS and access-control issues in the same cluster provide the credential-capture step. EBO WebStation is browser-based and frequently published to the corporate network for facilities staff, which is where the credentials get phished.","remediation":"Software upgrade of the EcoStruxure Building Operation server and WebReports to a fixed release per Schneider's SEVD advisories - a server-side change in a normal window, no controller firmware, no cooling downtime. Also enforce MFA on the path to WebStation, remove shared facilities accounts, and take the server off the general corporate segment. In a leased colo the EBO server is the landlord's and shared across tenants: ask for the version, and note that 'authenticated' is a weak barrier when a dozen contractor accounts exist on that server.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-7569","https://nvd.nist.gov/vuln/detail/CVE-2020-7572","https://www.se.com/ww/en/work/support/cybersecurity/security-notifications.jsp"],"status":"curated"},{"id":"CVE-2021-21974","cve":"CVE-2021-21974","aliases":[],"title":"VMware ESXi (OpenSLP): OpenSLP heap overflow - the ESXiArgs ransomware entry point that mass-encrypted thousands","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"VMware ESXi (OpenSLP)","year":"2021","cvss_score":8.8,"severity":"high","kev":true,"impact":"OpenSLP heap overflow - the ESXiArgs ransomware entry point that mass-encrypted thousands of hosts [KEV]","attack_vector":"Unauthenticated network on the same segment","remediation":"Patch + disable SLP. The canonical example of why the hypervisor management plane must never share a broadcast domain with tenants","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-21974"],"status":"curated","published":"2021-02-24"},{"id":"CVE-2021-25296","cve":"CVE-2021-25296","aliases":[],"title":"Nagios XI: OS command injection in the windowswmi config wizard (authenticated)","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Nagios XI","year":"2021","cvss_score":8.8,"severity":"high","kev":true,"impact":"OS command injection in the windowswmi config wizard (authenticated) -> shell on the Nagios host","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; management tooling should be VPN-only","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-25296"],"status":"curated","published":"2021-02-15"},{"id":"CVE-2021-25297","cve":"CVE-2021-25297","aliases":[],"title":"Nagios XI: OS command injection in the switch config wizard","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Nagios XI","year":"2021","cvss_score":8.8,"severity":"high","kev":true,"impact":"OS command injection in the switch config wizard -> server compromise","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-25297"],"status":"curated","published":"2021-02-15"},{"id":"CVE-2021-25298","cve":"CVE-2021-25298","aliases":[],"title":"Nagios XI: OS command injection in the cloud-vm config wizard","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Nagios XI","year":"2021","cvss_score":8.8,"severity":"high","kev":true,"impact":"OS command injection in the cloud-vm config wizard -> server compromise","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-25298"],"status":"curated","published":"2021-02-15"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-306"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2021-25312","cve":"CVE-2021-25312","aliases":["HTCONDOR-2021-0001"],"title":"HTCondor (IDTOKENS authentication): A flaw in IDTOKENS lets a user authenticate as another user or as the condor","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HTCondor (IDTOKENS authentication)","year":"2021","cvss_score":8.8,"severity":"high","kev":false,"impact":"A flaw in IDTOKENS lets a user authenticate as another user or as the condor service itself. Once you are the condor service you own the scheduler - you decide whose jobs run on which GPUs and you can submit work under any tenant's identity.","attack_vector":"A user who can authenticate to an HTCondor daemon with IDTOKENS, which is the recommended modern method and therefore widely enabled.","remediation":"Upgrade to HTCondor 8.9.11 or later and restart all daemons. Rotate the IDTOKEN signing keys after upgrading and reissue tokens - a token minted while the flaw was live cannot be distinguished from a legitimate one.","references":["https://htcondor.org/security/vulnerabilities/HTCONDOR-2021-0001.html","https://nvd.nist.gov/vuln/detail/CVE-2021-25312"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2021-25741","cve":"CVE-2021-25741","aliases":[],"title":"Kubernetes (kubelet): subPath volume mount symlink race gives access to host files and directories outside the volume","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kubelet)","year":"2021","cvss_score":8.8,"severity":"high","kev":false,"impact":"subPath volume mount symlink race gives access to host files and directories outside the volume; host escape","attack_vector":"Cluster user able to create a pod with a subPath mount","remediation":"Rolling kubelet upgrade with node drain; interim mitigation is an admission policy banning subPath","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-25741"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2021-09-20"},{"id":"CVE-2021-31215","cve":"CVE-2021-31215","aliases":[],"title":"Slurm: Environment mishandling in PrologSlurmctld/EpilogSlurmctld gives remote code execution as SlurmUser","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Slurm","year":"2021","cvss_score":8.8,"severity":"high","kev":false,"impact":"Environment mishandling in PrologSlurmctld/EpilogSlurmctld gives remote code execution as SlurmUser","attack_vector":"Any user who can submit a job","remediation":"Upgrade Slurm; audit prolog/epilog scripts","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-31215"],"status":"curated","published":"2021-05-13"},{"id":"CVE-2021-34824","cve":"CVE-2021-34824","aliases":[],"title":"Istio: Gateway/DestinationRule credentialName can read TLS secrets from other namespaces","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Istio","year":"2021","cvss_score":8.8,"severity":"high","kev":false,"impact":"Gateway/DestinationRule credentialName can read TLS secrets from other namespaces; cross-tenant private key theft","attack_vector":"Cluster user with namespace access who can create Gateways","remediation":"Rolling istiod upgrade; rotate all mesh TLS keys","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-34824"],"status":"curated","published":"2021-06-29"},{"id":"CVE-2021-3570","cve":"CVE-2021-3570","aliases":[],"title":"linuxptp / ptp4l (PTP message forwarding): A missing length check when ptp4l forwards a PTP message between ports leaks","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"linuxptp / ptp4l (PTP message forwarding)","year":"2021","cvss_score":8.8,"severity":"high","kev":false,"impact":"A missing length check when ptp4l forwards a PTP message between ports leaks memory contents to a remote attacker and can be pushed into a crash. PTP is the time source for AI clusters that do distributed tracing, RDMA telemetry correlation, or lockstep checkpointing — and ptp4l typically runs as root with raw socket access on every node. An information leak out of that process is a leak out of a root-privileged daemon reachable from the fabric.","attack_vector":"Remote — any host that can send PTP messages to a node running ptp4l as a boundary/transparent clock. PTP is unauthenticated by default, so no credentials are involved and any tenant on the same segment qualifies.","remediation":"Package upgrade of linuxptp and a restart of ptp4l — no reboot needed, and the restart costs a brief loss of clock discipline rather than a node outage. The durable control is to run PTP on a dedicated VLAN that tenant workloads cannot source traffic onto, which is a switch config change.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-3570"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2021-07-09"},{"id":"CVE-2021-36230","cve":"CVE-2021-36230","aliases":[],"title":"Terraform Enterprise: Missing authorization on a subset of run-token API requests","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Terraform Enterprise","year":"2021","cvss_score":8.8,"severity":"high","kev":false,"impact":"Missing authorization on a subset of run-token API requests -> privilege escalation to organization owner","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade TFE to v202107-1+; review org owner membership","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-36230"],"status":"curated","published":"2021-07-20"},{"id":"CVE-2021-39298","cve":"CVE-2021-39298","aliases":[],"title":"AMD System Management Mode (SMM) interrupt handler: A flaw in the AMD SMM interrupt handler lets a high-privilege","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD System Management Mode (SMM) interrupt handler","year":"2021","cvss_score":8.8,"severity":"high","kev":false,"impact":"A flaw in the AMD SMM interrupt handler lets a high-privilege attacker reach System Management Mode and execute arbitrary code there. SMM sits above the hypervisor and above the OS - code running in SMM can read and write all physical memory including SEV-protected regions in some configurations, and it is invisible to every security tool you run. This is the classic 'ring -2' compromise: persistent, undetectable from the OS, and it survives reinstalling everything above it.","attack_vector":"Local, requires high privilege (root) on the host first. Not a tenant-reachable bug, but the payoff for an attacker who already has root is enormous - it converts a recoverable host compromise into an unrecoverable one.","remediation":"Fixed in AMD reference firmware (AGESA) and delivered only as an OEM SBIOS package - the OEM rebuild and requalification means **one to six months of lag**, and on end-of-support platforms possibly never. Applying it is a drain plus full power cycle. There is no OS-level mitigation for an SMM handler bug. On a bare-metal fleet, the compensating control is firmware measurement between tenants: if you cannot attest that SMM code is unchanged, you cannot honestly claim a node was cleaned by reimaging.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-39298","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2022-02-16"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-285"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2021-41137","cve":"CVE-2021-41137","aliases":[],"title":"MinIO (IAM policy engine): A regular user can step outside the policy restrictions applied to them, reaching operations","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"MinIO (IAM policy engine)","year":"2021","cvss_score":8.8,"severity":"high","kev":false,"impact":"A regular user can step outside the policy restrictions applied to them, reaching operations and objects the policy was written to deny. The bucket policy you rely on to keep tenants apart stops being a boundary.","attack_vector":"Any authenticated MinIO user with network access to the S3 endpoint.","remediation":"Upgrade to RELEASE.2021-10-10T16-53-30Z or later and restart the cluster. Re-verify tenant separation with an explicit access test per bucket rather than assuming the policy document is being honoured.","references":["https://github.com/minio/minio/security/advisories/GHSA-v64v-g97p-577c","https://nvd.nist.gov/vuln/detail/CVE-2021-41137"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2021-43858","cve":"CVE-2021-43858","aliases":[],"title":"MinIO: Hand-crafted admin API call updates a user's policy","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"MinIO","year":"2021","cvss_score":8.8,"severity":"high","kev":false,"impact":"Hand-crafted admin API call updates a user's policy -> privilege escalation","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; rotate admin credentials","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-43858"],"status":"curated","published":"2021-12-27"},{"id":"CVE-2021-44142","cve":"CVE-2021-44142","aliases":[],"title":"Samba (SMB gateway): Out-of-bounds heap read/write in vfs_fruit","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Samba (SMB gateway)","year":"2021","cvss_score":8.8,"severity":"high","kev":false,"impact":"Out-of-bounds heap read/write in vfs_fruit -> code execution as root on the file server","attack_vector":"Network (remote)","remediation":"Data-plane: emergency smbd upgrade on every SMB gateway/file node","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-44142"],"status":"curated","published":"2022-02-21"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-863"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2021-45102","cve":"CVE-2021-45102","aliases":["HTCONDOR-2021-0004"],"title":"HTCondor (SciTokens authentication): A SciToken is granted more authorization than the token's scopes should permit.","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HTCondor (SciTokens authentication)","year":"2021","cvss_score":8.8,"severity":"high","kev":false,"impact":"A SciToken is granted more authorization than the token's scopes should permit. Federated sites use SciTokens precisely to bound what a remote submitter may do, so this turns a deliberately narrow delegation into a wide one.","attack_vector":"Anyone holding a valid SciToken accepted by the pool, including remote federation partners.","remediation":"Upgrade to HTCondor 9.0.4 or 9.1.2 and restart the daemons. Re-derive your authorization policy from the token issuer's scopes afterwards rather than assuming the previous mapping was enforced.","references":["https://htcondor.org/security/vulnerabilities/HTCONDOR-2021-0004.html","https://nvd.nist.gov/vuln/detail/CVE-2021-45102"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2021-47390","cve":"CVE-2021-47390","aliases":[],"title":"Linux KVM x86 - stack out-of-bounds in ioapic_write_indirect(): A guest write to the virtual IOAPIC causes a stack","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux KVM x86 - stack out-of-bounds in ioapic_write_indirect()","year":"2021","cvss_score":8.8,"severity":"high","kev":false,"impact":"A guest write to the virtual IOAPIC causes a stack out-of-bounds access in the host kernel, reported by KASAN. At CVSS 8.8 this is a guest-to-host memory corruption primitive reachable by writing to an emulated device every VM has - stack corruption in the hypervisor is the shortest path from one tenant's VM to owning the machine and everything else on it.","attack_vector":"From inside a guest VM, by writing to the emulated IOAPIC. Tenant-reachable with no privilege beyond running a VM.","remediation":"Fixed in the Linux kernel. Distro kernel update plus host reboot - no firmware, no VBIOS. Treat as top priority on any host running untrusted guest VMs.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-47390"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-05-21"},{"id":"CVE-2022-0811","cve":"CVE-2022-0811","aliases":[],"title":"CRI-O: \"cr8escape\": kernel sysctl injection via pod spec gives container escape and arbitrary code execution","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"CRI-O","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"\"cr8escape\": kernel sysctl injection via pod spec gives container escape and arbitrary code execution as root on the node","attack_vector":"Any cluster user who can deploy a pod","remediation":"Emergency CRI-O upgrade on all nodes; drain and recreate pods. Add admission policy blocking sysctl values with newlines","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-0811"],"status":"curated","fleet":{"ubiquity":"Common - CRI-O is the default runtime on OpenShift and on several K8s distros neoclouds resell; less common than containerd but not niche","remediation_pain":"`daemon-restart` - upgrading CRI-O restarts the node's container runtime, which on most configurations restarts all pods on that node, so effectively `node-drain` for GPU workloads mid-training","pain_class":"node-drain","why_fleet_wide":"Anyone who can create a pod (i.e. any tenant with namespace access) sets arbitrary host kernel parameters via `sysctls`, abuses `kernel.core_pattern` and gets root code execution on *any* node in the cluster"},"published":"2022-03-16"},{"id":"CVE-2022-1025","cve":"CVE-2022-1025","aliases":[],"title":"Argo CD: Improper access control lets any user escalate to admin-level","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"Improper access control lets any user escalate to admin-level","attack_vector":"Any authenticated Argo CD user","remediation":"Rolling Argo CD upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-1025"],"status":"curated","published":"2022-07-12"},{"id":"CVE-2022-1227","cve":"CVE-2022-1227","aliases":[],"title":"Podman: Malicious image causes privilege escalation when a user runs `podman top`","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Podman","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"Malicious image causes privilege escalation when a user runs `podman top`","attack_vector":"Malicious image","remediation":"Upgrade Podman","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-1227"],"status":"curated","published":"2022-04-29"},{"id":"CVE-2022-1552","cve":"CVE-2022-1552","aliases":[],"title":"PostgreSQL: Autovacuum, REINDEX, CLUSTER etc. apply protections too late","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"PostgreSQL","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"Autovacuum, REINDEX, CLUSTER etc. apply protections too late -> user code runs privileged","attack_vector":"Network (remote)","remediation":"Control-plane: minor-version upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-1552"],"status":"curated","published":"2022-08-31"},{"id":"CVE-2022-2031","cve":"CVE-2022-2031","aliases":[],"title":"Samba (AD DC): KDC and kpasswd share keys","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Samba (AD DC)","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"KDC and kpasswd share keys -> a user forced to change password can obtain tickets to other services","attack_vector":"Network (remote)","remediation":"Control-plane: DC-only upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-2031"],"status":"curated","published":"2022-08-25"},{"id":"CVE-2022-20824","cve":"CVE-2022-20824","aliases":[],"title":"Cisco NX-OS / FXOS (Cisco Discovery Protocol): Root code execution on the switch from a crafted CDP frame sent","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco NX-OS / FXOS (Cisco Discovery Protocol)","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"Root code execution on the switch from a crafted CDP frame sent by anything on an adjacent link. CDP is a layer-2 protocol that is on by default and is not authenticated, so a compromised server NIC — or a tenant's bare-metal node — can attack the leaf it is cabled to directly. This is the classic 'the switch trusts the host' failure and it is why CDP/LLDP should be off on tenant-facing ports.","attack_vector":"Unauthenticated, adjacent — layer-2 reachability to the switch port. Any host on the link, including a tenant's own machine.","remediation":"NX-OS upgrade plus reload. Immediate config mitigation: disable CDP globally or per-interface on all host-facing ports (`no cdp enable`), which is a live change with no reload and is good hygiene independent of the CVE. Same treatment applies to the older CDP RCEs (CVE-2020-3119, CVE-2020-3172, CVE-2018-0303).","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-20824","https://nvd.nist.gov/vuln/detail/CVE-2020-3119"],"status":"curated","published":"2022-08-25"},{"id":"CVE-2022-23182","cve":"CVE-2022-23182","aliases":[],"title":"Intel Data Center Manager: Improper access control in Data Center Manager lets an unauthenticated attacker","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel Data Center Manager","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"Improper access control in Data Center Manager lets an unauthenticated attacker with adjacent network access escalate privilege. DCM is the fleet-wide power and thermal management plane - it holds credentials to platform management across every node it monitors, so compromising it is a lateral-movement jackpot rather than a single-host problem.","attack_vector":"Unauthenticated attacker on the same network segment as the DCM server. Whether that is a realistic position depends entirely on your management network segmentation.","remediation":"Upgrade the Intel Data Center Manager software. This is a management-plane application, so the update is an application upgrade and service restart - no node drain, no firmware, no reboot of managed hosts. The real work is deciding what DCM is allowed to reach: it holds credentials for platform power and telemetry across the fleet, so its network exposure matters more than its version. Upgrade to 4.1 or later.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-23182","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00662.html"],"status":"curated","fleet":{"pain_class":"daemon-restart"},"published":"2022-08-18"},{"id":"CVE-2022-24842","cve":"CVE-2022-24842","aliases":[],"title":"MinIO: Non-admin user can create service accounts for root/admin users and assume their policies","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"MinIO","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"Non-admin user can create service accounts for root/admin users and assume their policies","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade + rotate all service-account credentials","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-24842"],"status":"curated","published":"2022-04-12"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-287"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2022-26110","cve":"CVE-2022-26110","aliases":["HTCONDOR-2022-0003"],"title":"HTCondor (CLAIMTOBE authentication method): Once a user has authenticated to a daemon with CLAIMTOBE - a method that","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HTCondor (CLAIMTOBE authentication method)","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"Once a user has authenticated to a daemon with CLAIMTOBE - a method that amounts to asserting your own identity - every subsequent command on that connection can claim to be anyone. That includes submitting and controlling jobs as other tenants and issuing administrative commands.","attack_vector":"Anyone able to reach a daemon that lists CLAIMTOBE among its accepted methods. CLAIMTOBE has historically been in the default READ method list, so many pools are exposed without having chosen to be.","remediation":"Upgrade to HTCondor 8.8.16, 9.0.10 or 9.6.0 and restart the daemons. Separately, strike CLAIMTOBE out of SEC_*_AUTHENTICATION_METHODS everywhere - there is no configuration in which it is a real authentication method, and doing this also closes the CLAIMTOBE half of CVE-2019-18823.","references":["https://htcondor.org/security/vulnerabilities/HTCONDOR-2022-0003.html","https://www.debian.org/security/2022/dsa-5144","https://nvd.nist.gov/vuln/detail/CVE-2022-26110"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2022-28639","cve":"CVE-2022-28639","aliases":["HPESBHF04365"],"title":"HPE iLO 5 (adjacent-network code execution / DoS): Arbitrary code execution on the iLO from an adjacent network","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HPE iLO 5 (adjacent-network code execution / DoS)","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"Arbitrary code execution on the iLO from an adjacent network position, with denial of service as the softer outcome. Code execution on the service processor means the attacker owns power, Virtual Media, console and firmware for that node, and can leave an implant that outlives any host reinstall. The DoS variant is its own operational problem on a GPU fleet: losing iLO means losing the only way to power-cycle or console into a wedged training node, so an outage turns into a truck roll.","attack_vector":"Adjacent network - an attacker already on the same management segment as the iLO. That is a compromised jump host, a monitoring collector, another node's BMC, or anything else sharing the OOB VLAN. Not internet-reachable by design, but flat management networks make 'adjacent' mean 'the entire datacenter'.","remediation":"Flash iLO 5 to v2.72 or later (v2.71 and earlier are affected). Out-of-band, per-node, no host reboot and no job drain. The durable control is segmentation: this class of bug is only exploitable from the management network, so the value of putting each rack's BMCs behind their own segment with an explicit allowlist is high and it is a network change rather than a per-node campaign.","references":["https://support.hpe.com/hpsc/doc/public/display?docLocale=en_US&docId=emr_na-hpesbhf04365en_us","https://nvd.nist.gov/vuln/detail/CVE-2022-28639"],"status":"curated","published":"2022-09-20"},{"id":"CVE-2022-29178","cve":"CVE-2022-29178","aliases":[],"title":"Cilium: Incorrect default permissions on Cilium-managed host paths allow privilege escalation","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"Incorrect default permissions on Cilium-managed host paths allow privilege escalation","attack_vector":"Any tenant workload on the node","remediation":"Rolling Cilium agent DaemonSet upgrade; brief per-node dataplane interruption but no GPU pod eviction","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-29178"],"status":"curated","published":"2022-05-20"},{"id":"CVE-2022-29500","cve":"CVE-2022-29500","aliases":[],"title":"Slurm: Incorrect access control leading to information disclosure across users' jobs","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Slurm","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"Incorrect access control leading to information disclosure across users' jobs","attack_vector":"Any user who can submit a job","remediation":"Upgrade Slurm","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-29500"],"status":"curated","fleet":{"ubiquity":"Very common - Slurm is the default scheduler on HPC-style GPU clouds and on most bare-metal H100/GB200 clusters sold to AI labs","remediation_pain":"`daemon-restart` fleet-wide - SchedMD is explicit that the cluster stays vulnerable until *every* slurmdbd, slurmctld and slurmd has restarted, i.e. a coordinated restart across all compute nodes","pain_class":"daemon-restart","why_fleet_wide":"Credential-handling flaw lets an unprivileged user impersonate SlurmUser and then run arbitrary processes as root; affects every Slurm release since 1.0.0, so one bug covers the whole scheduler estate"},"published":"2022-05-05"},{"id":"CVE-2022-29501","cve":"CVE-2022-29501","aliases":[],"title":"Slurm: Incorrect access control leading to privilege escalation and code execution","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Slurm","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"Incorrect access control leading to privilege escalation and code execution","attack_vector":"Any user who can submit a job","remediation":"Upgrade Slurm","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-29501"],"status":"curated","published":"2022-05-05"},{"id":"CVE-2022-30243","cve":"CVE-2022-30243","aliases":["CVE-2022-30242","CVE-2022-30244","CVE-2022-30245"],"title":"Honeywell Alerton Visual Logic, Ascent Control Module (ACM) and Compass 1.6.5: Unauthenticated program writes","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Honeywell Alerton Visual Logic, Ascent Control Module (ACM) and Compass 1.6.5","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"Unauthenticated program writes to the controller. Not configuration - program. An attacker on the network can push new control logic onto an Alerton controller and stop or replace the running program with no verification of who they are. This is the deepest form of BMS compromise available: the attacker is not sending a bad setpoint that an operator might notice and override, they are rewriting the control algorithm so the controller itself now does the wrong thing and reports the right thing. Applied to the controllers sequencing CRAHs or chilled water in a GPU hall, an attacker can write logic that holds fans low, ignores high-temperature alarms, or trips the plant on a delay so the failure looks like a mechanical fault. Thermal shutdown of a 40-140 kW rack row follows in minutes, and the forensics point at the HVAC contractor rather than at an intrusion. The companion issues let configuration be changed the same way, unauthenticated.","attack_vector":"A crafted packet from any host on the controller's network - no authentication exists on the programming path at all. Alerton gear sits on the building/facility VLAN, and the Alerton BACtalk ecosystem is BACnet-based, so anything that can route BACnet to the controller can do this. Physical access to a mechanical room panel is an equally valid path.","remediation":"Effectively unpatchable as a class - the vendor's guidance for this family is defensive positioning rather than a fix that adds authentication to the programming path, because the programming protocol was never designed with any. The only real control is to make the controllers unreachable: isolated VLAN carrying BACnet only, explicit allow-list from the Alerton supervisor (Compass/Envision) and nothing else, no internet path, port security on the switch ports feeding mechanical rooms, and physical locks on control panels. Where Honeywell offers a newer controller generation with authenticated programming, replacing the controllers is the actual fix and it is a capital project. In a leased site, you cannot touch these - require the landlord to attest that BACnet programming traffic is not routable from any tenant or corporate network.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-30243","https://nvd.nist.gov/vuln/detail/CVE-2022-30244","https://github.com/scadafence/Honeywell-Alerton-Vulnerabilities","https://www.honeywell.com/us/en/product-security"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-32552","cve":"CVE-2022-32552","aliases":[],"title":"Pure Storage Purity//FA and Purity//FB restricted shell (Python environment variables): A logged-in user manipulates","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Pure Storage Purity//FA and Purity//FB restricted shell (Python environment variables)","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"A logged-in user manipulates Python environment variables to break out of the restricted array shell into an unrestricted root shell. The restricted shell is the whole boundary between an array operator and the appliance operating system, and it does not hold.","attack_vector":"Any valid login to an affected FlashArray or FlashBlade shell - including a low-privilege operator account handed out for day-to-day array work.","remediation":"Apply Pure's opt-in or manual patch, or upgrade Purity to an unaffected release. Treat every account that had shell access on an affected array as having had root, and rotate anything reachable from the appliance.","references":["https://support.purestorage.com/Pure_Security/Security_Bundle_2022-04-04/Security_Advisory_for_%E2%80%9Csecurity-bundle-2022-04-04","https://nvd.nist.gov/vuln/detail/CVE-2022-32552"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-32553","cve":"CVE-2022-32553","aliases":[],"title":"Pure Storage Purity//FA and Purity//FB restricted shell (environment variables): A second route out of the restricted","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Pure Storage Purity//FA and Purity//FB restricted shell (environment variables)","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"A second route out of the restricted array shell to a root shell, this one through general environment variable manipulation rather than the Python-specific path. Same outcome: an array operator becomes root on the appliance.","attack_vector":"Any valid shell login on an affected FlashArray or FlashBlade.","remediation":"Apply the same Pure patch or Purity upgrade that fixes the sibling issue - they ship together in the 2022-04-04 security bundle. Verify after patching that the restricted shell actually rejects an environment override attempt.","references":["https://support.purestorage.com/Pure_Security/Security_Bundle_2022-04-04/Security_Advisory_for_%E2%80%9Csecurity-bundle-2022-04-04","https://nvd.nist.gov/vuln/detail/CVE-2022-32553"],"status":"curated"},{"id":"CVE-2022-32744","cve":"CVE-2022-32744","aliases":[],"title":"Samba (AD DC): KDC accepts kpasswd requests encrypted with any key it knows","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Samba (AD DC)","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"KDC accepts kpasswd requests encrypted with any key it knows -> change any user's password, full domain takeover","attack_vector":"Network (remote)","remediation":"Control-plane: DC upgrade; force org-wide credential reset","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-32744"],"status":"curated","published":"2022-08-25"},{"id":"CVE-2022-33183","cve":"CVE-2022-33183","aliases":[],"title":"Brocade Fabric OS CLI: A remote authenticated attacker can act beyond their role through the Fabric OS CLI","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Brocade Fabric OS CLI","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"A remote authenticated attacker can act beyond their role through the Fabric OS CLI. SAN switches are routinely given shared operator accounts across a storage team, so role separation on FOS is thin to begin with; this removes what is left. Companion privilege escalation: CVE-2022-33182.","attack_vector":"Authenticated remote user on FOS before 9.1.0 / 9.0.1e / 8.2.3c / 8.2.0cbn5 / 7.4.2.j.","remediation":"Fabric OS upgrade plus reboot per fabric. Move to individual named accounts with RBAC roles rather than a shared switch login — a config and process change worth doing regardless.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-33183","https://nvd.nist.gov/vuln/detail/CVE-2022-33182"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2022-10-25"},{"id":"CVE-2022-34669","cve":"CVE-2022-34669","aliases":[],"title":"NVIDIA GPU Display Driver - Windows user mode layer: The user mode driver layer lets an unprivileged user read","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows user mode layer","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"The user mode driver layer lets an unprivileged user read or modify files critical to the driver, which is a straightforward path to SYSTEM on the node. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5415. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34669","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-73"],"fleet":{"pain_class":"node-drain"},"published":"2022-12-30"},{"id":"CVE-2022-36348","cve":"CVE-2022-36348","aliases":[],"title":"Intel Server Platform Services (SPS) firmware: Active debug code left enabled in shipped SPS firmware lets","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Server Platform Services (SPS) firmware","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"Active debug code left enabled in shipped SPS firmware lets an authenticated user escalate privilege. SPS is the server variant of the management engine - it is the component running on your Xeon nodes, not the consumer CSME - so debug code in production firmware is squarely a datacenter problem.","attack_vector":"Authenticated local access on an affected server.","remediation":"Fixed in Intel SPS firmware, which reaches you as an OEM BIOS or firmware package - not as a microcode or OS update. That means: wait for your server vendor to ship it, drain the node, flash, and reboot. OEM availability is the long pole and routinely lags the Intel advisory by one or more quarters on server platforms. Track it per platform SKU, because vendors ship these unevenly across their own product lines.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-36348","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00718.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2023-02-16"},{"id":"CVE-2022-42309","cve":"CVE-2022-42309","aliases":["XSA-414"],"title":"Xen (xenstored): Guest can crash xenstored, taking down control-plane services for all guests on the host","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen (xenstored)","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"Guest can crash xenstored, taking down control-plane services for all guests on the host","attack_vector":"Tenant VM guest","remediation":"Patch xenstored + restart; noisy-neighbour availability risk rather than a confidentiality break","references":["https://xenbits.xen.org/xsa/advisory-414.html"],"status":"curated","published":"2022-11-01"},{"id":"CVE-2022-4886","cve":"CVE-2022-4886","aliases":[],"title":"ingress-nginx: `log_format` directive bypasses path sanitization","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"`log_format` directive bypasses path sanitization; read arbitrary files including the SA token","attack_vector":"Cluster user who can create Ingress objects","remediation":"Rolling controller upgrade; scope the controller's RBAC down from cluster-wide secret read","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-4886"],"status":"curated","published":"2023-10-25"},{"id":"CVE-2022-48864","cve":"CVE-2022-48864","aliases":["vdpa/mlx5 VIRTIO_NET_CTRL_MQ_VQ_PAIRS_SET missing validation"],"title":"Linux kernel drivers/vdpa/mlx5 (mlx5 vDPA net device): A guest with an assigned mlx5 vDPA net device sends an","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel drivers/vdpa/mlx5 (mlx5 vDPA net device)","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"A guest with an assigned mlx5 vDPA net device sends an unvalidated queue-pair-count control command and panics the host kernel. CVSS scope is Changed - this is a guest breaking out of its own blast radius into the hypervisor. On a multi-tenant node using mlx5 vDPA for accelerated guest networking, one tenant VM takes down every other VM on that host.","attack_vector":"A malicious virtio driver inside a guest VM with an mlx5 vDPA device, or any local process with access to /dev/vhost-vdpa (typically the qemu/kvm group). No host root required.","remediation":"Upgrade the host kernel to 5.17, or a stable backport (5.15.29, 5.16.15) or your distro's patched kernel. Kernel upgrade means a rolling reboot of every hypervisor node using mlx5 vDPA, with live migration or workload drain per node. If you cannot reboot soon, stop exposing vDPA devices to untrusted guests - fall back to SR-IOV VFs or software virtio.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-48864","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2022/CVE-2022-48864.json"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-07-16"},{"cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-49931","cve":"CVE-2022-49931","aliases":[],"title":"Linux kernel (drivers/infiniband/hw/hfi1): The node panics when the fabric link goes down while any sender is waiting","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/hw/hfi1)","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"The node panics when the fabric link goes down while any sender is waiting for send credits. A corrupted list move in the freeze path dereferences a bad pointer in a workqueue, so a single link event turns into a whole-node outage for every tenant on that host - and the CNA rates it as giving confidentiality and integrity impact, not just availability.","attack_vector":"Adjacent-network: anyone able to bounce the Omni-Path link reaches it - a port flap, a switch-side action, a peer resetting the port, or a cable event. No credentials on the host are required and no tenant device node is needed; the trigger is the link transition itself while send waiters are queued. Requires hfi1 hardware (Intel Omni-Path), so this matters on HPC-heritage fabric nodes rather than pure Ethernet/RoCE clusters.","remediation":"Update to 5.4.224, 5.10.154, 5.15 or later. There is no useful interim control - the trigger is a normal fabric event, so schedule the reboot rather than trying to gate access.","references":["https://git.kernel.org/stable/c/25760a41e3802f54aadcc31385543665ab349b8e","https://git.kernel.org/stable/c/7c4260f8f188df32414a5ecad63e8b934c2aa3f0","https://nvd.nist.gov/vuln/detail/CVE-2022-49931"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2023-0184","cve":"CVE-2023-0184","aliases":[],"title":"GPU Display Driver: Local privesc to host root (use-after-free)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"Local privesc to host root (use-after-free)","attack_vector":"Any tenant with a container; also vGPU guest","remediation":"Driver + vGPU Manager upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0184","https://github.com/NVIDIA/product-security/tree/main/2023/5452"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-822"],"fleet":{"pain_class":"node-reboot"},"published":"2023-04-22"},{"id":"CVE-2023-0189","cve":"CVE-2023-0189","aliases":[],"title":"GPU Display Driver: Local privesc to host root (use-after-free)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"Local privesc to host root (use-after-free)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node, evict tenants","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0189","https://github.com/NVIDIA/product-security/tree/main/2023/5452"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-822"],"fleet":{"pain_class":"node-reboot"},"published":"2023-04-01"},{"id":"CVE-2023-22612","cve":"CVE-2023-22612","aliases":["INSYDE-SA-2023019"],"title":"Insyde InsydeH2O (IhisiSmm SMI handler): A malicious host OS calls an Insyde SMI handler with malformed arguments","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (IhisiSmm SMI handler)","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"A malicious host OS calls an Insyde SMI handler with malformed arguments and corrupts SMM memory. IHISI is Insyde's own firmware-services interface, used by BIOS update and configuration tooling, so it is reachable by design from the OS - the bug is that it trusts what it is told. Successful exploitation is ring -2 code execution: firmware persistence, attestation you can no longer believe, and a foothold under the hypervisor.","attack_vector":"Local admin/root on the host OS invoking the IHISI SMI interface with crafted arguments.","remediation":"OEM BIOS update built on the fixed Insyde kernel (5.0-5.5 affected). Firmware flash plus one reboot per node. No config workaround - the interface is a supported firmware service and cannot be disabled. Reduce blast radius by restricting which host-side tools can issue SMIs and by not granting untrusted tenants root on bare metal running unpatched firmware. NCC Group published the underlying research, which is worth reading before assessing your exposure.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-22612","https://www.insyde.com/security-pledge/SA-2023019","https://research.nccgroup.com/2023/04/11/stepping-insyde-system-management-mode/"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-04-11"},{"id":"CVE-2023-23583","cve":"CVE-2023-23583","aliases":[],"title":"Intel CPU (Reptar): Redundant REX-prefix MOVSB causes unpredictable behaviour","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel CPU (Reptar)","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"Redundant REX-prefix MOVSB causes unpredictable behaviour - a guest can hang or crash the entire host, with possible privilege escalation","attack_vector":"Any tenant process in a container; tenant VM guest","remediation":"Microcode update + reboot. Pure availability risk for a neocloud: one tenant can take down a whole node","references":["https://access.redhat.com/security/cve/CVE-2023-23583"],"status":"curated","fleet":{"ubiquity":"very common - broad recent Intel Xeon coverage, the host CPU under most NVIDIA GPU nodes","remediation_pain":"microcode+reboot - Intel shipped microcode, but on many platforms it only lands via a BIOS/UEFI update from the OEM, adding an OEM-image dependency on top of the reboot","pain_class":"microcode + reboot","why_fleet_wide":"An unprivileged instruction sequence can hang or machine-check the host, or potentially escalate privilege - so a single tenant can crash a whole GPU node, and the same instruction works on every node of that Xeon generation."},"published":"2023-11-14"},{"id":"CVE-2023-25194","cve":"CVE-2023-25194","aliases":[],"title":"Apache Kafka Connect: Attacker able to create/modify a connector sets a SASL JAAS JndiLoginModule config","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Apache Kafka Connect","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"Attacker able to create/modify a connector sets a SASL JAAS JndiLoginModule config -> RCE / JNDI SSRF","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade Connect workers; block arbitrary connector configuration","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25194"],"status":"curated","published":"2023-02-07"},{"id":"CVE-2023-25528","cve":"CVE-2023-25528","aliases":[],"title":"DGX H100 BMC (openBMC): RCE on BMC (web server plugin stack overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX H100 BMC (openBMC)","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"RCE on BMC (web server plugin stack overflow)","attack_vector":"Network-adjacent attacker on the management LAN","remediation":"Flash BMC to 23.08.18 out-of-band; segregate BMC network; no tenant eviction required","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5473/5473.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-121"],"published":"2023-09-20"},{"id":"CVE-2023-25548","cve":"CVE-2023-25548","aliases":["SEVD-2023-101-04"],"title":"Schneider Electric StruxureWare Data Center Expert (V7.9.2 and prior) - device credential endpoints: Incorrect","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Schneider Electric StruxureWare Data Center Expert (V7.9.2 and prior) - device credential endpoints","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"Incorrect authorisation lets a low-privileged DCE user read device credentials from endpoints that were never meant to expose them. The read-only NOC account you gave a monitoring contractor becomes the credential set for the entire power and cooling estate.","attack_vector":"Any low-privileged authenticated user on the DCE appliance - including accounts issued to third-party monitoring and maintenance vendors.","remediation":"Upgrade past V7.9.2. Then rotate device credentials and audit who holds DCE accounts. In most operators this audit is the finding: DCE accounts accumulate for vendors, integrators and former staff, and nobody owns the list.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25548"],"status":"curated","published":"2023-04-18"},{"id":"CVE-2023-28410","cve":"CVE-2023-28410","aliases":[],"title":"Intel i915 graphics driver for Linux (kernel < 6.2.10): A memory-buffer bounds failure in the i915 kernel driver that","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel i915 graphics driver for Linux (kernel < 6.2.10)","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"A memory-buffer bounds failure in the i915 kernel driver that an authenticated local user can drive into privilege escalation. i915 is the driver behind Intel integrated and Data Center GPU nodes, and it is reachable from inside any container that has been granted /dev/dri - so this is a container-to-host kernel escape on Intel-GPU nodes.","attack_vector":"Any local user or container with access to the DRM render node. No special hardware access beyond having been scheduled a GPU.","remediation":"Update to a Linux kernel with the fix (6.2.10 or later upstream, or your distro's backport) and reboot. i915 cannot be reloaded under running GPU workloads, so drain the node. No BIOS, firmware or microcode component.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-28410","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00886.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2023-05-10"},{"id":"CVE-2023-28433","cve":"CVE-2023-28433","aliases":[],"title":"MinIO: Windows deployments fail to filter `\\`","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"MinIO","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"Windows deployments fail to filter `\\` -> arbitrary object placement across buckets by a low-privilege key","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; only relevant to Windows-hosted MinIO","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-28433"],"status":"curated","published":"2023-03-22"},{"id":"CVE-2023-28434","cve":"CVE-2023-28434","aliases":[],"title":"MinIO: Crafted request bypasses PostPolicyBucket metadata bucket-name check","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"MinIO","year":"2023","cvss_score":8.8,"severity":"high","kev":true,"impact":"Crafted request bypasses PostPolicyBucket metadata bucket-name check -> object write to any bucket","attack_vector":"Network (remote)","remediation":"Control-plane: gateway upgrade; audit for cross-tenant object writes","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-28434"],"status":"curated","published":"2023-03-22"},{"id":"CVE-2023-32460","cve":"CVE-2023-32460","aliases":["DSA-2023-361"],"title":"Dell PowerEdge Server BIOS (privilege management): An improper privilege-management flaw in PowerEdge BIOS","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell PowerEdge Server BIOS (privilege management)","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"An improper privilege-management flaw in PowerEdge BIOS that an unauthenticated local attacker can use to escalate. 'Unauthenticated local' is the sharp part: it does not require a valid OS login, so a tenant with any code execution path on the box, or someone with brief physical access during a rack move or RMA, can reach it. The prize is the firmware layer - an implant there sits below the hypervisor, survives reimaging and disk replacement, and is invisible to every host-level agent you run.","attack_vector":"Local to the host with no authentication required - a tenant workload that escapes its boundary, a technician with console access during maintenance, or anyone with the box open. Does not touch the management VLAN at all, so OOB network segmentation buys you nothing here.","remediation":"System BIOS update, which is materially more expensive than a BMC flash: the payload can be staged out-of-band through iDRAC or OME, but it only applies on the next host reboot. That means draining running training jobs, or waiting for a natural maintenance window - realistically a scheduled rolling campaign across the fleet, not a same-day fix. Version floors differ per platform; take them from the advisory's per-model table. No config-only mitigation.","references":["https://www.dell.com/support/kbdoc/en-us/000219550/dsa-2023-361-security-update-for-dell-poweredge-server-bios-for-an-improper-privilege-management-security-vulnerability","https://nvd.nist.gov/vuln/detail/CVE-2023-32460"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-12-08"},{"id":"CVE-2023-33412","cve":"CVE-2023-33412","aliases":[],"title":"Supermicro BMC web interface CGI endpoints on X11 and M11 based boards with BMC firmware before 3.17.02","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC web interface CGI endpoints on X11 and M11 based boards with BMC firmware before 3.17.02","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"An authenticated BMC user runs arbitrary commands on the controller. Because X11 boards are the ones most likely to be running years-behind firmware in a rented or secondhand fleet, this is a realistic path from 'we gave the customer an IPMI login' to full out-of-band ownership of the node: power, console, virtual media, and firmware persistence. Any operator who has ever handed a tenant a BMC account on X11 hardware should assume this is exploitable. The X11 generation is the bulk of the older Supermicro GPU and storage fleet still in production at neoclouds.","attack_vector":"An authenticated BMC session over the network, at ordinary user privilege rather than administrator. Reachable by anything routable to the OOB management VLAN that holds any valid credential.","remediation":"Firmware flash to BMC 3.17.02 or later, per board, from Supermicro's December 2023 BMC advisory. On X11 boards this is often a multi-hop upgrade because the fleet is far behind, and some very old X11 SKUs stopped receiving images - for those the honest answer is that there is no fix and the only control is network isolation plus never issuing tenant-facing BMC accounts. Revoke any BMC credentials previously issued to tenants or contractors as part of the same change.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-33412","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2023/33xxx/CVE-2023-33412.json"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2023-33413","cve":"CVE-2023-33413","aliases":[],"title":"Supermicro BMC configuration functionality on X11 and M11 based boards through firmware 3.17.02: Arbitrary command","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC configuration functionality on X11 and M11 based boards through firmware 3.17.02","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"Arbitrary command execution on the BMC from an authenticated session, reached through configuration rather than a web handler. The operational significance of it being a separate path is that disabling or firewalling one interface does not close both - an operator who patched around CVE-2023-33412 without flashing still has this one. Outcome is BMC-level control of the node: power, boot device, console, and persistent firmware residency. A second, distinct command-execution path from the same December 2023 disclosure, this one in the settings surface rather than the CGI endpoints.","attack_vector":"An authenticated remote BMC user reaching the controller's configuration surface over the management network, at ordinary user privilege rather than administrator. Any valid credential on an X11 or M11 BMC below firmware 3.17.02 is enough, which includes the read-only accounts operators hand to monitoring systems and the shared credentials that most whitebox fleets still use across every node.","remediation":"Firmware flash to 3.17.02 or later per board SKU from the December 2023 Supermicro BMC advisory - the same image that fixes CVE-2023-33411 and CVE-2023-33412, so treat all three as one flash campaign rather than three. For X11 boards past end of firmware support, network isolation of the management VLAN is the only remaining control, and that should be documented as an accepted permanent risk rather than a temporary workaround.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-33413","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2023/33xxx/CVE-2023-33413.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-269"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-36628","cve":"CVE-2023-36628","aliases":[],"title":"Pure Storage FlashArray VASA provider: A vSphere or ESXi administrator with VASA access to a FlashArray escalates to","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Pure Storage FlashArray VASA provider","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"A vSphere or ESXi administrator with VASA access to a FlashArray escalates to root on the array itself. A hypervisor admin, who should only be able to manage their own VMs' storage, ends up owning the array for every consumer of it.","attack_vector":"VMware admin rights against a FlashArray registered as a VASA provider. The trust path runs from the virtualization layer into the storage appliance.","remediation":"Apply Pure's VASA security bulletin fix. Separately, reconsider whether the VASA registration should use a scoped service account rather than a broad one, so a hypervisor compromise does not reach array root.","references":["https://support.purestorage.com/Pure_Storage_Technical_Services/Field_Bulletins/Security_Bulletins/Security_Bulletin_for_Privilege_Escalation_in_VASA_CVE-2023-36628","https://nvd.nist.gov/vuln/detail/CVE-2023-36628"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2023-3676","cve":"CVE-2023-3676","aliases":[],"title":"Kubernetes: Command injection via pod spec on Windows nodes","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"Command injection via pod spec on Windows nodes; escalation to node admin","attack_vector":"Cluster user able to create pods on a Windows node","remediation":"Rolling control-plane and kubelet upgrade; Windows node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-3676"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2023-10-31"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-384"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2023-38002","cve":"CVE-2023-38002","aliases":[],"title":"IBM Storage Scale session management: An authenticated user steals or fixates another user's active session and","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"IBM Storage Scale session management","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"An authenticated user steals or fixates another user's active session and inherits their rights. If the victim is a storage administrator, the attacker owns the management plane for the whole cluster.","attack_vector":"Any authenticated user of Storage Scale 5.1.0.0 through 5.1.9.2 who can reach the session-bearing interface. No admin role needed as a starting point.","remediation":"Upgrade to 5.1.9.3 or later. After upgrading, invalidate all existing sessions and force re-authentication - the fix does not retire sessions that were already fixated.","references":["https://www.ibm.com/support/pages/node/7149699","https://nvd.nist.gov/vuln/detail/CVE-2023-38002"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2023-3893","cve":"CVE-2023-3893","aliases":[],"title":"kubernetes-csi-proxy: Insufficient input sanitisation in csi-proxy leads to Windows node admin","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"kubernetes-csi-proxy","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"Insufficient input sanitisation in csi-proxy leads to Windows node admin","attack_vector":"Cluster user able to create pods with CSI volumes on Windows","remediation":"Upgrade csi-proxy; Windows node drain","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2023-11-03"},{"id":"CVE-2023-3955","cve":"CVE-2023-3955","aliases":[],"title":"Kubernetes: Second Windows-node input-sanitisation escalation to admin","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"Second Windows-node input-sanitisation escalation to admin","attack_vector":"Cluster user able to create pods on a Windows node","remediation":"Rolling upgrade; Windows node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-3955"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2023-10-31"},{"id":"CVE-2023-42130","cve":"CVE-2023-42130","aliases":["ZDI-23-1496"],"title":"A10 Thunder ADC (FileMgmtExport): An authenticated attacker can walk outside the intended export directory","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"A10 Thunder ADC (FileMgmtExport)","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"An authenticated attacker can walk outside the intended export directory via the FileMgmtExport component, reading or deleting arbitrary files on the ADC — enough to pull sensitive configuration data or sabotage the device by deleting files it depends on.","attack_vector":"Requires authentication to the management interface, then a crafted path sent to the file-export functionality.","remediation":"Software upgrade to the fixed ACOS release per A10's security advisory. Standard upgrade-and-reboot per Thunder ADC instance; coordinate with failover if this device is in the active traffic path for a cluster.","references":["https://support.a10networks.com/support/security_advisory/a10-acos-file-access-vulnerability/","https://www.zerodayinitiative.com/advisories/ZDI-23-1496/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-03"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-120"],"fleet":{"pain_class":"node-drain"},"id":"CVE-2023-44466","cve":"CVE-2023-44466","aliases":[],"title":"CephFS/RBD kernel client (libceph messenger v2): A signedness bug in net/ceph/messenger_v2.c turns an attacker-chosen","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"CephFS/RBD kernel client (libceph messenger v2)","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"A signedness bug in net/ceph/messenger_v2.c turns an attacker-chosen frame length into a buffer overflow inside the kernel, reachable through HELLO and other early control frames. That is remote code execution in kernel context on every node that mounts CephFS or maps RBD.","attack_vector":"Anything that can complete or spoof the start of a messenger v2 handshake with a client node - a rogue mon/OSD, or an attacker on the storage network able to answer a client's connection.","remediation":"Update to Linux 6.4.5 or later (or a vendor kernel with the backport) on all Ceph client nodes and reboot. Restrict which hosts can reach client nodes on the Ceph ports and keep the storage fabric off tenant-routable networks.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-44466"],"status":"curated"},{"id":"CVE-2023-45234","cve":"CVE-2023-45234","aliases":["PixieFail","VU#132380"],"title":"EDK II NetworkPkg (DHCPv6 DNS Servers option handling): A crafted DNS Servers option inside a DHCPv6 Advertise","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"EDK II NetworkPkg (DHCPv6 DNS Servers option handling)","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"A crafted DNS Servers option inside a DHCPv6 Advertise overflows a firmware buffer, giving memory corruption and a plausible route to pre-boot code execution. Same class of loss as the Server ID overflow: an attacker who lands here is executing inside DXE, above the OS and outside anything the tenant's EDR or attestation agent can see.","attack_vector":"Unauthenticated attacker able to answer DHCPv6 on the provisioning network during the node's PXE boot. No physical access, no host credentials.","remediation":"Firmware flash from the server OEM, not from Tianocore - the upstream edk2 patch has to be rebased by your IBV and then re-qualified by the OEM, which historically takes one to two BIOS release cycles. One reboot per node. Immediate config workaround: disable IPv6 in the UEFI network boot stack, or disable network boot on nodes that do not need it, and restrict who can emit DHCPv6/RA on the deployment VLAN.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-45234","https://kb.cert.org/vuls/id/132380","https://github.com/tianocore/edk2/security/advisories/GHSA-hc6x-cw6p-gj7h"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-01-16"},{"id":"CVE-2023-45235","cve":"CVE-2023-45235","aliases":["PixieFail","VU#132380"],"title":"EDK II NetworkPkg (DHCPv6 proxy Advertise, Server ID option): Buffer overflow in the proxy-DHCPv6 path","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"EDK II NetworkPkg (DHCPv6 proxy Advertise, Server ID option)","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"Buffer overflow in the proxy-DHCPv6 path - the exact path a PXE/HTTP-boot provisioning flow uses when the boot server and the address server are different boxes. Memory corruption in DXE means the attacker can potentially take the node before it has an OS, which for a multi-tenant bare-metal GPU fleet means a persistent foothold that outlives the tenant lease and the reimage.","attack_vector":"Unauthenticated, on-link attacker impersonating or racing the proxy DHCPv6 server on the provisioning segment. Pre-OS.","remediation":"OEM BIOS update, flash and reboot each node. There is no OS-level patch and no runtime mitigation - the vulnerable code runs before the OS. Until the OEM ships, the practical controls are network-side: segment the provisioning VLAN away from tenant traffic, enforce DHCPv6 guard on access ports, and disable UEFI network boot on any node whose boot source is local NVMe.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-45235","https://kb.cert.org/vuls/id/132380","https://github.com/tianocore/edk2/security/advisories/GHSA-hc6x-cw6p-gj7h"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-01-16"},{"id":"CVE-2023-46229","cve":"CVE-2023-46229","aliases":[],"title":"LangChain (recursive URL loader): SSRF — crawling proceeds to internal hosts","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LangChain (recursive URL loader)","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"SSRF — crawling proceeds to internal hosts","attack_vector":"Attacker-supplied or crawled external page","remediation":"Upgrade past 0.0.317; egress-restrict RAG ingestion workers","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-46229"],"status":"curated","published":"2023-10-19"},{"id":"CVE-2023-49935","cve":"CVE-2023-49935","aliases":[],"title":"Slurm: slurmd message-integrity bypass permits reuse of root-level authentication tokens","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Slurm","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"slurmd message-integrity bypass permits reuse of root-level authentication tokens","attack_vector":"Any user who can reach slurmd","remediation":"Upgrade Slurm; rotate munge keys","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-49935"],"status":"curated","published":"2023-12-14"},{"id":"CVE-2023-5178","cve":"CVE-2023-5178","aliases":[],"title":"Linux NVMe-oF (nvmet-tcp): Use-after-free/double-free in nvmet_tcp_free_crypto","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux NVMe-oF (nvmet-tcp)","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"Use-after-free/double-free in nvmet_tcp_free_crypto -> may permit remote code execution on the target","attack_vector":"Network (remote)","remediation":"Data-plane: kernel patch on NVMe-oF targets; storage-fabric maintenance window","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-5178"],"status":"curated","published":"2023-11-01"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-52801","cve":"CVE-2023-52801","aliases":[],"title":"Linux kernel (drivers/iommu/iommufd): Splitting a mapping area - which is what a partial unmap does - leaves the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/iommufd)","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"Splitting a mapping area - which is what a partial unmap does - leaves the domains interval tree pointing at the old node. Upstream states the outcome plainly as a use-after-free. That tree is what maps IOVA ranges to the IOMMU domains they are programmed into, so corrupting it also means invalidation and teardown target the wrong ranges, leaving live DMA windows behind.","attack_vector":"A holder of /dev/iommu issuing an unmap that falls inside an existing mapping on an IOAS with a domain attached - completely routine VMM behaviour, not a crafted edge case. No host root, no special hardware beyond iommufd being in use.","remediation":"The record lists no fixed release; boot a kernel carrying the stable fix commits below. Interim control: do not expose /dev/iommu to tenants.","references":["https://git.kernel.org/stable/c/836db2e7e4565d8218923b3552304a1637e2f28d","https://git.kernel.org/stable/c/fcb32111f01ddf3cbd04644cde1773428e31de6a","https://nvd.nist.gov/vuln/detail/CVE-2023-52801"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-190","CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-52910","cve":"CVE-2023-52910","aliases":[],"title":"Linux kernel (drivers/iommu): The IOVA allocator's retry path overflows, so the lower-bound check is made against zero","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu)","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"The IOVA allocator's retry path overflows, so the lower-bound check is made against zero and an allocation succeeds with an IOVA below the domain's start address. A DMA address gets handed out outside the range the domain was set up to cover - an address the IOMMU domain does not actually own, which is the shape of a DMA window pointing where it should not. Vendor scores it scope-changed.","attack_vector":"Reached when an IOVA allocation request exceeds the size of the domain's address space, or when the node with the largest address is removed and a subsequent oversized allocation runs the retry path. Driven by in-kernel DMA API users through dma-iommu, so it is the host's device drivers on the hot path rather than a tenant ioctl - a tenant influences it only indirectly, through I/O sizes and mapping churn.","remediation":"The record lists no fixed release; boot a kernel carrying the stable fix commits below. No useful interim control - the allocator is shared by every DMA API user on the node.","references":["https://git.kernel.org/stable/c/c929a230c84441e400c32e7b7b4ab763711fb63e","https://git.kernel.org/stable/c/61cbf790e7329ed78877560be7136f0b911bba7f","https://nvd.nist.gov/vuln/detail/CVE-2023-52910"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-787","CWE-125"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-53630","cve":"CVE-2023-53630","aliases":[],"title":"Linux kernel (drivers/iommu/iommufd): An unmap runs off the end of the pinned page list and drops pin counts on pages","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/iommufd)","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"An unmap runs off the end of the pinned page list and drops pin counts on pages that were never part of the tenant's mapping. Those pages belong to the host kernel or to another tenant, and unpinning them lets memory that is still in use be freed and reallocated. Scope-changed corruption driven from one tenant's ioctl.","attack_vector":"A holder of /dev/iommu issuing IOMMU_IOAS_UNMAP while an access object (an emulated/mediated user of the same IOAS) is present on the range - the ordinary VMM flow. syzkaller-reachable from userspace ioctls; no host root and no hardware precondition beyond iommufd being in use.","remediation":"The record lists no fixed release; boot a kernel carrying the stable fix commits below. Interim control: do not hand /dev/iommu to tenants; run passthrough through the host VMM.","references":["https://git.kernel.org/stable/c/70726ce4d898db57bfc4ae30ecd7da63b0dd0aa4","https://git.kernel.org/stable/c/727c28c1cef2bc013d2c8bb6c50e410a3882a04e","https://nvd.nist.gov/vuln/detail/CVE-2023-53630"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-54043","cve":"CVE-2023-54043","aliases":[],"title":"Linux kernel (drivers/iommu/iommufd): The same hardware page table gets linked into an address space's page-table list","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/iommufd)","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"The same hardware page table gets linked into an address space's page-table list twice, corrupting the list. Iteration over that list is what drives map, unmap and invalidation, so a corrupted list means those operations walk freed or wrong entries - stale IOMMU mappings and host memory corruption, scope-changed per the vendor score.","attack_vector":"A holder of /dev/iommu performing an explicit HWPT attach (attaching a device to a specific hardware page table rather than to an IOAS) - the normal flow a VMM uses for a passthrough device. No host root. The upstream note says the in-tree test suite could not cover this path, so it shipped unexercised.","remediation":"The record lists no fixed release; boot a kernel carrying the stable fix commits below. Interim control: keep /dev/iommu out of tenant containers and mediate device attach in the host VMM.","references":["https://git.kernel.org/stable/c/c44adefdcf472f946f0632f4e0ddcbf3e00b8516","https://git.kernel.org/stable/c/b4ff830eca097df51af10a9be29e8cc817327919","https://nvd.nist.gov/vuln/detail/CVE-2023-54043"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-476","CWE-125"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-54060","cve":"CVE-2023-54060","aliases":[],"title":"Linux kernel (drivers/iommu/iommufd): The pfn batch end index is left at zero after a carry, so the unpin path walks an","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/iommufd)","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"The pfn batch end index is left at zero after a carry, so the unpin path walks an entry that was never filled and dereferences garbage. This is inside the page-unpinning code that every tenant DMA unmap and every access-domain teardown goes through, so a tenant can crash or corrupt the host from a normal unmap sequence.","attack_vector":"A holder of /dev/iommu doing ordinary map/unmap plus access-domain destroy. Upstream reproduced it from the in-tree test binary hitting the production batch code; the code path is not test-only. No host root, no hardware precondition beyond iommufd being the passthrough path.","remediation":"The record lists no fixed release; boot a kernel carrying the stable fix commits below. Interim control: keep /dev/iommu out of tenant containers.","references":["https://git.kernel.org/stable/c/176f36a376c417b58d19f79edfce20db9317eaa2","https://git.kernel.org/stable/c/b7c822fa6b7701b17e139f1c562fc24135880ed4","https://nvd.nist.gov/vuln/detail/CVE-2023-54060"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-54262","cve":"CVE-2023-54262","aliases":[],"title":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core/en/tc): Hardware flow-offload rules are programmed from a stale","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core/en/tc)","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"Hardware flow-offload rules are programmed from a stale duplicate of the flow attribute, so when a neighbour update rewrites an encapsulation the driver pushes a freed object into the firmware flow-table command. The use-after-free lands in mlx5_cmd_set_fte, meaning freed kernel memory decides what steering rule the NIC installs - corrupted or attacker-influenced steering state on the device that forwards every tenant's traffic, plus node crashes.","attack_vector":"Driven by neighbour (ARP/ND) update events on the uplink, which any host on the adjacent L2 segment - including a tenant VM or container with its own IP on the fabric - can provoke by changing or churning its MAC-to-IP binding while tunnel-encapsulated TC flows are offloaded. Requires eswitch/switchdev mode with TC hardware offload and encapsulation rules in use, which is the normal configuration on a neocloud node running OVS offload.","remediation":"Update to a kernel carrying the fix on your stream. Interim: disable TC hardware offload on the mlx5 uplink (`ethtool -K <dev> hw-tc-offload off`) or stop using tunnel-encap offload rules, and keep the fabric segment free of untrusted L2 neighbours.","references":["https://git.kernel.org/stable/c/c382b693ffcb1f1ebf60d76ab9dedfe9ea13eedf","https://git.kernel.org/stable/c/8fd1dac646e6b08d03e3f1ad3c5b34255b1e08e8","https://nvd.nist.gov/vuln/detail/CVE-2023-54262"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-362","CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-54318","cve":"CVE-2023-54318","aliases":[],"title":"Linux kernel (net/smc): The IB port-up handler walks the global link-group list without holding its lock, so a fabric","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/smc)","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"The IB port-up handler walks the global link-group list without holding its lock, so a fabric port event racing tenant connect/disconnect activity dereferences a link group that is being added or freed underneath it. The published crash is a NULL dereference in smcr_port_add from the IB port-event worker - a node panic that a peer or a flapping fabric link can provoke while tenants are churning connections.","attack_vector":"Adjacent-network: the trigger is an RDMA port event (link up/flap on the RoCE/IB port), which any fabric-side disturbance produces, raced against link-group list mutation driven by tenant SMC connections. No tenant privilege is needed on either side - the tenants only have to be creating and tearing down AF_SMC connections, which is unprivileged and autoloads the module. Link flaps are routine in a large GPU-cluster fabric, so this is not a contrived race.","remediation":"Boot a kernel carrying the fix commits (takes smc_lgr_list.lock around the iteration). Interim: blacklist smc on nodes not using SMC-R; there is no way to suppress port events on a live fabric, so patching is the real control.","references":["https://git.kernel.org/stable/c/d1c6c93c27a4bf48006ab16cd9b38d85559d7645","https://git.kernel.org/stable/c/70c8d17007dc4a07156b7da44509527990e569b3","https://nvd.nist.gov/vuln/detail/CVE-2023-54318"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-0130","cve":"CVE-2024-0130","aliases":[],"title":"UFM Enterprise / UFM Appliance / UFM CyberAI: Fabric-manager privesc, data corruption, service disruption via improper","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"UFM Enterprise / UFM Appliance / UFM CyberAI","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"Fabric-manager privesc, data corruption, service disruption via improper authentication on the Ethernet mgmt interface","attack_vector":"Network-adjacent attacker on the fabric management network","remediation":"Upgrade UFM; isolate the UFM mgmt interface to a dedicated VLAN; no tenant eviction, but InfiniBand fabric control is at stake","references":["https://github.com/NVIDIA/product-security/tree/main/2024/5584","https://nvd.nist.gov/vuln/detail/CVE-2024-0130"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-287"],"published":"2024-12-06"},{"id":"CVE-2024-10979","cve":"CVE-2024-10979","aliases":[],"title":"PostgreSQL: PL/Perl lets an unprivileged DB user change process env vars (e.g. PATH)","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"PostgreSQL","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"PL/Perl lets an unprivileged DB user change process env vars (e.g. PATH) -> arbitrary code execution","attack_vector":"Network (remote)","remediation":"Control-plane: minor-version upgrade of the control-plane DB","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-10979"],"status":"curated","published":"2024-11-14"},{"id":"CVE-2024-20449","cve":"CVE-2024-20449","aliases":[],"title":"Cisco Nexus Dashboard Fabric Controller (path traversal to RCE via SCP): A low-privileged authenticated attacker","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Cisco Nexus Dashboard Fabric Controller (path traversal to RCE via SCP)","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"A low-privileged authenticated attacker uploads a malicious file over SCP and executes code on the fabric controller.","attack_vector":"Authenticated low-privilege remote access to NDFC.","remediation":"Upgrade NDFC per cisco-sa-ndfc-ptrce-BUSHLbp.","references":["https://sec.cloudapps.cisco.com/security/center/content/CiscoSecurityAdvisory/cisco-sa-ndfc-ptrce-BUSHLbp"],"status":"curated"},{"id":"CVE-2024-20536","cve":"CVE-2024-20536","aliases":[],"title":"Cisco Nexus Dashboard Fabric Controller (SQL injection): A read-only NDFC user executes arbitrary SQL on the controller","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Cisco Nexus Dashboard Fabric Controller (SQL injection)","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"A read-only NDFC user executes arbitrary SQL on the controller database - which holds the fabric's configuration and credentials.","attack_vector":"Authenticated read-only remote access to NDFC.","remediation":"Upgrade NDFC per cisco-sa-ndfc-sqli-CyPPAxrL.","references":["https://sec.cloudapps.cisco.com/security/center/content/CiscoSecurityAdvisory/cisco-sa-ndfc-sqli-CyPPAxrL"],"status":"curated"},{"id":"CVE-2024-21802","cve":"CVE-2024-21802","aliases":[],"title":"llama.cpp / GGUF library: Heap buffer overflow in GGUF `info","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"llama.cpp / GGUF library","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"Heap buffer overflow in GGUF `info->ne` parsing","attack_vector":"Customer-supplied GGUF model file","remediation":"Rebuild any llama.cpp-derived binary; no host patch exists because the parser is statically linked into each tenant build","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21802"],"status":"curated","published":"2024-02-26"},{"id":"CVE-2024-21807","cve":"CVE-2024-21807","aliases":[],"title":"Intel ice driver (Ethernet 800 Series, Linux kernel mode): Improper initialisation in the Linux kernel-mode driver for","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel ice driver (Ethernet 800 Series, Linux kernel mode)","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"Improper initialisation in the Linux kernel-mode driver for Intel 800-series Ethernet, reachable by an authenticated user for privilege escalation. The 800 series (E810) is the NIC under most RoCE/RDMA AI fabrics, so a kernel-mode driver escalation here is host compromise reached from whoever can talk to the network stack - and on nodes exposing SR-IOV VFs or RDMA verbs to tenants, that includes tenants.","attack_vector":"Authenticated local user; on nodes that expose VFs or RDMA devices into containers, that extends to tenant workloads.","remediation":"Fixed in the Intel out-of-tree ice driver (or the equivalent in-kernel version). Updating the driver requires unloading and reloading the module, which drops every link on that NIC - on a node whose RDMA fabric carries collective traffic, that is a job-killing event, so drain first. If you take it via a distro kernel update instead, it is a reboot. No firmware flash for the driver-side fixes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21807","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00918.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-08-14"},{"id":"CVE-2024-21810","cve":"CVE-2024-21810","aliases":[],"title":"Intel ice driver (Ethernet 800 Series, Linux kernel mode): Kernel-mode driver flaw in the Intel 800-series Ethernet","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel ice driver (Ethernet 800 Series, Linux kernel mode)","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"Kernel-mode driver flaw in the Intel 800-series Ethernet Linux driver (improper input validation) giving an authenticated user privilege escalation. Same family and same fix train as the other 2024 ice driver issues; the reason to care is that E810 is the fabric NIC on most GPU nodes and its driver runs in the kernel on the host.","attack_vector":"Authenticated local user, including tenant workloads on nodes that map VFs or RDMA devices into containers.","remediation":"Fixed in the Intel out-of-tree ice driver (or the equivalent in-kernel version). Updating the driver requires unloading and reloading the module, which drops every link on that NIC - on a node whose RDMA fabric carries collective traffic, that is a job-killing event, so drain first. If you take it via a distro kernel update instead, it is a reboot. No firmware flash for the driver-side fixes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21810","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00918.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-08-14"},{"id":"CVE-2024-21825","cve":"CVE-2024-21825","aliases":[],"title":"llama.cpp / GGUF: Heap overflow in `GGUF_TYPE_ARRAY`/`GGUF_TYPE_STRING` parsing","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"llama.cpp / GGUF","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"Heap overflow in `GGUF_TYPE_ARRAY`/`GGUF_TYPE_STRING` parsing","attack_vector":"Customer-supplied GGUF file","remediation":"Rebuild from a patched commit","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21825"],"status":"curated","published":"2024-02-26"},{"id":"CVE-2024-21836","cve":"CVE-2024-21836","aliases":[],"title":"llama.cpp / GGUF: Heap overflow in `header.n_tensors` handling","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"llama.cpp / GGUF","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"Heap overflow in `header.n_tensors` handling","attack_vector":"Customer-supplied GGUF file","remediation":"Rebuild","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21836"],"status":"curated","published":"2024-02-26"},{"id":"CVE-2024-23496","cve":"CVE-2024-23496","aliases":[],"title":"llama.cpp / GGUF: Heap overflow in `gguf_fread_str`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"llama.cpp / GGUF","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"Heap overflow in `gguf_fread_str`","attack_vector":"Customer-supplied GGUF file","remediation":"Rebuild","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-23496"],"status":"curated","published":"2024-02-26"},{"id":"CVE-2024-23497","cve":"CVE-2024-23497","aliases":[],"title":"Intel ice driver (Ethernet 800 Series, Linux kernel mode): Kernel-mode driver flaw in the Intel 800-series Ethernet","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel ice driver (Ethernet 800 Series, Linux kernel mode)","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"Kernel-mode driver flaw in the Intel 800-series Ethernet Linux driver (an out-of-bounds write) giving an authenticated user privilege escalation. Same family and same fix train as the other 2024 ice driver issues; the reason to care is that E810 is the fabric NIC on most GPU nodes and its driver runs in the kernel on the host.","attack_vector":"Authenticated local user, including tenant workloads on nodes that map VFs or RDMA devices into containers.","remediation":"Fixed in the Intel out-of-tree ice driver (or the equivalent in-kernel version). Updating the driver requires unloading and reloading the module, which drops every link on that NIC - on a node whose RDMA fabric carries collective traffic, that is a job-killing event, so drain first. If you take it via a distro kernel update instead, it is a reboot. No firmware flash for the driver-side fixes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-23497","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00918.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-08-14"},{"id":"CVE-2024-23605","cve":"CVE-2024-23605","aliases":[],"title":"llama.cpp / GGUF: Heap overflow in `header.n_kv`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"llama.cpp / GGUF","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"Heap overflow in `header.n_kv`","attack_vector":"Customer-supplied GGUF file","remediation":"Rebuild","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-23605"],"status":"curated","published":"2024-02-26"},{"id":"CVE-2024-23898","cve":"CVE-2024-23898","aliases":[],"title":"Jenkins: No origin validation on the CLI WebSocket endpoint","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Jenkins","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"No origin validation on the CLI WebSocket endpoint -> cross-site WebSocket hijacking, attacker runs CLI commands","attack_vector":"Network (remote)","remediation":"Control-plane: same patch window as CVE-2024-23897","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-23898"],"status":"curated","published":"2024-01-24"},{"id":"CVE-2024-23918","cve":"CVE-2024-23918","aliases":[],"title":"Intel Xeon memory controller configuration (with SGX): An improper conditions check in Xeon memory controller","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Xeon memory controller configuration (with SGX)","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"An improper conditions check in Xeon memory controller configuration under SGX gives a privileged local user privilege escalation. Highest-scored of the memory-controller-plus-SGX family and the one to prioritise if you run SGX on Xeon.","attack_vector":"Privileged local access on the host.","remediation":"OEM platform BIOS update plus TCB recovery and re-attestation. Drain and reboot per node; OEM availability is the long pole.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-23918","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01079.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2024-11-13"},{"id":"CVE-2024-23981","cve":"CVE-2024-23981","aliases":[],"title":"Intel ice driver (Ethernet 800 Series, Linux kernel mode): Kernel-mode driver flaw in the Intel 800-series Ethernet","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel ice driver (Ethernet 800 Series, Linux kernel mode)","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"Kernel-mode driver flaw in the Intel 800-series Ethernet Linux driver (a wrap-around (integer) error) giving an authenticated user privilege escalation. Same family and same fix train as the other 2024 ice driver issues; the reason to care is that E810 is the fabric NIC on most GPU nodes and its driver runs in the kernel on the host.","attack_vector":"Authenticated local user, including tenant workloads on nodes that map VFs or RDMA devices into containers.","remediation":"Fixed in the Intel out-of-tree ice driver (or the equivalent in-kernel version). Updating the driver requires unloading and reloading the module, which drops every link on that NIC - on a node whose RDMA fabric carries collective traffic, that is a job-killing event, so drain first. If you take it via a distro kernel update instead, it is a reboot. No firmware flash for the driver-side fixes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-23981","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00918.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-08-14"},{"id":"CVE-2024-24747","cve":"CVE-2024-24747","aliases":[],"title":"MinIO: Access keys inherit the parent's `admin:*` actions, not just `s3:*`","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"MinIO","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"Access keys inherit the parent's `admin:*` actions, not just `s3:*` -> silent admin rights","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; explicitly deny admin actions in the access-key hierarchy","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-24747"],"status":"curated","published":"2024-01-31"},{"id":"CVE-2024-24986","cve":"CVE-2024-24986","aliases":[],"title":"Intel ice driver (Ethernet 800 Series, Linux kernel mode): Kernel-mode driver flaw in the Intel 800-series Ethernet","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel ice driver (Ethernet 800 Series, Linux kernel mode)","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"Kernel-mode driver flaw in the Intel 800-series Ethernet Linux driver (an access-control failure) giving an authenticated user privilege escalation. Same family and same fix train as the other 2024 ice driver issues; the reason to care is that E810 is the fabric NIC on most GPU nodes and its driver runs in the kernel on the host.","attack_vector":"Authenticated local user, including tenant workloads on nodes that map VFs or RDMA devices into containers.","remediation":"Fixed in the Intel out-of-tree ice driver (or the equivalent in-kernel version). Updating the driver requires unloading and reloading the module, which drops every link on that NIC - on a node whose RDMA fabric carries collective traffic, that is a job-killing event, so drain first. If you take it via a distro kernel update instead, it is a reboot. No firmware flash for the driver-side fixes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-24986","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00918.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-08-14"},{"id":"CVE-2024-25744","cve":"CVE-2024-25744","aliases":["Heckler"],"title":"Linux guest kernel - hypervisor-injected int 0x80 on the 32-bit syscall path (SEV-SNP / SEV-ES, AMD-SB-3008): The","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux guest kernel - hypervisor-injected int 0x80 on the 32-bit syscall path (SEV-SNP / SEV-ES, AMD-SB-3008)","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"The higher-scoring half of Heckler. A malicious hypervisor injects interrupt 0x80 - the legacy 32-bit syscall gate - into a confidential guest at a chosen instruction, and the guest kernel services it as a real syscall. The researchers turned this into an OpenSSH authentication bypass and a sudo-to-root escalation *inside* the confidential VM, with the host never touching guest memory. At 8.8 with a changed scope this is the most severe confidential-guest issue in the set, and it also covers Intel TDX, so it is a property of the interrupt-delivery design rather than an AMD-only bug.","attack_vector":"Malicious or compromised hypervisor injecting interrupts into its own guest. No guest vulnerability and no tenant cooperation needed.","remediation":"**Guest kernel fix, not a host fix** - Linux 6.6.7 / 6.9 and later carry the int80 hardening series. That inverts the usual rollout: patching every hypervisor you own does nothing, because the protection has to live in the tenant's VM image. As the operator, publish a minimum guest kernel for confidential workloads and gate admission on it. A quicker guest-side control is to disable 32-bit x86 emulation in the guest kernel entirely, which removes the int 0x80 gate. The hardware answer - protected/restricted interrupt delivery - had no mainline Linux support at disclosure. No host reboot, no firmware, no BIOS.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-25744","https://ahoi-attacks.github.io/heckler/","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3008.html"],"status":"curated","tags":["tenant-isolation"],"published":"2024-02-12"},{"id":"CVE-2024-2961","cve":"CVE-2024-2961","aliases":[],"title":"glibc (iconv): Out-of-bounds write in the ISO-2022-CN-EXT iconv converter - turns PHP/app file-read bugs into RCE","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"glibc (iconv)","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"Out-of-bounds write in the ISO-2022-CN-EXT iconv converter - turns PHP/app file-read bugs into RCE","attack_vector":"Unauthenticated network (via an app) or local user","remediation":"Package update + restart all consuming services","references":["https://access.redhat.com/security/cve/CVE-2024-2961"],"status":"curated","published":"2024-04-17"},{"id":"CVE-2024-30368","cve":"CVE-2024-30368","aliases":["ZDI-24-524"],"title":"A10 Thunder ADC (CsrRequestView): An authenticated attacker can inject a system-call payload through the CsrRequestView","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"A10 Thunder ADC (CsrRequestView)","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"An authenticated attacker can inject a system-call payload through the CsrRequestView component (used for certificate-signing-request handling), running arbitrary code on the load balancer with the privileges of the vulnerable process.","attack_vector":"Requires authentication first — the advisory doesn't specify a high privilege tier, meaning even a lower-privileged operator account may be enough to trigger it.","remediation":"Software upgrade to the fixed ACOS release per A10's advisory for CVE-2024-30368/CVE-2024-30369 (the two ship together). Upgrade and reboot each Thunder ADC instance; if it's fronting inference traffic, plan for a failover to a standby unit during the upgrade rather than a hard outage.","references":["https://support.a10networks.com/support/security_advisory/cve-2024-30368-cve-2024-30369","https://www.zerodayinitiative.com/advisories/ZDI-24-524/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-06-06"},{"id":"CVE-2024-31856","cve":"CVE-2024-31856","aliases":[],"title":"CyberPower PowerPanel MQTT message handling: An attacker with MQTT publish permissions can craft messages","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"CyberPower PowerPanel MQTT message handling","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"An attacker with MQTT publish permissions can craft messages to all managed PowerPanel devices, achieving SQL injection, arbitrary file write and remote code execution. Combined with the shared-certificate flaw, obtaining those permissions is not a high bar. One compromised device becomes code execution across the power-management estate.","attack_vector":"Anyone able to publish on the PowerPanel MQTT broker - reachable via the shared device certificate, or from any device already on the management network.","remediation":"Vendor upgrade. Separately, put the MQTT broker on a segment reachable only by the devices that must use it, and monitor for publishers you do not recognise. MQTT wildcards should be blocked at the broker (see the companion issue CVE-2024-31409).","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-31856"],"status":"curated","published":"2024-05-15"},{"id":"CVE-2024-33659","cve":"CVE-2024-33659","aliases":[],"title":"AMI AptioV BIOS (improper input validation, SMM): A local attacker overwrites arbitrary memory and executes code at SMM","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI AptioV BIOS (improper input validation, SMM)","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"A local attacker overwrites arbitrary memory and executes code at SMM level with scope change. AptioV is the UEFI firmware under a very large fraction of x86 servers, so this is a cross-OEM firmware issue.","attack_vector":"Local low-privilege access to the server.","remediation":"AMI ships the fix to OEMs, not to you - obtain the updated BIOS from your board/server vendor (Supermicro, Gigabyte, ASRock Rack, Quanta, Tyan etc.) and flash it. Expect a lag of weeks to months between the AMI advisory and an OEM image for your exact SKU, and expect some SKUs never to get one. Cold reboot per node.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/2025/AMI-SA-2025002.pdf"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-35843","cve":"CVE-2024-35843","aliases":[],"title":"Linux kernel (drivers/iommu/intel): The VT-d I/O page-fault reporting path looks up the faulting device with no","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/intel)","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"The VT-d I/O page-fault reporting path looks up the faulting device with no synchronisation against the IOMMU release path, so a device can report a fault against a struct that is being freed underneath it. Use-after-free in host kernel context, with the timing of one side driven by a device a tenant controls and the other by a device release the tenant triggers when it tears down its VM.","attack_vector":"A tenant with an ATS/PRI-capable assigned device keeps the device emitting I/O page faults while releasing it - stopping the VM, unbinding, or letting the control plane reclaim the GPU/NIC. The fault report races the IOMMU's device-release path and lands on freed fault parameters. Requires Intel VT-d with PRI enabled on the assigned device; no host root. Note this is a teardown-window race, so it needs the tenant to control both the fault stream and the release, which a normal VM stop provides.","remediation":"No fixed release is listed in this record; apply the linked stable commits or run a current stable/LTS kernel. Interim: disable PRI on tenant-assigned devices, and quiesce the device (stop the guest driver, disable ATS) before unbinding it during reclaim rather than pulling it while faults are in flight.","references":["https://git.kernel.org/stable/c/3d39238991e745c5df85785604f037f35d9d1b15","https://git.kernel.org/stable/c/def054b01a867822254e1dda13d587f5c7a99e2a","https://nvd.nist.gov/vuln/detail/CVE-2024-35843"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-36324","cve":"CVE-2024-36324","aliases":[],"title":"AMD Graphics Driver - crafted pointer leading to arbitrary code execution: Improper input validation in the AMD","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD Graphics Driver - crafted pointer leading to arbitrary code execution","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"Improper input validation in the AMD graphics driver lets an attacker supply a crafted pointer and reach arbitrary code execution. At 8.8 this is the most severe of AMD's own graphics-driver advisories in the window: a pointer that crosses the driver boundary unvalidated means kernel-level execution from whatever context can issue the call.","attack_vector":"Local, via the graphics driver interface - reachable by a process holding the GPU device.","remediation":"Update the AMD graphics driver package, then reload the driver or reboot. Confirm the fixed version is present in the ROCm/amdgpu build you actually deploy, since the packaged AMD driver and the mainline kernel driver move on different schedules.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36324","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-02-11"},{"id":"CVE-2024-37032","cve":"CVE-2024-37032","aliases":["Probllama"],"title":"Ollama: Path traversal in the digest field","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ollama","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"Path traversal in the digest field → arbitrary file overwrite → RCE","attack_vector":"Unauthenticated network to the Ollama API (`/api/pull` from an attacker-controlled registry)","remediation":"Upgrade to 0.1.34+. Ollama binds 11434 with no auth by default — never expose to a tenant network","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-37032"],"status":"curated","fleet":{"ubiquity":"Common - Ollama is the standard single-node LLM server on GPU dev boxes and small inference tenants; Docker installs run the API as root","remediation_pain":"`daemon-restart` - upgrade to 0.1.34+; trivial on the daemon, but every compromised host is root-owned and needs rebuild","pain_class":"daemon-restart","why_fleet_wide":"Unvalidated `digest` in OCI manifests gives path traversal; a rogue registry writes `/etc/ld.so.preload` and gets unauthenticated RCE as root - a poisoned model pull compromises every node that pulled it"},"published":"2024-05-31"},{"id":"CVE-2024-37052","cve":"CVE-2024-37052","aliases":[],"title":"MLflow (model flavors): Deserialization RCE from a maliciously uploaded model (one of a family: 37052–37060)","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"MLflow (model flavors)","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"Deserialization RCE from a maliciously uploaded model (one of a family: 37052–37060)","attack_vector":"Customer-supplied model artifact loaded via `mlflow.pyfunc.load_model`","remediation":"No format fix — every MLflow model flavor wraps pickle. Restrict who can register models and sandbox model-load workers","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-37052"],"status":"curated","published":"2024-06-04"},{"id":"CVE-2024-37061","cve":"CVE-2024-37061","aliases":[],"title":"MLflow (recipes / pyfunc): RCE via a maliciously crafted MLproject","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"MLflow (recipes / pyfunc)","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"RCE via a maliciously crafted MLproject","attack_vector":"Customer-supplied MLproject/recipe run by the platform","remediation":"Upgrade. If the provider runs a managed MLflow, tenant projects execute in the provider plane","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-37061"],"status":"curated","published":"2024-06-04"},{"id":"CVE-2024-43044","cve":"CVE-2024-43044","aliases":[],"title":"Jenkins: Agent processes can read arbitrary controller files via ClassLoaderProxy#fetchJar","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Jenkins","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"Agent processes can read arbitrary controller files via ClassLoaderProxy#fetchJar","attack_vector":"Network (remote)","remediation":"Control-plane: controller upgrade; treat build agents as untrusted","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-43044"],"status":"curated","published":"2024-08-07"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-670","CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-44994","cve":"CVE-2024-44994","aliases":[],"title":"Linux kernel (drivers/iommu): A dropped return statement made the IOMMU fault handler process a partial PRI","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu)","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"A dropped return statement made the IOMMU fault handler process a partial PRI page-request group as if it were complete, instead of collecting it and waiting. The device decides when a group is partial, so a tenant driving a PRI-capable passthrough device can emit the right sequence of page requests and crash the host kernel from inside its own VM or container.","attack_vector":"The fault is generated by the device, and in a passthrough cluster the device is under tenant control. A tenant with an ATS/PRI-capable assigned device (SVM-capable GPU, PRI-capable NIC, DSA/IAA accelerator) issues page requests marked as part of a multi-request group; the host handler mishandles the partial group and eventually crashes. No host root, no fabric access needed - just the assigned device. Conditional on PRI/IOPF being enabled for devices handed to tenants.","remediation":"No fixed release is listed in this record; apply the linked stable commits or run a current stable/LTS kernel. Interim: disable PRI/ATS on tenant-assigned devices where the workload does not require demand paging, which removes the fault-reporting path entirely.","references":["https://git.kernel.org/stable/c/cc6bc2ab1663ec9353636416af22452b078510e9","https://git.kernel.org/stable/c/fca5b78511e98bdff2cdd55c172b23200a7b3404","https://nvd.nist.gov/vuln/detail/CVE-2024-44994"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-48013","cve":"CVE-2024-48013","aliases":[],"title":"Dell SmartFabric OS10 (execution with unnecessary privileges): A low-privileged attacker escalates through an OS10","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell SmartFabric OS10 (execution with unnecessary privileges)","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"A low-privileged attacker escalates through an OS10 component running with more privilege than it needs. OS10 is a Linux-based NOS, so escalation here is root on a box that programs the forwarding ASIC — arbitrary control over which tenant's traffic goes where. Sits alongside a long series of OS10 command-injection findings (CVE-2024-48830, CVE-2024-49557, CVE-2024-49560, CVE-2025-22472, CVE-2025-22473, CVE-2025-46427, CVE-2025-46428) that all give a low-privileged local or remote user a path to root.","attack_vector":"Low-privileged attacker with access to the switch, versions 10.5.4.x through 10.6.0.x.","remediation":"OS10 upgrade plus reload. Because so many of these share the same precondition — a low-privileged account on the switch — the highest-leverage control is eliminating low-privilege switch accounts entirely and driving all changes through an automation account on a bastion.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-48013","https://nvd.nist.gov/vuln/detail/CVE-2025-46427","https://nvd.nist.gov/vuln/detail/CVE-2025-46428"],"status":"curated","published":"2025-03-17"},{"id":"CVE-2024-49559","cve":"CVE-2024-49559","aliases":[],"title":"Dell SmartFabric OS10 (default password): A default password in SmartFabric OS10 across 10.5.4.x through 10.6.0.x","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell SmartFabric OS10 (default password)","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"A default password in SmartFabric OS10 across 10.5.4.x through 10.6.0.x, usable remotely by a low-privileged attacker to escalate. Default credentials on a datacenter switch are the same failure the BMC world has been fighting for a decade — the switch arrives with a working account nobody in the deployment checklist knows to remove.","attack_vector":"Low-privileged attacker with remote access to the switch.","remediation":"OS10 upgrade plus reload. Also add a switch-intake step that enumerates and disables every non-provisioned local account, the same way you rotate BMC credentials at rack intake — a process change that catches the next one of these before an advisory does.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49559"],"status":"curated","published":"2025-03-17"},{"id":"CVE-2024-50627","cve":"CVE-2024-50627","aliases":[],"title":"Digi ConnectPort LTS (before 1.4.12): An attacker who can reach the ConnectPort LTS's file-upload feature","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Digi ConnectPort LTS (before 1.4.12)","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"An attacker who can reach the ConnectPort LTS's file-upload feature with only limited/local-network privileges can upload and execute a malicious file, escalating straight to full control of the device — which puts them on the cellular/serial gateway path some fleets use for out-of-band access when the primary network is down.","attack_vector":"Requires local-area-network reach to the device's web upload feature and some baseline permission level (not full admin); no internet-facing exposure needed if the device is on a segmented OOB VLAN, but that's exactly the network this appliance is meant to be reachable from.","remediation":"Software upgrade to ConnectPort LTS firmware 1.4.12 or later. This is a firmware flash per unit; because ConnectPort LTS is frequently the fallback OOB path when the primary network is unavailable, schedule the upgrade during planned maintenance rather than waiting for an outage to force it.","references":["https://www.digi.com/getattachment/Resources/Security/Alerts/Digi-ConnectPort-LTS-Firmware-Update/ConnectPort-LTS-KB.pdf","https://www.digi.com/resources/security"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-12-09"},{"id":"CVE-2024-5187","cve":"CVE-2024-5187","aliases":[],"title":"ONNX (`download_model_with_test_data`): Arbitrary file overwrite from a crafted model archive","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"ONNX (`download_model_with_test_data`)","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"Arbitrary file overwrite from a crafted model archive","attack_vector":"Customer-supplied model URL/archive","remediation":"Upgrade; do not run model-zoo download helpers as root or with shared mounts","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-5187"],"status":"curated","published":"2024-06-06"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-53135","cve":"CVE-2024-53135","aliases":[],"title":"Linux kernel (arch/x86/kvm/vmx): KVM's guest/host-mode Intel PT virtualization was broken end to end and the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (arch/x86/kvm/vmx)","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"KVM's guest/host-mode Intel PT virtualization was broken end to end and the maintainers say the bugs put host stability and health at risk. KVM trusts the guest's CPUID configuration to decide which RTIT MSRs to save and load, so a VM enumerating more address ranges than hardware supports drives the host into passing through, saving and loading non-existent MSRs - WARN storms, ToPA errors on the host, and a potential host deadlock. The feature was disabled outright rather than fixed.","attack_vector":"Requires the host to run kvm_intel with pt_mode=host_guest, which is not the default, and the VMM to expose Intel PT to the guest. Given that, the guest's own CPUID and RTIT_CTL programming drives the broken host paths from inside the VM. Nodes left on the default system-wide PT mode are not exposed.","remediation":"Update to a kernel where PT guest/host mode is buried behind CONFIG_BROKEN. Interim and permanent control: never set kvm_intel.pt_mode=host_guest, and keep Intel PT out of tenant guest CPUID.","references":["https://git.kernel.org/stable/c/b8a1d572478b6f239061ee9578b2451bf2f021c2","https://git.kernel.org/stable/c/d4b42f926adcce4e5ec193c714afd9d37bba8e5b","https://nvd.nist.gov/vuln/detail/CVE-2024-53135"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-53224","cve":"CVE-2024-53224","aliases":[],"title":"Linux kernel mlx5_ib (pkey change notifier): A race between InfiniBand device deregistration and the pkey-change work","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_ib (pkey change notifier)","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"A race between InfiniBand device deregistration and the pkey-change work item leaves the handler running after the device is gone, causing a NULL dereference. Pkey changes are pushed by the subnet manager, so a fabric-side event - benign reconfiguration or a hostile SM - can crash hosts that are cycling their RDMA devices.","attack_vector":"Adjacent, unauthenticated: an actor able to cause partition-key change events on the subnet (a compromised or spoofed subnet manager) combined with device teardown on the target.","remediation":"Upgrade the host kernel to 6.13 or a stable backport (6.6.64, 6.11.11, 6.12.2). Rolling reboot of RDMA hosts. Complementary control: lock down who can run a subnet manager on the fabric and set SM priority so a rogue SM cannot take over.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53224","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2024/CVE-2024-53224.json"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-12-27"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-56669","cve":"CVE-2024-56669","aliases":[],"title":"Linux kernel (drivers/iommu/intel): Use-after-free of VT-d cache-tag objects. Device-TLB cache tags outlive the IOMMU","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/intel)","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"Use-after-free of VT-d cache-tag objects. Device-TLB cache tags outlive the IOMMU domain they belong to, so a later IOTLB flush walks freed kernel memory. The upstream report reproduces it exactly as an operator would hit it - several VFs from different PFs passed through to one userspace process - and the crash arrives via the tenant's own DMA unmap ioctl. Scope-changed: a tenant's ordinary unmap corrupts host kernel state.","attack_vector":"A tenant process or VM holding /dev/vfio/* with one or more ATS-capable VFs assigned. The freed cache tag is reached from vfio_iommu_type1's VFIO_IOMMU_UNMAP_DMA path, i.e. an ioctl the tenant issues itself. Conditional on Intel VT-d with device-TLB/ATS enabled on the assigned functions. No host root required.","remediation":"The record lists no fixed release; boot a kernel carrying the stable fix commits below. Interim controls: avoid assigning VFs from multiple different PFs to a single tenant, and disable ATS on assigned functions where the device allows it.","references":["https://git.kernel.org/stable/c/9a0a72d3ed919ebe6491f527630998be053151d8","https://git.kernel.org/stable/c/1f2557e08a617a4b5e92a48a1a9a6f86621def18","https://nvd.nist.gov/vuln/detail/CVE-2024-56669"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-5948","cve":"CVE-2024-5948","aliases":["CVE-2024-5947","CVE-2024-5949","CVE-2024-5950","CVE-2024-5951","CVE-2024-5952","ZDI-24-672"],"title":"Deep Sea Electronics DSE855 generator communications gateway: Six unauthenticated flaws in one device: two stack-based","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Deep Sea Electronics DSE855 generator communications gateway","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"Six unauthenticated flaws in one device: two stack-based buffer overflows giving remote code execution, an infinite-loop DoS, a configuration-backup disclosure that leaks the device's stored settings and credentials, and missing authentication on both factory-reset and restart. The reset and restart issues deserve particular attention because they need no exploitation skill at all - a network-adjacent attacker simply asks the gateway to factory-reset itself, and the link between the generator controllers and the monitoring system is gone along with its configuration. Code execution gives persistence on a device sitting inside the electrical infrastructure segment. For an operator running GPU racks, the consequence is the same class as any standby-power monitoring failure: you find out your generators did not start when the hall goes dark and the cooling stops, and the accelerators cross thermal shutdown minutes later. The configuration disclosure additionally hands over credentials that are usually reused across the site's other DSE and BMS gear.","attack_vector":"Network-adjacent and unauthenticated for all six - no credentials required for any of them, including the destructive ones. The device's web service on the facility or electrical-infrastructure VLAN is the entire attack surface. As with the newer DSE855 issue, the generator contractor's remote-monitoring path is the most likely route in from outside.","remediation":"Firmware update from Deep Sea Electronics. The unit is small and the flash is fast, but ownership is the friction: these are usually specified, installed and maintained by the generator vendor, not by the datacenter's own team, so the change has to go through that contract. Given that a factory reset can be triggered by anyone who can reach the device, an ACL restricting the gateway's web port to the monitoring server is a same-day control worth taking regardless of patch status. Keep an offline copy of the gateway configuration so a triggered factory reset is a ten-minute restore rather than a contractor callout.","references":["https://www.zerodayinitiative.com/advisories/ZDI-24-672/","https://www.zerodayinitiative.com/advisories/ZDI-24-675/","https://nvd.nist.gov/vuln/detail/CVE-2024-5948","https://nvd.nist.gov/vuln/detail/CVE-2024-5951"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-6983","cve":"CVE-2024-6983","aliases":[],"title":"LocalAI: RCE — the backend accepts inputs beyond the config path","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LocalAI","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"RCE — the backend accepts inputs beyond the config path","attack_vector":"Unauthenticated/low-privilege network to the LocalAI API","remediation":"Upgrade past 2.17.1","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-6983"],"status":"curated","published":"2024-09-27"},{"id":"CVE-2024-7348","cve":"CVE-2024-7348","aliases":[],"title":"PostgreSQL: TOCTOU race in pg_dump","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"PostgreSQL","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"TOCTOU race in pg_dump -> an object creator runs arbitrary SQL as the (often superuser) dump operator","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; stop running scheduled pg_dump as superuser","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-7348"],"status":"curated","published":"2024-08-08"},{"id":"CVE-2024-7646","cve":"CVE-2024-7646","aliases":[],"title":"ingress-nginx: Annotation validation bypass reaching config injection and cluster-wide secret access","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"Annotation validation bypass reaching config injection and cluster-wide secret access","attack_vector":"Cluster user who can create Ingress objects","remediation":"Rolling controller upgrade, no GPU drain","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated","published":"2024-08-16"},{"id":"CVE-2025-0657","cve":"CVE-2025-0657","aliases":["CVE-2025-0658"],"title":"Automated Logic / Carrier i-Vu Gen5 BACnet router (drv_gen5_106-01-2380) and i-Vu Zone Controller: Malformed BACnet","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Automated Logic / Carrier i-Vu Gen5 BACnet router (drv_gen5_106-01-2380) and i-Vu Zone Controller","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"Malformed BACnet MS/TP frames put the router and the zone controllers into a fault state, and the vendor is explicit that recovery requires a manual power cycle - on the zone controller a second packet after reset leaves it permanently unresponsive until someone physically touches it. That is the important operator detail: this is not a reboot-and-recover DoS, it is a bricking-until-truck-roll DoS against the devices that command air handling in the hall. Lose the zone controllers and the affected zone stops modulating; depending on the failsafe wiring you either get fans stuck at last-known state or dampers closed. Either way you have lost closed-loop thermal control over a GPU hall and you are now running on whatever the mechanical failsafe does, with a technician on the way. For a 40 kW+ rack density that is a race you can lose. Expect this to hit during the worst possible moment because an attacker will fire it while the plant is already at high load.","attack_vector":"An attacker on the BACnet MS/TP segment, or on any BACnet/IP segment that routes onto it. MS/TP is an RS-485 serial bus, so pure MS/TP access means physical proximity to the field wiring - but the Gen5 router exists precisely to bridge IP to MS/TP, and that is the exposed side. Anyone who can send BACnet/IP to the router can reach the serial devices behind it. On the facility VLAN this needs no credentials because BACnet has no authentication to begin with.","remediation":"Driver/firmware update from Automated Logic or Carrier for the Gen5 router and controllers. Realistically this means the controls contractor on site with a laptop touching every controller, during a maintenance window, on live cooling - which is exactly the work most operators defer. Because the fix is slow, the compensating control is the one that matters: no arbitrary host should be able to originate BACnet/IP toward the router. Put the BACnet segment behind an allow-list at the switch or a firewall that permits only the BMS supervisor's address, and disable BACnet routing between building zones that do not need to talk. If you lease, you cannot flash the landlord's controllers - get the driver version in writing and require them to schedule it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0657","https://nvd.nist.gov/vuln/detail/CVE-2025-0658","https://www.corporate.carrier.com/product-security/advisories-resources/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-1097","cve":"CVE-2025-1097","aliases":[],"title":"ingress-nginx: Config injection via unsanitized auth-tls-match-cn annotation","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"Config injection via unsanitized auth-tls-match-cn annotation","attack_vector":"Cluster user who can create Ingress objects","remediation":"Rolling controller upgrade","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated","published":"2025-03-25"},{"id":"CVE-2025-1098","cve":"CVE-2025-1098","aliases":[],"title":"ingress-nginx: Config injection via unsanitized mirror annotations","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"Config injection via unsanitized mirror annotations","attack_vector":"Cluster user who can create Ingress objects","remediation":"Rolling controller upgrade","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated","published":"2025-03-25"},{"id":"CVE-2025-15566","cve":"CVE-2025-15566","aliases":[],"title":"ingress-nginx: Config injection via the auth-proxy-set-headers annotation","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"Config injection via the auth-proxy-set-headers annotation","attack_vector":"Cluster user who can create Ingress objects","remediation":"Rolling controller upgrade","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated","published":"2026-02-06"},{"id":"CVE-2025-22086","cve":"CVE-2025-22086","aliases":["RDMA/mlx5 fix mlx5_poll_one() cur_qp update flow"],"title":"Linux kernel mlx5_ib (InfiniBand/RoCE completion queue polling): mlx5_poll_one() compares the firmware's QP number","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel mlx5_ib (InfiniBand/RoCE completion queue polling)","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"mlx5_poll_one() compares the firmware's QP number against the wrong structure's QP number, so the wrong queue pair is used to handle a completion, leading to a NULL dereference. Anyone already on the InfiniBand subnet can drive it: unsolicited SMP/GMP/CM management datagrams land on QP0/QP1 and generate the completions, and MAD reception on an IB fabric is entirely unauthenticated. On a shared InfiniBand fabric this is a way for any attached node to crash other nodes' RDMA stacks.","attack_vector":"Any node on the same InfiniBand subnet, unauthenticated - no login on the target, just fabric attachment. Adjacent-network attack vector.","remediation":"Upgrade the host kernel to 6.15 or a stable backport (5.4.292, 5.10.236, 5.15.180, 6.1.134, 6.6.87, 6.12.23, 6.13.11, 6.14.2). Rolling reboot of every InfiniBand/RoCE host. Complementary control: enforce partition keys and restrict which nodes can attach to the subnet - this bug is a strong argument for not treating an IB fabric as a trusted flat network.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-22086","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2025/CVE-2025-22086.json"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-04-16"},{"id":"CVE-2025-23120","cve":"CVE-2025-23120","aliases":[],"title":"Veeam Backup & Replication: Remote code execution reachable by any domain user on a domain-joined backup server","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Veeam Backup & Replication","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"Remote code execution reachable by any domain user on a domain-joined backup server","attack_vector":"Network (remote)","remediation":"Control-plane: patch; take the backup server off the production AD domain","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23120"],"status":"curated","published":"2025-03-20"},{"id":"CVE-2025-23121","cve":"CVE-2025-23121","aliases":[],"title":"Veeam Backup & Replication: Authenticated domain user achieves remote code execution on the Backup Server","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Veeam Backup & Replication","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"Authenticated domain user achieves remote code execution on the Backup Server","attack_vector":"Network (remote)","remediation":"Control-plane: patch; workgroup-isolate the backup infrastructure","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23121"],"status":"curated","published":"2025-06-19"},{"id":"CVE-2025-23253","cve":"CVE-2025-23253","aliases":[],"title":"NVIDIA App: Local privesc","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA App","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"Local privesc","attack_vector":"Local Windows user","remediation":"Consumer app; no DC action","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23253","https://github.com/NVIDIA/product-security/tree/main/2025/5644"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/UI:N/PR:L/S:C/C:H/I:H/A:H","cwe":["CWE-547"],"published":"2025-04-22"},{"id":"CVE-2025-23254","cve":"CVE-2025-23254","aliases":[],"title":"TensorRT-LLM: RCE via unsafe pickle deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT-LLM","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"RCE via unsafe pickle deserialization","attack_vector":"Malicious model / untrusted engine file","remediation":"Bump TensorRT-LLM; rebuild serving images; enforce signed model artifacts","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23254","https://github.com/NVIDIA/product-security/tree/main/2025/5648"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2025-05-01"},{"id":"CVE-2025-24325","cve":"CVE-2025-24325","aliases":[],"title":"Intel ice driver (Ethernet 800 Series, Linux kernel mode): Improper input validation in the 800-series Linux kernel","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel ice driver (Ethernet 800 Series, Linux kernel mode)","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"Improper input validation in the 800-series Linux kernel driver allowing an authenticated user to escalate. Highest-scored of the 2025 ice driver batch.","attack_vector":"Authenticated local user; extends to tenants on VF/RDMA-exposing nodes.","remediation":"Fixed in the Intel out-of-tree ice driver (or in-kernel equivalent). Updating the driver requires unloading and reloading the module, which drops every link on that NIC - on a node whose RDMA fabric carries collective traffic, that is a job-killing event, so drain first. If you take it via a distro kernel update instead, it is a reboot. No firmware flash for the driver-side fixes. Target ice 1.17.2 or later.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-24325","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01296.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-08-12"},{"id":"CVE-2025-24514","cve":"CVE-2025-24514","aliases":[],"title":"ingress-nginx: Config injection via unsanitized auth-url annotation (part of the IngressNightmare set)","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"Config injection via unsanitized auth-url annotation (part of the IngressNightmare set)","attack_vector":"Cluster user who can create Ingress objects","remediation":"Rolling controller upgrade","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated","published":"2025-03-25"},{"id":"CVE-2025-32431","cve":"CVE-2025-32431","aliases":[],"title":"Traefik: Path matcher flaw in PathPrefix/Path/PathRegex routing enables route and authorization bypass","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Traefik","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"Path matcher flaw in PathPrefix/Path/PathRegex routing enables route and authorization bypass","attack_vector":"Unauthenticated network","remediation":"Rolling Traefik upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-32431"],"status":"curated","published":"2025-04-21"},{"id":"CVE-2025-33186","cve":"CVE-2025-33186","aliases":[],"title":"NVIDIA AIStore - AuthN: A flaw in the AIStore authentication component reaches privilege escalation, information","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA AIStore - AuthN","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"A flaw in the AIStore authentication component reaches privilege escalation, information disclosure and data tampering - effectively an authentication bypass in front of your training data store. Scored 8.8 network with no privileges required.","attack_vector":"Network, unauthenticated, one user-interaction step. Anyone with a route to AIStore's AuthN service.","remediation":"Upgrade AIStore per bulletin 5724 as a priority and rotate AuthN tokens afterwards - the patch does not invalidate credentials an attacker may already hold. Cost: rolling control-plane restart.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33186","https://github.com/NVIDIA/product-security/tree/main/2025/5724"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-798"],"published":"2025-11-11"},{"id":"CVE-2025-33208","cve":"CVE-2025-33208","aliases":[],"title":"NVIDIA TAO Toolkit: An uncontrolled search path loads an attacker-planted resource, reaching privilege escalation","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA TAO Toolkit","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"An uncontrolled search path loads an attacker-planted resource, reaching privilege escalation and code execution. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5730 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33208","https://github.com/NVIDIA/product-security/tree/main/2025/5730"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-427"],"published":"2025-12-03"},{"id":"CVE-2025-33213","cve":"CVE-2025-33213","aliases":[],"title":"NVIDIA Merlin Transformers4Rec: The Trainer component deserializes untrusted data, reaching code execution","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Merlin Transformers4Rec","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"The Trainer component deserializes untrusted data, reaching code execution. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5739 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33213","https://github.com/NVIDIA/product-security/tree/main/2025/5739"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2025-12-09"},{"id":"CVE-2025-33214","cve":"CVE-2025-33214","aliases":[],"title":"NVIDIA NVTabular: The Workflow component deserializes untrusted data, reaching code execution","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NVTabular","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"The Workflow component deserializes untrusted data, reaching code execution. Feature-engineering workflows are commonly shared as artifacts, which is the delivery path. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5739 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33214","https://github.com/NVIDIA/product-security/tree/main/2025/5739"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2025-12-09"},{"id":"CVE-2025-36593","cve":"CVE-2025-36593","aliases":[],"title":"Dell OpenManage Network Integration (RADIUS auth bypass): An attacker on the local network forges a valid RADIUS Accept","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Dell OpenManage Network Integration (RADIUS auth bypass)","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"An attacker on the local network forges a valid RADIUS Accept in response to a failed authentication - so a rejected login becomes an accepted one. This is the Blast-RADIUS protocol weakness landing in Dell's fabric management tool.","attack_vector":"Local network position between OMNI and the RADIUS server.","remediation":"Upgrade OMNI to 3.8. Beyond the patch, the structural fix is running RADIUS over an authenticated transport (RADSEC/TLS) or moving fabric admin auth off RADIUS entirely.","references":["https://www.dell.com/support/kbdoc/en-us/000337238/dsa-2025-257-security-update-for-dell-openmanage-network-integration-omni-vulnerabilities"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-37885","cve":"CVE-2025-37885","aliases":[],"title":"Linux kernel (arch/x86/kvm/svm): When a GSI route changed to something that cannot be posted, KVM only fixed up the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (arch/x86/kvm/svm)","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"When a GSI route changed to something that cannot be posted, KVM only fixed up the interrupt remapping table entry if the new route was an MSI, leaving IRTEs still posting interrupts directly to a vCPU. An assigned device's interrupts then get delivered into a guest that should no longer receive them, and if the VM is torn down while the entry still points at it the hardware writes posted-interrupt state into freed host memory - a device-driven use-after-free that outlives the VM.","attack_vector":"Requires device assignment - VFIO passthrough, which is the normal shape of a GPU tenant node - plus AMD AVIC or Intel posted interrupts. The stale entry is created by an interrupt-routing update on a running VM and is then exercised by the physical device itself, so the fallout lands on the host and on whoever gets that device next, not only on the VM that created it.","remediation":"Update to a stable kernel carrying the linked fix (no fixed release enumerated; take the branch with commit 023816bd5fa4). Interim controls: avoid IRQ-routing changes on running passthrough VMs, scrub and re-probe assigned devices between tenants, or disable interrupt posting (kvm_amd.avic=0 / kvm_intel.enable_apicv=0) at a performance cost.","references":["https://git.kernel.org/stable/c/023816bd5fa46fab94d1e7917fe131b79ed1fb41","https://git.kernel.org/stable/c/116c7d35b8f72eac383b9fd371d7c1a8ffc2968b","https://nvd.nist.gov/vuln/detail/CVE-2025-37885"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-37957","cve":"CVE-2025-37957","aliases":[],"title":"Linux kernel (arch/x86/kvm): A guest that is in SMM and then triple-faults makes SVM take the SHUTDOWN intercept and","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (arch/x86/kvm)","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"A guest that is in SMM and then triple-faults makes SVM take the SHUTDOWN intercept and reset the vCPU without first forcing it out of SMM, leaving host-side vCPU state in an architecturally impossible configuration. The same omission in the nested path was previously a use-after-free, and the kernel CNA rates this one as a scope-changing confidentiality, integrity and availability break.","attack_vector":"Guest-driven on AMD hosts and reproduced by syzkaller with nothing but a VM and one vCPU: enter SMM (a guest can direct an SMI at itself through the emulated local APIC, or the VMM's KVM_SMI path is used) and then execute instructions that cascade into a triple fault. No passthrough device and no host privilege required.","remediation":"Update to a stable kernel with the linked fix; the record points at the 5.16 and 6.1 lines, so take the point release on your branch that contains commit e9b28bc65fd3. No meaningful interim control - SMM emulation cannot be disabled per tenant.","references":["https://git.kernel.org/stable/c/e9b28bc65fd3a56755ba503258024608292b4ab1","https://git.kernel.org/stable/c/d362b21fefcef7eda8f1cd78a5925735d2b3287c","https://nvd.nist.gov/vuln/detail/CVE-2025-37957"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-696","CWE-665"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-38216","cve":"CVE-2025-38216","aliases":[],"title":"Linux kernel (drivers/iommu/intel): VT-d switched from set-and-check to clear-and-reset when programming device-table","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/intel)","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"VT-d switched from set-and-check to clear-and-reset when programming device-table context entries on domain attach. For PCI functions that share a requester ID through a PCIe-to-PCI bridge, attaching one function blows away the context entry the sibling is using, so devices in the same IOMMU group end up with a wrong or absent translation and start taking DMA faults. Attaching a device to a tenant's domain therefore disturbs the translation state of every other function aliased to it - the shared entry is the boundary, and the attach path stops treating it as shared.","attack_vector":"Triggered by domain attach on Intel VT-d hosts whenever devices are PCI-aliased behind a PCIe-to-PCI bridge - exactly the case where those functions must be passed through as one IOMMU group. Reached through the normal VFIO/iommufd attach a tenant's VMM performs when claiming a device, so a tenant's start/stop cycle drives it; no host root required. Conditional on Intel VT-d and on the node actually having aliased functions (upstream reproduced it with an Apple SPI controller, but the pattern is generic to bridge-aliased devices).","remediation":"No fixed release is listed in this record; apply the linked stable commits or run a current stable/LTS kernel on Intel passthrough nodes. Interim: avoid assigning devices that live in an aliased IOMMU group behind a PCIe-to-PCI bridge, and audit dmesg for DMAR PTE-read faults on nodes where a tenant attach coincides with another device going quiet.","references":["https://git.kernel.org/stable/c/fb5873b779dd5858123c19bbd6959566771e2e83","https://git.kernel.org/stable/c/d43c81b691813e16a2d08208ce8947aebdab83cd","https://nvd.nist.gov/vuln/detail/CVE-2025-38216"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-38283","cve":"CVE-2025-38283","aliases":[],"title":"Linux kernel (drivers/vfio/pci/hisilicon): The guest decides whether the host's VFIO migration code has a valid queue","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/vfio/pci/hisilicon)","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"The guest decides whether the host's VFIO migration code has a valid queue address to work with. A guest that simply never loads the VF driver leaves the migration payload empty, and the destination host dereferences the resulting null address while restoring device state. A tenant crashes the host it is being migrated onto - scored scope-changed by the kernel CNA, and on a shared node that is everyone else's outage.","attack_vector":"A tenant VM assigned a HiSilicon accelerator VF (Kunpeng ZIP/SEC/HPRE class) under hisi_acc_vfio_pci is live-migrated while the guest has not loaded the VF driver. The guest controls the condition entirely - it just does nothing - and the fault lands in host kernel context on the destination. Conditional on the hisi_acc_vfio_pci variant driver being bound, SR-IOV VFs assigned to guests, and live migration being enabled; not reachable on an NVIDIA/AMD GPU fleet that never loads this driver.","remediation":"No fixed release is listed in this record; apply the linked stable commits or run a current stable/LTS kernel on any node using hisi_acc_vfio_pci. Interim: disable live migration for HiSilicon accelerator VFs, or bind plain vfio-pci instead of the migration-capable variant driver.","references":["https://git.kernel.org/stable/c/b5ef128926cd34dffa2a66607b9c82b902581ef8","https://git.kernel.org/stable/c/59a834592dd200969fdf3c61be1cb0615c647e45","https://nvd.nist.gov/vuln/detail/CVE-2025-38283"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-226","CWE-908"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-38511","cve":"CVE-2025-38511","aliases":[],"title":"Linux kernel (drivers/gpu/drm/xe): On SR-IOV-partitioned Intel GPUs, the local-memory translation tables handed to a VF","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/xe)","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"On SR-IOV-partitioned Intel GPUs, the local-memory translation tables handed to a VF are allocated without being zeroed, so page-table entries outside the range actually provisioned to that VF still point at whatever was there before - another tenant's VRAM allocations or the host PF's own pages. A malicious guest that walks past its provisioned LMEM window reads, and can write, memory belonging to a different tenant or to the host.","attack_vector":"A guest that owns an SR-IOV VF of the GPU (i.e. a VM tenant given a virtual function rather than a whole card) reaches this by simply addressing local memory beyond its provisioned range. Requires SR-IOV to be enabled and VFs provisioned by the xe PF driver; a plain container holding /dev/dri/renderD* on a non-virtualized card is not affected. No host root or physical access needed - the boundary broken is exactly the VF/PF partitioning boundary the operator is selling.","remediation":"Boot a kernel carrying the fix commits below (no fixed_in version published by the kernel CNA - track the stable trees the commits landed in). Interim: stop provisioning GPU VFs to untrusted tenants, or disable SR-IOV on the xe cards and hand out whole devices instead, until the patched kernel is rolled out.","references":["https://git.kernel.org/stable/c/ff4b8c9ade1b82979fdd01e6f45b60f92eed26d8","https://git.kernel.org/stable/c/5d21892c2e15b6a27f8bc907693eca7c6b7cc269","https://nvd.nist.gov/vuln/detail/CVE-2025-38511"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-190"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-38688","cve":"CVE-2025-38688","aliases":[],"title":"Linux kernel (drivers/iommu/iommufd): The IOVA allocator's alignment arithmetic wraps near ULONG_MAX and yields a","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/iommufd)","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"The IOVA allocator's alignment arithmetic wraps near ULONG_MAX and yields a corrupted IOVA. The upstream fix states it plainly - userspace can create a mapping that overlaps an existing mapping or a reserved range. That is an over-broad DMA window: a tenant's passthrough device gets translations into memory the IOMMU was supposed to be fencing off, including ranges reserved for MSI and for other domains. This is the direct tenant-to-host DMA escape shape.","attack_vector":"A tenant holding /dev/iommu issues IOAS map/allocate requests with a length and alignment chosen so the candidate range sits near the top of the address space and the alignment rounds past ULONG_MAX. No host root, no special hardware - just the iommufd ioctls a passthrough tenant already uses to program its own DMA mappings. Conditional on iommufd being the passthrough path on the node.","remediation":"Apply the linked stable commits or run a current stable/LTS kernel on every node that exposes iommufd; no fixed version string is carried in this record, so track the commits into your own kernel build. Interim: do not expose /dev/iommu to untrusted tenants, keep the VMM as the only iommufd client, and prefer legacy VFIO type1 containers where the platform still supports it until the kernel is patched.","references":["https://git.kernel.org/stable/c/d19b817540c0abe84854a64ee9ee34cecc3bbeef","https://git.kernel.org/stable/c/ebb6021560b94649bec6b8faba6fe0dca2218e81","https://nvd.nist.gov/vuln/detail/CVE-2025-38688"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-3928","cve":"CVE-2025-3928","aliases":[],"title":"Commvault Web Server: Remote authenticated attacker creates and executes webshells","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Commvault Web Server","year":"2025","cvss_score":8.8,"severity":"high","kev":true,"impact":"Remote authenticated attacker creates and executes webshells","attack_vector":"Network (remote)","remediation":"Control-plane: patch; sweep the web server for webshells","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-3928"],"status":"curated","published":"2025-04-25"},{"id":"CVE-2025-39961","cve":"CVE-2025-39961","aliases":[],"title":"Linux iommu/amd - race while increasing host page table level: The AMD IOMMU host page table implementation supports","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux iommu/amd - race while increasing host page table level","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"The AMD IOMMU host page table implementation supports growing the page table dynamically, and the code that increases the level races with concurrent users. At CVSS 8.8 this is the most severe AMD IOMMU issue in the set. The AMD IOMMU is what constrains device DMA on a GPU host - it is the boundary that stops a passed-through or SR-IOV accelerator from reading memory belonging to another tenant - so a race that corrupts its page tables is a direct threat to device-level isolation, and a corruption primitive in host kernel memory besides.","attack_vector":"Local, triggered by concurrent DMA mapping activity. On a GPU node with high-rate accelerator and RDMA NIC DMA, the concurrency needed to hit this arises from normal workload behaviour, and a tenant can drive it deliberately by hammering mapping operations.","remediation":"Fixed in the Linux kernel AMD IOMMU driver. Take the distro kernel update and reboot the host - no firmware, VBIOS or AGESA step. Prioritise this on nodes doing GPU passthrough, SR-IOV or heavy RDMA: those are the configurations that exercise the dynamic page-table growth path hardest and depend most on IOMMU correctness for tenant isolation.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-39961"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-10-09"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-40058","cve":"CVE-2025-40058","aliases":[],"title":"Linux kernel (drivers/iommu/intel): VT-d advertised IOMMU dirty-page tracking on units whose page walk is not coherent","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/intel)","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"VT-d advertised IOMMU dirty-page tracking on units whose page walk is not coherent with the CPU. Per the Intel spec, the IOMMU then takes a non-recoverable fault the moment it tries to atomically set an A/D bit, so a tenant enabling dirty tracking on its own passthrough domain wedges the IOMMU unit that serves other devices on that node. The kernel CNA scored this scope-changed, which matches: the damage lands outside the requesting domain. Separately, dirty bits that never land quietly break live-migration correctness for any workload relying on them.","attack_vector":"A tenant (or the migration control plane acting on a tenant's request) allocates an iommufd hardware page table with dirty tracking enabled - IOMMU_HWPT_ALLOC with the dirty-tracking flag - for a device behind a VT-d unit that reports scalable-mode A/D support without snooped page-walk coherency. Reachable from /dev/iommu with no host root. Conditional on Intel VT-d, scalable mode, and hardware that reports ecap_slads without ecap_smpwc.","remediation":"No fixed release is listed in this record; apply the linked stable commits or run a current stable/LTS kernel. Interim: do not enable IOMMU dirty tracking (and therefore IOMMU-assisted live migration of passthrough devices) on hosts whose VT-d units lack snooped page-walk coherency; gate the dirty-tracking flag in your VMM/control plane rather than letting tenants request it.","references":["https://git.kernel.org/stable/c/ebe16d245a00626bb87163862a1b07daf5475a3e","https://git.kernel.org/stable/c/8d096ce0e87bdc361f0b25d7943543bc53aa0b9e","https://nvd.nist.gov/vuln/detail/CVE-2025-40058"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-787","CWE-125"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-40336","cve":"CVE-2025-40336","aliases":[],"title":"Linux kernel (drivers/gpu/drm): The shared GPU SVM layer mis-computes the mapping order when an HMM range only","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm)","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"The shared GPU SVM layer mis-computes the mapping order when an HMM range only partially covers a huge page, so the driver programs GPU page-table entries for memory outside the range the tenant asked for - including pages not mapped by that process's mm at all. The GPU then reads and writes host memory the tenant was never granted, which is a direct route to another tenant's or the kernel's pages.","attack_vector":"A tenant container holding /dev/dri/renderD* on a driver using the shared GPU SVM/userptr path (xe, and drivers moving onto drm_gpusvm) reaches this from normal SVM/userptr binds - it only has to arrange a userspace range that straddles a transparent-huge-page boundary. No special capability, no display access, no host root.","remediation":"Update to a kernel carrying the drm_gpusvm fix commits below (the kernel CNA published no fixed_in list; follow the stable branches these landed on). Interim: for untrusted tenants, disable transparent hugepages for the workload or drop /dev/dri/renderD* from containers that do not need GPU SVM.","references":["https://git.kernel.org/stable/c/08e9fd78ba1b9e95141181c69cc51795c9888157","https://git.kernel.org/stable/c/c50729c68aaf93611c855752b00e49ce1fdd1558","https://nvd.nist.gov/vuln/detail/CVE-2025-40336"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","fleet":{"pain_class":"node-drain"},"id":"CVE-2025-40362","cve":"CVE-2025-40362","aliases":[],"title":"CephFS kernel client (ceph.ko, MDS auth caps): In a multi-FS Ceph cluster the kernel client applies an MDS auth cap","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"CephFS kernel client (ceph.ko, MDS auth caps)","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"In a multi-FS Ceph cluster the kernel client applies an MDS auth cap from one file system to another because it never compares fsname. A key granted read-only on fsname1 and read-write on fsname2 ends up with write rights on fsname1 - so a tenant scoped to their own file system can write into another one.","attack_vector":"Any user on a compute node that kernel-mounts CephFS in a cluster running more than one file system, using a CephX key with differing caps per file system.","remediation":"Update the kernel on all CephFS client nodes to one carrying the multifs auth-caps fix, drain and reboot each node. Until then avoid granting one CephX principal different caps across file systems in the same cluster - use a distinct key per file system.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-40362","https://git.kernel.org/pub/scm/linux/security/vulns.git"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-41225","cve":"CVE-2025-41225","aliases":[],"title":"VMware vCenter Server (authenticated command execution via alarms): A user with permission to create or modify alarms","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"VMware vCenter Server (authenticated command execution via alarms)","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"A user with permission to create or modify alarms and run script actions executes arbitrary commands on vCenter. The alarm/script-action feature is a legitimate automation path that doubles as a privilege-escalation route to root on the management server.","attack_vector":"Authenticated vCenter user holding alarm-management privileges - a role commonly granted to monitoring integrations.","remediation":"Apply the Broadcom fix per advisory 25717. Also audit which service accounts hold alarm/script-action rights; most monitoring integrations do not need them.","references":["https://support.broadcom.com/web/ecx/support-content-notification/-/external/content/SecurityAdvisories/0/25717"],"status":"curated"},{"id":"CVE-2025-46427","cve":"CVE-2025-46427","aliases":[],"title":"Dell SmartFabric OS10 (command injection): Second command-injection path in the same OS10 advisory, giving","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell SmartFabric OS10 (command injection)","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"Second command-injection path in the same OS10 advisory, giving a low-privileged remote user command execution on the switch.","attack_vector":"Authenticated low-privilege access to the switch.","remediation":"Upgrade SmartFabric OS10 to 10.6.1.0. Requires switch reboot.","references":["https://www.dell.com/support/kbdoc/en-us/000391062/dsa-2025-407-security-update-for-dell-networking-os10-vulnerabilities"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-46428","cve":"CVE-2025-46428","aliases":[],"title":"Dell SmartFabric OS10 (command injection): A low-privileged remote attacker executes code on the switch OS","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell SmartFabric OS10 (command injection)","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"A low-privileged remote attacker executes code on the switch OS. On a GPU cluster these switches carry the storage and east-west traffic, so switch compromise means traffic interception across the fabric.","attack_vector":"Authenticated low-privilege access to the switch management plane.","remediation":"Upgrade SmartFabric OS10 to 10.6.1.0. Switch OS upgrade with a reboot - do it leaf-by-leaf on a redundant topology, or take a fabric maintenance window.","references":["https://www.dell.com/support/kbdoc/en-us/000391062/dsa-2025-407-security-update-for-dell-networking-os10-vulnerabilities"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-49847","cve":"CVE-2025-49847","aliases":[],"title":"llama.cpp (vocab): Attacker-supplied GGUF vocabulary triggers memory corruption","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"llama.cpp (vocab)","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"Attacker-supplied GGUF vocabulary triggers memory corruption","attack_vector":"Customer-supplied model file","remediation":"Rebuild past b5662","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-49847"],"status":"curated","published":"2025-06-17"},{"id":"CVE-2025-51480","cve":"CVE-2025-51480","aliases":[],"title":"ONNX (`save_external_data`): Path traversal","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"ONNX (`save_external_data`)","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"Path traversal → overwrite arbitrary files","attack_vector":"Customer-supplied ONNX model with crafted external-data entries","remediation":"Upgrade past 1.17.0","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-51480"],"status":"curated","published":"2025-07-22"},{"id":"CVE-2025-62164","cve":"CVE-2025-62164","aliases":[],"title":"vLLM (multimodal embeddings): Memory corruption","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (multimodal embeddings)","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"Memory corruption → crash and possible RCE","attack_vector":"Unauthenticated request to an exposed OpenAI-compatible serving port carrying crafted embeddings","remediation":"Upgrade to 0.11.1+. If the provider hosts the endpoint, this is provider-owned; if the tenant runs their own vLLM, it is tenant-owned and the provider can only enforce network policy","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-62164"],"status":"curated","fleet":{"ubiquity":"Very common - vLLM is the default open-source inference server on nearly every neocloud's serverless-inference and BYO-endpoint product","remediation_pain":"`daemon-restart` - upgrade to 0.11.1+ and roll the serving fleet; every replica of every model endpoint must be restarted","pain_class":"daemon-restart","why_fleet_wide":"The Completions API `torch.load`s user-supplied prompt embeddings; with PyTorch 2.8's sparse-tensor checks off by default a crafted tensor causes an out-of-bounds write - network-only, unauthenticated-adjacent, on every serving replica"},"published":"2025-11-21"},{"id":"CVE-2025-62623","cve":"CVE-2025-62623","aliases":[],"title":"AMD Pensando ionic cloud driver for VMware ESXi (heap overflow): Second heap overflow in the ionic ESXi driver","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"AMD Pensando ionic cloud driver for VMware ESXi (heap overflow)","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"Second heap overflow in the ionic ESXi driver with the same privilege-escalation and code-execution outcome.","attack_vector":"Local low-privilege access on the ESXi host.","remediation":"Apply the AMD driver update per AMD-SB-2001; VIB update plus host reboot.","references":["https://www.amd.com/en/resources/product-security/bulletin/AMD-SB-2001.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-62624","cve":"CVE-2025-62624","aliases":[],"title":"AMD Pensando ionic cloud driver for VMware ESXi (heap overflow): Heap-based buffer overflow in the ionic SmartNIC","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"AMD Pensando ionic cloud driver for VMware ESXi (heap overflow)","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"Heap-based buffer overflow in the ionic SmartNIC driver on ESXi allowing privilege escalation and arbitrary code execution. The DPU driver sits in the hypervisor's network datapath, so compromise reaches all tenant traffic on the host.","attack_vector":"Local low-privilege access on the ESXi host; high attack complexity.","remediation":"Apply the AMD-provided driver update per AMD-SB-2001. Driver VIB update plus host reboot - coordinate with any DPU firmware update in the same bulletin so you take one outage rather than two.","references":["https://www.amd.com/en/resources/product-security/bulletin/AMD-SB-2001.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-6685","cve":"CVE-2025-6685","aliases":["ZDI-25-650"],"title":"ATEN eco DC (DCIM/environmental management platform): The web interface doesn't check a user's assigned role","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ATEN eco DC (DCIM/environmental management platform)","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"The web interface doesn't check a user's assigned role before acting on their requests, so an authenticated low-privileged user can escalate to actions normally reserved for administrators on the datacenter-infrastructure-management platform — which typically has visibility and control hooks into PDUs and environmental sensors across the facility.","attack_vector":"Requires a valid but low-privileged account on ATEN eco DC; the attacker sends requests for admin-level functions that the server fails to gate on role.","remediation":"Software upgrade to the patched eco DC release per ATEN's advisory. This is a server-side application (not per-rack firmware), so it's a single upgrade rather than a fleet-wide rollout, but audit who has any account on it since the bug turns any low-privileged login into an admin one.","references":["https://www.aten.com/global/en/supportcenter/info/security-advisory/25/","https://www.zerodayinitiative.com/advisories/ZDI-25-650/"],"status":"curated","published":"2025-09-02"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-71202","cve":"CVE-2025-71202","aliases":[],"title":"Linux kernel (drivers/iommu): This is the substantive fix for stale IOMMU translations of the kernel address space","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu)","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"This is the substantive fix for stale IOMMU translations of the kernel address space under SVA. Without it the IOMMU keeps cached paging-cache entries for kernel page-table pages that have already been freed and handed out for other use, so a device can DMA through a translation that now points at attacker-placed data - arbitrary physical memory access and privilege escalation from a device a tenant drives. The kernel CNA scored it scope-changed, which is the right read: the blast radius is the host, not the tenant.","attack_vector":"Local unprivileged, on x86 hosts with IOMMU SVA and an SVA/PASID-capable device the tenant can bind (Intel DSA/IAA, SVM-capable GPUs, PRI-capable NICs). The attacker recycles kernel page-table pages through ordinary means - the series calls out vfree() as the common, unprivileged trigger - while an SVA-bound device still has the stale entry cached, then grooms the reallocated page. Requires CONFIG_IOMMU_SVA and a driver exposing SVA binding to userspace.","remediation":"No fixed release is listed in this record; take the whole series from the linked stable commits (this plus CVE-2025-71089) or run a current stable/LTS kernel that carries it. Interim: disable SVA/PASID on tenant-facing devices and keep SVA-capable accelerator nodes out of untrusted containers - the upstream series itself disables x86 SVA outright until this invalidation path exists.","references":["https://git.kernel.org/stable/c/9f0a7ab700f8620e433b05c57fbd26c92ea186d9","https://git.kernel.org/stable/c/e37d5a2d60a338c5917c45296bac65da1382eda5","https://nvd.nist.gov/vuln/detail/CVE-2025-71202"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-8557","cve":"CVE-2025-8557","aliases":[],"title":"Lenovo XClarity Orchestrator (alternate communication channel): An attacker on the LXCO network segment manipulates","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lenovo XClarity Orchestrator (alternate communication channel)","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"An attacker on the LXCO network segment manipulates a local device to create an alternate communication channel into the management stack - a network-position attack against the fleet controller.","attack_vector":"Access to a device on the LXCO local network segment. Unauthenticated.","remediation":"Apply the LXCO update per LEN-201014, and put the orchestrator on a dedicated management segment rather than a shared server VLAN.","references":["https://support.lenovo.com/us/en/product_security/LEN-201014"],"status":"curated"},{"id":"CVE-2025-8876","cve":"CVE-2025-8876","aliases":[],"title":"N-able N-central: Improper input validation","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"N-able N-central","year":"2025","cvss_score":8.8,"severity":"high","kev":true,"impact":"Improper input validation -> OS command injection","attack_vector":"Network (remote)","remediation":"Control-plane: patch to 2025.3.1+ immediately","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-8876"],"status":"curated","published":"2025-08-14"},{"id":"CVE-2026-13622","cve":"CVE-2026-13622","aliases":[],"title":"KubeVirt virt-handler (symlink following in migration proxy): During live migration virt-handler dials Unix sockets","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"KubeVirt virt-handler (symlink following in migration proxy)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"During live migration virt-handler dials Unix sockets inside the target virt-launcher pod through /proc/<pid>/root/ without symlink protection, and those paths sit in qemu-owned directories the launcher user can write. A tenant who controls a VM's launcher redirects the handler into arbitrary host paths, escalating out of the pod with scope change.","attack_vector":"A tenant with control inside a virt-launcher pod, triggering or riding a live migration.","remediation":"Apply the Red Hat OpenShift Virtualization errata (RHSA-2026:51031 / RHSA-2026:53655) or upgrade upstream KubeVirt. Operator-driven rolling update of virt-handler across nodes; VMs keep running but migrations should be paused during the rollout.","references":["https://access.redhat.com/errata/RHSA-2026:51031","https://access.redhat.com/errata/RHSA-2026:53655"],"status":"curated"},{"id":"CVE-2026-14371","cve":"CVE-2026-14371","aliases":[],"title":"Lenovo XClarity Integrator for Windows Admin Center (PowerShell command injection): PowerShell command injection","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lenovo XClarity Integrator for Windows Admin Center (PowerShell command injection)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"PowerShell command injection on the Windows Admin Center gateway when establishing remote commands - code execution on a host that holds administrative reach into the managed estate.","attack_vector":"Authenticated low-privilege access to the WAC gateway, with user interaction.","remediation":"Upgrade the XClarity Integrator WAC plugin past 5.1.1. Plugin update on the gateway host.","references":["https://pcsupport.lenovo.com/us/en/product_security/home"],"status":"curated"},{"id":"CVE-2026-1580","cve":"CVE-2026-1580","aliases":[],"title":"ingress-nginx: Config injection via the auth-method annotation","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Config injection via the auth-method annotation","attack_vector":"Cluster user who can create Ingress objects","remediation":"Rolling controller upgrade","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated","published":"2026-02-03"},{"id":"CVE-2026-16793","cve":"CVE-2026-16793","aliases":[],"title":"Lenovo XClarity Orchestrator (OS command injection): An authenticated attacker executes arbitrary OS commands","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lenovo XClarity Orchestrator (OS command injection)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"An authenticated attacker executes arbitrary OS commands as a privileged user on LXCO. Orchestrator sits above XClarity Administrator and drives multi-site fleet management, so compromise there reaches a very large number of servers.","attack_vector":"Authenticated low-privilege access to LXCO 2.2.0.","remediation":"Apply the Lenovo LXCO update. Appliance upgrade with a service restart; rotate the credentials LXCO uses to reach managed XCC endpoints afterwards.","references":["https://support.lenovo.com/my/en/solutions/ht509976-lenovo-xclarity-orchestrator"],"status":"curated","fleet":{"pain_class":"daemon-restart"}},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-269"],"fleet":{"pain_class":"hot-patch"},"id":"CVE-2026-18951","cve":"CVE-2026-18951","aliases":[],"title":"Kubeflow Training Operator (RHOAI overlay, trainjobs aggregated into the edit ClusterRole): The RHOAI overlay","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Kubeflow Training Operator (RHOAI overlay, trainjobs aggregated into the edit ClusterRole)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"The RHOAI overlay aggregates trainjobs create/update/delete into Kubernetes' built-in edit ClusterRole, so every namespace editor silently gains full control over training jobs. Chained with the companion flaw that allows arbitrary pod configuration in a TrainJob, a tenant with routine edit rights escalates to arbitrary code execution on the GPU nodes their jobs land on.","attack_vector":"Any principal bound to the standard edit ClusterRole in a namespace where the training operator is installed - a very common default for application teams.","remediation":"Apply the Red Hat OpenShift AI errata (RHSA-2026:53262/53263). Independently, audit which ClusterRoles aggregate trainjobs verbs in your own overlays and split TrainJob management into a dedicated role instead of folding it into edit.","references":["https://access.redhat.com/security/cve/CVE-2026-18951","https://nvd.nist.gov/vuln/detail/CVE-2026-18951"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-288"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-22049","cve":"CVE-2026-22049","aliases":[],"title":"NetApp ONTAP WebAuthn multi-factor authentication (Relying Party ID): An attacker who already has valid credentials","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"NetApp ONTAP WebAuthn multi-factor authentication (Relying Party ID)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"An attacker who already has valid credentials sidesteps the WebAuthn second factor because the Relying Party ID is not bound correctly. The hardware-key requirement protecting storage admin logins stops being a barrier.","attack_vector":"A remote attacker in possession of valid ONTAP credentials, against a system running 9.16.1 or later with WebAuthn MFA configured.","remediation":"Upgrade to the ONTAP release NetApp names in the advisory. Until then, do not count WebAuthn as the control that stops credential reuse - rotate any password suspected of exposure rather than relying on the second factor.","references":["https://security.netapp.com/advisory/NTAP-20260722-0001/","https://nvd.nist.gov/vuln/detail/CVE-2026-22049"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-22622","cve":"CVE-2026-22622","aliases":["eaton-va-2026-1005"],"title":"Eaton Tripp Lite series PADM firmware, session management interface: A low-privilege authenticated user escalates","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Eaton Tripp Lite series PADM firmware, session management interface","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"A low-privilege authenticated user escalates to unrestricted device access. In practice that means a read-only monitoring account - the kind operators hand to a DCIM tool or an NOC vendor - becomes full control of rack power.","attack_vector":"Any authenticated user on the PDU, including the shared read-only accounts typically configured for monitoring integrations.","remediation":"PADM firmware update or replacement for EOL SKUs. Separately, audit which third parties hold PDU accounts - monitoring integrations are the usual source of the low-privilege credential this bug needs.","references":["https://www.eaton.com/content/dam/eaton/company/news-insights/cybersecurity/security-bulletins/eaton-va-2026-1005.pdf"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2026-07-30"},{"id":"CVE-2026-22807","cve":"CVE-2026-22807","aliases":[],"title":"vLLM (HF `auto_map`): Loads Hugging Face dynamic modules during model resolution","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (HF `auto_map`)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Loads Hugging Face dynamic modules during model resolution → code execution","attack_vector":"Customer-supplied or poisoned Hub model repo","remediation":"Upgrade to 0.14.0+. Any `trust_remote_code`-adjacent path means the model repo is executable content","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-22807"],"status":"curated","published":"2026-01-21"},{"id":"CVE-2026-24054","cve":"CVE-2026-24054","aliases":[],"title":"Kata Containers: Malformed or layer-less container image breaks Kata's handling","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kata Containers","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Malformed or layer-less container image breaks Kata's handling","attack_vector":"Malicious image","remediation":"Upgrade Kata to 3.26.0+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24054"],"status":"curated","published":"2026-01-29"},{"id":"CVE-2026-24164","cve":"CVE-2026-24164","aliases":[],"title":"BioNeMo Framework: Remote RCE via untrusted serialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"BioNeMo Framework","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Remote RCE via untrusted serialization","attack_vector":"Malicious model artifact / network payload","remediation":"Bump BioNeMo; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24164","https://github.com/NVIDIA/product-security/tree/main/2026/5808"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2026-03-31"},{"id":"CVE-2026-24186","cve":"CVE-2026-24186","aliases":[],"title":"NVIDIA FLARE SDK: RCE via unsafe deserialization in message handling","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA FLARE SDK","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"RCE via unsafe deserialization in message handling","attack_vector":"Network peer in the federation","remediation":"Upgrade FLARE; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24186","https://github.com/NVIDIA/product-security/tree/main/2026/5819"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2026-04-28"},{"id":"CVE-2026-24187","cve":"CVE-2026-24187","aliases":[],"title":"GPU Display Driver: Local privesc to host root (use-after-free in context handling)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Local privesc to host root (use-after-free in context handling)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node, evict all tenant workloads","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24187","https://github.com/NVIDIA/product-security/tree/main/2026/5821"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"published":"2026-05-26"},{"id":"CVE-2026-24217","cve":"CVE-2026-24217","aliases":[],"title":"BioNeMo Framework: Arbitrary file read/write via path traversal","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"BioNeMo Framework","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Arbitrary file read/write via path traversal","attack_vector":"Malicious model archive","remediation":"Bump BioNeMo; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24217","https://github.com/NVIDIA/product-security/tree/main/2026/5831"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-29"],"published":"2026-05-20"},{"id":"CVE-2026-24512","cve":"CVE-2026-24512","aliases":[],"title":"ingress-nginx: Config injection via rules.http.paths.path","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Config injection via rules.http.paths.path","attack_vector":"Cluster user who can create Ingress objects","remediation":"Rolling controller upgrade","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated","published":"2026-02-03"},{"id":"CVE-2026-24747","cve":"CVE-2026-24747","aliases":[],"title":"PyTorch (`weights_only` unpickler): Bypass of the `weights_only` allowlist","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"PyTorch (`weights_only` unpickler)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Bypass of the `weights_only` allowlist → arbitrary code execution","attack_vector":"Customer-supplied checkpoint; defeats the mitigation shipped for CVE-2025-32434","remediation":"Rebuild images on torch >= 2.10.0. Proves the format itself, not the flag, is the control — push checkpoint-format policy (safetensors-only ingest) rather than version-chasing","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24747"],"status":"curated","published":"2026-01-27"},{"id":"CVE-2026-27893","cve":"CVE-2026-27893","aliases":[],"title":"vLLM (hardcoded `trust_remote_code`, second instance): Same class, two more model files, through 0.18.0","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (hardcoded `trust_remote_code`, second instance)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Same class, two more model files, through 0.18.0","attack_vector":"Customer-supplied model repo","remediation":"Upgrade to 0.18.0+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-27893"],"status":"curated","published":"2026-03-27"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-31588","cve":"CVE-2026-31588","aliases":[],"title":"Linux kernel (arch/x86/kvm): An emulated MMIO write that straddles a page boundary onto a second MMIO page is split","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (arch/x86/kvm)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"An emulated MMIO write that straddles a page boundary onto a second MMIO page is split into two userspace exits, with the second exit still pointing at an on-stack variable from the first. If the second KVM_RUN comes from a different task, the host kernel reads a freed kernel stack - KASAN caught exactly that. Freed host kernel stack contents flow into data the VM side can observe.","attack_vector":"Started by the guest: it issues a store that splits a page and lands on emulated MMIO on both halves. The freed-stack condition needs the completing KVM_RUN to be issued from a different task, which the VMM process controls - so full weaponisation wants /dev/kvm access (nested-virt tenant or local user), while the guest alone controls the trigger.","remediation":"Update to a kernel with the referenced stable commits. Interim: keep /dev/kvm out of tenant containers and disable nested virtualization for tenants on unpatched nodes.","references":["https://git.kernel.org/stable/c/019d0bd32b9a4646ba35d904907452039e2db700","https://git.kernel.org/stable/c/4569c66dd9e94a22cd0796b6514a8b25ffff16a1","https://nvd.nist.gov/vuln/detail/CVE-2026-31588"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-459","CWE-672"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-31735","cve":"CVE-2026-31735","aliases":[],"title":"Linux kernel (drivers/iommu/generic_pt): When an unmap lands in the middle of a large or contiguous page-table entry","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/generic_pt)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"When an unmap lands in the middle of a large or contiguous page-table entry the kernel unmaps more than was requested but only invalidates the range that was requested. The extra IOVAs stay live in the IOMMU TLB, so a passthrough device keeps a working DMA window into pages the kernel has already handed back to the allocator. That is a direct tenant-to-host DMA escape: the device can read and write whatever lands in those pages next. Note the upstream author's own caveat that he knows of nothing that currently unmaps a large entry this way, so he believes it is not triggerable in practice - the vendor still scores it 8.8 scope-changed.","attack_vector":"Whichever side drives the unmap: a VMM's IOMMU_IOAS_UNMAP on /dev/iommu, or an in-kernel driver DMA unmap. Conditional on the generic page-table (iommupt) backend being in use and on the mapping having been built with large or contiguous IOPTEs, and on the unmap boundary falling inside one of them. Not reachable from the network fabric.","remediation":"The record lists no fixed release; boot a kernel carrying the stable fix commits below. There is no useful interim control for a stale-TLB bug other than the patched kernel - do not rely on unmapping as a security boundary until the node is patched.","references":["https://git.kernel.org/stable/c/50ecd96a28f712f8b682c0441f4cb9b086d28816","https://git.kernel.org/stable/c/ee6e69d032550687a3422504bfca3f834c7b5061","https://nvd.nist.gov/vuln/detail/CVE-2026-31735"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-3288","cve":"CVE-2026-3288","aliases":[],"title":"ingress-nginx: rewrite-target annotation injects nginx config","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"rewrite-target annotation injects nginx config; arbitrary code execution in the controller context","attack_vector":"Cluster user who can create Ingress objects","remediation":"Rolling controller upgrade, no GPU drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-3288"],"status":"curated","published":"2026-03-09"},{"id":"CVE-2026-33175","cve":"CVE-2026-33175","aliases":[],"title":"JupyterHub OAuthenticator: Authenticated user bypasses the intended identity check","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"JupyterHub OAuthenticator","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Authenticated user bypasses the intended identity check","attack_vector":"Notebook user","remediation":"Upgrade to 17.4.0+; the auth plugin is the tenant boundary on a managed notebook service","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-33175"],"status":"curated","published":"2026-04-03"},{"id":"CVE-2026-34197","cve":"CVE-2026-34197","aliases":[],"title":"Apache ActiveMQ: Improper input validation and code injection in the broker","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Apache ActiveMQ","year":"2026","cvss_score":8.8,"severity":"high","kev":true,"impact":"Improper input validation and code injection in the broker","attack_vector":"Network (remote)","remediation":"Control-plane: broker upgrade; ActiveMQ is a recurring ransomware entry point","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-34197"],"status":"curated","published":"2026-04-07"},{"id":"CVE-2026-35029","cve":"CVE-2026-35029","aliases":[],"title":"LiteLLM (`/config/update`): Endpoint does not enforce admin authorization","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LiteLLM (`/config/update`)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Endpoint does not enforce admin authorization","attack_vector":"Any authenticated proxy user","remediation":"Upgrade to 1.83.0+; a tenant key can rewrite the gateway config","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-35029"],"status":"curated","published":"2026-04-06"},{"id":"CVE-2026-35044","cve":"CVE-2026-35044","aliases":[],"title":"BentoML (Dockerfile generation): Injection into generated Dockerfile","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"BentoML (Dockerfile generation)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Injection into generated Dockerfile → build-time code execution","attack_vector":"Customer-supplied `bentofile.yaml` built by a provider-run build service","remediation":"Provider-owned if the provider offers managed builds: a tenant's build manifest executes in the provider's builder","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-35044"],"status":"curated","published":"2026-04-06"},{"id":"CVE-2026-35397","cve":"CVE-2026-35397","aliases":[],"title":"Jupyter Server: Path traversal in the REST API","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Jupyter Server","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Path traversal in the REST API","attack_vector":"Notebook user or anyone reaching an exposed notebook port","remediation":"Upgrade past 2.17.0. Notebook servers on GPU nodes are frequently exposed with token auth only","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-35397"],"status":"curated","published":"2026-05-05"},{"id":"CVE-2026-3821","cve":"CVE-2026-3821","aliases":[],"title":"Supermicro SMASH service (X14DBG-DAP, X14DBI): An attacker with any authorised BMC login escalates through the SMASH","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Supermicro SMASH service (X14DBG-DAP, X14DBI)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"An attacker with any authorised BMC login escalates through the SMASH shell to arbitrary code execution against the BMC, or knocks the controller offline entirely. On the DBG/DBI platform boards this is a current-generation GPU node, so the payoff is out-of-band control of live accelerator hardware: power cycling to disrupt long training runs, virtual-media boot into an attacker image, and a firmware-resident implant that persists across the node being returned to the pool. The CLI management shell exposed over SSH on the BMC of Supermicro's newest GPU-platform boards.","attack_vector":"An authenticated low-privilege BMC account with SSH reachability to the controller. Read-only or operator-tier BMC accounts handed to monitoring systems, support staff or tenants are enough - the privilege bar is low, and SMASH-over-SSH is enabled by default on these boards.","remediation":"Firmware flash from Supermicro's July 2026 BMC/IPMI advisory batch. There is a genuine config-only mitigation here that most operators should apply regardless of patch state: disable the SSH/SMASH service on the BMC entirely if your tooling uses Redfish or IPMI-over-LAN, which removes this and the whole SMASH overflow family from your attack surface at zero rollout cost. Otherwise restrict SSH to the BMC to a management-host allowlist and audit every non-admin BMC account you have handed out.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-3821","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2026/3xxx/CVE-2026-3821.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2026-40217","cve":"CVE-2026-40217","aliases":[],"title":"LiteLLM (`/guardrails/test_custom_code`): RCE via bytecode rewriting","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LiteLLM (`/guardrails/test_custom_code`)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"RCE via bytecode rewriting","attack_vector":"Authenticated network user of the proxy","remediation":"Upgrade; the guardrail-test endpoint is an eval sink","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-40217"],"status":"curated","published":"2026-04-10"},{"cvss_vector":"CVSS:4.0/AV:N/AC:L/AT:N/PR:N/UI:N/VC:N/VI:H/VA:L/SC:N/SI:N/SA:N","cwe":["CWE-287","CWE-306"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2026-40344","cve":"CVE-2026-40344","aliases":[],"title":"MinIO (S3 API, Snowball auto-extract): The Snowball auto-extract path skips signature verification entirely, so an","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"MinIO (S3 API, Snowball auto-extract)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"The Snowball auto-extract path skips signature verification entirely, so an unauthenticated caller uploads a TAR that MinIO unpacks into arbitrary object keys. That is unauthenticated write into any bucket - dataset poisoning and checkpoint tampering across tenant boundaries.","attack_vector":"Any client with network reach to the MinIO S3 endpoint. Pre-authentication.","remediation":"Upgrade MinIO to the fixed release from GHSA-9c4q-hq6p-c237 and roll a restart across the cluster. If you do not use Snowball ingestion, block the x-minio-extract / snowball request path at the proxy as an interim control, and review object write history on shared buckets.","references":["https://github.com/minio/minio/security/advisories/GHSA-9c4q-hq6p-c237","https://nvd.nist.gov/vuln/detail/CVE-2026-40344"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:4.0/AV:N/AC:L/AT:N/PR:N/UI:N/VC:N/VI:H/VA:L/SC:N/SI:N/SA:N","cwe":["CWE-287"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2026-41145","cve":"CVE-2026-41145","aliases":[],"title":"MinIO (S3 API, unsigned-trailer uploads): The signature on a query-string-credential unsigned-trailer upload is not","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"MinIO (S3 API, unsigned-trailer uploads)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"The signature on a query-string-credential unsigned-trailer upload is not properly verified, so an attacker with no valid secret key writes objects into buckets they have no rights to. Anyone who can reach the endpoint can overwrite a checkpoint, a dataset shard, or a model artifact belonging to another tenant.","attack_vector":"Any client that can reach the MinIO S3 endpoint over the network. No valid credential is needed.","remediation":"Upgrade MinIO to the release named in GHSA-hv4r-mvr4-25vw and restart every node in the erasure set (rolling restart is supported). Afterwards audit object versions and modification times on buckets that were internet- or tenant-reachable, and turn on versioning plus object lock for artifacts you cannot afford to have silently rewritten.","references":["https://github.com/minio/minio/security/advisories/GHSA-hv4r-mvr4-25vw","https://nvd.nist.gov/vuln/detail/CVE-2026-41145"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-41486","cve":"CVE-2026-41486","aliases":[],"title":"Ray Data (Arrow extension types): Custom Arrow extension types deserialized unsafely","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ray Data (Arrow extension types)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Custom Arrow extension types deserialized unsafely","attack_vector":"Customer-supplied Arrow/Parquet dataset","remediation":"Upgrade to 2.55.0+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-41486"],"status":"curated","published":"2026-05-08"},{"id":"CVE-2026-42271","cve":"CVE-2026-42271","aliases":[],"title":"LiteLLM proxy: Two endpoints allow privilege escalation / unauthorized action","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LiteLLM proxy","year":"2026","cvss_score":8.8,"severity":"high","kev":true,"impact":"Two endpoints allow privilege escalation / unauthorized action","attack_vector":"Authenticated low-privilege proxy user","remediation":"**[KEV]** Patch to 1.83.7+ immediately","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-42271"],"status":"curated","published":"2026-05-08"},{"id":"CVE-2026-4342","cve":"CVE-2026-4342","aliases":[],"title":"ingress-nginx: Comment-based nginx configuration injection","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Comment-based nginx configuration injection","attack_vector":"Cluster user who can create Ingress objects","remediation":"Rolling controller upgrade","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated","published":"2026-03-19"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-78"],"fleet":{"pain_class":"hot-patch"},"id":"CVE-2026-44345","cve":"CVE-2026-44345","aliases":["GHSA-78f9-r8mh-4xm2"],"title":"BentoML (Dockerfile template, docker.base_image interpolation): A multi-line docker.base_image value in bento.yaml","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"BentoML (Dockerfile template, docker.base_image interpolation)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"A multi-line docker.base_image value in bento.yaml smuggles arbitrary directives into the generated Dockerfile, and bentoml containerize then executes them via docker build on the machine doing the build. On a shared build node or CI runner that is code execution with the builder's credentials - registry push tokens, cloud roles, and whatever else the runner holds.","attack_vector":"Anyone who can get a victim to run bentoml containerize against an attacker-supplied bento.yaml - a pull request, a shared model repo, or a marketplace bento.","remediation":"Upgrade BentoML to 1.4.39 or later. Run bentoml build and containerize for untrusted bentos in a throwaway sandbox with no registry or cloud credentials mounted, since this is one of a series of Dockerfile-template injection bugs in the same code path.","references":["https://github.com/bentoml/BentoML/security/advisories/GHSA-78f9-r8mh-4xm2","https://nvd.nist.gov/vuln/detail/CVE-2026-44345"],"status":"curated"},{"id":"CVE-2026-44346","cve":"CVE-2026-44346","aliases":[],"title":"BentoML (`bentofile.yaml`): Malicious build manifest","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"BentoML (`bentofile.yaml`)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Malicious build manifest → code execution in the build pipeline","attack_vector":"Customer-supplied build config","remediation":"Upgrade to 1.4.39+; isolate build workers per tenant","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-44346"],"status":"curated","published":"2026-05-27"},{"id":"CVE-2026-45831","cve":"CVE-2026-45831","aliases":[],"title":"ChromaDB (SimpleRBAC): Authorization provider evaluates permissions incorrectly","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"ChromaDB (SimpleRBAC)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Authorization provider evaluates permissions incorrectly","attack_vector":"Authenticated low-privilege tenant","remediation":"Upgrade past 0.5.0","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-45831"],"status":"curated","published":"2026-06-12"},{"id":"CVE-2026-45832","cve":"CVE-2026-45832","aliases":[],"title":"ChromaDB (V1 endpoints): Tenant/database passed as `None` to the authz layer","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"ChromaDB (V1 endpoints)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Tenant/database passed as `None` to the authz layer → cross-tenant access","attack_vector":"Authenticated tenant on a shared instance","remediation":"Upgrade. Direct cross-tenant data access — the multi-tenancy control simply does not apply","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-45832"],"status":"curated","published":"2026-06-12"},{"id":"CVE-2026-45833","cve":"CVE-2026-45833","aliases":[],"title":"ChromaDB: Authenticated code injection","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"ChromaDB","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Authenticated code injection","attack_vector":"Any authenticated tenant of a shared Chroma instance","remediation":"Upgrade; do not multi-tenant a single Chroma","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-45833"],"status":"curated","published":"2026-06-12"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-362","CWE-665"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-45945","cve":"CVE-2026-45945","aliases":[],"title":"Linux kernel (drivers/iommu/intel): A live 512-bit VT-d PASID entry is replaced with a single structure copy, so the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/intel)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"A live 512-bit VT-d PASID entry is replaced with a single structure copy, so the IOMMU can fetch a half-old, half-new entry and briefly translate through a page table that belongs to neither the old nor the new domain. Domain replacement is exactly what happens when a passthrough device is moved between address spaces or between tenants, which is where a torn entry turns into DMA against the wrong tenant's memory.","attack_vector":"Triggered on the domain-replacement path, which a tenant or its VMM drives through iommufd (attach/replace a HWPT) or through vfio device rebinding while the device is actively issuing DMA. Conditional on VT-d scalable mode with PASID; no host root. Timing-dependent, but the attacker controls both the replacement and the DMA traffic that races it.","remediation":"Update to a stable kernel carrying commits 47180078 / 66a7aff4. Interim: quiesce the device (stop tenant DMA) around any domain attach/replace, and avoid runtime HWPT replacement for tenant devices on unpatched hosts.","references":["https://git.kernel.org/stable/c/4718007870547e1efebbdd6745d9fce58f008fef","https://git.kernel.org/stable/c/66a7aff480a82b8642b3991fed5fdc9780022157","https://nvd.nist.gov/vuln/detail/CVE-2026-45945"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-46113","cve":"CVE-2026-46113","aliases":[],"title":"Linux kernel (arch/x86/kvm/mmu): The shadow MMU derives GFNs for direct shadow pages arithmetically, which breaks if","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (arch/x86/kvm/mmu)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"The shadow MMU derives GFNs for direct shadow pages arithmetically, which breaks if the guest page tables change behind KVM's back. A leaf SPTE and rmap entry then land outside the range the parent shadow page covers, survive the zap, and are walked after the shadow page is freed. Dirty logging or an MMU-notifier invalidation then dereferences a freed kvm_mmu_page - host kernel use-after-free.","attack_vector":"Same reachability as its follow-up CVE-2026-53359: shadow paging in use, the PDE modified from outside the guest, and a memslot deletion. That is a process holding /dev/kvm behaving as its own VMM - a tenant with nested virt, or a local unprivileged user with /dev/kvm - after which normal host memory reclaim (MADV_DONTNEED, THP, swap) triggers the stale rmap walk without further attacker action.","remediation":"Update to a kernel with the referenced stable commits, alongside CVE-2026-53359. Interim: remove /dev/kvm from tenant containers and disable nested virtualization on unpatched nodes.","references":["https://git.kernel.org/stable/c/e9d4ea13aa2b6400bb10ec64b370ba3dadcd22f0","https://git.kernel.org/stable/c/488e386484ec8c0e558be6e156edf34ed9f4d5c8","https://nvd.nist.gov/vuln/detail/CVE-2026-46113"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-47101","cve":"CVE-2026-47101","aliases":[],"title":"LiteLLM (key generation): internal_user can mint keys with routes their role forbids","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LiteLLM (key generation)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"internal_user can mint keys with routes their role forbids","attack_vector":"Authenticated tenant user","remediation":"Upgrade to 1.83.14+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47101"],"status":"curated","published":"2026-05-21"},{"id":"CVE-2026-47102","cve":"CVE-2026-47102","aliases":[],"title":"LiteLLM (`/user/update`): User can self-elevate `user_role`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LiteLLM (`/user/update`)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"User can self-elevate `user_role`","attack_vector":"Authenticated tenant user","remediation":"Upgrade to 1.83.10+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47102"],"status":"curated","published":"2026-05-21"},{"id":"CVE-2026-4944","cve":"CVE-2026-4944","aliases":[],"title":"vLLM (hardcoded `trust_remote_code=True`): Two model implementation files force remote code execution regardless","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (hardcoded `trust_remote_code=True`)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Two model implementation files force remote code execution regardless of operator setting","attack_vector":"Customer-supplied model repo","remediation":"Upgrade past 0.14.1. Operator-level `--trust-remote-code=false` does not protect you — no config fix","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-4944"],"status":"curated","published":"2026-05-28"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-52952","cve":"CVE-2026-52952","aliases":[],"title":"Linux kernel (drivers/iommu): The IOMMU group's domain pointer is left stale when a device reset races a detach, and","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"The IOMMU group's domain pointer is left stale when a device reset races a detach, and the reset-completion path then re-attaches a domain that has already been freed. That is a use-after-free on the object that defines which host memory a passed-through device may DMA to - the strongest primitive a tenant device can get for reaching outside its own address space.","attack_vector":"A tenant holding /dev/vfio/* triggers a device reset - VFIO_DEVICE_RESET, or simply closing the device fd, which resets on release - while another thread detaches or replaces the domain. Multi-device IOMMU groups and PCI DMA-alias quirks widen the window, and those are common in GPU topologies behind PCIe switches. No host root required.","remediation":"Update to a stable kernel carrying commits 8fc289e8 / 5474e6e1. Interim: put each tenant's passthrough devices in a single-device IOMMU group where the topology allows it (ACS on the upstream switch ports), and avoid handing tenants devices whose group contains functions belonging to other workloads.","references":["https://git.kernel.org/stable/c/8fc289e809f3eb7e36cadc4684ab6fad747a5a93","https://git.kernel.org/stable/c/5474e6e17a262db45c60575c73f70210f5c7001f","https://nvd.nist.gov/vuln/detail/CVE-2026-52952"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-863","CWE-670"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-53053","cve":"CVE-2026-53053","aliases":[],"title":"Linux kernel (drivers/iommu/amd): On AMD hosts the Device Table Entry copied to a DMA-alias device is looked up using","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/amd)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"On AMD hosts the Device Table Entry copied to a DMA-alias device is looked up using the wrong source device ID, so an alias can be programmed with another device's - or a stale - translation table. The aliased function then DMAs through an IOMMU domain that does not belong to it, which is exactly the cross-tenant DMA the IOMMU is there to prevent.","attack_vector":"No tenant action is required - the wrong Device Table Entry is installed when the device is set up. It bites on any AMD-Vi host where a passed-through function has a PCI DMA alias: devices behind PCIe-to-PCI bridges and multi-function endpoints covered by alias quirks, which includes a lot of real GPU and NIC topologies. From then on the tenant's device translates through the wrong domain.","remediation":"Update to a stable kernel carrying commits dbd76a53 / 20b3c566 (the CNA's fixed-version field on this record is not a usable target - confirm the commits in your distro kernel). Interim: audit which passthrough devices have DMA aliases (check the IOMMU group membership against the PCI topology) and avoid handing out aliased functions to tenants until patched.","references":["https://git.kernel.org/stable/c/dbd76a537d8cb814e7f5b795ab21ecb7949c821d","https://git.kernel.org/stable/c/20b3c566e2702e5d4d0545be8a97029a2eebcc0e","https://nvd.nist.gov/vuln/detail/CVE-2026-53053"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-672","CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-53057","cve":"CVE-2026-53057","aliases":[],"title":"Linux kernel (drivers/iommu/riscv): The RISC-V IOMMU driver updated device-directory and process-directory entries","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/riscv)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"The RISC-V IOMMU driver updated device-directory and process-directory entries without issuing the invalidations the specification requires, so the hardware keeps using cached device and PASID context after the kernel has moved or freed it. A device therefore keeps translating through a context the kernel believes is gone - a stale mapping that survives detach, which is a DMA window into whatever that memory becomes.","attack_vector":"Reached on any domain attach/detach or PASID setup for a device behind a RISC-V IOMMU - a tenant closing or rebinding a passthrough device is enough. Hardware-conditional: this affects RISC-V platforms only and is not reachable on the x86 or Arm nodes that make up essentially all GPU fleets today. Track it only if RISC-V hosts are in the estate.","remediation":"Update to a stable kernel carrying commits 3f917d9b / d99d1c13 on RISC-V hosts. No action needed on x86 (VT-d/AMD-Vi) or Arm SMMU nodes.","references":["https://git.kernel.org/stable/c/3f917d9bff68600f77561900f3145bd4706dc840","https://git.kernel.org/stable/c/d99d1c13faa793ff1abab0d20ab6473c838081b3","https://nvd.nist.gov/vuln/detail/CVE-2026-53057"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-863","CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-53188","cve":"CVE-2026-53188","aliases":[],"title":"Linux kernel (drivers/infiniband/core): The RDMA user-capability check identified the capability file only by device","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/core)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"The RDMA user-capability check identified the capability file only by device number, and character and block device numbers alias. A tenant that opens a block device with a colliding major/minor can pass that fd off as an authentic ucap file and be granted an RDMA user capability it was never issued - for example privileged mlx5 control or raw-QP access. The CNA scored it scope-changed, i.e. the privilege gained reaches beyond the caller's own container.","attack_vector":"A tenant container that holds /dev/infiniband/* and can open any block device node whose dev_t matches the ucap character device. Both are things a container with device access routinely has. No fabric peer and no host root needed.","remediation":"Update to a stable kernel carrying 96b6e98ff12d (or aa181287ebdc / 4a1b1ac27446) and reboot. Interim: do not expose block device nodes and /dev/infiniband/* to the same untrusted container, and drop the RDMA ucap devices from tenant containers that do not need privileged verbs.","references":["https://git.kernel.org/stable/c/96b6e98ff12d50ed5817230c6f1188e1150d225d","https://git.kernel.org/stable/c/aa181287ebdcc53ee0ba5c2f8243e2d541ebc19b","https://nvd.nist.gov/vuln/detail/CVE-2026-53188"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-53240","cve":"CVE-2026-53240","aliases":[],"title":"Linux kernel (net/xfrm): An unlocked read of the IPTFS reassembly state lets two CPUs disagree about who owns a socket","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"An unlocked read of the IPTFS reassembly state lets two CPUs disagree about who owns a socket buffer, so the receive path trims and frees an skb that reassembly (or the drop timer) has already released. A remote peer that can keep two reassembly contexts in flight gets a use-after-free in the skbuff slab on the decrypt path - kernel heap corruption reachable from the encrypted fabric.","attack_vector":"Driven entirely by inbound ESP traffic on an IPTFS SA plus concurrency: the attacker only has to send interleaved partial IPTFS payloads so reassembly completes on one CPU while another is still in __input_process_payload. Reachable by any peer holding the SA - a peer node, a compromised node in the fleet, or the remote end of a tenant overlay. Conditional on IPTFS mode being configured; no tenant device node and no local privilege on the victim node needed.","remediation":"Boot a kernel carrying the fix commits below (no fixed stable version published). Interim control: disable IP-TFS mode on cluster SAs, or blacklist xfrm_iptfs where it is not in use, so the reassembly path is not reachable at all.","references":["https://git.kernel.org/stable/c/8d9a79fbf5172d9c4c0146057af2360913265a11","https://git.kernel.org/stable/c/ff2ee35b6ce5fa8a8e24ea50b15733d5c8780198","https://nvd.nist.gov/vuln/detail/CVE-2026-53240"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-416","CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-53281","cve":"CVE-2026-53281","aliases":[],"title":"Linux kernel (drivers/iommu/intel): When the PASID is not found on the device list, VT-d runs the teardown anyway and","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/intel)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"When the PASID is not found on the device list, VT-d runs the teardown anyway and decrements a reference count it never took, so the domain's refcount can reach zero while other devices are still attached to it. The IOMMU domain is then freed under live devices - a use-after-free on the translation state of unrelated passthrough devices that share that domain, plus a NULL dereference on the simpler path.","attack_vector":"A tenant detaching or closing a PASID-attached device through /dev/vfio/* plus /dev/iommu drives the teardown; hitting the not-found case takes a detach that races another detach or a domain replacement. Conditional on VT-d scalable mode with PASID in use. The blast radius is other devices sharing the same domain, which is what makes this cross-tenant rather than self-inflicted.","remediation":"Update to a stable kernel carrying commits 9022cb9a / cdfe3c9f (the record's fixed-version list predates the fix and is not a usable target). Interim: avoid sharing one IOMMU domain across devices belonging to different tenants, and disable PASID/scalable mode where SVA is not required.","references":["https://git.kernel.org/stable/c/9022cb9ac0c2a72a57fa8ebf92ac74f953ca0153","https://git.kernel.org/stable/c/cdfe3c9f2c9e28a8651ee463c88ad191ced2f840","https://nvd.nist.gov/vuln/detail/CVE-2026-53281"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-53322","cve":"CVE-2026-53322","aliases":[],"title":"Linux kernel (drivers/vfio/pci): When a tenant closes its passed-through PCI device, vfio disables the function before","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/vfio/pci)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"When a tenant closes its passed-through PCI device, vfio disables the function before it revokes the dma-buf exports over that device's BARs. In the window between the two, the BAR pages stay mapped while the underlying resources are released back to the kernel, so the departing tenant - or anyone it handed the dma-buf fd to - keeps read/write access to MMIO that a different driver, and potentially another tenant's device, now owns.","attack_vector":"A tenant holding /dev/vfio/<group> exports a dma-buf over its device BARs, keeps or passes that fd, then closes the device fd. Entirely reachable from inside the container with only the passthrough device node; requires a kernel with vfio-pci dma-buf export support and that feature in use.","remediation":"Update to a stable kernel carrying commits 4f1000a3 / d9770870. Interim: do not enable or permit the vfio-pci dma-buf export feature for tenant-held devices, and do not allow dma-buf fds to be passed out of the tenant's namespace.","references":["https://git.kernel.org/stable/c/4f1000a30f67cf7d328059242776a858611d5ef9","https://git.kernel.org/stable/c/d97708701434ce72968e771976aaf9d3438fcafd","https://nvd.nist.gov/vuln/detail/CVE-2026-53322"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-53359","cve":"CVE-2026-53359","aliases":[],"title":"Linux kernel (arch/x86/kvm/mmu): Shadow-page lookup reuses a page without comparing its role, so a direct (2MB) shadow","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (arch/x86/kvm/mmu)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Shadow-page lookup reuses a page without comparing its role, so a direct (2MB) shadow page gets reused for a 4KB indirect mapping. The rmap entry recorded under the walked GFN is never removed when the page is zapped, and after the memslot is dropped the freed shadow page is still reachable through that stale rmap. Any later rmap walk - dirty logging, MMU notifier invalidation - dereferences freed host kernel memory.","attack_vector":"Needs the shadow MMU plus a writer that KVM's write tracking does not see: the guest's PDE is changed from outside the guest and then a memslot is deleted. In practice that means a process holding /dev/kvm acting as its own VMM - a tenant with nested virtualization enabled, or any local user with /dev/kvm - not a container tenant with no KVM access. The MMU-notifier trigger (e.g. MADV_DONTNEED, host memory reclaim) then fires on its own.","remediation":"Update to a kernel with the referenced stable commits; take it together with CVE-2026-46113, which is the first half of the same hole. Interim: do not hand /dev/kvm to tenants and disable nested virtualization on unpatched nodes.","references":["https://git.kernel.org/stable/c/b1337aae5e194324e4810d561764e7793f8b3864","https://git.kernel.org/stable/c/9291654d69e08542de37755cebe4d5b02c3170d1","https://nvd.nist.gov/vuln/detail/CVE-2026-53359"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-53360","cve":"CVE-2026-53360","aliases":[],"title":"Linux KVM - GHCB v2+ scratch area location enforcement: KVM did not require the GHCB software scratch area to live","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux KVM - GHCB v2+ scratch area location enforcement","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"KVM did not require the GHCB software scratch area to live inside the GHCB's own shared buffer when GHCB v2+ is in use, as the spec demands. A guest can therefore point the scratch area at memory outside the shared region and get the host to read or write there on its behalf - a confused-deputy path from a confidential guest into host memory. Guest-to-host escape shape, and the CVSS 8.8 reflects it.","attack_vector":"From inside an SEV-ES/SNP guest via the GHCB protocol - tenant-reachable with no host privilege.","remediation":"Fixed in the Linux kernel - KVM/x86 SEV code or the ccp/PSP driver. Take the distro kernel update (RHEL/Rocky, Ubuntu, SLES) and **reboot the host**; SEV/SNP hypervisor paths cannot be live-patched in any meaningful way, and SNP platform init/shutdown is not safe to cycle under running guests. Drain confidential-VM tenants, reboot, then re-admit. No firmware, VBIOS or AGESA step needed, which makes this one of the cheaper classes of SEV fix to roll out. Top-of-queue for any node hosting tenant-supplied confidential VMs.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53360"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-07-04"},{"id":"CVE-2026-53374","cve":"CVE-2026-53374","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): Memory is handed to a consumer without being","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Memory is handed to a consumer without being initialised or cleared in the amdgpu GEM/VM/command-submission ioctl surface. Whatever the previous owner left behind is readable - and on a GPU node the previous owner is very often a different tenant's job. This is the classic residual-data leak between workloads sharing a card: model weights, activations, keys or tokens from the prior tenant can surface in a fresh allocation. Upstream fix: drm/amdgpu: zero-initialize GART table on allocation","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53374","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-07-19"},{"id":"CVE-2026-53375","cve":"CVE-2026-53375","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu/vce): A correctness defect in the amdgpu kernel driver core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu/vce)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu kernel driver core reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/vce: Prevent partial address patches","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53375","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-07-19"},{"id":"CVE-2026-56340","cve":"CVE-2026-56340","aliases":[],"title":"vLLM (sparse tensor validation): Missing sparse-tensor invariant checks in multimodal embeddings","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (sparse tensor validation)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Missing sparse-tensor invariant checks in multimodal embeddings → memory corruption","attack_vector":"Unauthenticated request to the serving port","remediation":"Upgrade past 0.13.0; PyTorch disables sparse invariant checks by default, so the fix must be in vLLM","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-56340"],"status":"curated","published":"2026-06-20"},{"id":"CVE-2026-57516","cve":"CVE-2026-57516","aliases":[],"title":"Ray (WebDataset reader): Unsafe deserialization","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ray (WebDataset reader)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Unsafe deserialization → RCE from a malicious tar archive","attack_vector":"Customer-supplied dataset consumed by a Ray Data pipeline","remediation":"Upgrade to 2.56.0+. Dataset files are now an RCE vector, not just model files","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-57516"],"status":"curated","published":"2026-07-01"},{"id":"CVE-2026-59093","cve":"CVE-2026-59093","aliases":[],"title":"Weaviate: RBAC role assignment does not verify the assigner holds the granted permissions","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Weaviate","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"RBAC role assignment does not verify the assigner holds the granted permissions → privilege escalation","attack_vector":"Authenticated tenant of a shared Weaviate","remediation":"Upgrade to 1.38.0+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-59093"],"status":"curated","published":"2026-07-02"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-22","CWE-639"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2026-59733","cve":"CVE-2026-59733","aliases":[],"title":"rclone (serve restic --private-repos): --private-repos is meant to confine each authenticated user to their own","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"rclone (serve restic --private-repos)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"--private-repos is meant to confine each authenticated user to their own repository, but a .. in the URL path walks out of it. An authenticated tenant reads, overwrites and deletes other tenants' backup repositories - so the isolation flag you deployed specifically for multi-user backup does not hold.","attack_vector":"Any authenticated user of an rclone serve restic endpoint running with --private-repos.","remediation":"Upgrade rclone and restart the serve restic instance. Audit repository contents and object timestamps for cross-user writes. Where possible back the separation with per-user storage credentials or separate buckets instead of trusting the path prefix.","references":["https://github.com/rclone/rclone/security/advisories/GHSA-fqj9-69pf-6pjg","https://nvd.nist.gov/vuln/detail/CVE-2026-59733"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-125"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-63807","cve":"CVE-2026-63807","aliases":[],"title":"Linux kernel (arch/x86/kvm/mmu): A guest that creates a hugepage mapping extending below the bounds of a memslot makes","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (arch/x86/kvm/mmu)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"A guest that creates a hugepage mapping extending below the bounds of a memslot makes KVM link a shadow page whose GFN is outside the slot; hugepage recovery then indexes the slot's lpage_info array out of bounds. The observed result is a host page fault on a vmalloc address - a guest-driven out-of-bounds access in host kernel memory that takes the whole node with it.","attack_vector":"Driven from inside the guest by arranging its own page tables: the guest installs a hugepage mapping whose range crosses the edge of a memslot, then the host's hugepage recovery worker walks it. Requires the shadow MMU (guest without EPT/NPT-backed TDP, nested guests, or dirty-logging-forced 4K mappings). Any tenant VM on an affected node can set this up.","remediation":"Update to a kernel with the referenced stable commits. Interim: avoid running tenants on the shadow-MMU path where possible - keep TDP enabled (kvm.tdp_mmu=Y / ept=1 / npt=1) and avoid long dirty-logging windows on unpatched nodes.","references":["https://git.kernel.org/stable/c/7b52008023b7facf40fba3ebe92449bda8ea53b9","https://git.kernel.org/stable/c/5cab1c989f938f5e1b9a0de66486f1fc2c28479b","https://nvd.nist.gov/vuln/detail/CVE-2026-63807"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-63919","cve":"CVE-2026-63919","aliases":[],"title":"Linux kernel (net/xfrm): Transport-mode reinjection stashes a network-namespace pointer in the socket buffer's control","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Transport-mode reinjection stashes a network-namespace pointer in the socket buffer's control block and dereferences it later from a deferred workqueue callback, without ever taking a reference. If the namespace is destroyed between queueing and the callback, the IPsec receive path runs against freed namespace state - a use-after-free driven by inbound encrypted traffic on a node where tenant namespaces come and go, which is every hour of every day in a container fleet.","attack_vector":"Two ingredients, both cheap on a multi-tenant node: inbound ESP traffic in transport mode that takes the deferred reinjection path (async crypto), and a network namespace being torn down. A tenant supplies both itself - keep a peer sending encrypted traffic into its namespace, then exit the namespace - and normal pod churn supplies the second half by accident. Reachable from the fabric side by any peer that can send ESP transport-mode packets to the node.","remediation":"Boot a kernel carrying the fix commits below (no fixed stable version published). Interim control: prefer tunnel mode over transport mode for tenant-facing SAs, and disable async/offloaded crypto so reinjection is not deferred, until nodes can be rebooted.","references":["https://git.kernel.org/stable/c/7ee59eda8820b758ed29e1cd3222359c7b97302c","https://git.kernel.org/stable/c/2df7059a18afb7d3aee6c36cad5d371c198111d4","https://nvd.nist.gov/vuln/detail/CVE-2026-63919"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-367","CWE-787"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-63937","cve":"CVE-2026-63937","aliases":[],"title":"Linux kernel (arch/x86/kvm/svm): KVM read Page State Change entries and indices out of a guest-writable buffer more","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (arch/x86/kvm/svm)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"KVM read Page State Change entries and indices out of a guest-writable buffer more than once, so a guest that rewrites the buffer between the bounds check and the use defeats the validation and steers the host past the buffer. This is what makes the neighbouring PSC bounds fixes bypassable, turning them back into a guest-to-host memory-safety break.","attack_vector":"A SEV-SNP guest races a second vCPU against the vCPU that submitted the PSC request, mutating the shared GHCB buffer while the host processes it. Entirely guest-side; requires only that the node runs SEV-SNP guests under kvm_amd.","remediation":"Update to a kernel carrying the referenced stable commits (no fixed release string published). This fix travels with the other GHCB scratch-area fixes in the same series - take the whole set, not just one. Interim: move SEV-SNP tenants off unpatched hosts.","references":["https://git.kernel.org/stable/c/bd232801ef1d1fd985d2d4ca3cd1d888303ca86f","https://git.kernel.org/stable/c/b1dfaa6f7a957726a6800135be3659fbe4bbf2a4","https://nvd.nist.gov/vuln/detail/CVE-2026-63937"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-20","CWE-863"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-64042","cve":"CVE-2026-64042","aliases":[],"title":"Linux kernel (drivers/vfio/pci): Vfio-pci exports a dma-buf over BAR memory without confirming those BAR resources were","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/vfio/pci)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Vfio-pci exports a dma-buf over BAR memory without confirming those BAR resources were ever actually reserved, so a tenant can obtain a mapping to physical MMIO the driver does not own. That is a handle onto whatever else lives at those addresses - another device's registers, or host-owned MMIO - reachable from inside the container.","attack_vector":"A tenant holding /dev/vfio/<group> calls the vfio-pci dma-buf feature ioctl on a device whose BAR resources were not reserved at probe (BAR sizing quirks, resource conflicts, or a device where the reservation silently failed). Requires the vfio-pci dma-buf export feature to be present and enabled; no host root.","remediation":"Update to a stable kernel carrying commits 8443cd44 / 702809da. Interim: do not enable the vfio-pci dma-buf export feature for tenant-held devices, and verify at bind time that each passthrough device's BARs were successfully reserved before handing the group to a tenant.","references":["https://git.kernel.org/stable/c/8443cd4497a4498c4b01058d76a92116244cb605","https://git.kernel.org/stable/c/702809dabdecca807bdd50cfdcc1c980feb2ba62","https://nvd.nist.gov/vuln/detail/CVE-2026-64042"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-416","CWE-459"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-64475","cve":"CVE-2026-64475","aliases":[],"title":"Linux kernel (drivers/vfio/pci): If vfio-pci device registration fails after the device joined the VGA arbiter, the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/vfio/pci)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"If vfio-pci device registration fails after the device joined the VGA arbiter, the arbiter keeps a callback registered against a vfio device cookie that is about to be freed. The VGA arbiter then calls into freed state on the next arbitration - a use-after-free on a node where a GPU failed to bind for passthrough, and the upstream note is explicit that it becomes exploitable again as soon as the callback follows drvdata.","attack_vector":"Requires a vfio-pci registration failure on a VGA-class device, so the trigger is host-side provisioning: binding a GPU to vfio-pci where a later registration step fails. Not tenant-initiated. Reachable only on devices that participate in VGA arbitration, which in practice means display-capable GPUs rather than headless datacenter accelerators.","remediation":"Update to a stable kernel carrying commits 0f2a35a0 / 8d65decd (5.10.261 / 5.12 / 5.13 are listed by the CNA for older branches). Interim: treat any vfio-pci bind failure on a VGA-capable GPU as requiring a node reboot before the device is offered to a tenant.","references":["https://git.kernel.org/stable/c/0f2a35a0c7ea7da347b814750eaa78adf3582381","https://git.kernel.org/stable/c/8d65decde9afd2bd78bcfffdc0df73b82a0b5509","https://nvd.nist.gov/vuln/detail/CVE-2026-64475"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-64516","cve":"CVE-2026-64516","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vce1): An out-of-bounds access in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vce1)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu firmware, ACPI and IP-block initialisation - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu/vce1: Fix VCE 1 firmware size and offsets","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-64516","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-07-25"},{"id":"CVE-2026-64522","cve":"CVE-2026-64522","aliases":["net/mlx5e eswitch mode block underflow on IPsec acquire SA"],"title":"Linux kernel mlx5_core IPsec offload / eswitch mode interlock: The acquire-SA path unconditionally calls","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel mlx5_core IPsec offload / eswitch mode interlock","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"The acquire-SA path unconditionally calls mlx5_eswitch_unblock_mode() without a matching block, underflowing the counter that is supposed to prevent eswitch mode transitions while IPsec offload is live. Once that interlock is broken, switchdev/legacy mode can be flipped out from under active offloads - and eswitch mode is what defines VF steering and isolation on the NIC. A remote packet is enough to start unwinding the enforcement point for tenant separation.","attack_vector":"Network-reachable, unauthenticated: a remote TCP SYN routed through an administrator-configured outbound IPsec policy reaches the vulnerable acquire-SA callback. No account on the host.","remediation":"Upgrade the host kernel to 7.1 or a stable backport (6.18.34, 7.0.11). Rolling reboot of every node running mlx5 IPsec full offload. Interim: if you are not depending on hardware IPsec offload, disable it on the mlx5 interfaces (config change, no reboot) to take the path out of reach.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-64522","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-64522.json"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-07-25"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-64561","cve":"CVE-2026-64561","aliases":[],"title":"Linux kernel (arch/x86/kvm/mmu): If reclaiming shadow pages invalidates the root a fault is being serviced against, KVM","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (arch/x86/kvm/mmu)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"If reclaiming shadow pages invalidates the root a fault is being serviced against, KVM maps into that invalid root and creates child shadow pages that inherit the invalid role, putting invalid pages on the active MMU list. That breaks the invariant the zapping code relies on, leaving live shadow-page-table state pointing at pages KVM believes are gone - the classic setup for a host-side use-after-free reachable from guest page faults.","attack_vector":"Reachable from an ordinary guest: fault in enough memory to push the shadow MMU into reclaiming pages while a root is being used, on any node where the shadow MMU is active (nested guests, or guests running without TDP). No host access required.","remediation":"Update to a kernel with the referenced stable commits. Interim: keep TDP MMU enabled and avoid exposing nested virtualization to tenants on unpatched nodes; consider raising kvm.mmu_shadow_page limits so reclaim is not hit under normal tenant load.","references":["https://git.kernel.org/stable/c/65c4f7a1028cf01a93a2762d679c289810ede990","https://git.kernel.org/stable/c/35e77467610c4a37cb0ff54ee56b85f73b1f5700","https://nvd.nist.gov/vuln/detail/CVE-2026-64561"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-64562","cve":"CVE-2026-64562","aliases":[],"title":"Linux kernel (arch/x86/kvm/vmx): Nested teardown freed the shadow VMCS page while vmcs01 still referenced it, and","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (arch/x86/kvm/vmx)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Nested teardown freed the shadow VMCS page while vmcs01 still referenced it, and because the clear is asynchronous a vCPU that migrates afterwards makes the CPU execute VMCLEAR against a page already returned to the allocator. That is hardware writing into reallocated host kernel memory - a use-after-free in the host driven by ordinary nested-VMX activity in a guest.","attack_vector":"Guest-driven: a tenant uses VMX inside its VM and then tears the nested state down (VMCLEAR/VMXOFF, or nested state teardown), and the freed page is written when the vCPU is next scheduled on a different physical CPU. Requires nested VMX exposed to the guest (Intel host, kvm_intel nested=1, VMX in guest CPUID). No host privilege needed.","remediation":"Update to a stable kernel with the linked fix (no fixed release enumerated; take the branch carrying commit dc3eecfa219e). Interim control: disable nested virtualization for tenant guests (kvm_intel.nested=0) - that removes the whole path.","references":["https://git.kernel.org/stable/c/dc3eecfa219ebc9d01eaf7d1abd1441efe884dab","https://git.kernel.org/stable/c/4f50e6aec16f69627dbad5704d1e90a255d766a7","https://nvd.nist.gov/vuln/detail/CVE-2026-64562"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-68107","cve":"CVE-2026-68107","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn4): A correctness defect in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn4)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/vcn4: avoid rereading IB param length","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68107","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-10"},{"id":"CVE-2026-68108","cve":"CVE-2026-68108","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vce): An out-of-bounds access in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vce)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu firmware, ACPI and IP-block initialisation - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu/vce: fix integer overflow in image size","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68108","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-10"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-68329","cve":"CVE-2026-68329","aliases":[],"title":"Linux kernel (drivers/iommu/amd): Iommu_completion_wait() returned without waiting whenever another CPU had already","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/amd)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Iommu_completion_wait() returned without waiting whenever another CPU had already queued a completion-wait, so a CPU could free page-table pages while the AMD IOMMU was still walking the translations those pages describe. The tenant's device then DMAs through a page table backed by memory the host has reallocated - a stale mapping that is a direct tenant-to-host DMA read/write escape, and the most serious bug in this batch.","attack_vector":"Any high-rate DMA map/unmap workload on an AMD-Vi host reaches it: a tenant hammering unmap through a passed-through NIC or GPU, or issuing VFIO_IOMMU_UNMAP_DMA in a loop from /dev/vfio/*, while a second CPU does concurrent IOMMU work. No host root, no fabric access; the race gets easier the busier the node is, so a co-tenant generating IOMMU traffic helps the attacker.","remediation":"Update to a stable kernel carrying commits ab7faf5a / 93494bd4 on every AMD-Vi node. Interim controls are weak: reducing unmap rate or pinning tenants to fewer sockets only narrows the window. Treat this as a mandatory reboot on AMD hosts that run passthrough for tenants.","references":["https://git.kernel.org/stable/c/ab7faf5a172ebfdc423ebb3eea4d472740de82f9","https://git.kernel.org/stable/c/93494bd446396c257fb589f59894577e96e406e2","https://nvd.nist.gov/vuln/detail/CVE-2026-68329"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-68432","cve":"CVE-2026-68432","aliases":[],"title":"Linux VXLAN driver (CAP_NET_ADMIN check on changelink across netns): A VXLAN tunnel's `changelink()` operates across","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux VXLAN driver (CAP_NET_ADMIN check on changelink across netns)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"A VXLAN tunnel's `changelink()` operates across two network namespaces — the device's namespace and the sticky underlay namespace — but the capability check only covers the device's. Once a VXLAN device has been created in or moved to a different namespace, a caller with CAP_NET_ADMIN in only one of them can reconfigure the tunnel's underlay side. In a container platform, network namespaces are the tenant boundary and CAP_NET_ADMIN inside a namespace is something you grant routinely; this turns namespace-local privilege into control over the underlay encapsulation that other tenants share.","attack_vector":"A container or tenant holding CAP_NET_ADMIN in its own network namespace, against a VXLAN device whose underlay namespace differs from its device namespace.","remediation":"Kernel upgrade plus host reboot across container hosts. Interim: do not grant CAP_NET_ADMIN to tenant containers — a container-runtime policy change and one of the highest-value single restrictions available on a shared GPU host, since it also closes a long tail of similar netlink-reachable issues.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68432"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-08-12"},{"id":"CVE-2026-72045","cve":"CVE-2026-72045","aliases":[],"title":"Linux octeontx2-af (Marvell OCTEON CN10K, LMTLINE mailbox handler): The OCTEON CN10K admin-function mailbox handler","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux octeontx2-af (Marvell OCTEON CN10K, LMTLINE mailbox handler)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"The OCTEON CN10K admin-function mailbox handler uses a caller-supplied `base_pcifunc` as a direct index into the LMT map table, reading *another* PCI function's LMTLINE physical base address and copying it into the caller's own map-table entry. The mailbox dispatcher authenticates the requesting function, then ignores that authentication for the field that selects whose memory window you get. A VF assigned to one tenant can therefore point itself at another function's doorbell region on a shared OCTEON DPU. This is a textbook SR-IOV isolation break: the hardware isolation exists, the software hands out the key.","attack_vector":"A tenant holding an OCTEON VF — an SR-IOV virtual function passed into a VM or container — sending a crafted mailbox request to the admin function.","remediation":"Kernel upgrade plus host reboot on every node with Marvell OCTEON CN10K networking. Rolling drain across the fleet; nothing to flash. Until patched, do not assign OCTEON VFs to untrusted tenants — the isolation you are relying on is not being enforced.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-72045"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-08-15"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-863"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-72136","cve":"CVE-2026-72136","aliases":[],"title":"Linux kernel (net/xfrm): The rtnetlink changelink path for xfrm interfaces checked CAP_NET_ADMIN only against the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"The rtnetlink changelink path for xfrm interfaces checked CAP_NET_ADMIN only against the namespace the request came from, not the namespace the interface actually lives in. A caller privileged in its own network namespace can therefore rewrite an xfrm interface belonging to a different namespace - changing its if_id, link, or the netns it is bound to. That reassigns which SA the interface's traffic is encrypted under, which is a direct route to another tenant's traffic being decrypted with the wrong key, steered to the wrong endpoint, or emitted in clear. The CNA marks the scope as changed, which is the right read.","attack_vector":"A tenant container that holds CAP_NET_ADMIN in its own user+network namespace and can see an xfrm interface whose link netns is elsewhere - the standard situation when the host or an orchestrator creates an xfrm device and moves it into a pod namespace, or shares one across namespaces. No fabric access, no host root, and no device node beyond an rtnetlink socket are needed.","remediation":"Boot a kernel carrying the fix commits below (no fixed stable version published in the record). Interim control: stop granting CAP_NET_ADMIN in tenant user namespaces; do not create xfrm interfaces in one namespace and expose them into another; keep the IPsec control plane entirely on the host side of the boundary.","references":["https://git.kernel.org/stable/c/04c1aa57d08471b1953bf27c84ac9b3d78d71831","https://git.kernel.org/stable/c/bdfd1c21d90e628a58a9de79e024cdfcbedfa15c","https://nvd.nist.gov/vuln/detail/CVE-2026-72136"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-72283","cve":"CVE-2026-72283","aliases":[],"title":"Linux kernel (arch/x86/kvm): When KVM failed to program the interrupt remapping table for irq bypass, it left a","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (arch/x86/kvm)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"When KVM failed to program the interrupt remapping table for irq bypass, it left a dangling pointer to the interrupt producer. For VFIO PCI the producer lives inside the per-vector context and is freed when the vector is disabled, so KVM later dereferences freed memory - a host kernel use-after-free in the interrupt path of a passthrough device.","attack_vector":"Requires device passthrough with interrupt bypass - exactly the configuration used for GPU and NIC passthrough in a GPU cloud. A tenant (or the VMM acting on tenant-controlled MSI-X configuration) disables an MSI-X vector on the assigned device while KVM's routing still references the producer; an IRTE update failure leaves the stale pointer behind. Needs /dev/vfio access plus an assigned device, i.e. any tenant VM with a passed-through GPU or NIC.","remediation":"Update to a kernel with the referenced stable commits. Interim: there is no clean workaround short of dropping IRQ bypass (disable posted interrupts / kvm halt_poll and IRQ bypass on the node) or not passing devices through, both of which cost GPU-VM performance - patch and reboot is the real fix.","references":["https://git.kernel.org/stable/c/d1379888cc4230bac647ec24ab83306afbd03e88","https://git.kernel.org/stable/c/d5560b6569cd05ba72c6b33427fbabc6ec46b8cf","https://nvd.nist.gov/vuln/detail/CVE-2026-72283"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-72286","cve":"CVE-2026-72286","aliases":[],"title":"Linux KVM - intra-host migration/mirroring of SEV-SNP VMs: KVM allowed intra-host migration and mirroring of SEV-SNP","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux KVM - intra-host migration/mirroring of SEV-SNP VMs","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"KVM allowed intra-host migration and mirroring of SEV-SNP VMs even though the feature was never fully implemented for SNP - SNP-specific state such as the guest request mutex and message counters is not carried across. Migrating or mirroring an SNP VM through this path lands it in an inconsistent confidential state, which is a route to breaking the guest's isolation and to host-side memory corruption. At 8.8 this is the most severe SEV-related kernel issue in the current set.","attack_vector":"Reachable through the KVM ioctl surface used for intra-host migration/mirroring - so a process with access to /dev/kvm and the ability to drive migration, i.e. the VMM. In a multi-tenant control plane that is your orchestrator, which makes control-plane compromise the realistic path in.","remediation":"Fixed in the Linux kernel - KVM/x86 SEV code or the ccp/PSP driver. Take the distro kernel update (RHEL/Rocky, Ubuntu, SLES) and **reboot the host**; SEV/SNP hypervisor paths cannot be live-patched in any meaningful way, and SNP platform init/shutdown is not safe to cycle under running guests. Drain confidential-VM tenants, reboot, then re-admit. No firmware, VBIOS or AGESA step needed, which makes this one of the cheaper classes of SEV fix to roll out. The upstream fix simply rejects the operation for SNP VMs. Until you are patched, disable intra-host migration and VM mirroring for SEV-SNP guests in your VMM configuration - that removes the reachable path without a reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-72286"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-08-15"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-20","CWE-190"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-72497","cve":"CVE-2026-72497","aliases":[],"title":"Linux kernel (drivers/infiniband/hw/bnxt_re): The variable-WQE send-queue slot count came straight from userspace with","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/hw/bnxt_re)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"The variable-WQE send-queue slot count came straight from userspace with no upper or lower bound, so a tenant could program a queue geometry the hardware cannot represent (above the 64K maximum, or zero). The CNA scored it scope-changed with full confidentiality, integrity and availability loss - the resulting adapter state escapes the requesting container's own boundary.","attack_vector":"A tenant container holding /dev/infiniband/uverbs* on a Broadcom bnxt_re NIC issues ibv_create_qp with a crafted slot count in variable-WQE mode. No fabric peer, no host root. Only bnxt_re nodes are affected.","remediation":"No fixed version is listed in the record - take the stable kernel carrying a59d815cbe66 (or dc95931b7e13) and reboot. Interim: remove /dev/infiniband/* from untrusted containers on bnxt_re nodes, or blacklist bnxt_re where RoCE is not needed.","references":["https://git.kernel.org/stable/c/a59d815cbe667929b693b5fa6716a074e6a31c5b","https://git.kernel.org/stable/c/dc95931b7e1326dacae547874bf38c092e5960d8","https://nvd.nist.gov/vuln/detail/CVE-2026-72497"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-72499","cve":"CVE-2026-72499","aliases":[],"title":"Linux bnxt_re RoCE driver (CQ toggle page use-after-free): The completion-queue variant of the toggle-page","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux bnxt_re RoCE driver (CQ toggle page use-after-free)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"The completion-queue variant of the toggle-page use-after-free in the Broadcom RoCE driver. Same reachability — the RDMA object lifecycle, which tenants exercise every time they set up or tear down a connection — and the same concern that a stray firmware interrupt writes into memory the kernel has already handed back.","attack_vector":"Local user with RDMA verbs access, racing completion-queue destruction against a notification-queue interrupt.","remediation":"Kernel/driver upgrade plus host reboot. Same rolling-drain cost as its SRQ sibling; both land in the same patch, so plan one reboot, not two.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-72499"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-08-15"},{"id":"CVE-2026-72500","cve":"CVE-2026-72500","aliases":[],"title":"Linux bnxt_re RoCE driver (SRQ toggle page use-after-free): A use-after-free in the Broadcom RoCE driver — the toggle","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux bnxt_re RoCE driver (SRQ toggle page use-after-free)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"A use-after-free in the Broadcom RoCE driver — the toggle page backing a shared receive queue is freed before firmware teardown completes, so a notification-queue interrupt arriving mid-destroy writes into freed memory. This is in the RDMA path, which means it is reachable from the queue-pair lifecycle that tenants drive directly when they open and close RDMA connections. A write-after-free driven by an interrupt in the RDMA control path is the shape of bug that turns into cross-tenant memory access on a shared RoCE fabric. Companion issue CVE-2026-72499 is the completion-queue equivalent.","attack_vector":"A local user with RDMA verbs access — i.e. any tenant running RoCE workloads — creating and destroying shared receive queues to race the teardown against an incoming NQ interrupt.","remediation":"Kernel/driver upgrade plus host reboot. On a RoCE cluster this is a full rolling reboot of every node using Broadcom RDMA, and jobs must be drained first. If you cannot patch quickly, restricting which containers get RDMA device access (verbs char devices) narrows who can drive the vulnerable lifecycle — a config change in your container runtime.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-72500","https://nvd.nist.gov/vuln/detail/CVE-2026-72499"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-08-15"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-787","CWE-682"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-74277","cve":"CVE-2026-74277","aliases":[],"title":"Linux kernel (drivers/iommu): Every peer-to-peer segment in a scatter-gather list inherits the length of the first","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Every peer-to-peer segment in a scatter-gather list inherits the length of the first segment, so the IOMMU is programmed with DMA lengths that do not match the actual buffers. Devices read and write past the end of short segments and silently truncate long ones - cross-buffer memory corruption on exactly the peer-to-peer path that GPUDirect and GPU-to-NIC traffic runs over.","attack_vector":"Any workload that drives PCI peer-to-peer DMA through dma_map_sg reaches this: GPUDirect RDMA from a tenant holding /dev/infiniband/uverbs* plus /dev/dri/renderD*, NVMe peer-to-peer, or GPU-to-NIC staging. No elevated privilege is required and no unusual configuration beyond PCI P2PDMA being in use with multi-segment scatterlists, which is the normal case for large transfers.","remediation":"Update to a stable kernel carrying commits 8646f00c / db50fb87 (no fixed-version list was published by the kernel CNA - confirm the backport with your distro). Interim: disable PCI P2PDMA / GPUDirect peer-to-peer for tenant workloads if your stack allows it, since the corruption only occurs on the P2PDMA branch of iommu_dma_map_sg().","references":["https://git.kernel.org/stable/c/8646f00ce021e49f4f05bc1d4060a0c27b25d0e1","https://git.kernel.org/stable/c/db50fb87015b955a5a0c155293b2dd40d63a3b9e","https://nvd.nist.gov/vuln/detail/CVE-2026-74277"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-74328","cve":"CVE-2026-74328","aliases":[],"title":"Linux kernel (drivers/iommu/iommufd): Iommufd tears down the page-tracking state behind a dma-buf backed IOAS mapping","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/iommufd)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Iommufd tears down the page-tracking state behind a dma-buf backed IOAS mapping while the exporter can still fire invalidation callbacks into it. A tenant can race that teardown into a use-after-free on the structure that decides which host pages its device is allowed to touch, which is a step toward host memory disclosure or a host kernel crash that takes the whole node down.","attack_vector":"A tenant container holding /dev/iommu together with a device fd under /dev/vfio/* sets up an IOAS mapping over a dma-buf (the vfio-pci BAR dma-buf export is the usual source), then races close/detach against the exporter's invalidation callback. No host root and no fabric access are needed - only the passthrough device nodes inside the container, on a kernel new enough to have iommufd dma-buf support.","remediation":"Boot a stable kernel carrying the fix commits (0507fced / f2d70dbd); the kernel CNA published no fixed-version list for this one, so track the commits into your distro kernel. Interim: do not expose /dev/iommu directly to tenant containers - keep passthrough behind a VMM the operator controls - and do not enable the vfio-pci dma-buf export path for tenant-held devices.","references":["https://git.kernel.org/stable/c/0507fcedbdcc87281ef8639c045fc9980363bbb6","https://git.kernel.org/stable/c/f2d70dbd3dcefa8e3c380beff9c31f5f033a4221","https://nvd.nist.gov/vuln/detail/CVE-2026-74328"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-74520","cve":"CVE-2026-74520","aliases":[],"title":"Linux kernel (drivers/iommu): An I/O page-fault group is handed to userspace through iommufd while still sitting on the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"An I/O page-fault group is handed to userspace through iommufd while still sitting on the generic IOPF pending list, so a detach or hardware-page-table replacement frees the group out from under the fault fd. A later read, response, or cleanup touches freed memory - a use-after-free in the very path that arbitrates what a tenant's device is allowed to translate.","attack_vector":"A tenant container holding /dev/iommu with a PRI/IOPF-capable device generates page faults and concurrently detaches the device or replaces its HWPT. Reachable with a GPU or accelerator using SVA, or any passthrough device with ATS+PRI enabled. Conditional on PRI/IOPF being enabled on the device; no host root.","remediation":"Update to a stable kernel carrying commits 6da8f374 / 4e74a369. Interim: disable ATS/PRI (and therefore IOPF) for tenant passthrough devices where the workload does not need demand paging, and do not expose /dev/iommu directly to untrusted containers.","references":["https://git.kernel.org/stable/c/6da8f37419dd4c456f26fc203f04e000186f4b3d","https://git.kernel.org/stable/c/4e74a369236424114b94cf6a9f5ff9e848b430b4","https://nvd.nist.gov/vuln/detail/CVE-2026-74520"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-74527","cve":"CVE-2026-74527","aliases":[],"title":"Linux octeontx2-af (VF clobbering shared CGX PKIND state): PF and VF NIX logical functions that share a CGX MAC reuse","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux octeontx2-af (VF clobbering shared CGX PKIND state)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"PF and VF NIX logical functions that share a CGX MAC reuse the same hardware packet-parse (PKIND) programming, and a VF allocating a NIX LF could reset the MAC's RX PKIND and default TX parse configuration set up by the PF. Parse configuration determines how the adapter interprets every frame on that MAC — so a tenant VF can change packet interpretation for everyone sharing the physical port, including the operator. The fix adds an explicit permission check that was simply absent.","attack_vector":"A tenant VF on an OCTEON adapter allocating a NIX logical function on a CGX MAC shared with the PF or with other tenants.","remediation":"Kernel upgrade plus host reboot on OCTEON nodes. Rolling drain. Structurally, avoid sharing a single CGX MAC between an operator PF and tenant VFs where the platform allows dedicating MACs instead — a provisioning-layout decision rather than a patch.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-74527"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-08-15"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-269"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2026-75481","cve":"CVE-2026-75481","aliases":[],"title":"SkyPilot (API server, service account role update authorization): SkyPilot never checks whether the caller is entitled","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"SkyPilot (API server, service account role update authorization)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"SkyPilot never checks whether the caller is entitled to grant the administrator role when updating service account permissions. Any authenticated user creates a service account, promotes it to admin, then authenticates with its bearer token and owns every user and every workspace - which on SkyPilot means control of the clusters and cloud accounts it provisions GPUs in.","attack_vector":"Any authenticated SkyPilot user with API server access. No admin rights needed to start.","remediation":"Upgrade SkyPilot past the fixed commit (8a3e0025) and restart the API server. Then enumerate existing service accounts and their roles - an attacker who already ran this leaves a legitimate-looking admin service account behind that the upgrade does not remove.","references":["https://github.com/skypilot-org/skypilot/issues/9846","https://nvd.nist.gov/vuln/detail/CVE-2026-75481"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-8828","cve":"CVE-2026-8828","aliases":[],"title":"ChromaDB (Rust): Missing authorization validation","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"ChromaDB (Rust)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Missing authorization validation → arbitrary read/write across collections","attack_vector":"Authenticated tenant","remediation":"Upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-8828"],"status":"curated","published":"2026-06-12"},{"id":"NCVD-2021-006-infiniband-subnet-management-sub","cve":null,"aliases":["InfiniBand P_Key enforcement gap","Q_Key weakness","SM MAD trust","subnet manager spoofing","GUID spoofing","M_Key optional"],"title":"InfiniBand subnet management - Subnet Management Packets (SMPs), P_Key/Q_Key partition enforcement, port/node GUIDs","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"InfiniBand subnet management - Subnet Management Packets (SMPs), P_Key/Q_Key partition enforcement, port/node GUIDs","year":"2021","cvss_score":8.8,"severity":"high","kev":false,"impact":"InfiniBand's tenant boundary is the partition key, and its fabric control plane is the subnet manager speaking unauthenticated management datagrams. Three weaknesses compound. First, P_Key enforcement is a per-port switch capability that must be explicitly turned on; where it is left off - a common default on smaller fabrics - partition membership is advisory and any node can talk to any other. Second, Q_Keys guarding unreliable-datagram traffic are frequently left at well-known or predictable values, so UD traffic including management traffic is forgeable. Third, subnet management packets are protected only by the optional M_Key, and if M_Key is unset or set to zero any node on the subnet can issue SMPs - reprogramming LIDs and routing tables, or standing up a rogue subnet manager that takes over the fabric. A node with a spoofed GUID can inherit another node's partition membership outright.","attack_vector":"From any host attached to the IB subnet, the attacker sends SMPs on QP0 (unauthenticated when M_Key is unset) to read and rewrite switch forwarding tables and port configuration, or announces a higher-priority subnet manager and wins the SM election. Partition membership is then whatever the attacker says it is, and traffic can be mirrored, redirected, or blackholed. GUID spoofing is a driver/firmware-level parameter on many adapters. None of this requires exploiting a software defect - the specification permits all of it when the optional protections are not configured.","remediation":"Config change, and it is cheap relative to the exposure - do it this quarter. Set a non-zero M_Key with lease protection on every port so SMPs from unauthorised nodes are rejected; enable P_Key enforcement on all switch ports facing tenant hosts; assign a distinct, non-default Q_Key per tenant; and pin the subnet manager by priority with SM handover disabled, ideally running OpenSM or NVIDIA UFM on a management node tenants cannot reach. Applying M_Key and P_Key enforcement is pushed through the SM configuration and takes effect on the next sweep - no switch reload and no host reboot, though a botched M_Key rollout can lock you out of your own fabric, so stage it. Audit with ibnetdiscover/saquery that enforcement is actually on, since it silently defaults off.","references":["https://www.usenix.org/conference/usenixsecurity21/presentation/rothenberger","https://netsec.ethz.ch/publications/papers/sec21summer-redmark.pdf","https://arxiv.org/abs/2202.08080"],"status":"curated","tags":["tenant-isolation"]},{"id":"NCVD-2021-012-infiniband-subnet-management-sub","cve":null,"aliases":["InfiniBand P_Key enforcement gap","Q_Key weakness","SM MAD trust","subnet manager spoofing","GUID spoofing","M_Key optional"],"title":"InfiniBand subnet management - Subnet Management Packets (SMPs), P_Key/Q_Key partition enforcement, port/node GUIDs","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"InfiniBand subnet management - Subnet Management Packets (SMPs), P_Key/Q_Key partition enforcement, port/node GUIDs","year":"2021","cvss_score":8.8,"severity":"high","kev":false,"impact":"InfiniBand's tenant boundary is the partition key, and its fabric control plane is the subnet manager speaking unauthenticated management datagrams. Three weaknesses compound. First, P_Key enforcement is a per-port switch capability that must be explicitly turned on; where it is left off - a common default on smaller fabrics - partition membership is advisory and any node can talk to any other. Second, Q_Keys guarding unreliable-datagram traffic are frequently left at well-known or predictable values, so UD traffic including management traffic is forgeable. Third, subnet management packets are protected only by the optional M_Key, and if M_Key is unset or set to zero any node on the subnet can issue SMPs - reprogramming LIDs and routing tables, or standing up a rogue subnet manager that takes over the fabric. A node with a spoofed GUID can inherit another node's partition membership outright.","attack_vector":"From any host attached to the IB subnet, the attacker sends SMPs on QP0 (unauthenticated when M_Key is unset) to read and rewrite switch forwarding tables and port configuration, or announces a higher-priority subnet manager and wins the SM election. Partition membership is then whatever the attacker says it is, and traffic can be mirrored, redirected, or blackholed. GUID spoofing is a driver/firmware-level parameter on many adapters. None of this requires exploiting a software defect - the specification permits all of it when the optional protections are not configured.","remediation":"Config change, and it is cheap relative to the exposure - do it this quarter. Set a non-zero M_Key with lease protection on every port so SMPs from unauthorised nodes are rejected; enable P_Key enforcement on all switch ports facing tenant hosts; assign a distinct, non-default Q_Key per tenant; and pin the subnet manager by priority with SM handover disabled, ideally running OpenSM or NVIDIA UFM on a management node tenants cannot reach. Applying M_Key and P_Key enforcement is pushed through the SM configuration and takes effect on the next sweep - no switch reload and no host reboot, though a botched M_Key rollout can lock you out of your own fabric, so stage it. Audit with ibnetdiscover/saquery that enforcement is actually on, since it silently defaults off.","references":["https://www.usenix.org/conference/usenixsecurity21/presentation/rothenberger","https://netsec.ethz.ch/publications/papers/sec21summer-redmark.pdf","https://arxiv.org/abs/2202.08080"],"status":"curated","tags":["tenant-isolation"]},{"id":"NCVD-2022-001-infiniband-roce-local-rnic-kerne","cve":null,"aliases":["NeVerMore","RDMA local packet injection","Taranov et al., arXiv:2202.08080"],"title":"InfiniBand/RoCE local RNIC - kernel bypass path shared by all local processes: NeVerMore showed that an unprivileged","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"InfiniBand/RoCE local RNIC - kernel bypass path shared by all local processes","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"NeVerMore showed that an unprivileged local user can inject packets into any RDMA connection created on the same local network controller, bypassing the operating system and kernel entirely. On a multi-tenant node this is a complete break of the host's process boundary at the fabric layer: a low-privilege container that has been given a /dev/infiniband device can forge traffic belonging to a co-resident tenant's queue pairs, and from there acquire unauthorized block access to NVMe-oF targets that trusted the RDMA connection as the authenticator. In a GPU cluster this converts one compromised training pod into read/write access to the whole tenant's remote datasets.","attack_vector":"The attacker needs only ordinary access to the local RDMA verbs device - which is how RDMA is normally exposed to containers, since kernel bypass is the whole point. They construct raw queue pairs and emit packets whose transport headers name another local process's QP, because the RNIC does not verify that the submitting context owns the source identity it stamps on the wire. The paper implements four RDMA-protocol attacks and seven NVMe-oF attacks and verifies them against both SPDK and the Linux kernel NVMe-oF implementations.","remediation":"No patch closes this in general. Config change with real cost: stop sharing one RNIC/PF across trust boundaries - give each tenant a dedicated SR-IOV VF or a dedicated physical NIC, and never mount /dev/infiniband into an untrusted container. On BlueField DPUs, terminate the RDMA connection on the DPU so the host cannot forge fabric identity - that is a hardware/topology change. Layer real authentication above the transport: enable NVMe-oF in-band DH-HMAC-CHAP so block access does not rest on the RDMA connection alone (config change on target and initiator, no reboot). Treat 'RDMA device in an untrusted container' as equivalent to root on the fabric.","references":["https://arxiv.org/abs/2202.08080","https://arxiv.org/abs/1903.09355"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"],"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2025-017-vllm-openai-compatible-server-qw","cve":null,"aliases":["GHSA-79j6-g2m3-jgfw","CVE-2025-9141 (reserved)"],"title":"vLLM OpenAI-compatible server (qwen3_coder tool-call parser): Code execution inside the serving process, which on a GPU","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM OpenAI-compatible server (qwen3_coder tool-call parser)","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"Code execution inside the serving process, which on a GPU node means execution next to the model weights, the GPU, and whatever credentials the inference pod holds. vLLM's Qwen3-Coder tool parser falls back to Python eval() when converting a tool-call parameter whose type it does not recognise. Any authenticated API user who can steer the model into emitting a crafted argument gets that string evaluated on the server. For an inference provider this is the worst shape of the multi-tenant problem: one customer's prompt reaches the interpreter on a box that is simultaneously serving other customers' requests, giving access to in-flight prompts and completions, the model artifacts on local disk, and the node's service-account token and cloud identity. Prompt content is attacker-controlled by definition in a serving product, so the 'requires authentication' qualifier buys very little.","attack_vector":"Network, any authenticated API client. Requires the server to be started with --enable-auto-tool-choice and --tool-call-parser qwen3_coder; the payload arrives as an ordinary completion request whose tool-call parameters carry an unrecognised type.","remediation":"Upgrade vLLM to 0.10.1.1 or later and restart the serving processes. If you cannot upgrade immediately, stop the server with the qwen3_coder parser or disable automatic tool choice — that removes the code path entirely. Longer term, run inference pods with a scoped service account and no ambient cloud credentials so an eval() foothold does not become a cluster foothold.","references":["https://github.com/vllm-project/vllm/security/advisories/GHSA-79j6-g2m3-jgfw","https://github.com/vllm-project/vllm/pull/21396"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-20","CWE-123","CWE-502","CWE-787"],"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2026-043-vllm-multimodal-prompt-embedding","cve":null,"aliases":["GHSA-mcmc-2m55-j8jj"],"title":"vLLM (multimodal prompt embeddings, sparse tensor validation): This is the advisory saying the earlier fix did not","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (multimodal prompt embeddings, sparse tensor validation)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"This is the advisory saying the earlier fix did not hold. The previous remediation for the prompt-embeds denial of service only flipped the feature off by default; the underlying flaw, missing sparse-tensor validation on caller-supplied embeddings, was never addressed. PyTorch disables sparse tensor invariant checks by default for performance, so a malformed tensor with out-of-range or negative indices is accepted and acted on, giving an out-of-bounds write primitive alongside the crash. Any operator who re-enables prompt embeds — which multimodal and embedding-serving deployments routinely do, since that is the feature — is exposed again on a supposedly patched version. Impact runs from killing a shared serving replica to memory corruption in the process holding other tenants' in-flight requests on the GPU node.","attack_vector":"Network, authenticated API client, on any deployment where the prompt-embeds feature is enabled. The attacker submits a crafted sparse tensor as multimodal embedding input in an ordinary request.","remediation":"Upgrade to vLLM 0.11.1 or later and restart the servers. Leave prompt embeds disabled unless you genuinely need them; if you need them, do not treat the earlier default-off change as a fix, since it only removed the default exposure. Enforce pod memory limits and restart policy so a corrupted or wedged worker is recycled.","references":["https://github.com/vllm-project/vllm/security/advisories/GHSA-mcmc-2m55-j8jj","https://github.com/vllm-project/vllm/pull/30649"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-78"],"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2026-051-kubeedge-configupdatejob-handler","cve":null,"aliases":["GHSA-m3c6-2p7h-cfr3","CVE-2026-62182 (reserved)"],"title":"KubeEdge (ConfigUpdateJob handler, updateFields): REMOTE CODE EXECUTION ON EDGE NODES via a normal Kubernetes API","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"KubeEdge (ConfigUpdateJob handler, updateFields)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"REMOTE CODE EXECUTION ON EDGE NODES via a normal Kubernetes API object. The ConfigUpdateJob handler concatenated caller-supplied updateFields values into a command string and ran it through a system shell, so shell metacharacters in a configuration field execute as commands on every targeted node with the privileges of the KubeEdge process. For anyone running distributed inference on edge or far-edge GPU boxes, the practical picture is that a namespaced RBAC grant intended to let a team push configuration turns into node-level execution across the fleet — and the fleet is the part of the estate with the least physical oversight and the weakest chance of anyone noticing.","attack_vector":"Network, authenticated: the attacker needs RBAC permission to create or update ConfigUpdateJob resources and an enrolled target edge node. No user interaction beyond the job being processed.","remediation":"Upgrade to KubeEdge 1.23.1, 1.22.2 or 1.21.2, where keadm config-update is invoked with structured arguments and the whole --set value is passed as a single literal. Before that lands, restrict RBAC on ConfigUpdateJob to trusted administrators, avoid ConfigUpdateJob where the input cannot be fully trusted, and monitor edge-node process activity for unexpected commands.","references":["https://github.com/kubeedge/kubeedge/security/advisories/GHSA-m3c6-2p7h-cfr3"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-78"],"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2026-052-kubeedge-nodeupgradejob-handler","cve":null,"aliases":["GHSA-5jpj-293f-rhvj","CVE-2026-62371 (reserved)"],"title":"KubeEdge (NodeUpgradeJob handler, v1alpha2 API): REMOTE CODE EXECUTION ON EDGE NODES through the upgrade path. The","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"KubeEdge (NodeUpgradeJob handler, v1alpha2 API)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"REMOTE CODE EXECUTION ON EDGE NODES through the upgrade path. The handler built its 'keadm upgrade edge' invocation by concatenating the user-controlled spec.version and spec.image fields into a shell command, so metacharacters in either field run as commands on the targeted node with the upgrade process's privileges. The affected range goes back to 1.12.0, far wider than the sibling ConfigUpdateJob issue. What makes this the sharper of the two operationally is the field involved: spec.image is exactly the kind of value operators parameterise in GitOps repos and let platform teams or automation set, so the injection point sits in a value that routinely flows from a less-trusted source into a cluster-wide job.","attack_vector":"Network, authenticated: RBAC permission to create or update NodeUpgradeJob resources through the v1alpha2 API, targeting an enrolled edge node.","remediation":"Upgrade to KubeEdge 1.23.1, 1.22.2 or 1.21.2, which invoke keadm through exec.Command with version and image as separate literal arguments. In the interim restrict NodeUpgradeJob create/update to trusted administrators, never let untrusted users or tenants influence spec.version or spec.image, and avoid NodeUpgradeJob-based upgrades where those fields come from automation you do not fully control.","references":["https://github.com/kubeedge/kubeedge/security/advisories/GHSA-5jpj-293f-rhvj"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-502"],"fleet":{"pain_class":"unpatchable / mitigate-only"},"id":"NCVD-2026-054-mlflow-statsmodels-flavor-mlflow","cve":null,"aliases":["GHSA-gqvg-gmmx-x4hm"],"title":"MLflow (statsmodels flavor, MLFLOW_ALLOW_PICKLE_DESERIALIZATION guard): SECURITY CONTROL BYPASS LEADING TO RCE: the","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"MLflow (statsmodels flavor, MLFLOW_ALLOW_PICKLE_DESERIALIZATION guard)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"SECURITY CONTROL BYPASS LEADING TO RCE: the switch operators flip to stop pickle execution does not cover the statsmodels flavor. MLFLOW_ALLOW_PICKLE_DESERIALIZATION=False is the documented control for the 2024 pickle-RCE family, and mlflow/sklearn implements the guard as the reference pattern; mlflow/statsmodels has no such check anywhere in the file and calls smio.load_pickle() straight through. So an attacker who can place a crafted MLmodel artifact anywhere in a reachable artifact store executes code in any process that later calls mlflow.pyfunc.load_model() against it — even on a deployment the operator has explicitly hardened. In a shared GPU cluster the loading process is usually a training or serving pod with GPU access, registry credentials and a cluster identity, and artifact stores are frequently writable by more tenants than the set allowed to deploy models. The danger is the false assurance: teams that adopted the flag believe this class is closed.","attack_vector":"Network / artifact-store write. The attacker needs write access to any artifact store a victim will load from, and a victim process that calls mlflow.pyfunc.load_model() on the malicious model. No MLflow credentials are required if the store is writable through another path.","remediation":"Do not rely on MLFLOW_ALLOW_PICKLE_DESERIALIZATION as the boundary — treat model artifacts as executable code and control who can write to artifact stores with the same rigour as who can push container images. Restrict artifact-store write access per tenant, load models only from stores whose writers you trust, and run model-loading pods with a scoped service account and no ambient cloud credentials. Track the MLflow fix for the statsmodels flavor and upgrade when it ships.","references":["https://github.com/mlflow/mlflow/security/advisories/GHSA-gqvg-gmmx-x4hm"],"status":"curated"},{"id":"CVE-2022-1798","cve":"CVE-2022-1798","aliases":[],"title":"KubeVirt: Path traversal lets a user who can configure KubeVirt read arbitrary host files","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"KubeVirt","year":"2022","cvss_score":8.7,"severity":"high","kev":false,"impact":"Path traversal lets a user who can configure KubeVirt read arbitrary host files","attack_vector":"Cluster user with VM configuration rights","remediation":"Upgrade KubeVirt; restart virt-handler DaemonSet","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-1798"],"status":"curated","published":"2022-09-15"},{"id":"CVE-2023-20514","cve":"CVE-2023-20514","aliases":[],"title":"AMD Secure Processor - TEE parameter handling: A privileged attacker can hand an arbitrary memory value to functions","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor - TEE parameter handling","year":"2023","cvss_score":8.7,"severity":"high","kev":false,"impact":"A privileged attacker can hand an arbitrary memory value to functions inside the trusted execution environment, reaching arbitrary code execution in the ASP. At CVSS 8.7 this is one of the more direct host-root-to-secure-processor escalations in the set: the OS administrator, who is supposed to be outside the ASP trust boundary, gets inside it.","attack_vector":"Local, privileged (host root). No physical access and no signed-TA requirement.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. Because this sits inside the SEV-SNP trust boundary, the update also moves the platform's reported TCB version: after patching you must refresh VCEK certificates from AMD's KDS and update whatever attestation policy your tenants (or your own confidential-VM control plane) pin against, or every guest launch will start failing validation.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-20514","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2026-02-11"},{"id":"CVE-2024-0106","cve":"CVE-2024-0106","aliases":[],"title":"BlueField / ConnectX firmware: Improper certificate validation / access control","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"BlueField / ConnectX firmware","year":"2024","cvss_score":8.7,"severity":"high","kev":false,"impact":"Improper certificate validation / access control","attack_vector":"Network-adjacent","remediation":"Flash NIC/DPU firmware; node reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0106","https://github.com/NVIDIA/product-security/tree/main/2024/5562"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:L/I:H/A:H","cwe":["CWE-274"],"fleet":{"pain_class":"node-reboot"},"published":"2024-11-01"},{"id":"CVE-2024-0108","cve":"CVE-2024-0108","aliases":[],"title":"Jetson (Xavier/TX/Nano): Improper error handling","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Jetson (Xavier/TX/Nano)","year":"2024","cvss_score":8.7,"severity":"high","kev":false,"impact":"Improper error handling -> privesc","attack_vector":"Local attacker on the device","remediation":"Flash JetPack; edge fleet only","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0108","https://github.com/NVIDIA/product-security/tree/main/2024/5555"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:L","cwe":["CWE-755"],"published":"2024-08-08"},{"id":"CVE-2024-23651","cve":"CVE-2024-23651","aliases":[],"title":"BuildKit: Race between parallel build steps sharing cache mounts with subpaths","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"BuildKit","year":"2024","cvss_score":8.7,"severity":"high","kev":false,"impact":"Race between parallel build steps sharing cache mounts with subpaths; host file access","attack_vector":"Two concurrent builds on a shared builder","remediation":"Upgrade BuildKit; isolate build cache per tenant","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-23651"],"status":"curated","published":"2024-01-31"},{"id":"CVE-2024-58310","cve":"CVE-2024-58310","aliases":[],"title":"APC Network Management Card 4 (NMC4): An unauthenticated attacker can manipulate URL parameters to walk out of the web","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"APC Network Management Card 4 (NMC4)","year":"2024","cvss_score":8.7,"severity":"high","kev":false,"impact":"An unauthenticated attacker can manipulate URL parameters to walk out of the web root and read arbitrary system files off the card, including files like /etc/passwd — enough to harvest system account information and plan a follow-on attack against the UPS/PDU's management plane.","attack_vector":"Fully remote and unauthenticated — a crafted HTTP request with encoded directory-traversal sequences is enough.","remediation":"Firmware flash of the NMC4 card to the fixed release. Roll out per card; the UPS/PDU keeps serving power to its load during the flash, but remote monitoring/management of that unit drops briefly.","references":["https://www.exploit-db.com/exploits/51897","https://www.vulncheck.com/advisories/apc-network-management-card-path-traversal-via-directory-traversal"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2025-12-11"},{"cvss_vector":"CVSS:4.0/AV:N/AC:L/AT:N/PR:N/UI:N/VC:H/VI:N/VA:N/SC:N/SI:N/SA:N","cwe":["CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-0051","cve":"CVE-2025-0051","aliases":[],"title":"Pure Storage FlashArray authentication input validation: Malformed input during authentication takes the FlashArray","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Pure Storage FlashArray authentication input validation","year":"2025","cvss_score":8.7,"severity":"high","kev":false,"impact":"Malformed input during authentication takes the FlashArray into a denial of service, before any credential is checked. Losing the array means every GPU node that mounts it loses its data path at once.","attack_vector":"Network reach to the FlashArray authentication surface. Pre-authentication, so no account is required.","remediation":"Upgrade Purity//FA to the release named in Pure's security bulletin. Until patched, limit which networks can reach the array's login endpoints - this is reachable from anywhere the login prompt is.","references":["https://support.purestorage.com/bundle/m_security_bulletins/page/Pure_Security/topics/concept/c_security_bulletins.html","https://nvd.nist.gov/vuln/detail/CVE-2025-0051"],"status":"curated"},{"id":"CVE-2025-20105","cve":"CVE-2025-20105","aliases":["INTEL-SA-01234","CVE-2025-20064","CVE-2025-20068","CVE-2025-20027"],"title":"UEFI firmware SMM modules in Intel reference platform firmware (SMM handler, FlashUcAcmSmm, ImcErrorHandler, WheaERST","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"UEFI firmware SMM modules in Intel reference platform firmware","year":"2025","cvss_score":8.7,"severity":"high","kev":false,"impact":"Improper input validation in SMM modules that Intel ships as reference code and that OEMs build into their server BIOS. SMM is the most privileged execution mode on the platform - above the hypervisor - so an escalation here gives an attacker control of the platform beneath every isolation boundary the fleet relies on, including access to the SPI flash write path. The result is a firmware implant that survives reimage and crosses tenant handoff, and that can neutralise measured boot from underneath. Because this is Intel reference code, the same defect propagates identically across every OEM that consumed that code drop, so exposure is fleet-wide across mixed vendors rather than isolated to one supplier.","attack_vector":"A privileged local user on the host - local root or an existing kernel foothold triggering the SMI. On bare-metal GPU nodes, that is the tenant.","remediation":"BIOS update from each OEM once they pick up Intel's fixed reference code - Dell, HPE, Supermicro, Lenovo, Gigabyte, Quanta and Wiwynn all ship independently, and for reference-code advisories the lag from Intel's disclosure to a shipped server BIOS routinely runs one to two quarters, longer for ODM whitebox. Requires host reboot and job drain. Track this by OEM BIOS version rather than by CVE, because OEM release notes often reference only their own advisory ID. No runtime mitigation exists for SMM defects.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-20105","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01234.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-20163","cve":"CVE-2025-20163","aliases":[],"title":"Cisco Nexus Dashboard Fabric Controller (SSH host key validation): NDFC does not validate the SSH host keys","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Cisco Nexus Dashboard Fabric Controller (SSH host key validation)","year":"2025","cvss_score":8.7,"severity":"high","kev":false,"impact":"NDFC does not validate the SSH host keys of the switches it manages, so anyone positioned on the management network can impersonate a managed switch and machine-in-the-middle the controller's sessions with it — harvesting the device credentials NDFC presents and feeding back forged state. Affects all NDFC versions regardless of configuration, which means every NDFC deployment ever built has been trusting its management network implicitly.","attack_vector":"Unauthenticated attacker with a position on the path between NDFC and its managed devices — a compromised management-network host, a rogue device, or ARP/route manipulation on the OOB VLAN.","remediation":"Upgrade NDFC to a release that pins host keys. Controller software upgrade. Also rotate every switch credential NDFC holds, since they may have been captured — that credential rotation across a fabric is the expensive part, not the upgrade.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-20163"],"status":"curated","published":"2025-06-04"},{"id":"CVE-2025-23256","cve":"CVE-2025-23256","aliases":[],"title":"BlueField DPU: Access-control bypass on the DPU","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"BlueField DPU","year":"2025","cvss_score":8.7,"severity":"high","kev":false,"impact":"Access-control bypass on the DPU -> control of the tenant network path","attack_vector":"Network-adjacent attacker / tenant on the DPU-served host","remediation":"Flash DPU firmware + upgrade DOCA; DPU reset drops tenant networking, drain node","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23256","https://github.com/NVIDIA/product-security/tree/main/2025/5655"],"status":"curated","fleet":{"ubiquity":"very common - ConnectX NICs are in essentially every GPU node; BlueField DPUs increasingly own the tenant network and storage path","remediation_pain":"firmware-flash of NIC/DPU firmware per node (BF-2/BF-3 images 45.1020, 35.4554 LTS22, 39.5050 LTS23, 43.3608 LTS24), often requiring a host reboot to activate","pain_class":"firmware-flash","why_fleet_wide":"The DPU is a full independent computer with DMA to the host and control of the tenant's network and storage offload - compromise is below the hypervisor and invisible to the host OS, on every node carrying the card."},"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:L/I:H/A:H","cwe":["CWE-863"],"published":"2025-09-04"},{"id":"CVE-2025-23293","cve":"CVE-2025-23293","aliases":[],"title":"NVIDIA License System - Delegated Licensing Service (DLS): An unauthorised action against the DLS reaches high","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA License System - Delegated Licensing Service (DLS)","year":"2025","cvss_score":8.7,"severity":"high","kev":false,"impact":"An unauthorised action against the DLS reaches high integrity and availability impact with a changed scope - scored 8.7, the most serious of the licensing-appliance issues. An attacker on the management network can take the licensing service down or corrupt it, which eventually strands every vGPU guest.","attack_vector":"Adjacent network, low privileges, no user interaction. Any account on the management network reaches it.","remediation":"Update the DLS appliance per bulletin 5705 as a priority, and segment the licensing appliance off the general management VLAN. Cost: appliance restart; plan it against your licence lease window so guests do not notice.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23293","https://github.com/NVIDIA/product-security/tree/main/2025/5705"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:L/UI:N/S:C/C:N/I:H/A:H","cwe":["CWE-306"],"published":"2025-09-30"},{"cvss_vector":"CVSS:4.0/AV:N/AC:L/AT:N/PR:N/UI:N/VC:N/VI:H/VA:N/SC:N/SI:N/SA:N","cwe":["CWE-347"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2025-31489","cve":"CVE-2025-31489","aliases":[],"title":"MinIO (S3 API, unsigned-trailer uploads): Signature validation on unsigned-trailer uploads is incomplete, so knowing","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"MinIO (S3 API, unsigned-trailer uploads)","year":"2025","cvss_score":8.7,"severity":"high","kev":false,"impact":"Signature validation on unsigned-trailer uploads is incomplete, so knowing only an access key ID - not the secret - is enough to write objects as that identity. Any tenant whose access key ID is visible in logs or config can be impersonated for writes.","attack_vector":"Any network client that can reach the S3 endpoint and has seen an access key ID.","remediation":"Upgrade to MinIO RELEASE.2025-04-03T14-56-28Z or later and restart the cluster. Treat access key IDs as semi-sensitive going forward, and audit writes on shared buckets over the exposure window.","references":["https://github.com/minio/minio/security/advisories/GHSA-wg47-6jq2-q2hh","https://nvd.nist.gov/vuln/detail/CVE-2025-31489"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-37101","cve":"CVE-2025-37101","aliases":[],"title":"HPE OneView for VMware vCenter (vertical privilege escalation): A read-only user performs administrative actions","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HPE OneView for VMware vCenter (vertical privilege escalation)","year":"2025","cvss_score":8.7,"severity":"high","kev":false,"impact":"A read-only user performs administrative actions through the OneView vCenter plugin - so view-only access to vCenter becomes administrative control over HPE hardware management.","attack_vector":"Authenticated read-only user of the OV4VC plugin, with user interaction.","remediation":"Apply the HPE update per HPESBGN04876. Plugin update inside vCenter; no server firmware work.","references":["https://support.hpe.com/hpesc/public/docDisplay?docId=hpesbgn04876en_us&docLocale=en_US"],"status":"curated"},{"id":"CVE-2025-54412","cve":"CVE-2025-54412","aliases":[],"title":"skops (scikit-learn model sharing): Inconsistency in the `Operator` handling lets an untrusted model bypass the safe","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"skops (scikit-learn model sharing)","year":"2025","cvss_score":8.7,"severity":"high","kev":false,"impact":"Inconsistency in the `Operator` handling lets an untrusted model bypass the safe loader","attack_vector":"Customer-supplied skops model file","remediation":"Upgrade past 0.11.0; skops is the \"safe alternative to pickle\" and it too has loader bypasses","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-54412"],"status":"curated","published":"2025-07-26"},{"id":"CVE-2025-54413","cve":"CVE-2025-54413","aliases":[],"title":"skops: Method-handling inconsistency","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"skops","year":"2025","cvss_score":8.7,"severity":"high","kev":false,"impact":"Method-handling inconsistency → safe-loader bypass","attack_vector":"Customer-supplied model file","remediation":"Upgrade past 0.11.0","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-54413"],"status":"curated","published":"2025-07-26"},{"id":"CVE-2025-61958","cve":"CVE-2025-61958","aliases":["K000154647"],"title":"F5 BIG-IP (iHealth command / tmsh restricted shell): An authenticated attacker with at least a resource-administrator","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"F5 BIG-IP (iHealth command / tmsh restricted shell)","year":"2025","cvss_score":8.7,"severity":"high","kev":false,"impact":"An authenticated attacker with at least a resource-administrator role can use the iHealth command to break out of the restricted tmsh shell and get a full bash shell on the device — this specifically defeats BIG-IP's Appliance mode, the hardened mode operators use to lock admins out of the underlying OS on shared/regulated deployments.","attack_vector":"Requires an authenticated account with resource-administrator role (not full root/admin) — the attack is a privilege-escalation/shell-escape from a role that was supposed to be constrained.","remediation":"Software upgrade to the fixed BIG-IP release per F5 K000154647. Part of the same October 2025 remediation batch as CVE-2025-53521 — apply in the same maintenance window. This specifically matters for shared/managed BIG-IP deployments that rely on Appliance mode to keep administrators out of shell access.","references":["https://my.f5.com/manage/s/article/K000154647"],"status":"curated","published":"2025-10-15"},{"id":"CVE-2026-31837","cve":"CVE-2026-31837","aliases":[],"title":"Istio: When JWKS resolution fails, istiod falls back to hardcoded defaults, weakening JWT validation","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Istio","year":"2026","cvss_score":8.7,"severity":"high","kev":false,"impact":"When JWKS resolution fails, istiod falls back to hardcoded defaults, weakening JWT validation","attack_vector":"Unauthenticated network, exploitable by first disrupting JWKS reachability","remediation":"Rolling istiod upgrade to 1.29.1/1.28.5/1.27.8+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-31837"],"status":"curated","published":"2026-03-10"},{"id":"CVE-2026-34940","cve":"CVE-2026-34940","aliases":[],"title":"KubeAI (Ollama engine controller): Injection in `ollamaStartupProbeScript()`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"KubeAI (Ollama engine controller)","year":"2026","cvss_score":8.7,"severity":"high","kev":false,"impact":"Injection in `ollamaStartupProbeScript()`","attack_vector":"Tenant-supplied model name in a Kubernetes AI operator","remediation":"Upgrade to 0.23.2+; operator-plane compromise from tenant input","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-34940"],"status":"curated","published":"2026-04-06"},{"id":"CVE-2026-53230","cve":"CVE-2026-53230","aliases":["net/mlx5 slab-out-of-bounds in mlx5_query_nic_vport_mac_list"],"title":"Linux kernel mlx5_core eswitch / vport (SR-IOV): Mlx5_core sizes a firmware command buffer from the physical function's","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel mlx5_core eswitch / vport (SR-IOV)","year":"2026","cvss_score":8.7,"severity":"high","kev":false,"impact":"Mlx5_core sizes a firmware command buffer from the physical function's MAC-list capability, but a virtual function can be configured with a larger max. Querying that VF makes the firmware response overflow the PF's buffer - a slab out-of-bounds in the host kernel driven by a value the VF side controls. CVSS scope is Changed. This is the clean VF-to-PF memory-safety break: a tenant holding an SR-IOV VF corrupts host kernel memory belonging to the NIC that serves everyone on the node.","attack_vector":"A tenant or container with local control of an mlx5 SR-IOV VF, in combination with a VF max-MAC-list setting larger than the PF's capability. Triggered from the PF-side eswitch worker, so no host root is needed on the attacker side.","remediation":"Upgrade the host kernel to 7.1 or a stable backport (6.6.143, 6.12.94, 6.18.36, 7.0.13). Rolling reboot of every SR-IOV host - drain GPU jobs per node, this is not live-patchable in practice. Interim: audit devlink VF MAC-list max settings and keep them at or below the PF capability.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53230","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-53230.json"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-06-25"},{"id":"CVE-2026-55732","cve":"CVE-2026-55732","aliases":["CVE-2026-12496","CVE-2026-12504","CVE-2026-55731"],"title":"Loytec L-INX, L-GATE, L-ROC, L-IOB, L-DALI, L-VIS, L-PAD and LIP-ME201C (through 8.4.18, LINX-A64): An out-of-bounds","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Loytec L-INX, L-GATE, L-ROC, L-IOB, L-DALI, L-VIS, L-PAD and LIP-ME201C (through 8.4.18, LINX-A64)","year":"2026","cvss_score":8.7,"severity":"high","kev":false,"impact":"An out-of-bounds read in BACnet packet parsing lets an unauthenticated attacker crash the main control process and reboot the device with a single malformed TimeSynchronization message - and BACnet TimeSynchronization is a broadcast service, so one packet can take out every Loytec device on the segment at once. Loytec controllers and gateways are the BACnet/LonWorks integration layer in a lot of European-designed and mixed-vendor datacenters, sitting between the IP supervisory network and the field bus that runs air handling. Rebooting them all simultaneously severs supervisory control of cooling for as long as the attacker keeps sending, which against 40-140 kW GPU racks is a straightforward path to thermal shutdown across a hall. The companion issues in the same disclosure set (unauthenticated stored XSS in the OPC XML-DA statistics page, a PAM misconfiguration allowing authentication as an unintended uid, and an SNMP agent loop-condition bug) mean the same devices also offer credential-theft and persistence paths, not just DoS.","attack_vector":"Unauthenticated, remote, over BACnet on the facility network - and via broadcast, so it does not even need to know device addresses. The SNMP and web issues are likewise reachable from anywhere on that segment. Anyone who can put a frame on the building VLAN can do this.","remediation":"Firmware update from Loytec above 8.4.18 for the device family, plus LWEB-802 5.0.8+ for the management side. This is a per-device firmware flash across every gateway and controller, done by the controls integrator, with each device offline during the flash - a maintenance window on live cooling. Because a broadcast packet is the trigger, the compensating control has to actually block broadcast BACnet from untrusted hosts, which usually means putting the BACnet segment on its own VLAN with no untrusted hosts on it at all rather than trying to filter by address. Disable the SNMP agent and the OPC XML-DA interface if you are not using them.","references":["https://www.loytec.com/support/product-security/advisories","https://nvd.nist.gov/vuln/detail/CVE-2026-55732","https://nvd.nist.gov/vuln/detail/CVE-2026-12504"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2026-5747","cve":"CVE-2026-5747","aliases":[],"title":"Firecracker: Out-of-bounds write in the virtio PCI transport","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Firecracker","year":"2026","cvss_score":8.7,"severity":"high","kev":false,"impact":"Out-of-bounds write in the virtio PCI transport; guest root can crash or potentially compromise the VMM process","attack_vector":"Any tenant guest VM with root inside it","remediation":"Upgrade Firecracker; restart microVMs, which for a neocloud means terminating tenant instances","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-5747"],"status":"curated","published":"2026-04-08"},{"id":"CVE-2026-6445","cve":"CVE-2026-6445","aliases":[],"title":"Pure Storage FlashArray Purity (data path information exposure): Insufficient filtering on certain data paths exposes","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Pure Storage FlashArray Purity (data path information exposure)","year":"2026","cvss_score":8.7,"severity":"high","kev":false,"impact":"Insufficient filtering on certain data paths exposes sensitive information to an authenticated low-privileged user of the array.","attack_vector":"Authenticated low-privilege array user.","remediation":"Apply the Purity update referenced in Pure's security bulletins. Non-disruptive array software upgrade.","references":["https://support.purestorage.com/bundle/m_security_bulletins/page/Pure_Security/topics/concept/c_security_bulletins.html"],"status":"curated"},{"id":"CVE-2017-3883","cve":"CVE-2017-3883","aliases":[],"title":"Cisco FXOS / NX-OS AAA: AAA implementation flaw enabling remote DoS via brute-force login attempts against the switch","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco FXOS / NX-OS AAA","year":"2017","cvss_score":8.6,"severity":"high","kev":false,"impact":"AAA implementation flaw enabling remote DoS via brute-force login attempts against the switch management plane","attack_vector":"Network, unauthenticated","remediation":"NX-OS/FXOS upgrade; also a reason to keep switch management auth off any tenant-reachable path","references":["https://nvd.nist.gov/vuln/detail/CVE-2017-3883"],"status":"curated","published":"2017-10-19"},{"id":"CVE-2018-0378","cve":"CVE-2018-0378","aliases":[],"title":"Cisco NX-OS PTP feature (Nexus 5500/5600/6000): An unauthenticated remote attacker takes down a Nexus switch through","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco NX-OS PTP feature (Nexus 5500/5600/6000)","year":"2018","cvss_score":8.6,"severity":"high","kev":false,"impact":"An unauthenticated remote attacker takes down a Nexus switch through the PTP subsystem. PTP is one of the few unauthenticated protocols that switches process in the control plane by design, so it is a reliable path to the CPU on every device that has it enabled. The Cisco IOS equivalent is CVE-2018-0473.","attack_vector":"Unauthenticated, remote — PTP packets reaching the switch's PTP-enabled interfaces.","remediation":"NX-OS upgrade plus reload. Immediate config mitigation: disable PTP on interfaces that face tenant workloads and keep it only on the links that actually need it — live change, no reload.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-0378","https://nvd.nist.gov/vuln/detail/CVE-2018-0473"],"status":"curated","tags":["fabric-dos"],"published":"2018-10-17"},{"id":"CVE-2018-7093","cve":"CVE-2018-7093","aliases":[],"title":"HPE iLO3/4/5: Remote unauthenticated denial of service against the management controller","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HPE iLO3/4/5","year":"2018","cvss_score":8.6,"severity":"high","kev":false,"impact":"Remote unauthenticated denial of service against the management controller — loses out-of-band access to the node during an incident","attack_vector":"Network, unauthenticated","remediation":"iLO firmware update; DoS on the BMC is an availability problem specifically because it removes the recovery path","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-7093"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2018-08-14"},{"id":"CVE-2019-5736","cve":"CVE-2019-5736","aliases":[],"title":"runc: Host runc binary overwritten from inside a container","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"runc","year":"2019","cvss_score":8.6,"severity":"high","kev":false,"impact":"Host runc binary overwritten from inside a container; full host root. The canonical container-escape","attack_vector":"Any tenant workload that can exec as root in its own container, or a malicious image","remediation":"Replace the runc binary on every node; already-running containers keep the vulnerable fd, so a full drain and pod restart is required","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-5736"],"status":"curated","fleet":{"ubiquity":"Universal - the original runc escape; affected Docker, containerd and CRI-O simultaneously","remediation_pain":"`node-drain` - the host runc binary itself is overwritten by the exploit, so remediation is binary replacement plus recreation of every container; the standing mitigation is making runc immutable","pain_class":"node-drain","why_fleet_wide":"A container process rewrites `/proc/self/exe` and overwrites the *host* runc binary, so every subsequent container start on that host executes attacker code as root - the canonical fleet-wide container-runtime emergency"},"published":"2019-02-11"},{"id":"CVE-2021-1587","cve":"CVE-2021-1587","aliases":[],"title":"Cisco NX-OS (VXLAN OAM / NGOAM): A crafted VXLAN OAM packet reloads a VTEP. In a VXLAN/EVPN GPU fabric every leaf is a","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco NX-OS (VXLAN OAM / NGOAM)","year":"2021","cvss_score":8.6,"severity":"high","kev":false,"impact":"A crafted VXLAN OAM packet reloads a VTEP. In a VXLAN/EVPN GPU fabric every leaf is a VTEP, so an attacker with a foothold in any tenant overlay can knock out leaves one at a time and stall collectives cluster-wide. Reachable from inside a tenant's own overlay, which is what makes it interesting — it does not need underlay access.","attack_vector":"Unauthenticated, remote — the attacker needs to be able to land a crafted VXLAN packet on the switch's VTEP address. In practice that means a compromised workload or a tenant that can source arbitrary UDP.","remediation":"NX-OS upgrade plus reload. If NGOAM is not in use, disabling the feature is a live config change with no reload and removes the exposure entirely — do that first, patch on the next window.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1587"],"status":"curated","tags":["fabric-dos"],"published":"2021-08-25"},{"id":"CVE-2021-32777","cve":"CVE-2021-32777","aliases":[],"title":"Envoy: ext-authz header handling flaw allows bypassing the external authorization service","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy","year":"2021","cvss_score":8.6,"severity":"high","kev":false,"impact":"ext-authz header handling flaw allows bypassing the external authorization service","attack_vector":"Unauthenticated network","remediation":"Upgrade Envoy","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-32777"],"status":"curated","published":"2021-08-24"},{"id":"CVE-2021-32779","cve":"CVE-2021-32779","aliases":[],"title":"Envoy: URI fragment treated as part of the path","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy","year":"2021","cvss_score":8.6,"severity":"high","kev":false,"impact":"URI fragment treated as part of the path; authorization bypass","attack_vector":"Unauthenticated network","remediation":"Upgrade Envoy","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-32779"],"status":"curated","published":"2021-08-24"},{"id":"CVE-2021-32781","cve":"CVE-2021-32781","aliases":[],"title":"Envoy: Processing continues after a local reply, causing undefined behaviour","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy","year":"2021","cvss_score":8.6,"severity":"high","kev":false,"impact":"Processing continues after a local reply, causing undefined behaviour","attack_vector":"Unauthenticated network","remediation":"Upgrade Envoy","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-32781"],"status":"curated","published":"2021-08-24"},{"id":"CVE-2021-33141","cve":"CVE-2021-33141","aliases":[],"title":"Intel Ethernet Adapter manageability firmware (NC-SI / sideband path): Improper input validation in the *manageability*","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Ethernet Adapter manageability firmware (NC-SI / sideband path)","year":"2021","cvss_score":8.6,"severity":"high","kev":false,"impact":"Improper input validation in the *manageability* firmware of Intel Ethernet adapters, exploitable by an unauthenticated user. Manageability firmware is the NC-SI sideband engine — the path that carries BMC traffic over the same physical NIC as production data. A flaw there is the bridge between the data network and out-of-band management: an attacker on the fabric reaches the sideband channel, and the sideband channel reaches the BMC, which controls power and virtual media for the node. This is the highest-value shape of NIC firmware bug for a multi-tenant operator, because it crosses the boundary between 'tenant network' and 'operator management plane'.","attack_vector":"Unauthenticated attacker with network access to the adapter. No host account required.","remediation":"Flash adapter manageability firmware via the OEM firmware bundle (this is separate from the main NVM image on some platforms); cold power cycle. Architecturally, the durable control is to stop sharing the production NIC with BMC traffic — use a dedicated BMC NIC on a physically separate OOB network rather than NC-SI sideband. That is a hardware/topology decision, so it applies to your next buildout, not this one. Companion issues in the same advisory family: CVE-2021-33162, CVE-2021-33161, CVE-2021-33158, CVE-2021-33157, CVE-2022-37341.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-33141","https://nvd.nist.gov/vuln/detail/CVE-2021-33162"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-02-23"},{"id":"CVE-2021-47136","cve":"CVE-2021-47136","aliases":[],"title":"Linux kernel mlx5_core representor TC path + net/sched tc extension: The TC_SKB_EXT skb extension is not zeroed","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_core representor TC path + net/sched tc extension","year":"2021","cvss_score":8.6,"severity":"high","kev":false,"impact":"The TC_SKB_EXT skb extension is not zeroed on allocation and mlx5's representor restore path never initialized the newer fields, so Open vSwitch reads uninitialized kernel memory as a boolean. This leaks host kernel memory contents into the OVS datapath - and it is triggered remotely, by sending packets that miss hardware offload. On a switchdev host running OVS over ConnectX, that is remote kernel-memory disclosure into the software datapath.","attack_vector":"Remote, unauthenticated: send traffic crafted to miss the hardware offload path on an mlx5 switchdev host running OVS.","remediation":"Upgrade the host kernel to 5.13 or a stable backport (5.10.42, 5.12.9). Rolling reboot. Note NVD scores this 5.5 while the kernel CNA scores it 8.6 - if you triage from NVD feeds you will under-prioritize it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-47136","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2021/CVE-2021-47136.json"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-03-25"},{"id":"CVE-2022-2601","cve":"CVE-2022-2601","aliases":[],"title":"GRUB2 (font engine, grub_font_construct_glyph): Buffer overflow when constructing a glyph from a crafted GRUB font","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (font engine, grub_font_construct_glyph)","year":"2022","cvss_score":8.6,"severity":"high","kev":false,"impact":"Buffer overflow when constructing a glyph from a crafted GRUB font file. Fonts are unsigned data sitting in the boot partition on virtually every install, which makes this one of the cheapest Secure Boot bypasses in the family.","attack_vector":"Anyone who can write a font file to the boot partition - local root, prior tenant, or a poisoned image build.","remediation":"grub2 package update + reboot. Fonts are rarely needed on a headless server image; dropping the graphical GRUB theme removes this surface outright.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-2601","https://access.redhat.com/security/cve/CVE-2022-2601"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2022-12-14"},{"id":"CVE-2022-36392","cve":"CVE-2022-36392","aliases":[],"title":"Intel AMT / Standard Manageability firmware: Improper input validation in AMT/ISM firmware, scored high because","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel AMT / Standard Manageability firmware","year":"2022","cvss_score":8.6,"severity":"high","kev":false,"impact":"Improper input validation in AMT/ISM firmware, scored high because of network reach. Another entry in the recurring pattern: the out-of-band management engine keeps producing network-reachable memory-safety bugs, and each one needs an OEM firmware cycle to fix.","attack_vector":"Network access to the AMT interface on affected firmware versions.","remediation":"Fixed in Intel CSME/SPS firmware, which reaches you as an OEM BIOS or firmware package - not as a microcode or OS update. That means: wait for your server vendor to ship it, drain the node, flash, and reboot. OEM availability is the long pole and routinely lags the Intel advisory by one or more quarters on server platforms. Track it per platform SKU, because vendors ship these unevenly across their own product lines.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-36392","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00783.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2023-08-11"},{"id":"CVE-2022-3775","cve":"CVE-2022-3775","aliases":[],"title":"GRUB2 (font engine, blit_comb): Integer underflow when rendering certain unicode sequences writes out of bounds","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (font engine, blit_comb)","year":"2022","cvss_score":8.6,"severity":"high","kev":false,"impact":"Integer underflow when rendering certain unicode sequences writes out of bounds. Same class as the glyph-construction bug and shipped in the same advisory wave - if you patched one you probably need both.","attack_vector":"Crafted font or text rendered by GRUB, reachable by anyone who can write boot-partition content.","remediation":"grub2 package update + reboot. Verify your distro's package covers both font CVEs, not just the first.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-3775","https://access.redhat.com/security/cve/CVE-2022-3775"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2022-12-19"},{"id":"CVE-2023-35941","cve":"CVE-2023-35941","aliases":[],"title":"Envoy: Malicious client constructs permanently valid credentials in the OAuth filter","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy","year":"2023","cvss_score":8.6,"severity":"high","kev":false,"impact":"Malicious client constructs permanently valid credentials in the OAuth filter","attack_vector":"Unauthenticated network","remediation":"Upgrade Envoy; rotate OAuth HMAC secrets","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-35941"],"status":"curated","published":"2023-07-25"},{"id":"CVE-2024-20267","cve":"CVE-2024-20267","aliases":[],"title":"Cisco NX-OS (MPLS traffic handling / netstack): Crafted MPLS traffic restarts netstack, which stops the switch","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco NX-OS (MPLS traffic handling / netstack)","year":"2024","cvss_score":8.6,"severity":"high","kev":false,"impact":"Crafted MPLS traffic restarts netstack, which stops the switch forwarding or reloads it. Applies even where you are not consciously running MPLS, because the parse path is reachable regardless. One packet stream, one leaf down.","attack_vector":"Unauthenticated, remote — an attacker able to send MPLS-labelled frames toward the device.","remediation":"NX-OS upgrade plus reload. Filter MPLS ethertype at the fabric edge as a live mitigation if you do not use MPLS, which most GPU-cluster fabrics do not.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-20267"],"status":"curated","tags":["fabric-dos"],"published":"2024-02-29"},{"id":"CVE-2024-20321","cve":"CVE-2024-20321","aliases":[],"title":"Cisco NX-OS (eBGP implementation): An unauthenticated remote attacker can wedge the switch through the eBGP","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco NX-OS (eBGP implementation)","year":"2024","cvss_score":8.6,"severity":"high","kev":false,"impact":"An unauthenticated remote attacker can wedge the switch through the eBGP implementation. In a BGP-underlay leaf/spine — the standard GPU-cluster design — the routing daemon going down means the rack loses reachability, and a coordinated attack against several leaves partitions the fabric mid-training-run.","attack_vector":"Unauthenticated, remote to the BGP process. Requires the ability to reach the switch's BGP listener.","remediation":"NX-OS upgrade plus reload. Harden with strict neighbor ACLs and CoPP in the meantime — live config. Because this affects the underlay control plane, schedule the reload per MLAG/ECMP pair so the fabric never loses both paths.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-20321"],"status":"curated","tags":["fabric-dos"],"published":"2024-02-29"},{"id":"CVE-2024-20446","cve":"CVE-2024-20446","aliases":[],"title":"Cisco NX-OS (DHCPv6 relay agent): A crafted DHCPv6 packet takes the switch out. Relevant because DHCP relay is normally","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco NX-OS (DHCPv6 relay agent)","year":"2024","cvss_score":8.6,"severity":"high","kev":false,"impact":"A crafted DHCPv6 packet takes the switch out. Relevant because DHCP relay is normally configured on exactly the SVIs that face tenant workloads, so this is reachable from inside a tenant network with no credentials at all.","attack_vector":"Unauthenticated, remote — a host on a VLAN where DHCPv6 relay is configured.","remediation":"NX-OS upgrade plus reload. If IPv6 is unused in your cluster (still common), disabling the DHCPv6 relay is a live config change that removes the exposure with no downtime.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-20446"],"status":"curated","fleet":{"pain_class":"hot-patch"},"tags":["fabric-dos"],"published":"2024-08-28"},{"id":"CVE-2024-21626","cve":"CVE-2024-21626","aliases":[],"title":"runc: \"Leaky Vessels\": internal file descriptor leak lets a container process start with cwd in the host filesystem","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"runc","year":"2024","cvss_score":8.6,"severity":"high","kev":false,"impact":"\"Leaky Vessels\": internal file descriptor leak lets a container process start with cwd in the host filesystem; full host escape","attack_vector":"Malicious image (a crafted WORKDIR is enough) or any tenant workload","remediation":"Replace runc on all nodes; running containers stay vulnerable, so drain and recreate all GPU pods","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21626"],"status":"curated","fleet":{"ubiquity":"Universal - runc is the default OCI runtime under Docker, containerd, CRI-O and therefore under essentially every GPU container on every neocloud","remediation_pain":"`node-drain` - the runc binary is replaced on the host, and while the swap itself is atomic, already-running containers keep the vulnerable runtime semantics; safe remediation means evacuating and recreating every container, i.e. draining paying GPU jobs off each node","pain_class":"node-drain","why_fleet_wide":"A leaked file descriptor lets any customer-supplied container image (or `docker exec`) land its working directory in the host filesystem namespace, giving host root from an untrusted tenant workload on every node in the fleet"},"published":"2024-01-31"},{"id":"CVE-2024-23324","cve":"CVE-2024-23324","aliases":["GHSA-gq3v-vvhj-96j6"],"title":"Envoy proxy (ext_authz filter): When Envoy's ext_authz filter is configured with failure_mode_allow set to true","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy proxy (ext_authz filter)","year":"2024","cvss_score":8.6,"severity":"high","kev":false,"impact":"When Envoy's ext_authz filter is configured with failure_mode_allow set to true, a downstream client can force an invalid gRPC request that circumvents the external-authorization check entirely — meaning traffic that should have been checked against an auth service (a common pattern for gating access to inference or cluster-management endpoints) sails through unauthenticated.","attack_vector":"Remote — a downstream client crafts a malformed gRPC request against an Envoy instance where ext_authz is set to fail open.","remediation":"Software upgrade to Envoy 1.29.1, 1.28.1, 1.27.3, 1.26.7, or later. No config workaround exists (the vendor advisory states none), so this is a binary/image upgrade and restart across every Envoy instance doing ext_authz-based access control in the cluster ingress path — a rolling restart avoids a hard outage if you run multiple replicas.","references":["https://github.com/envoyproxy/envoy/security/advisories/GHSA-gq3v-vvhj-96j6"],"status":"curated","fleet":{"pain_class":"daemon-restart"},"published":"2024-02-09"},{"id":"CVE-2024-4325","cve":"CVE-2024-4325","aliases":[],"title":"Gradio (`/queue/join`): SSRF","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Gradio (`/queue/join`)","year":"2024","cvss_score":8.6,"severity":"high","kev":false,"impact":"SSRF","attack_vector":"Unauthenticated network to the demo","remediation":"Upgrade past 4.21.0; blocks IMDS access from the GPU node","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-4325"],"status":"curated","published":"2024-06-06"},{"id":"CVE-2024-48248","cve":"CVE-2024-48248","aliases":[],"title":"NAKIVO Backup & Replication: Unauthenticated absolute path traversal via getImageByPath","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"NAKIVO Backup & Replication","year":"2024","cvss_score":8.6,"severity":"high","kev":true,"impact":"Unauthenticated absolute path traversal via getImageByPath -> arbitrary file read incl. cleartext credentials","attack_vector":"Network (remote)","remediation":"Control-plane: patch + rotate all credentials in the backup product database","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-48248"],"status":"curated","published":"2025-03-04"},{"id":"CVE-2024-48882","cve":"CVE-2024-48882","aliases":["CVE-2024-49572","CVE-2025-20085","CVE-2025-23417","CVE-2025-26858","CVE-2025-55221","CVE-2025-54848","TALOS-2024-2119"],"title":"Socomec DIRIS Digiware M-70 1.6.9 (Modbus TCP and Modbus RTU-over-TCP): A large cluster of unauthenticated Modbus","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Socomec DIRIS Digiware M-70 1.6.9 (Modbus TCP and Modbus RTU-over-TCP)","year":"2024","cvss_score":8.6,"severity":"high","kev":false,"impact":"A large cluster of unauthenticated Modbus denial-of-service and buffer-overflow issues in a device that sits on the electrical monitoring backbone - and two of them do something worse than crash: the DoS also weakens credentials such that the documented default credentials become valid on the device again. That converts a crash into an authentication bypass, which is a genuinely nasty combination on gear that monitors and in some deployments controls electrical distribution. Losing the Digiware monitoring layer blinds the operator to power conditions across the hall during whatever else the attacker is doing, and a device that has silently reverted to default credentials is a persistent foothold in the electrical segment. The Modbus angle is the general lesson here: these are unauthenticated packets to TCP 502, and the same class of embedded Modbus stack fragility exists across chiller, CDU and ATS controllers throughout the facility.","attack_vector":"Unauthenticated network packets to the device's Modbus TCP service - Talos confirms a single crafted packet suffices for several of these. No credentials, no interaction. The device lives on the facility/electrical VLAN, typically polled by the BMS or a DCIM collector, and is often reachable from any host on that segment because Modbus deployments almost never carry ACLs.","remediation":"Firmware update from Socomec for the DIRIS Digiware M-70. That is a per-device flash on live electrical monitoring gear, coordinated with the electrical contractor - not a cooling outage, but still a scheduled window and a technician per device. Because credential reversion is in scope, after patching you must re-verify that default credentials no longer work on every unit, not just assume the update handled it. The durable control is the Modbus one: default-deny on TCP 502, permit only the poller, and alert on any Modbus write function code appearing on a segment that should only see reads.","references":["https://talosintelligence.com/vulnerability_reports/TALOS-2024-2119","https://talosintelligence.com/vulnerability_reports/TALOS-2024-2118","https://nvd.nist.gov/vuln/detail/CVE-2024-48882","https://nvd.nist.gov/vuln/detail/CVE-2024-49572"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-22108","cve":"CVE-2025-22108","aliases":[],"title":"Linux bnxt_en driver (TX BD bd_cnt field masking): The 5-bit bd_cnt field in the transmit buffer descriptor","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux bnxt_en driver (TX BD bd_cnt field masking)","year":"2025","cvss_score":8.6,"severity":"high","kev":false,"impact":"The 5-bit bd_cnt field in the transmit buffer descriptor is not masked, so an out-of-range value corrupts transmit descriptors and produces TX timeouts. Practically: the NIC stops transmitting. On a training node that is a hung collective, and the failure looks like a network problem rather than a driver bug, so it burns debugging time.","attack_vector":"Triggered by transmit paths that produce a descriptor count above the representable range. Local, driven by workload traffic patterns.","remediation":"Kernel/driver upgrade plus host reboot. No firmware change required.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-22108"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-04-16"},{"id":"CVE-2025-32008","cve":"CVE-2025-32008","aliases":["INTEL-SA-01315","CVE-2025-20080","CVE-2026-20715","INTEL-SA-01427"],"title":"Intel AMT and Intel Standard Manageability firmware (current CSME generations): Out-of-bounds write in AMT/ISM firmware","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel AMT and Intel Standard Manageability firmware (current CSME generations)","year":"2025","cvss_score":8.6,"severity":"high","kev":false,"impact":"Out-of-bounds write in AMT/ISM firmware reachable by an unauthenticated network adversary, with a companion null-pointer dereference in the same advisory and a further unauthenticated network denial-of-service in the August 2026 batch. The immediate effect is that anyone on the management path can knock the manageability engine of a node over; memory corruption in the ME is also the standard precursor to code execution in it. The operator-facing point is that AMT is still, in 2026, an unauthenticated network attack surface on the management VLAN - and losing the ME on a node loses your out-of-band recovery path exactly when you need it.","attack_vector":"Network adversary, unauthenticated, reaching the AMT/ISM listener on the node. Same exposure as the 2017-era AMT bugs: management VLAN, or tenant space if the manageability path is not fully isolated.","remediation":"CSME firmware flash from the OEM (Dell, HPE, Supermicro, Lenovo, Gigabyte, Quanta) with a host reboot and job drain. The lasting control is the same one operators keep skipping: unprovision AMT on every SKU where you do not actively use it, disable it in the BIOS profile, and block 16992/16993/623/664/5900 anywhere a tenant-reachable segment could touch it. If you do use AMT, put it behind TLS with mutual auth and a dedicated segment.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-32008","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01315.html","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01427.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:L/A:L","cwe":["CWE-908","CWE-200"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-38608","cve":"CVE-2025-38608","aliases":[],"title":"Linux kernel (net/tls): When a BPF socket policy shrinks the plaintext after the ciphertext length was computed, kTLS","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/tls)","year":"2025","cvss_score":8.6,"severity":"high","kev":false,"impact":"When a BPF socket policy shrinks the plaintext after the ciphertext length was computed, kTLS encrypts and transmits the stale tail - uninitialized kernel memory is appended to a complete Application Data record and sent to the peer. Kernel heap bytes leave the node on the wire, and the receiver sees a malformed record.","attack_vector":"Requires an sk_msg/BPF policy calling bpf_msg_pop_data() attached to a kTLS TX socket - the shape of a service mesh or CNI sidecar doing L7 policy over encrypted traffic. The party who receives the leak is whoever is on the far end of the connection, which for a tenant-facing proxy can be the tenant itself. Attaching the policy needs CAP_BPF; reading the leaked bytes needs nothing.","remediation":"Boot a kernel carrying the linked stable commits. Interim: stop running sk_msg programs that shorten payloads (bpf_msg_pop_data) over kTLS sockets, or terminate TLS in userspace on affected nodes.","references":["https://git.kernel.org/stable/c/6ba20ff3cdb96a908b9dc93cf247d0b087672e7c","https://git.kernel.org/stable/c/849d24dc5aed45ebeb3490df429356739256ac40","https://nvd.nist.gov/vuln/detail/CVE-2025-38608"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-59887","cve":"CVE-2025-59887","aliases":["ETN-VA-2025-1026"],"title":"Eaton UPS Companion (EUC) software installer: The installer does not properly authenticate the library files it loads","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Eaton UPS Companion (EUC) software installer","year":"2025","cvss_score":8.6,"severity":"high","kev":false,"impact":"The installer does not properly authenticate the library files it loads, so an attacker who can place a file alongside the installation package gets arbitrary code execution during install - at whatever privilege the installer runs with, which is administrative. The exposure window is your own deployment process.","attack_vector":"An attacker with write access to wherever the installation package is staged - a shared drive, a downloads folder, an imaging share.","remediation":"Use the fixed EUC version from Eaton's download centre and stage installers somewhere with restricted write access. Verify package hashes before running. Companion issues CVE-2025-59888 (unquoted search path) and CVE-2025-67450 (insecure library loading) have the same fix and the same mitigation.","references":["https://www.eaton.com/content/dam/eaton/company/news-insights/cybersecurity/security-bulletins/etn-va-2025-1026.pdf"],"status":"curated","published":"2025-12-26"},{"id":"CVE-2026-1603","cve":"CVE-2026-1603","aliases":[],"title":"Ivanti Endpoint Manager (EPM): Auth bypass via alternate path","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ivanti Endpoint Manager (EPM)","year":"2026","cvss_score":8.6,"severity":"high","kev":true,"impact":"Auth bypass via alternate path -> unauthenticated leak of stored credential data","attack_vector":"Network (remote)","remediation":"Control-plane: patch to 2024 SU5+; rotate every credential EPM stored","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-1603"],"status":"curated","published":"2026-02-10"},{"id":"CVE-2026-22620","cve":"CVE-2026-22620","aliases":["eaton-va-2026-1005"],"title":"Eaton Tripp Lite series PADM firmware (rack PDU / ATS management): Unauthenticated authentication bypass gives","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Eaton Tripp Lite series PADM firmware (rack PDU / ATS management)","year":"2026","cvss_score":8.6,"severity":"high","kev":false,"impact":"Unauthenticated authentication bypass gives privileged access to the PDU management firmware. From a privileged session on a switched PDU an attacker controls outlet state for the rack - power-cycling nodes, or holding outlets off. Note the vendor has published an end-of-life notice for this product line alongside the advisory, which means for some deployed units the fix is replacement, not a patch.","attack_vector":"Unauthenticated, remote, against the PDU's management interface on the facility or OOB network.","remediation":"Update PADM firmware where a fixed build exists. For SKUs covered by the EOL notice there is no forward-fix path and the remediation is hardware replacement - a capex line and a rack-by-rack electrical swap, not a maintenance window. Until then, isolate the PDU management network and disable remote outlet switching.","references":["https://www.eaton.com/content/dam/eaton/company/news-insights/cybersecurity/security-bulletins/eaton-va-2026-1005.pdf"],"status":"curated","tags":["physical-impact"],"published":"2026-07-30"},{"id":"CVE-2026-24222","cve":"CVE-2026-24222","aliases":[],"title":"NemoClaw: Sensitive info exposure in logs","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NemoClaw","year":"2026","cvss_score":8.6,"severity":"high","kev":false,"impact":"Sensitive info exposure in logs","attack_vector":"Anyone with log access (incl. shared logging backends)","remediation":"Upgrade the service; purge/rotate leaked secrets from log stores","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24222","https://github.com/NVIDIA/product-security/tree/main/2026/5837"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:C/C:H/I:N/A:N","cwe":["CWE-497"],"published":"2026-04-28"},{"id":"CVE-2026-28500","cve":"CVE-2026-28500","aliases":[],"title":"ONNX: Security-control bypass through 1.20.1","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"ONNX","year":"2026","cvss_score":8.6,"severity":"high","kev":false,"impact":"Security-control bypass through 1.20.1","attack_vector":"Customer-supplied ONNX model","remediation":"Upgrade; ONNX external-data handling has no durable fix — sandbox model parsing","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-28500"],"status":"curated","published":"2026-03-18"},{"id":"CVE-2026-34445","cve":"CVE-2026-34445","aliases":[],"title":"ONNX (`ExternalDataInfo`): Security control bypass in external-data path handling","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"ONNX (`ExternalDataInfo`)","year":"2026","cvss_score":8.6,"severity":"high","kev":false,"impact":"Security control bypass in external-data path handling","attack_vector":"Customer-supplied ONNX model","remediation":"Upgrade to 1.21.0+; fourth iteration of the same external-data traversal class","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-34445"],"status":"curated","published":"2026-04-01"},{"id":"CVE-2026-59707","cve":"CVE-2026-59707","aliases":[],"title":"LocalAI (`/models/apply`): Unauthenticated SSRF fetching arbitrary internal URLs","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LocalAI (`/models/apply`)","year":"2026","cvss_score":8.6,"severity":"high","kev":false,"impact":"Unauthenticated SSRF fetching arbitrary internal URLs","attack_vector":"Unauthenticated network to the model-install endpoint","remediation":"Upgrade; block egress to internal ranges and IMDS","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-59707"],"status":"curated","published":"2026-07-07"},{"id":"CVE-2026-63086","cve":"CVE-2026-63086","aliases":[],"title":"Text Generation Inference (TGI): SSRF in the OpenAI-compatible multimodal chat endpoint","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Text Generation Inference (TGI)","year":"2026","cvss_score":8.6,"severity":"high","kev":false,"impact":"SSRF in the OpenAI-compatible multimodal chat endpoint","attack_vector":"Unauthenticated request supplying an image URL","remediation":"Upgrade past 3.3.7 and block metadata/internal egress from serving pods","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63086"],"status":"curated","published":"2026-07-16"},{"id":"CVE-2026-6444","cve":"CVE-2026-6444","aliases":[],"title":"Pure Storage FlashArray Purity (management interface privilege bypass): An authenticated low-privileged user reaches","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Pure Storage FlashArray Purity (management interface privilege bypass)","year":"2026","cvss_score":8.6,"severity":"high","kev":false,"impact":"An authenticated low-privileged user reaches functionality beyond their assigned privileges in the array management interface - horizontal and vertical RBAC bypass on the storage control plane.","attack_vector":"Authenticated low-privilege user of the Purity management interface.","remediation":"Apply the Purity update per Pure's security bulletins; review array RBAC assignments afterwards.","references":["https://support.purestorage.com/bundle/m_security_bulletins/page/Pure_Security/topics/concept/c_security_bulletins.html"],"status":"curated"},{"cvss_vector":"CVSS:4.0/AV:N/AC:L/AT:N/PR:L/UI:N/VC:H/VI:H/VA:N/SC:N/SI:N/SA:N","cwe":["CWE-22"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2026-71309","cve":"CVE-2026-71309","aliases":[],"title":"rclone (serve restic): Path validation in serve restic is incomplete, so an authenticated caller escapes the configured","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"rclone (serve restic)","year":"2026","cvss_score":8.6,"severity":"high","kev":false,"impact":"Path validation in serve restic is incomplete, so an authenticated caller escapes the configured backend root and reaches data outside the served subtree. On a shared backup or artifact endpoint that is one tenant reading and writing another's files.","attack_vector":"Any authenticated user of an rclone serve restic endpoint.","remediation":"Upgrade rclone to the release in GHSA-45pq-889g-fcgh and restart the serve process. Check for reads and writes outside each user's expected prefix, and scope the underlying storage credential to the served subtree so an escape at the rclone layer still hits a storage-side denial.","references":["https://github.com/rclone/rclone/security/advisories/GHSA-45pq-889g-fcgh","https://nvd.nist.gov/vuln/detail/CVE-2026-71309"],"status":"curated","tags":["tenant-isolation"]},{"id":"NCVD-2021-003-infiniband-rocev2-transport-rnic","cve":null,"aliases":["ReDMArk","ReDMArk QP/PSN predictability","RDMA packet injection by impersonation","Rothenberger et al., USENIX Security 2021"],"title":"InfiniBand / RoCEv2 transport - RNIC connection state (QP number, PSN) on Mellanox ConnectX-class and compatible RNICs","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"InfiniBand / RoCEv2 transport - RNIC connection state (QP number, PSN) on Mellanox ConnectX-class and compatible RNICs","year":"2021","cvss_score":8.6,"severity":"high","kev":false,"impact":"RDMA has no cryptographic binding between a packet and the connection it claims to belong to. A Reliable Connected queue pair is identified only by destination QP number plus packet sequence number, and ReDMArk showed both are highly predictable on real RNICs - QP numbers are handed out near-sequentially by the firmware and initial PSNs are drawn from a weak generator. Anyone who can put a frame on the fabric with the victim's source GID/LID can inject a valid-looking RDMA WRITE or SEND into an established connection between two other tenants. On a shared GPU cluster that means a neighbouring tenant, or a compromised node anywhere in the same partition, can write into another tenant's registered memory - which in an AI cluster is model weights, KV cache, gradient buffers, or NCCL communication buffers - with no software on the victim host ever seeing the write.","attack_vector":"Attacker needs one host with an RNIC on the same L2/L3 RoCE domain or IB subnet as the victims (a rented bare-metal node, a container with a VF or an SR-IOV VF, or a compromised storage/management node). They enumerate the QPN space by opening their own connections to the victim host to learn the allocator's current position, then brute-force or predict PSN and emit crafted BTH/RETH headers with a spoofed source. On RoCEv2 the outer frame is ordinary UDP/4791, so spoofing is as easy as a raw socket if the switch does not enforce source MAC/IP filtering. No exploit of a software bug is required - this is the protocol working as designed.","remediation":"No patch exists; this is architectural. Config change first: put every tenant in its own InfiniBand partition (P_Key) or its own RoCE VLAN/VRF, and turn on switch-side source-address enforcement (port security / IP source guard / MAC learning locks) so a node cannot emit frames claiming another node's GID - that is a switch config push, no reload needed on most platforms. Where the NIC supports it, enable RoCEv2 link-layer encryption/authentication (IPsec or PSP offload on ConnectX-6 Dx and later, or NVIDIA BlueField DPU-terminated crypto) - this is a firmware flash plus driver upgrade on the NIC fleet and costs a rolling host reboot per node. Hard-partitioning tenants onto separate physical fabrics or separate IB subnets is the only complete answer and is a capacity/cost decision, not a patch.","references":["https://www.usenix.org/conference/usenixsecurity21/presentation/rothenberger","https://netsec.ethz.ch/publications/papers/sec21summer-redmark.pdf","https://www.usenix.org/conference/usenixsecurity22/presentation/xing"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"tags":["tenant-isolation"]},{"id":"NCVD-2021-009-infiniband-rocev2-transport-rnic","cve":null,"aliases":["ReDMArk","ReDMArk QP/PSN predictability","RDMA packet injection by impersonation","Rothenberger et al., USENIX Security 2021"],"title":"InfiniBand / RoCEv2 transport - RNIC connection state (QP number, PSN) on Mellanox ConnectX-class and compatible RNICs","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"InfiniBand / RoCEv2 transport - RNIC connection state (QP number, PSN) on Mellanox ConnectX-class and compatible RNICs","year":"2021","cvss_score":8.6,"severity":"high","kev":false,"impact":"RDMA has no cryptographic binding between a packet and the connection it claims to belong to. A Reliable Connected queue pair is identified only by destination QP number plus packet sequence number, and ReDMArk showed both are highly predictable on real RNICs - QP numbers are handed out near-sequentially by the firmware and initial PSNs are drawn from a weak generator. Anyone who can put a frame on the fabric with the victim's source GID/LID can inject a valid-looking RDMA WRITE or SEND into an established connection between two other tenants. On a shared GPU cluster that means a neighbouring tenant, or a compromised node anywhere in the same partition, can write into another tenant's registered memory - which in an AI cluster is model weights, KV cache, gradient buffers, or NCCL communication buffers - with no software on the victim host ever seeing the write.","attack_vector":"Attacker needs one host with an RNIC on the same L2/L3 RoCE domain or IB subnet as the victims (a rented bare-metal node, a container with a VF or an SR-IOV VF, or a compromised storage/management node). They enumerate the QPN space by opening their own connections to the victim host to learn the allocator's current position, then brute-force or predict PSN and emit crafted BTH/RETH headers with a spoofed source. On RoCEv2 the outer frame is ordinary UDP/4791, so spoofing is as easy as a raw socket if the switch does not enforce source MAC/IP filtering. No exploit of a software bug is required - this is the protocol working as designed.","remediation":"No patch exists; this is architectural. Config change first: put every tenant in its own InfiniBand partition (P_Key) or its own RoCE VLAN/VRF, and turn on switch-side source-address enforcement (port security / IP source guard / MAC learning locks) so a node cannot emit frames claiming another node's GID - that is a switch config push, no reload needed on most platforms. Where the NIC supports it, enable RoCEv2 link-layer encryption/authentication (IPsec or PSP offload on ConnectX-6 Dx and later, or NVIDIA BlueField DPU-terminated crypto) - this is a firmware flash plus driver upgrade on the NIC fleet and costs a rolling host reboot per node. Hard-partitioning tenants onto separate physical fabrics or separate IB subnets is the only complete answer and is a capacity/cost decision, not a patch.","references":["https://www.usenix.org/conference/usenixsecurity21/presentation/rothenberger","https://netsec.ethz.ch/publications/papers/sec21summer-redmark.pdf","https://www.usenix.org/conference/usenixsecurity22/presentation/xing"],"status":"curated","tags":["tenant-isolation"]},{"id":"NCVD-2022-002-nvme-over-fabrics-protocol-over","cve":null,"aliases":["NeVerMore NVMe-oF attacks","NQN spoofing","NVMe-oF controller hijack over RDMA"],"title":"NVMe-over-Fabrics protocol over RDMA - SPDK NVMe-oF target and Linux kernel nvmet: NeVerMore implemented seven attacks","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVMe-over-Fabrics protocol over RDMA - SPDK NVMe-oF target and Linux kernel nvmet","year":"2022","cvss_score":8.6,"severity":"high","kev":false,"impact":"NeVerMore implemented seven attacks against the NVMe-oF protocol itself and confirmed them on the two implementations that matter operationally - SPDK and Linux nvmet. The core finding is that NVMe-oF's access control leans on the RDMA connection and on the host NQN string, neither of which is authenticated by default. A host NQN is just a text identifier the initiator asserts; there is nothing stopping a tenant from claiming another tenant's NQN and being handed their namespaces. For a neocloud selling disaggregated NVMe to GPU tenants, this means dataset and checkpoint volumes belonging to one customer can be attached read-write by another.","attack_vector":"The attacker connects to the target's discovery and I/O controllers over RDMA (or TCP) and presents a forged Host NQN, or hijacks an existing connection using the RDMA injection primitives above. Because NVMe-oF allow-lists are typically written as 'NQN X may see subsystem Y', spoofing the NQN is sufficient. Discovery controllers make reconnaissance trivial by listing every subsystem NQN and transport address on the fabric to any peer that asks.","remediation":"Config change, and it is the single highest-value one in this slice: enable NVMe-oF in-band authentication (DH-HMAC-CHAP, supported in Linux nvmet since 6.0 and in SPDK) so the NQN is proven rather than asserted, and enable TLS (NVMe/TCP) or fabric-level crypto where available. No reboot; nvmet accepts this via configfs at runtime, though initiators must be reconfigured in step so plan a rolling attach/detach. Additionally restrict the discovery controller to an authenticated management network rather than the tenant fabric, and put storage traffic on its own VLAN/P_Key. For SPDK, upgrade to a current release (process restart, brief I/O pause).","references":["https://arxiv.org/abs/2202.08080","https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2026/CVE-2026-64320.json"],"status":"curated","fleet":{"pain_class":"hot-patch"},"tags":["tenant-isolation"]},{"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2023-011-harbor-harbor-helm-default-core","cve":null,"aliases":["GHSA-j7jh-fmcm-xxwv"],"title":"Harbor (harbor-helm, default core.secretName JWT signing key): SUPPLY CHAIN, UNAUTHENTICATED REGISTRY ACCESS: Harbor","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Harbor (harbor-helm, default core.secretName JWT signing key)","year":"2023","cvss_score":8.6,"severity":"high","kev":false,"impact":"SUPPLY CHAIN, UNAUTHENTICATED REGISTRY ACCESS: Harbor installed via harbor-helm without core.secretName set falls back to a default public/private keypair to sign the JWT tokens that authorize image push and pull. The key is public, so anyone can forge a token and push or pull images in that Harbor with no authentication at all. For a GPU cluster whose registry is the source of every container image the fleet runs, forging push tokens means planting an image that nodes will pull and execute; forging pull tokens means reading every tenant's private images. Two details make this nastier than a normal default-credential bug: upgrading harbor-helm does not fix an already-installed instance, since the key is baked into the existing deployment, and robot accounts derive their tokens from the same key, so remediation forces regeneration of every robot token — legacy-marked robot accounts cannot be rotated at all and must be deleted and recreated.","attack_vector":"Network, fully unauthenticated. Anyone who can reach the Harbor API and knows the public default key can mint valid push/pull tokens. Applies to instances installed with affected harbor-helm versions where core.secretName was left unset; docker-compose, harbor-tile and TKG/Carvel installs are not affected.","remediation":"Do not treat the harbor-helm upgrade as the fix for a running instance. Set core.secretName to a generated secret and apply it to the existing deployment, then restart the core component. Regenerate every robot account token afterwards, and delete and recreate any robot account marked Legacy since it cannot be rotated. Audit the registry for images pushed during the exposure window and re-verify digests of anything running in the fleet. Fixed harbor-helm versions that remove the insecure default: 1.3.18, 1.9.6, 1.10.4, 1.11.1.","references":["https://github.com/goharbor/harbor/security/advisories/GHSA-j7jh-fmcm-xxwv"],"status":"curated"},{"id":"CVE-2020-11013","cve":"CVE-2020-11013","aliases":[],"title":"Helm: The `lookup` template function discloses in-cluster resources, including Secrets, to a chart author","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Helm","year":"2020","cvss_score":8.5,"severity":"high","kev":false,"impact":"The `lookup` template function discloses in-cluster resources, including Secrets, to a chart author","attack_vector":"Anyone who can get an operator to install their chart","remediation":"Upgrade Helm; review third-party charts before install","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-11013"],"status":"curated","published":"2020-04-24"},{"id":"CVE-2021-30465","cve":"CVE-2021-30465","aliases":[],"title":"runc: Container filesystem breakout via directory traversal in mount handling","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"runc","year":"2021","cvss_score":8.5,"severity":"high","kev":false,"impact":"Container filesystem breakout via directory traversal in mount handling; host root","attack_vector":"Any tenant workload able to create multiple pods with specific mount config","remediation":"Replace runc binary; drain node","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-30465"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2021-05-27"},{"id":"CVE-2022-28181","cve":"CVE-2022-28181","aliases":[],"title":"NVIDIA GPU Display Driver - kernel mode layer (Windows nvlddmkm.sys and Linux nvidia.ko): A specially crafted shader","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - kernel mode layer (Windows nvlddmkm.sys and Linux nvidia.ko)","year":"2022","cvss_score":8.5,"severity":"high","kev":false,"impact":"A specially crafted shader causes an out-of-bounds write in the kernel mode layer, reaching code execution with a changed scope - and NVIDIA scores this AV:Network, meaning a remotely delivered shader (WebGL, a remote render session, a shared compute service) can reach it. This is the most dangerous display-driver bug in the 2022 set for anyone running remote rendering or browser-facing GPU workloads. Both the Windows and Linux datacenter drivers are affected, so a mixed fleet needs two separate rollouts.","attack_vector":"Local and unprivileged on either OS. On Linux it is reachable from any GPU container via /dev/nvidia*; on Windows from any session holding a GPU handle.","remediation":"Upgrade both the Linux and the Windows datacenter driver branches listed in bulletin 5353. Cost: Linux needs a drain and nvidia.ko reload per node; Windows needs a reboot per node. Two change windows unless your fleet is homogeneous.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28181","https://github.com/NVIDIA/product-security/tree/main/2022/5353"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2022-05-17"},{"id":"CVE-2022-28182","cve":"CVE-2022-28182","aliases":[],"title":"NVIDIA GPU Display Driver - Windows DirectX 11 user mode driver (nvwgf2um.dll / nvwgf2umx.dll): A crafted shader causes","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows DirectX 11 user mode driver (nvwgf2um.dll / nvwgf2umx.dll)","year":"2022","cvss_score":8.5,"severity":"high","kev":false,"impact":"A crafted shader causes an out-of-bounds write in the DX11 user mode driver, reaching code execution with a changed scope. NVIDIA scores it AV:Network with no privileges required, which is the signature of a shader delivered over a remote rendering or browser path rather than by a local user. For a cloud-gaming, VDI or remote-workstation operator this is the bug in the 2022 set that actually crosses a customer boundary.","attack_vector":"Network, no privileges. The attacker supplies a shader that the victim's GPU compiles and runs - via a web page, a streamed application, or a shared render pipeline. On a multi-session Windows host this reaches other users' sessions.","remediation":"Install the fixed Windows driver from bulletin 5353. Cost: node reboot after driver replacement, so drain sessions. Until patched, the only real compensating control is not accepting untrusted shader input, which is not an option for a cloud-gaming or remote-workstation product.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28182","https://github.com/NVIDIA/product-security/tree/main/2022/5353"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2022-05-17"},{"id":"CVE-2022-34671","cve":"CVE-2022-34671","aliases":[],"title":"GPU Display Driver (kernel): Local privesc to host root (kernel buffer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver (kernel)","year":"2022","cvss_score":8.5,"severity":"high","kev":false,"impact":"Local privesc to host root (kernel buffer overflow)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34671","https://github.com/NVIDIA/product-security/tree/main/2023/5468"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-reboot"},"published":"2022-12-30"},{"id":"CVE-2023-22736","cve":"CVE-2023-22736","aliases":[],"title":"Argo CD: Authorization bypass lets an Application be synced to a destination it is not permitted to reach","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2023","cvss_score":8.5,"severity":"high","kev":false,"impact":"Authorization bypass lets an Application be synced to a destination it is not permitted to reach","attack_vector":"Any authenticated Argo CD user","remediation":"Rolling Argo CD upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-22736"],"status":"curated","published":"2023-01-26"},{"id":"CVE-2024-22280","cve":"CVE-2024-22280","aliases":[],"title":"VMware Aria Automation (SQL injection): An authenticated user injects SQL and performs unauthorized read/write","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"VMware Aria Automation (SQL injection)","year":"2024","cvss_score":8.5,"severity":"high","kev":false,"impact":"An authenticated user injects SQL and performs unauthorized read/write against the Aria Automation database, which drives automated provisioning across the estate.","attack_vector":"Authenticated low-privilege Aria Automation user.","remediation":"Apply the Broadcom fix per advisory 24598. Appliance patch and restart.","references":["https://support.broadcom.com/web/ecx/support-content-notification/-/external/content/SecurityAdvisories/0/24598"],"status":"curated"},{"id":"CVE-2025-14459","cve":"CVE-2025-14459","aliases":[],"title":"KubeVirt CDI: PVCs can be cloned from unauthorized namespaces via DataImportCron","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"KubeVirt CDI","year":"2025","cvss_score":8.5,"severity":"high","kev":false,"impact":"PVCs can be cloned from unauthorized namespaces via DataImportCron; cross-tenant data theft","attack_vector":"Cluster user with namespace access","remediation":"Upgrade CDI","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-14459"],"status":"curated","published":"2026-01-26"},{"id":"CVE-2025-22218","cve":"CVE-2025-22218","aliases":[],"title":"VMware Aria Operations for Logs (credential disclosure): A View Only Admin reads the credentials of other VMware","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"VMware Aria Operations for Logs (credential disclosure)","year":"2025","cvss_score":8.5,"severity":"high","kev":false,"impact":"A View Only Admin reads the credentials of other VMware products integrated with Aria Operations for Logs - so the lowest-privilege admin role yields credentials for vCenter and friends.","attack_vector":"Authenticated View Only Admin on Aria Operations for Logs.","remediation":"Apply the Broadcom fix per advisory 25329, then rotate the integration credentials that were stored - anyone holding that role could already have read them.","references":["https://support.broadcom.com/web/ecx/support-content-notification/-/external/content/SecurityAdvisories/0/25329"],"status":"curated"},{"id":"CVE-2025-23267","cve":"CVE-2025-23267","aliases":[],"title":"Container Toolkit: Container escape / host file write via symlink following","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Container Toolkit","year":"2025","cvss_score":8.5,"severity":"high","kev":false,"impact":"Container escape / host file write via symlink following","attack_vector":"Any tenant with a container","remediation":"Bump nvidia-container-toolkit + restart runtime; upgrade GPU Operator; evict tenants","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23267","https://github.com/NVIDIA/product-security/tree/main/2025/5659"],"status":"curated","fleet":{"ubiquity":"Universal - default hook path in every install","remediation_pain":"`daemon-restart` to 1.17.8","pain_class":"daemon-restart","why_fleet_wide":"Link-following (symlink) in the privileged `update-ldcache` hook lets a crafted image tamper with host files or DoS the node, taking down every tenant scheduled on it"},"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:C/C:N/I:L/A:H","cwe":["CWE-59"],"published":"2025-07-17"},{"id":"CVE-2025-41250","cve":"CVE-2025-41250","aliases":[],"title":"VMware vCenter (SMTP header injection via scheduled tasks): A non-administrative user with scheduled-task permissions","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"VMware vCenter (SMTP header injection via scheduled tasks)","year":"2025","cvss_score":8.5,"severity":"high","kev":false,"impact":"A non-administrative user with scheduled-task permissions manipulates vCenter notification emails - useful for phishing operators with mail that genuinely originates from vCenter.","attack_vector":"Authenticated low-privilege vCenter user able to create scheduled tasks.","remediation":"Apply the Broadcom fix per advisory 36150. vCenter patch; low urgency relative to the RCE issues but it enables convincing internal phishing.","references":["https://support.broadcom.com/web/ecx/support-content-notification/-/external/content/SecurityAdvisories/0/36150"],"status":"curated"},{"id":"CVE-2025-53547","cve":"CVE-2025-53547","aliases":[],"title":"Helm: Crafted Chart.yaml plus a symlinked Chart.lock gives local code execution when dependencies are updated","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Helm","year":"2025","cvss_score":8.5,"severity":"high","kev":false,"impact":"Crafted Chart.yaml plus a symlinked Chart.lock gives local code execution when dependencies are updated; compromises the CD runner","attack_vector":"Malicious chart pulled by a GitOps pipeline","remediation":"Upgrade Helm to 3.18.4+ everywhere charts are rendered, including Argo CD and Flux images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-53547"],"status":"curated","published":"2025-07-08"},{"id":"CVE-2025-61972","cve":"CVE-2025-61972","aliases":[],"title":"AMD NBIO register lock bits - System Management Network access: NBIO registers that should be locked after boot are","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD NBIO register lock bits - System Management Network access","year":"2025","cvss_score":8.5,"severity":"high","kev":false,"impact":"NBIO registers that should be locked after boot are not, so a local admin-privileged attacker gets arbitrary access to the System Management Network - the internal bus that reaches the ASP and the SMU. From SMN access the attacker executes code in the AMD Secure Processor itself, which collapses both confidentiality and integrity for every SEV-SNP guest on the machine. At 8.5 this is the most severe ASP-reachable issue in the current batch.","attack_vector":"Local, host administrator. Exactly the threat model SEV-SNP claims to defend against, which is what makes it serious rather than routine.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. Because this sits inside the SEV-SNP trust boundary, the update also moves the platform's reported TCB version: after patching you must refresh VCEK certificates from AMD's KDS and update whatever attestation policy your tenants (or your own confidential-VM control plane) pin against, or every guest launch will start failing validation. Treat any confidential-computing SLA you offer as void on unpatched nodes: the host operator - or anyone who compromises the host operator's tooling - can reach guest memory.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-61972","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2026-05-13"},{"cvss_vector":"CVSS:4.0/AV:N/AC:L/AT:N/PR:H/UI:N/VC:H/VI:H/VA:N/SC:N/SI:N/SA:N","cwe":["CWE-522"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2025-62157","cve":"CVE-2025-62157","aliases":["GHSA-c2hv-4pfj-mm2r"],"title":"Argo Workflows (workflow-controller, artifact repository credential logging): Workflow-controller writes artifact","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo Workflows (workflow-controller, artifact repository credential logging)","year":"2025","cvss_score":8.5,"severity":"high","kev":false,"impact":"Workflow-controller writes artifact repository credentials into its own logs in plaintext. Anyone who can read pod logs in the Argo namespace - a monitoring sidecar, an on-call engineer, a tenant with over-broad RBAC - gets full read/write/delete on the shared artifact bucket that holds every tenant's inputs, checkpoints and outputs.","attack_vector":"A principal with get on pods/log for the workflow-controller pod, or read access to whatever log pipeline collects it.","remediation":"Upgrade to 3.6.12 or 3.7.3 and restart the controller, then rotate the artifact repository credentials. Note the sibling CVE-2026-42295 covers the same leak in the executor, so patch both before declaring the credentials clean.","references":["https://github.com/argoproj/argo-workflows/security/advisories/GHSA-c2hv-4pfj-mm2r","https://nvd.nist.gov/vuln/detail/CVE-2025-62157"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-64324","cve":"CVE-2025-64324","aliases":[],"title":"KubeVirt: hostDisk feature mounts host files into a VM with insufficient restriction","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"KubeVirt","year":"2025","cvss_score":8.5,"severity":"high","kev":false,"impact":"hostDisk feature mounts host files into a VM with insufficient restriction; host data exposure","attack_vector":"Cluster user able to create a VM with hostDisk","remediation":"Upgrade KubeVirt to 1.6.1/1.7.0+; disable the hostDisk feature gate for tenants","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-64324"],"status":"curated","published":"2025-11-18"},{"cvss_vector":"CVSS:4.0/AV:N/AC:L/AT:N/PR:H/UI:N/VC:H/VI:N/VA:N/SC:H/SI:H/SA:H","cwe":["CWE-532"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-0207","cve":"CVE-2026-0207","aliases":[],"title":"Pure Storage FlashBlade logging: Sensitive material ends up in FlashBlade logs under certain conditions, and the scored","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Pure Storage FlashBlade logging","year":"2026","cvss_score":8.5,"severity":"high","kev":false,"impact":"Sensitive material ends up in FlashBlade logs under certain conditions, and the scored impact extends beyond the array itself - so whatever leaks is useful against adjacent systems, not just the storage.","attack_vector":"An account with high privileges on the array, or anyone able to read logs and support bundles collected from it.","remediation":"Upgrade Purity//FB to the fixed release from Pure's bulletin, purge affected logs, and rotate any secret that could have been written into them.","references":["https://support.purestorage.com/bundle/m_security_bulletins/page/Pure_Security/topics/concept/c_security_bulletins.html","https://nvd.nist.gov/vuln/detail/CVE-2026-0207"],"status":"curated"},{"id":"CVE-2026-20898","cve":"CVE-2026-20898","aliases":["INTEL-SA-01439","CVE-2025-20004","CVE-2025-24305","INTEL-SA-01273","INTEL-SA-01313"],"title":"Alias Checking Trusted Module (ACTM) firmware for Intel Xeon processors, including Xeon 6: Improper access control","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Alias Checking Trusted Module (ACTM) firmware for Intel Xeon processors, including Xeon 6","year":"2026","cvss_score":8.5,"severity":"high","kev":false,"impact":"Improper access control in ACTM, the Intel-signed module that validates memory-aliasing configuration as part of the platform's trusted boot and confidential-computing plumbing on current Xeon parts. An adversary positioned in startup code or SMM can escalate through it, which undermines the platform integrity guarantee that ACTM exists to provide. On the newest Xeon generations this is the layer that a confidential-computing or attestation story is built on, so a defect here means the node's integrity claims to a tenant cannot be relied upon. Persistence obtained at this level is firmware-resident, survives host reimage, and carries into the next tenant occupying the node.","attack_vector":"A privileged adversary already executing in startup code or SMM - reached from local root plus an SMM or early-boot bug, or from a supply-chain-modified BIOS image.","remediation":"Platform firmware/BIOS update carrying the fixed ACTM, from the OEM (Dell, HPE, Supermicro, Lenovo, Gigabyte, Quanta, Wiwynn). Host reboot and job drain. This is current-generation silicon, so it directly affects the Xeon 6 head nodes and CPU hosts under new GPU deployments - budget for it in the same maintenance window as GPU firmware. There is no way to disable ACTM, and no host-side mitigation.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-20898","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01439.html","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01273.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2026-24260","cve":"CVE-2026-24260","aliases":[],"title":"Container Toolkit: Container escape to host root via TOCTOU race","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Container Toolkit","year":"2026","cvss_score":8.5,"severity":"high","kev":false,"impact":"Container escape to host root via TOCTOU race","attack_vector":"Any tenant that can start a container on a GPU node","remediation":"Bump nvidia-container-toolkit + restart container runtime on every GPU node; upgrade GPU Operator chart; evict and re-admit tenant workloads","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24260","https://github.com/NVIDIA/product-security/tree/main/2026/5850"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-367"],"fleet":{"pain_class":"node-drain"},"published":"2026-07-01"},{"id":"CVE-2026-25628","cve":"CVE-2026-25628","aliases":[],"title":"Qdrant (`/logger`): Append to arbitrary files via the logger endpoint","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Qdrant (`/logger`)","year":"2026","cvss_score":8.5,"severity":"high","kev":false,"impact":"Append to arbitrary files via the logger endpoint","attack_vector":"Network user of the Qdrant API","remediation":"Upgrade to 1.16.0+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-25628"],"status":"curated","published":"2026-02-06"},{"cvss_vector":"CVSS:4.0/AV:N/AC:L/AT:N/PR:H/UI:N/VC:H/VI:H/VA:N/SC:N/SI:N/SA:N","cwe":["CWE-522"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2026-42295","cve":"CVE-2026-42295","aliases":["GHSA-7vf8-2cr6-54mf"],"title":"Argo Workflows (workflow executor, artifact driver logging): The executor logs the whole artifact driver struct, so S3","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo Workflows (workflow executor, artifact driver logging)","year":"2026","cvss_score":8.5,"severity":"high","kev":false,"impact":"The executor logs the whole artifact driver struct, so S3 access keys, GCS service account keys, Azure account keys and Git passwords land in plaintext workflow pod logs. Any tenant with pod-log read in that namespace walks off with the artifact repository credentials and can read, overwrite or delete every other tenant's datasets and model artifacts. Incomplete fix of CVE-2025-62157, which covered the controller but not the executor.","attack_vector":"A user or service account with get/list on pods/log in a namespace where workflows run. No workflow submission rights needed.","remediation":"Upgrade to 4.0.5 and restart controller and server so new executors ship the fixed logging. Then rotate the artifact repository credentials and tighten pods/log RBAC - the credentials are already in whatever log store scraped those pods.","references":["https://github.com/argoproj/argo-workflows/security/advisories/GHSA-7vf8-2cr6-54mf","https://nvd.nist.gov/vuln/detail/CVE-2026-42295"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:4.0/AV:N/AC:L/AT:N/PR:L/UI:N/VC:L/VI:H/VA:H/SC:N/SI:H/SA:H","cwe":["CWE-862"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2026-42297","cve":"CVE-2026-42297","aliases":["GHSA-xchc-cqwg-g76q"],"title":"Argo Workflows (Argo Server, ConfigMap-backed sync limit provider): The Sync Service's ConfigMap provider runs no","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo Workflows (Argo Server, ConfigMap-backed sync limit provider)","year":"2026","cvss_score":8.5,"severity":"high","kev":false,"impact":"The Sync Service's ConfigMap provider runs no auth.CanI check on any CRUD path, so any authenticated caller - including one presenting a bogus bearer token - can create, edit or delete the ConfigMaps that hold workflow synchronization limits. Those limits are the concurrency gates on shared resources, so a tenant can raise their own ceiling or zero out someone else's and starve or stampede the GPU pool.","attack_vector":"Any client that can reach the Argo Server API and present any bearer token. Effectively unauthenticated in deployments that accept client-mode tokens.","remediation":"Upgrade Argo Server to 4.0.5 and restart the deployment. Until patched, block the sync endpoints at the ingress and review the sync-limit ConfigMaps in each namespace for unexpected edits.","references":["https://github.com/argoproj/argo-workflows/security/advisories/GHSA-xchc-cqwg-g76q","https://nvd.nist.gov/vuln/detail/CVE-2026-42297"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-44944","cve":"CVE-2026-44944","aliases":["CVE-2026-44943","CVE-2026-55995"],"title":"open-iscsi / open-isns - iscsiuio control socket authorization and iSNS record handling: Three related defects","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"open-iscsi / open-isns - iscsiuio control socket authorization and iSNS record handling","year":"2026","cvss_score":8.5,"severity":"high","kev":false,"impact":"Three related defects in the initiator stack that ships on every Linux host doing iSCSI. Unprivileged local users can use the iscsiuio control socket, which controls the iSCSI offload/boot path on the NIC. A path-traversal issue lets a machine-in-the-middle attacker cause root-owned files to be created outside the iSCSI node database and inject lines into a record, which means an on-path attacker on the storage network can influence which target the host connects to on subsequent boots. And a double free in open-isns lets an unauthenticated MITM crash the daemon. For a GPU fleet where nodes boot or mount datasets over iSCSI, the middle one is the sharp end: it turns a passive network position into control over what a node believes its storage is.","attack_vector":"Local unprivileged user on the initiator host for the socket issue; machine-in-the-middle on the unauthenticated, unencrypted storage network for the path traversal and the open-isns double free.","remediation":"Package update for open-iscsi and open-isns on every initiator - i.e. every GPU node. It is a userspace daemon update, so a restart of iscsiuio/iscsid rather than a reboot, though sessions established through the offload path may need to be re-established. Disable iscsiuio entirely on hosts that do not use hardware iSCSI offload - most do not, and that removes the local-privilege surface outright. The MITM issues are an argument for CHAP mutual authentication or IPsec on the storage VLAN, since the underlying protocol offers no integrity otherwise.","references":["https://bugzilla.suse.com/show_bug.cgi?id=CVE-2026-44944","https://github.com/open-iscsi/open-iscsi/commit/668ca1df9c9a1e9bdd5c999ae1d67c9c8909237e","https://nvd.nist.gov/vuln/detail/CVE-2026-44944"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-07-29"},{"cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-327"],"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2026-040-ceph-rgw-sts-session-tokens","cve":null,"aliases":["GHSA-j73r-qrgx-jvq2","CVE-2026-39944 (reserved)"],"title":"Ceph RGW (STS session tokens): Any tenant holding one ordinary STS session token can edit it into RGW superuser. RGW","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ceph RGW (STS session tokens)","year":"2026","cvss_score":8.5,"severity":"high","kev":false,"impact":"Any tenant holding one ordinary STS session token can edit it into RGW superuser. RGW STS tokens ride the same unauthenticated AES-128-CBC handler as CephX, so the ciphertext carries no integrity protection. The attacker bit-flips the acct_type, perm_type and is_admin fields inside a token they already legitimately hold; a forged is_admin trips the global override in rgw_process_authenticated(), which short-circuits every check_caps() decision. The result is full RGW administrative access: read and write on every bucket belonging to every other tenant on the object store, plus the ability to mint and manage users. Unlike the CephX flaw this needs no encryption oracle and no packet capture — it is offline arithmetic on a token the attacker was legitimately issued, then a normal S3 request. For a neocloud selling shared object storage next to GPU capacity, this collapses the tenant boundary on the data plane customers stage their training corpora in.","attack_vector":"Network, reachable from the ordinary RGW S3 endpoint. Requires STS to be enabled on the cluster and one valid STS token of any privilege level. The token does not need to carry elevated rights and no network observation is needed.","remediation":"Upgrade to Ceph 20.2.4 or 19.2.6 and restart radosgw. If STS is not used, disable it to remove the surface. After patching, invalidate outstanding STS sessions and review RGW admin-level operations in the logs for the exposure window, since forged tokens verify cleanly and leave no distinguishing trace.","references":["https://github.com/ceph/ceph/security/advisories/GHSA-j73r-qrgx-jvq2","https://docs.ceph.com/en/latest/security/CVE-2026-39944"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-256","CWE-522"],"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2026-049-cloudnativepg-role-password-hand","cve":null,"aliases":["GHSA-w3gf-xc94-wvmj","CVE-2026-55765 (reserved)"],"title":"CloudNativePG (role password handling, pg_stat_statements exposure): CREDENTIAL DISCLOSURE ACROSS THE TENANT BOUNDARY","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"CloudNativePG (role password handling, pg_stat_statements exposure)","year":"2026","cvss_score":8.5,"severity":"high","kev":false,"impact":"CREDENTIAL DISCLOSURE ACROSS THE TENANT BOUNDARY: CloudNativePG sets role passwords by interpolating the cleartext into an ALTER ROLE utility statement. Utility statements are not parameterised, so pg_stat_statements records the literal when track_utility is on. A tenant role able to read those statistics harvests the platform-managed superuser and application-owner passwords on every rotation — meaning the credential refresh cycle keeps re-leaking rather than healing — reconnects as superuser, and where superuser TCP access is enabled reaches OS command execution in the database pod via COPY ... PROGRAM. The vendor is explicit that a default CNPG install is not affected: it needs pg_stat_statements preloaded, a tenant role granted pg_monitor or pg_read_all_stats, and enableSuperuserAccess: true. That combination is the normal shape of a managed-database product built on CNPG with a query-insights UI, which is exactly what a neocloud offering hosted Postgres alongside GPU capacity ends up building. An earlier fix suppressed log_statement only; pg_stat_statements is a separate subsystem and was untouched by it.","attack_vector":"Network, low privileges: a tenant-facing role with pg_monitor or pg_read_all_stats on a cluster that preloads pg_stat_statements. Escalation to OS execution additionally needs superuser TCP access enabled.","remediation":"Upgrade to CloudNativePG 1.28.4, 1.29.2 or 1.30.0 and roll the instances. Rotate the superuser and application-owner passwords after upgrading, since anything captured is already captured. Clusters that supply a SCRAM-SHA-256 verifier in the managed-role Secret store a non-replayable hash and were never exposed — moving to that is the durable fix. Otherwise revoke pg_monitor/pg_read_all_stats from tenant roles, or turn off enableSuperuserAccess.","references":["https://github.com/cloudnative-pg/cloudnative-pg/security/advisories/GHSA-w3gf-xc94-wvmj","https://github.com/cloudnative-pg/cloudnative-pg/pull/10724"],"status":"curated"},{"id":"CVE-2019-13139","cve":"CVE-2019-13139","aliases":[],"title":"Docker / moby: Command execution via crafted remote git build path in `docker build`","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2019","cvss_score":8.4,"severity":"high","kev":false,"impact":"Command execution via crafted remote git build path in `docker build`","attack_vector":"Anyone who can submit a build to a shared builder","remediation":"Upgrade Docker Engine; isolate tenant build runners","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-13139"],"status":"curated","published":"2019-08-22"},{"id":"CVE-2021-1051","cve":"CVE-2021-1051","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): A local user gets elevated enough to rewrite display configuration","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2021","cvss_score":8.4,"severity":"high","kev":false,"impact":"A local user gets elevated enough to rewrite display configuration through the escape handler, taking the display subsystem out. On a headless compute node the display path matters less, but the underlying escape-handler privilege gap is the same one that produces worse bugs.","attack_vector":"Any local user with GPU device access on a Windows host.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1051"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-01-08"},{"id":"CVE-2021-33162","cve":"CVE-2021-33162","aliases":[],"title":"Intel Ethernet Adapter manageability firmware (access control): Improper access control in Intel Ethernet adapter","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Ethernet Adapter manageability firmware (access control)","year":"2021","cvss_score":8.4,"severity":"high","kev":false,"impact":"Improper access control in Intel Ethernet adapter manageability firmware, allowing an authenticated user to escalate privilege. Same NC-SI sideband surface as its sibling — the difference is it needs an account, which in a bare-metal rental means the tenant. A tenant escalating through the manageability engine is a tenant reaching toward the BMC of the machine they rented, and from there toward the operator's management network.","attack_vector":"Authenticated user — on bare metal, the tenant with host access.","remediation":"OEM firmware bundle update for adapter manageability firmware plus cold power cycle. Verify NC-SI is actually needed on your platform; where the server has a dedicated BMC NIC, disabling the sideband channel in BIOS/BMC config removes this path entirely and is a config change rather than a flash.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-33162"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-02-23"},{"id":"CVE-2022-21163","cve":"CVE-2022-21163","aliases":[],"title":"Crypto API Toolkit for Intel SGX: Improper access control in the SGX Crypto API Toolkit lets an authenticated user","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Crypto API Toolkit for Intel SGX","year":"2022","cvss_score":8.4,"severity":"high","kev":false,"impact":"Improper access control in the SGX Crypto API Toolkit lets an authenticated user escalate privilege. The toolkit is what many deployments use to put HSM-style key operations inside an enclave, so a break here reaches the keys the enclave was protecting.","attack_vector":"Authenticated user of a system running the Crypto API Toolkit.","remediation":"Upgrade to Crypto API Toolkit 2.0 (commit 91ee496) or later and rotate any keys the toolkit held. Userspace/enclave update - requires re-signing and re-attesting the enclave; no reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-21163","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00746.html"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2023-02-16"},{"id":"CVE-2022-28627","cve":"CVE-2022-28627","aliases":["HPESBHF04333"],"title":"HPE iLO 5 (local privilege escalation to code execution): An unprivileged user can execute arbitrary code in the iLO","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HPE iLO 5 (local privilege escalation to code execution)","year":"2022","cvss_score":8.4,"severity":"high","kev":false,"impact":"An unprivileged user can execute arbitrary code in the iLO context. This is one of a batch of a dozen-plus iLO 5 defects HPE fixed in the same firmware release, which is itself the useful signal: the iLO 5 codebase had a cluster of memory-safety and access-control problems reachable without high privilege. Any one of them lands the attacker on the service processor, with the usual consequences - power, Virtual Media boot, console, and persistence below the hypervisor.","attack_vector":"An unprivileged actor with local access to the iLO's own execution environment - in practice, someone who already has a low-privilege iLO account or a foothold reached through one of the sibling defects in the same batch. Not an unauthenticated internet-facing entry point.","remediation":"Flash iLO 5 to v2.71 or later - and note that the later CVE-2022-28639 batch requires v2.72, so go straight to the highest available iLO 5 build rather than patching to the floor. Out-of-band, per-node, no host reboot and no drain. Treat this as one flash covering a whole batch of CVEs, which makes the per-node rollout cost-effective.","references":["https://support.hpe.com/hpsc/doc/public/display?docLocale=en_US&docId=emr_na-hpesbhf04333en_us","https://nvd.nist.gov/vuln/detail/CVE-2022-28627"],"status":"curated","published":"2022-08-12"},{"id":"CVE-2022-41736","cve":"CVE-2022-41736","aliases":[],"title":"IBM Spectrum Scale Container Native Storage Access: A local user obtains root privileges through the Spectrum Scale","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"IBM Spectrum Scale Container Native Storage Access","year":"2022","cvss_score":8.4,"severity":"high","kev":false,"impact":"A local user obtains root privileges through the Spectrum Scale container-native storage layer. Root on a node that mounts the shared training filesystem means access to whatever that node can see — which on a GPFS cluster is typically a very large namespace shared across tenants. Companion issue CVE-2022-43831 is the same shape via missing security-context settings.","attack_vector":"Local user on a node running Container Native Storage Access 5.1.2.1 through 5.1.6.0.","remediation":"Upgrade past 5.1.6.0 (rolling). Independently, enforce restrictive Kubernetes security contexts on the storage-access pods — a manifest change you can apply immediately and that closes the CVE-2022-43831 variant on its own.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-41736","https://nvd.nist.gov/vuln/detail/CVE-2022-43831"],"status":"curated","published":"2023-04-29"},{"id":"CVE-2022-42271","cve":"CVE-2022-42271","aliases":[],"title":"DGX servers (BMC firmware < 2.09.00): RCE on BMC (buffer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX servers (BMC firmware < 2.09.00)","year":"2022","cvss_score":8.4,"severity":"high","kev":false,"impact":"RCE on BMC (buffer overflow)","attack_vector":"Network-adjacent mgmt-LAN attacker","remediation":"Flash BMC to 2.09.00+ out-of-band","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42271","https://github.com/NVIDIA/product-security/tree/main/2024/5513"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-120"],"published":"2023-01-11"},{"id":"CVE-2023-0208","cve":"CVE-2023-0208","aliases":[],"title":"NVIDIA DCGM - nv-hostengine: A heap-based buffer overflow reachable through the bound socket gives denial of service","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DCGM - nv-hostengine","year":"2023","cvss_score":8.4,"severity":"high","kev":false,"impact":"A heap-based buffer overflow reachable through the bound socket gives denial of service and data tampering with a changed CVSS scope, against a root-privileged daemon on every GPU node.","attack_vector":"Local or network depending on how you bound the socket. If nv-hostengine is listening on a routable interface, any host on that network can reach it.","remediation":"Update DCGM per bulletin 5453 and restart nv-hostengine. Cost: telemetry gap of seconds, no GPU job impact, no drain. Restrict the listening socket to localhost as a standing control.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0208","https://github.com/NVIDIA/product-security/tree/main/2023/5453"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:N/I:H/A:H","cwe":["CWE-122"],"published":"2023-04-01"},{"id":"CVE-2023-22649","cve":"CVE-2023-22649","aliases":[],"title":"Rancher: Sensitive data leaked into Rancher audit logs","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Rancher","year":"2023","cvss_score":8.4,"severity":"high","kev":false,"impact":"Sensitive data leaked into Rancher audit logs","attack_vector":"Anyone with audit-log read access","remediation":"Upgrade Rancher; scrub and re-secure audit logs","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-22649"],"status":"curated","published":"2024-10-16"},{"id":"CVE-2023-31100","cve":"CVE-2023-31100","aliases":[],"title":"Phoenix SecureCore Technology 4 (SMI handler, improper access control): An SMI handler with missing access control lets","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Phoenix SecureCore Technology 4 (SMI handler, improper access control)","year":"2023","cvss_score":8.4,"severity":"high","kev":false,"impact":"An SMI handler with missing access control lets an attacker modify the SPI flash. That is the direct route to a permanent firmware implant: rewrite the boot flash and the compromise survives OS reinstall, disk replacement and node reimaging between tenants, while sitting below Secure Boot and below anything attestation can honestly measure. Highest-scored Phoenix advisory in this set.","attack_vector":"Local attacker on the host with the ability to invoke the SMI handler - in practice admin/root, or a tenant with kernel access on bare metal.","remediation":"OEM BIOS update on the fixed SecureCore Technology 4 build (affected from 4.3.0.0). Firmware flash, one reboot per node. No config workaround for the handler itself, but this is the case where platform SPI protections earn their keep: confirm BIOS Lock Enable, SMM BIOS Write Protect and protected range registers are actually set on your platform - many OEM defaults leave at least one of them off, and a properly locked flash blunts the primitive even on unpatched firmware.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31100"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-11-15"},{"id":"CVE-2023-31323","cve":"CVE-2023-31323","aliases":[],"title":"AMD Secure Processor - XGMI Trusted Agent (type confusion): Type confusion in the ASP's XGMI Trusted Agent means a","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor - XGMI Trusted Agent (type confusion)","year":"2023","cvss_score":8.4,"severity":"high","kev":false,"impact":"Type confusion in the ASP's XGMI Trusted Agent means a malformed argument is interpreted as the wrong kind of object, producing a memory-safety violation inside the secure processor. On a multi-GPU Instinct node this is a path from host-privileged code to controlling the trusted agent that governs the GPU interconnect - the highest-scoring of the XGMI pair at 8.4.","attack_vector":"Local, host-privileged, on multi-GPU XGMI-connected platforms.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. Same campaign as the XGMI TOCTOU issue - patch both in one OEM BIOS pass on MI-series nodes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31323","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2026-02-12"},{"id":"CVE-2024-1708","cve":"CVE-2024-1708","aliases":[],"title":"ConnectWise ScreenConnect: Path traversal enabling remote code execution","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"ConnectWise ScreenConnect","year":"2024","cvss_score":8.4,"severity":"high","kev":true,"impact":"Path traversal enabling remote code execution; chained with CVE-2024-1709","attack_vector":"Network (remote)","remediation":"Control-plane: same patch; audit installed extensions","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-1708"],"status":"curated","published":"2024-02-21"},{"id":"CVE-2024-36352","cve":"CVE-2024-36352","aliases":[],"title":"AMD Graphics Driver - crafted pointer leading to arbitrary writes: A specially crafted pointer passed to the AMD","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD Graphics Driver - crafted pointer leading to arbitrary writes","year":"2024","cvss_score":8.4,"severity":"high","kev":false,"impact":"A specially crafted pointer passed to the AMD graphics driver produces arbitrary writes or a crash. Arbitrary kernel writes from a GPU driver call is a privilege-escalation primitive available to anything with the device open - on a shared GPU node, that is the tenant.","attack_vector":"Local, via the graphics driver interface.","remediation":"Update the AMD graphics driver and reload or reboot. Driver-level fix, no firmware step.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36352","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-09-06"},{"id":"CVE-2024-48831","cve":"CVE-2024-48831","aliases":[],"title":"Dell SmartFabric OS10 (hard-coded password): A hard-coded password in OS10 10.5.6.x gives an unauthenticated attacker","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell SmartFabric OS10 (hard-coded password)","year":"2024","cvss_score":8.4,"severity":"high","kev":false,"impact":"A hard-coded password in OS10 10.5.6.x gives an unauthenticated attacker with local access full unauthorized access to the switch. Hard-coded means you cannot rotate it - only the patch removes it.","attack_vector":"Local access to the switch (console, or a compromised management host).","remediation":"Upgrade OS10 per DSA-2025-068. Because the credential is hard-coded, there is no config-level mitigation - the firmware upgrade and reboot is the only fix.","references":["https://www.dell.com/support/kbdoc/en-us/000295014/dsa-2025-068-security-update-for-dell-networking-os10-vulnerabilities"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:N/A:H","cwe":["CWE-125"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-50115","cve":"CVE-2024-50115","aliases":[],"title":"Linux kernel (arch/x86/kvm/svm): Hardware ignores the low five bits of CR3 when loading PDPTEs, but KVM's nested SVM","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (arch/x86/kvm/svm)","year":"2024","cvss_score":8.4,"severity":"high","kev":false,"impact":"Hardware ignores the low five bits of CR3 when loading PDPTEs, but KVM's nested SVM path did not, so an L1 guest can offset where the host reads the page-directory-pointer table from. Upstream states the worst case is an out-of-bounds read - if the target page sits at the end of a memslot and the VMM is not using guard pages, the host reads past it and feeds the result into the nested page-table walker.","attack_vector":"Guest-driven: a tenant with nested virtualization exposed sets nCR3 with bits 4:0 non-zero and executes VMRUN with PAE paging in L2. Requires an AMD host with kvm_amd nested=1 and SVM advertised in the guest's CPUID; not reachable if nested virt is withheld from tenants.","remediation":"Update to a stable kernel with the linked fix (no fixed release enumerated; take the branch carrying commit 6876793907cb). Interim control: disable nested virtualization for tenant guests (kvm_amd.nested=0, drop SVM from guest CPUID).","references":["https://git.kernel.org/stable/c/6876793907cbe19d42e9edc8c3315a21e06c32ae","https://git.kernel.org/stable/c/58cb697d80e669c56197f703e188867c8c54c494","https://nvd.nist.gov/vuln/detail/CVE-2024-50115"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-12007","cve":"CVE-2025-12007","aliases":[],"title":"Supermicro BMC firmware update signature/validation logic on the X13SEM-F motherboard family: The operator loses","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC firmware update signature/validation logic on the X13SEM-F motherboard family","year":"2025","cvss_score":8.4,"severity":"high","kev":false,"impact":"The operator loses the ability to trust or verify what firmware is actually running on the node. An attacker who can push an update writes their own BMC image to flash, and from that moment the BMC reports whatever the attacker wants it to report - version strings, attestation values, health telemetry. Reimaging the host does nothing, replacing the NVMe does nothing, and a fleet-wide firmware inventory will show the node as compliant. For a neocloud reselling bare metal this is the worst class of finding, because the implant persists across tenant boundaries and the next tenant has no way to detect it. The code path that is supposed to prove a firmware image came from Supermicro before writing it to the BMC's flash.","attack_vector":"Local access to the node's firmware update path - a host-side root process reaching the BMC over the KCS/in-band interface, or an operator-adjacent process with permission to invoke the update. This is the realistic post-exploitation move after a tenant escapes to host root on bare metal, or after any compromise of the provisioning tooling that flashes firmware during node turnup.","remediation":"Firmware flash with the fixed BMC image from Supermicro's January 2026 BMC/IPMI advisory batch. Note the ordering problem: an already-implanted BMC can lie about accepting the update, so for any node you suspect was touched you need an out-of-band SPI reflash with a hardware programmer rather than a software update, which is a hands-on-metal job per node. Going forward, restrict who can call the firmware update path at all - remove host-side IPMI/KCS access from tenant-facing bare metal images, and treat firmware updates as a privileged provisioning operation rather than something any admin session can trigger.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-12007","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2025/12xxx/CVE-2025-12007.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-23356","cve":"CVE-2025-23356","aliases":[],"title":"NVIDIA Isaac Lab (Isaac Sim): SB3 configuration parsing reaches code execution with no privileges and no user","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Isaac Lab (Isaac Sim)","year":"2025","cvss_score":8.4,"severity":"high","kev":false,"impact":"SB3 configuration parsing reaches code execution with no privileges and no user interaction required. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5708 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23356","https://github.com/NVIDIA/product-security/tree/main/2025/5708"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-306"],"published":"2025-10-14"},{"id":"CVE-2025-30479","cve":"CVE-2025-30479","aliases":[],"title":"Dell CloudLink (command injection): Command injection giving a privileged user full control of the CloudLink system","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Dell CloudLink (command injection)","year":"2025","cvss_score":8.4,"severity":"high","kev":false,"impact":"Command injection giving a privileged user full control of the CloudLink system.","attack_vector":"Adjacent network with a known privileged password.","remediation":"Upgrade CloudLink to 8.2.","references":["https://www.dell.com/support/kbdoc/en-us/000384363/dsa-2025-374-security-update-for-dell-cloudlink-multiple-security-vulnerabilities"],"status":"curated"},{"id":"CVE-2025-33225","cve":"CVE-2025-33225","aliases":[],"title":"NVIDIA Resiliency Extension: Predictable log-file names in the log-aggregation path let an attacker pre-create","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Resiliency Extension","year":"2025","cvss_score":8.4,"severity":"high","kev":false,"impact":"Predictable log-file names in the log-aggregation path let an attacker pre-create or hijack the target file, reaching privilege escalation and code execution. Scored 8.4 with no privileges required. The Resiliency Extension is what restarts failed large training jobs, so it runs with broad access across the job's nodes.","attack_vector":"Local, no privileges required, no user interaction. Any account on a node participating in a resilient training job.","remediation":"Update the Resiliency Extension per bulletin 5746 and rebuild training images. Cost: image rebuild and job restart; no host driver or firmware change.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33225","https://github.com/NVIDIA/product-security/tree/main/2025/5746"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-61"],"published":"2025-12-16"},{"id":"CVE-2025-45379","cve":"CVE-2025-45379","aliases":[],"title":"Dell CloudLink (console command injection): Command injection from the console giving shell access","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Dell CloudLink (console command injection)","year":"2025","cvss_score":8.4,"severity":"high","kev":false,"impact":"Command injection from the console giving shell access to the key-management appliance.","attack_vector":"Adjacent network plus a known privileged credential.","remediation":"Upgrade CloudLink to 8.2. Appliance upgrade with a maintenance window.","references":["https://www.dell.com/support/kbdoc/en-us/000384363/dsa-2025-374-security-update-for-dell-cloudlink-multiple-security-vulnerabilities"],"status":"curated"},{"id":"CVE-2025-52565","cve":"CVE-2025-52565","aliases":[],"title":"runc: Insufficient checks when bind-mounting /dev/console allow writes to arbitrary host procfs paths","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"runc","year":"2025","cvss_score":8.4,"severity":"high","kev":false,"impact":"Insufficient checks when bind-mounting /dev/console allow writes to arbitrary host procfs paths; container escape","attack_vector":"Any tenant workload / malicious image","remediation":"Replace runc on all nodes; drain required to restart containers","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-52565"],"status":"curated","fleet":{"ubiquity":"Universal - same runc version range","remediation_pain":"`node-drain` - same as above; the fix ships in runc 1.2.8 / 1.3.3 / 1.4.0-rc.3 and only applies to newly created containers","pain_class":"node-drain","why_fleet_wide":"`/dev/console` bind-mount race/symlink lets runc mount an unexpected target before LSM/mount protections apply, granting write access to procfs and a breakout"},"published":"2025-11-06"},{"id":"CVE-2025-54886","cve":"CVE-2025-54886","aliases":[],"title":"skops (`Card.get_model`): Model card loading has no trusted-types check","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"skops (`Card.get_model`)","year":"2025","cvss_score":8.4,"severity":"high","kev":false,"impact":"Model card loading has no trusted-types check → code execution","attack_vector":"Customer-supplied model card / repo","remediation":"Upgrade past 0.12.0","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-54886"],"status":"curated","published":"2025-08-08"},{"id":"CVE-2026-24233","cve":"CVE-2026-24233","aliases":[],"title":"TensorRT-LLM: RCE via unsafe deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT-LLM","year":"2026","cvss_score":8.4,"severity":"high","kev":false,"impact":"RCE via unsafe deserialization","attack_vector":"Malicious model artifact","remediation":"Bump TensorRT-LLM; rebuild and redeploy serving images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24233","https://github.com/NVIDIA/product-security/tree/main/2026/5840"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2026-07-14"},{"id":"CVE-2026-33747","cve":"CVE-2026-33747","aliases":[],"title":"BuildKit: Custom frontend can craft an API message causing daemon compromise","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"BuildKit","year":"2026","cvss_score":8.4,"severity":"high","kev":false,"impact":"Custom frontend can craft an API message causing daemon compromise","attack_vector":"Anyone who can supply a build frontend","remediation":"Upgrade BuildKit to 0.28.1+; restrict custom frontends","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-33747"],"status":"curated","published":"2026-03-27"},{"id":"CVE-2026-35204","cve":"CVE-2026-35204","aliases":[],"title":"Helm: Crafted plugin writes its contents to an arbitrary filesystem location on install or update","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Helm","year":"2026","cvss_score":8.4,"severity":"high","kev":false,"impact":"Crafted plugin writes its contents to an arbitrary filesystem location on install or update","attack_vector":"Malicious Helm plugin","remediation":"Upgrade Helm to 4.1.4+; restrict plugin sources","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-35204"],"status":"curated","published":"2026-04-09"},{"id":"CVE-2026-35205","cve":"CVE-2026-35205","aliases":[],"title":"Helm: Helm installs plugins with no provenance file even when signature verification is required","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Helm","year":"2026","cvss_score":8.4,"severity":"high","kev":false,"impact":"Helm installs plugins with no provenance file even when signature verification is required; supply-chain gate silently fails open","attack_vector":"Malicious Helm plugin","remediation":"Upgrade Helm to 4.1.4+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-35205"],"status":"curated","published":"2026-04-09"},{"id":"CVE-2026-53492","cve":"CVE-2026-53492","aliases":[],"title":"containerd: CRI trusts CDI annotations inside untrusted checkpoint image metadata","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2026","cvss_score":8.4,"severity":"high","kev":false,"impact":"CRI trusts CDI annotations inside untrusted checkpoint image metadata; can inject device access (directly relevant to GPU CDI devices)","attack_vector":"Malicious checkpoint image","remediation":"Rolling containerd upgrade with node drain; disable checkpoint import. High priority for GPU hosts since CDI is how GPUs are injected","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53492"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2026-07-01"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-125","CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-63860","cve":"CVE-2026-63860","aliases":[],"title":"Linux kernel (drivers/infiniband/core): IWARP port-mapper netlink attributes were accepted as plain strings with no","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/core)","year":"2026","cvss_score":8.4,"severity":"high","kev":false,"impact":"IWARP port-mapper netlink attributes were accepted as plain strings with no guarantee of a NUL terminator, then handed to strcmp and %s. A crafted attribute makes the kernel read past the attribute into adjacent memory - kernel memory disclosure into log output and comparisons, or an oops. The CNA scored it as requiring no privileges.","attack_vector":"Requires the ability to send RDMA_NL_IWPM netlink messages, normally the host's iwpmd port-mapper daemon. Relevant on iWARP-capable nodes (irdma in iWARP mode, cxgb4, siw); a tenant with CAP_NET_ADMIN in a non-user-namespaced net namespace, or anything that can impersonate the port mapper, reaches it. Not reachable from the fabric.","remediation":"Update to a stable kernel carrying fcd07d3b8ee7 (or 87111356d58d / abda65bdd130) and reboot. Interim: do not grant CAP_NET_ADMIN over the host netlink namespace to tenant workloads, and disable iWARP mode where the fabric does not require it.","references":["https://git.kernel.org/stable/c/fcd07d3b8ee7a39b344d73aed69c1a68cd9eacdf","https://git.kernel.org/stable/c/87111356d58d86edb221ba144d261ed83a5b8bbe","https://nvd.nist.gov/vuln/detail/CVE-2026-63860"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:N/A:H","cwe":["CWE-125","CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-64247","cve":"CVE-2026-64247","aliases":[],"title":"Linux kernel (arch/x86/kvm): A nested guest can put an out-of-range virtual-processor ID into an enlightened VMCS and","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (arch/x86/kvm)","year":"2026","cvss_score":8.4,"severity":"high","kev":false,"impact":"A nested guest can put an out-of-range virtual-processor ID into an enlightened VMCS and have KVM index a sparse-bank set with it unchecked, producing an out-of-bounds / use-after-free read in host kernel memory during a paravirtual TLB flush hypercall. A guest gets to read host memory it should never see, and can crash the node.","attack_vector":"Reachable from a guest that has Hyper-V enlightenments and nested virtualization available: the guest runs an L2 vCPU and copies an unbounded VP ID into the enlightened VMCS, then issues an HvFlush hypercall. KASAN caught it from an ordinary unprivileged process running a guest. Conditional on nested virt being exposed to tenants and Hyper-V enlightenments enabled.","remediation":"Update to a kernel with the referenced stable commits. Interim: turn off nested virtualization for tenant VMs (kvm_intel nested=0 / kvm_amd nested=0) and disable Hyper-V enlightenments in the guest CPU model until nodes are patched.","references":["https://git.kernel.org/stable/c/d18756b12aab30d07794446445c93112e5c69a2e","https://git.kernel.org/stable/c/83c2f52c6a78b1590034e955cff3fe0b052fe4ae","https://nvd.nist.gov/vuln/detail/CVE-2026-64247"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2019-9507","cve":"CVE-2019-9507","aliases":[],"title":"Vertiv Avocent UMG-4000 universal management gateway: Every command the UMG-4000's web interface runs executes as root","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Vertiv Avocent UMG-4000 universal management gateway","year":"2019","cvss_score":8.3,"severity":"high","kev":false,"impact":"Every command the UMG-4000's web interface runs executes as root on the underlying OS. An admin-authenticated attacker who can inject shell syntax into a web form gets root on the gateway — which sits between the operator and every server it's providing KVM/serial access to.","attack_vector":"Requires an authenticated administrator session on the web interface; the app fails to neutralize shell metacharacters before executing commands.","remediation":"Software/firmware upgrade from Vertiv; download the fixed build from Vertiv's Avocent UMG support page and flash each gateway. Since the UMG-4000 is the aggregation point for KVM access to many downstream nodes, schedule the update in a maintenance window and expect KVM sessions through that gateway to drop during the flash.","references":["https://www.vertiv.com/en-us/support/software-download/it-management/avocent-universal-management-gateway-appliance--software-downloads/"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2020-03-30"},{"id":"CVE-2021-23277","cve":"CVE-2021-23277","aliases":[],"title":"Eaton Intelligent Power Manager (IPM) prior to 1.69 - dynamic eval: Unauthenticated eval injection: user-controlled","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Eaton Intelligent Power Manager (IPM) prior to 1.69 - dynamic eval","year":"2021","cvss_score":8.3,"severity":"high","kev":false,"impact":"Unauthenticated eval injection: user-controlled code syntax reaches a dynamic evaluation path. Second unauthenticated route to code execution on the same power-management server, from the same advisory batch - which is the point worth taking away. Patching one CVE in this batch and not the rest leaves the door open.","attack_vector":"Unauthenticated, remote, to the IPM server.","remediation":"Upgrade to IPM 1.69 or later - the whole CVE-2021-23276 through -23281 batch lands in one release, so treat it as a single action.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-23277"],"status":"curated","published":"2021-04-13"},{"id":"CVE-2021-39155","cve":"CVE-2021-39155","aliases":[],"title":"Istio: Case-sensitivity mismatch in host matching bypasses authorization policy","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Istio","year":"2021","cvss_score":8.3,"severity":"high","kev":false,"impact":"Case-sensitivity mismatch in host matching bypasses authorization policy","attack_vector":"Unauthenticated network","remediation":"Rolling istiod upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-39155"],"status":"curated","published":"2021-08-24"},{"id":"CVE-2022-25987","cve":"CVE-2022-25987","aliases":["Trojan Source"],"title":"Intel C++ Compiler Classic / oneAPI toolkits (Unicode source handling): Improper handling of Unicode bidirectional","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel C++ Compiler Classic / oneAPI toolkits (Unicode source handling)","year":"2022","cvss_score":8.3,"severity":"high","kev":false,"impact":"Improper handling of Unicode bidirectional and homoglyph characters in source code means the compiler can build something materially different from what a reviewer reads. This is the Trojan Source class - a supply-chain problem for anyone compiling third-party or contributor-supplied kernels and operators into their inference stack.","attack_vector":"Anyone who can get source into your build - an internal contributor, a vendored dependency, or a model-op repo you compile from.","remediation":"Upgrade the compiler to 2021.6 / oneAPI 2022.2 or later, and add a CI check that rejects bidirectional control characters in source. Build-toolchain change only - no node reboot, no firmware. Rebuild any artefact compiled with an affected compiler if provenance matters.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-25987","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00674.html"],"status":"curated","published":"2023-02-16"},{"id":"CVE-2022-26843","cve":"CVE-2022-26843","aliases":[],"title":"Intel oneAPI DPC++/C++ compiler (homoglyph rendering): Homoglyph characters are not visually distinguished","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel oneAPI DPC++/C++ compiler (homoglyph rendering)","year":"2022","cvss_score":8.3,"severity":"high","kev":false,"impact":"Homoglyph characters are not visually distinguished by the toolchain, so two different identifiers can look identical in review. Companion to the bidirectional-override issue and part of the same supply-chain risk when compiling contributed GPU kernels.","attack_vector":"Anyone whose source reaches your compiler.","remediation":"Upgrade to oneAPI 2022.1 or later and screen source for confusable identifiers in CI. Toolchain-only, no reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-26843","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00674.html"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2023-02-16"},{"id":"CVE-2022-26872","cve":"CVE-2022-26872","aliases":[],"title":"AMI MegaRAC: Password reset interception via the API — attacker takes over an admin BMC account","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC","year":"2022","cvss_score":8.3,"severity":"high","kev":false,"impact":"Password reset interception via the API — attacker takes over an admin BMC account","attack_vector":"Network / BMC API","remediation":"BMC firmware update; interim mitigation is disabling the self-service password reset flow","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-26872"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-01-30"},{"id":"CVE-2022-31034","cve":"CVE-2022-31034","aliases":[],"title":"Argo CD: Predictable SSO state values allow authentication bypass during login","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2022","cvss_score":8.3,"severity":"high","kev":false,"impact":"Predictable SSO state values allow authentication bypass during login","attack_vector":"Unauthenticated network in a position to observe or race a login","remediation":"Rolling Argo CD upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31034"],"status":"curated","published":"2022-06-27"},{"id":"CVE-2022-31105","cve":"CVE-2022-31105","aliases":[],"title":"Argo CD: Improper certificate validation lets Argo CD be tricked into trusting a hostile endpoint","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2022","cvss_score":8.3,"severity":"high","kev":false,"impact":"Improper certificate validation lets Argo CD be tricked into trusting a hostile endpoint","attack_vector":"Unauthenticated network in a MITM position","remediation":"Rolling Argo CD upgrade; no GPU drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31105"],"status":"curated","published":"2022-07-12"},{"id":"CVE-2022-40259","cve":"CVE-2022-40259","aliases":[],"title":"AMI MegaRAC: Default credentials — Redfish API accessible with shipped account","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC","year":"2022","cvss_score":8.3,"severity":"high","kev":false,"impact":"Default credentials — Redfish API accessible with shipped account; full BMC control","attack_vector":"Network / Redfish API","remediation":"Credential rotation at rack intake; treat any node received from an ODM as compromised-by-default until BMC accounts are reset","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-40259"],"status":"curated","published":"2022-12-05"},{"id":"CVE-2023-0683","cve":"CVE-2023-0683","aliases":["LEN-99936"],"title":"Lenovo XClarity Controller (XCC) - API privilege escalation: A read-only XCC user gains elevated privileges through","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lenovo XClarity Controller (XCC) - API privilege escalation","year":"2023","cvss_score":8.3,"severity":"high","kev":false,"impact":"A read-only XCC user gains elevated privileges through a specifically crafted API call. Same shape as the later 2023 XCC batch and the same consequence: a monitoring-tier credential becomes control of the service processor, which means out-of-band power, Virtual Media boot of an attacker image, console into whatever the tenant is running, and firmware-level persistence that survives reimaging. Taken together with CVE-2023-4606 and CVE-2023-4607, this is a pattern rather than an isolated defect - XCC's API-side authorisation checks were repeatedly incomplete across 2023.","attack_vector":"An authenticated XCC account with read-only access, reaching the XCC API over the out-of-band management VLAN.","remediation":"Flash XCC to the version listed for your model in LEN-99936. Out-of-band, per-node, no host reboot and no job drain. Sequence it with the other 2023 XCC advisories so each node is touched once. Given three independent authorisation bypasses in the same year, the durable posture is to stop treating XCC read-only accounts as low-risk and to gate the XCC management segment tightly.","references":["https://support.lenovo.com/us/en/product_security/LEN-99936","https://nvd.nist.gov/vuln/detail/CVE-2023-0683"],"status":"curated","published":"2023-05-01"},{"id":"CVE-2023-25533","cve":"CVE-2023-25533","aliases":[],"title":"DGX H100 BMC: Privesc + code execution (web UI input validation)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX H100 BMC","year":"2023","cvss_score":8.3,"severity":"high","kev":false,"impact":"Privesc + code execution (web UI input validation)","attack_vector":"Authenticated BMC user on mgmt network","remediation":"Flash BMC 23.08.18 out-of-band","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5473/5473.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:H/UI:N/S:C/C:H/I:H/A:L","cwe":["CWE-20"],"published":"2023-09-20"},{"id":"CVE-2023-28083","cve":"CVE-2023-28083","aliases":["HPESBHF04456"],"title":"HPE iLO 4 / iLO 5 / iLO 6 (remote cross-site scripting): Cross-site scripting in the iLO web interface across all three","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HPE iLO 4 / iLO 5 / iLO 6 (remote cross-site scripting)","year":"2023","cvss_score":8.3,"severity":"high","kev":false,"impact":"Cross-site scripting in the iLO web interface across all three current generations. The reason a browser bug scores this high on a BMC is that the victim is an operator with an authenticated iLO session: script running in that session can drive the same actions the operator can - mount Virtual Media, change boot order, power-cycle, create accounts - from the operator's own browser and credentials. It converts a phishing link into out-of-band control of a node without the attacker ever needing network reachability to the iLO themselves.","attack_vector":"Requires tricking an authenticated operator into loading attacker-controlled content while they hold an iLO session. The attacker does not need to reach the management VLAN at all - the operator's browser is the bridge, which is precisely why 'the BMCs are on an isolated network' is not a complete answer.","remediation":"Flash iLO 4 to v2.82, iLO 5 to v2.78, or iLO 6 to v1.20 or later. Out-of-band, per-node, no host reboot and no job drain. Operational controls that help independently: use a dedicated browser profile or a privileged access workstation for BMC administration, and do not leave iLO sessions open in a browser that also handles general web traffic and email.","references":["https://support.hpe.com/hpesc/public/docDisplay?docLocale=en_US&docId=hpesbhf04456en_us","https://nvd.nist.gov/vuln/detail/CVE-2023-28083"],"status":"curated","published":"2023-03-22"},{"id":"CVE-2023-31009","cve":"CVE-2023-31009","aliases":[],"title":"DGX H100 BMC (REST): Code execution + privesc","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX H100 BMC (REST)","year":"2023","cvss_score":8.3,"severity":"high","kev":false,"impact":"Code execution + privesc","attack_vector":"Network-adjacent BMC REST client","remediation":"Flash BMC 23.08.18 out-of-band","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5473/5473.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:N/UI:N/S:U/C:L/I:H/A:H","cwe":["CWE-20"],"published":"2023-09-20"},{"id":"CVE-2023-37294","cve":"CVE-2023-37294","aliases":["AMI-SA-2023010"],"title":"AMI MegaRAC SPx 12 / SPx 13 (BMC network service): Heap corruption in the BMC reachable without credentials","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx 12 / SPx 13 (BMC network service)","year":"2023","cvss_score":8.3,"severity":"high","kev":false,"impact":"Heap corruption in the BMC reachable without credentials. A reliable exploit gives BMC code execution and therefore persistent, below-the-OS control of the node; a sloppy one just crashes the BMC, which on a GPU node means losing remote power control and console right when you need it - the node keeps running the training job but becomes un-manageable until someone walks the row.","attack_vector":"Adjacent network, unauthenticated, but high attack complexity - the attacker needs heap grooming or a race to win, so this is a targeted-effort bug rather than a spray. Precondition is still just L2 reachability to the BMC NIC.","remediation":"Firmware flash to SPx_12.7 / SPx_13.6, out-of-band and per node, subject to ODM rebase. Same rollout cost as the rest of the AMI-SA-2023010 batch, so treat all eight CVEs in that advisory as one flash campaign rather than eight tickets. Interim control is network segmentation of the BMC plane.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2023010.pdf","https://nvd.nist.gov/vuln/detail/CVE-2023-37294"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-01-09"},{"id":"CVE-2023-37295","cve":"CVE-2023-37295","aliases":[],"title":"AMI MegaRAC SPx (BMC heap memory corruption): Further unauthenticated heap corruption in the MegaRAC BMC reachable","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx (BMC heap memory corruption)","year":"2023","cvss_score":8.3,"severity":"high","kev":false,"impact":"Further unauthenticated heap corruption in the MegaRAC BMC reachable from an adjacent network.","attack_vector":"Adjacent-network access to the BMC.","remediation":"Obtain and flash updated BMC firmware from your board OEM.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2023010.pdf"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-37296","cve":"CVE-2023-37296","aliases":["AMI-SA-2023010"],"title":"AMI MegaRAC SPx 12 / SPx 13 (BMC network service): Stack memory corruption in the same unauthenticated BMC parsing","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx 12 / SPx 13 (BMC network service)","year":"2023","cvss_score":8.3,"severity":"high","kev":false,"impact":"Stack memory corruption in the same unauthenticated BMC parsing surface. Best case for the attacker is BMC code execution and a firmware-level foothold on the node; worst case for the operator without an exploit is a BMC that wedges and needs a physical AC cycle to recover, which on a dense GPU rack means a hands-on trip and possibly draining neighbouring nodes.","attack_vector":"Adjacent network, no credentials, high complexity. Anything on the management VLAN - including a compromised BMC on a neighbouring node - is close enough.","remediation":"Firmware flash to SPx_12.7 / SPx_13.6. Out-of-band, per node, ODM-gated. No config-only fix inside the BMC; the compensating control is to make the management VLAN unreachable from tenant and general corporate networks and to keep the BMC off any routable address space.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2023010.pdf","https://nvd.nist.gov/vuln/detail/CVE-2023-37296"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-01-09"},{"id":"CVE-2023-37297","cve":"CVE-2023-37297","aliases":[],"title":"AMI MegaRAC SPx (BMC heap memory corruption): Heap corruption in the BMC reachable from an adjacent network","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx (BMC heap memory corruption)","year":"2023","cvss_score":8.3,"severity":"high","kev":false,"impact":"Heap corruption in the BMC reachable from an adjacent network without authentication, with scope change.","attack_vector":"Adjacent-network access to the BMC.","remediation":"Obtain and flash updated BMC firmware from your board OEM.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2023010.pdf"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-39266","cve":"CVE-2023-39266","aliases":[],"title":"ArubaOS-Switch web management interface: Unauthenticated stored cross-site scripting against the ArubaOS-Switch web UI","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ArubaOS-Switch web management interface","year":"2023","cvss_score":8.3,"severity":"high","kev":false,"impact":"Unauthenticated stored cross-site scripting against the ArubaOS-Switch web UI. Stored XSS in a switch management interface is a credential-theft and config-change path aimed at your network operators: an attacker plants the payload without logging in, and it fires the next time an admin opens the page.","attack_vector":"Unauthenticated, remote to the switch's web management interface; the payload executes in an administrator's browser session.","remediation":"ArubaOS-Switch firmware upgrade plus reload. Immediate mitigation is a config change: disable the web management interface and manage via SSH/CLI, which most datacenter operators should be doing anyway. Related ArubaOS issue: CVE-2023-35971.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-39266"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-08-29"},{"id":"CVE-2023-40284","cve":"CVE-2023-40284","aliases":[],"title":"Supermicro BMC (IPMI web interface, XSS): Stored/reflected script injection in the BMC web UI","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC (IPMI web interface, XSS)","year":"2023","cvss_score":8.3,"severity":"high","kev":false,"impact":"Stored/reflected script injection in the BMC web UI. On its own it is 'just XSS', but on a BMC the session it hijacks is the one that can mount virtual media, power-cycle the node and flash firmware - so it is the entry step of a full out-of-band takeover chain rather than a cosmetic web bug.","attack_vector":"Requires an operator to load an attacker-influenced BMC page. Any engineer who administers BMCs from a browser is the target, and the attacker only needs to be able to plant content the BMC will render.","remediation":"BMC firmware flash per board. Until then, treat BMC web access as a privileged action: dedicated browser profile or jump host, never the same browser session used for general web browsing, and no BMC on a routable network.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-40284"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-03-27"},{"id":"CVE-2023-40287","cve":"CVE-2023-40287","aliases":[],"title":"Supermicro BMC (IPMI web interface, XSS): Script injection in the BMC management UI, scope-changing because","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC (IPMI web interface, XSS)","year":"2023","cvss_score":8.3,"severity":"high","kev":false,"impact":"Script injection in the BMC management UI, scope-changing because the compromised session controls power, console and firmware on the physical node.","attack_vector":"Network reach to the BMC web interface plus operator interaction.","remediation":"BMC firmware flash per board. Same rollout as the rest of the batch; the interim control is network isolation of the BMC plane, not browser hygiene alone.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-40287"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-03-27"},{"id":"CVE-2023-40288","cve":"CVE-2023-40288","aliases":[],"title":"Supermicro BMC (IPMI web interface, XSS): Further injection point in the same BMC web stack","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC (IPMI web interface, XSS)","year":"2023","cvss_score":8.3,"severity":"high","kev":false,"impact":"Further injection point in the same BMC web stack; the payoff is a hijacked administrative session with out-of-band control of the server.","attack_vector":"Network reach to the BMC web interface plus operator interaction.","remediation":"BMC firmware flash per board, bundled with the batch.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-40288"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-03-27"},{"id":"CVE-2023-40290","cve":"CVE-2023-40290","aliases":[],"title":"Supermicro BMC (IPMI web interface, XSS via IE11): Injection that fires specifically through Internet Explorer 11","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC (IPMI web interface, XSS via IE11)","year":"2023","cvss_score":8.3,"severity":"high","kev":false,"impact":"Injection that fires specifically through Internet Explorer 11. Worth keeping on the list precisely because BMC web UIs are the last place in a datacenter where an ancient browser is still in the loop - legacy Java KVM clients and jump hosts frozen on old images keep IE alive long after everything else moved on.","attack_vector":"An operator administering the BMC from IE11 on Windows, with network reach to the BMC.","remediation":"BMC firmware flash per board. The cheaper immediate action is retiring IE11 from your management jump hosts entirely, which also kills a class of legacy KVM client exposure at the same time.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-40290"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-03-27"},{"id":"CVE-2023-40547","cve":"CVE-2023-40547","aliases":[],"title":"shim (HTTP boot): Out-of-bounds write from a crafted HTTP response during network boot","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"shim (HTTP boot)","year":"2023","cvss_score":8.3,"severity":"high","kev":false,"impact":"Out-of-bounds write from a crafted HTTP response during network boot — full system compromise at the pre-boot stage, before any OS security control exists","attack_vector":"Adjacent network, MITM on the boot server or a compromised PXE/HTTP boot server","remediation":"New signed shim rollout plus dbx revocation of the old one. Directly relevant to neoclouds because netboot provisioning is the standard bare-metal reimaging path — the provisioning network is the attack surface","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-40547"],"status":"curated","published":"2024-01-25"},{"id":"CVE-2023-45230","cve":"CVE-2023-45230","aliases":["PixieFail","AMI-SA-2024001"],"title":"AMI AptioV UEFI BIOS (EDK II network stack, DHCPv6 client): Buffer overflow in the firmware's DHCPv6 client, triggered","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI AptioV UEFI BIOS (EDK II network stack, DHCPv6 client)","year":"2023","cvss_score":8.3,"severity":"high","kev":false,"impact":"Buffer overflow in the firmware's DHCPv6 client, triggered by an over-long server ID option. The attacker gets code execution in the pre-boot firmware environment - before the OS, before Secure Boot has finished mattering, with full access to the platform. For a GPU fleet the sharp edge is network provisioning: PXE and HTTP boot are how most clusters bring nodes up, and the firmware talks IPv6 DHCP during that window whether or not you intended to use IPv6. An attacker on the provisioning VLAN answers first and owns the node before it has an operating system.","attack_vector":"Adjacent network, unauthenticated, no interaction - the attacker just needs to be on the same segment as the booting node and respond to its DHCPv6 solicitation faster than your real server. Every node reboot is a fresh opportunity, so on a fleet that autoscales or recovers nodes constantly the window is effectively always open. Note that IPv6 is exercised even in IPv4-only deployments because the firmware stack solicits regardless.","remediation":"BIOS update carrying the patched EDK II network package - firmware flash plus host reboot per node, gated on your server vendor rebasing AMI's AptioV build; AMI's advisory names 'AptioV' generically rather than a BKC version, so confirm the specific BIOS release with your OEM. There is a real config-only mitigation and you should apply it regardless: disable PXE and network boot in BIOS on nodes that do not need it, and where you do need it, isolate the provisioning VLAN so nothing untrusted can answer DHCP on it. That is a BIOS setup change plus one reboot per node, far cheaper than the flash.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/2024/AMI-SA-2024001.pdf","https://github.com/tianocore/edk2/security/advisories/GHSA-hc6x-cw6p-gj7h","https://nvd.nist.gov/vuln/detail/CVE-2023-45230"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-01-16"},{"id":"CVE-2024-22424","cve":"CVE-2024-22424","aliases":[],"title":"Argo CD: CSRF against the Argo CD API allows deploying arbitrary workloads with Argo's cluster-admin rights","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2024","cvss_score":8.3,"severity":"high","kev":false,"impact":"CSRF against the Argo CD API allows deploying arbitrary workloads with Argo's cluster-admin rights","attack_vector":"An attacker who gets a logged-in operator to visit a hostile page","remediation":"Rolling Argo CD upgrade; no GPU drain. Also scope Argo's cluster credentials down","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-22424"],"status":"curated","published":"2024-01-19"},{"id":"CVE-2024-47084","cve":"CVE-2024-47084","aliases":[],"title":"Gradio: CORS origin validation bypass","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Gradio","year":"2024","cvss_score":8.3,"severity":"high","kev":false,"impact":"CORS origin validation bypass → cross-origin access to the demo API","attack_vector":"Malicious page visited by the demo user","remediation":"Upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-47084"],"status":"curated","published":"2024-10-10"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:C/C:H/I:H/A:H","cwe":["CWE-250","CWE-1220"],"fleet":{"pain_class":"hot-patch"},"id":"CVE-2024-52799","cve":"CVE-2024-52799","aliases":["GHSA-fgrf-2886-4q7m"],"title":"Argo Workflows Helm chart (argo-helm, workflow-role RBAC): The chart's workflow-role grants create on pods/exec to","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo Workflows Helm chart (argo-helm, workflow-role RBAC)","year":"2024","cvss_score":8.3,"severity":"high","kev":false,"impact":"The chart's workflow-role grants create on pods/exec to every workflow pod, which since Argo 3.4 and the Emissary executor is no longer needed. Any workflow can therefore exec into any other pod in the same namespace and run commands there - a tenant who gets a colleague to run a malicious template owns the whole namespace.","attack_vector":"A user who can get a workflow template executed in the namespace. Requires the argo-workflows Helm chart below 0.44.0 with appVersion 3.4 or above; upstream plain manifests are not affected.","remediation":"Upgrade the argo-workflows Helm chart to 0.44.0 or later and apply it. This is an RBAC Role change, so it takes effect on apply without restarting the controller. If you templated the chart into your own manifests, strip pods/exec from workflow-role directly.","references":["https://github.com/argoproj/argo-helm/security/advisories/GHSA-fgrf-2886-4q7m","https://nvd.nist.gov/vuln/detail/CVE-2024-52799"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:4.0/AV:N/AC:H/AT:N/PR:N/UI:N/VC:N/VI:L/VA:H/SC:N/SI:N/SA:N","cwe":["CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-0052","cve":"CVE-2025-0052","aliases":[],"title":"Pure Storage FlashBlade authentication input validation: The FlashBlade equivalent of the FlashArray pre-authentication","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Pure Storage FlashBlade authentication input validation","year":"2025","cvss_score":8.3,"severity":"high","kev":false,"impact":"The FlashBlade equivalent of the FlashArray pre-authentication denial of service: malformed authentication input stops the array serving data. For a fleet using FlashBlade as the shared training filesystem, that is a full stall.","attack_vector":"Network reach to the FlashBlade authentication surface, no credentials needed. Attack complexity is rated high, so it is not trivially repeatable, but it needs no account.","remediation":"Upgrade Purity//FB to the fixed release in Pure's security bulletin, and restrict network exposure of the login endpoints in the meantime.","references":["https://support.purestorage.com/bundle/m_security_bulletins/page/Pure_Security/topics/concept/c_security_bulletins.html","https://nvd.nist.gov/vuln/detail/CVE-2025-0052"],"status":"curated"},{"id":"CVE-2025-23359","cve":"CVE-2025-23359","aliases":[],"title":"Container Toolkit: Container escape to host filesystem (bypass of the CVE-2024-0132 fix)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Container Toolkit","year":"2025","cvss_score":8.3,"severity":"high","kev":false,"impact":"Container escape to host filesystem (bypass of the CVE-2024-0132 fix)","attack_vector":"Any tenant that can run an arbitrary container image","remediation":"Bump nvidia-container-toolkit to 1.17.4+ and restart the runtime on every GPU node; upgrade GPU Operator to 24.9.2+; evict tenant workloads","references":["https://github.com/NVIDIA/product-security/tree/main/2025/5616","https://nvd.nist.gov/vuln/detail/CVE-2025-23359"],"status":"curated","fleet":{"ubiquity":"Universal - all versions <= 1.17.3, i.e. everyone who thought they had already patched 0132","remediation_pain":"`daemon-restart` to 1.17.4 / GPU Operator 24.9.2 - a second forced patch cycle across the same fleet within five months","pain_class":"daemon-restart","why_fleet_wide":"Proves the class is not one-and-done: a mount-path TOCTOU bypass re-opens host filesystem access from a crafted container, so every operator who patched in Sept 2024 had to re-patch the whole fleet in Feb 2025"},"cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:R/S:C/C:H/I:H/A:H","cwe":["CWE-367"],"published":"2025-02-12"},{"id":"CVE-2025-26336","cve":"CVE-2025-26336","aliases":[],"title":"Dell Chassis Management Controller (PowerEdge FX2 / VRTX): Unauthenticated remote attacker overflows a stack buffer","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Dell Chassis Management Controller (PowerEdge FX2 / VRTX)","year":"2025","cvss_score":8.3,"severity":"high","kev":false,"impact":"Unauthenticated remote attacker overflows a stack buffer in the CMC and gains control of the chassis manager - the component that owns power, identity and console for every sled in the enclosure.","attack_vector":"Network access to the CMC management interface, no credentials required.","remediation":"Update CMC firmware to 2.40.200.202101130302 (FX2) or 3.41.200.202209300499 (VRTX). Firmware flash on the chassis controller; sleds keep running but management is interrupted. If these chassis are reachable from anything but a locked-down OOB VLAN, fix that first - it is the actual exposure.","references":["https://www.dell.com/support/kbdoc/en-us/000297463/dsa-2025-123-security-update-for-dell-chassis-management-controller-firmware-for-dell-poweredge-fx2-and-vrtx-vulnerabilities"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-67601","cve":"CVE-2025-67601","aliases":[],"title":"Rancher: CLI login with -skip-verify and no --cacert silently accepts any certificate","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Rancher","year":"2025","cvss_score":8.3,"severity":"high","kev":false,"impact":"CLI login with -skip-verify and no --cacert silently accepts any certificate; MITM on the management API","attack_vector":"Unauthenticated network in a MITM position","remediation":"Upgrade Rancher CLI; ban -skip-verify in runbooks","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-67601"],"status":"curated","published":"2026-02-25"},{"id":"CVE-2026-20751","cve":"CVE-2026-20751","aliases":[],"title":"Intel Data Center GPU driver for VMware ESXi (out-of-bounds read): Out-of-bounds read in the ESXi GPU driver exposing","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel Data Center GPU driver for VMware ESXi (out-of-bounds read)","year":"2026","cvss_score":8.3,"severity":"high","kev":false,"impact":"Out-of-bounds read in the ESXi GPU driver exposing data and enabling denial of service.","attack_vector":"Local privileged access on the ESXi host.","remediation":"Update the Intel Data Center Graphics Driver for ESXi to 2.0.2 or later; VIB update plus host reboot.","references":["https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01402.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-20879","cve":"CVE-2026-20879","aliases":[],"title":"Intel Data Center GPU driver for VMware ESXi (out-of-bounds write): Out-of-bounds write in the ESXi GPU driver causing","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel Data Center GPU driver for VMware ESXi (out-of-bounds write)","year":"2026","cvss_score":8.3,"severity":"high","kev":false,"impact":"Out-of-bounds write in the ESXi GPU driver causing data corruption and denial of service on the hypervisor host.","attack_vector":"Local privileged access on the ESXi host.","remediation":"Update the Intel Data Center Graphics Driver for ESXi to 2.0.2 or later; VIB update plus host reboot.","references":["https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01402.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-22621","cve":"CVE-2026-22621","aliases":["eaton-va-2026-1005"],"title":"Eaton Tripp Lite series PADM firmware, session management interface: An authenticated administrator can break out","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Eaton Tripp Lite series PADM firmware, session management interface","year":"2026","cvss_score":8.3,"severity":"high","kev":false,"impact":"An authenticated administrator can break out of the restricted shell and run arbitrary commands on the PDU. That turns a device you thought was an appliance into a persistent Linux foothold sitting on your out-of-band network, below every server it powers and outside any endpoint tooling you run.","attack_vector":"Requires administrator credentials on the PDU - which, given the companion authentication bypass, an unauthenticated attacker can obtain first. Chain the two and this is unauthenticated remote code execution on rack power infrastructure.","remediation":"PADM firmware update, or hardware replacement for EOL SKUs. Rotate PDU admin credentials, which are very commonly shared fleet-wide from the original commissioning.","references":["https://www.eaton.com/content/dam/eaton/company/news-insights/cybersecurity/security-bulletins/eaton-va-2026-1005.pdf"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2026-07-30"},{"id":"CVE-2026-24148","cve":"CVE-2026-24148","aliases":[],"title":"Jetson Xavier / Orin: Authentication bypass in a network service","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Jetson Xavier / Orin","year":"2026","cvss_score":8.3,"severity":"high","kev":false,"impact":"Authentication bypass in a network service","attack_vector":"Network-adjacent unauthenticated","remediation":"Flash JetPack; edge fleet","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24148","https://github.com/NVIDIA/product-security/tree/main/2026/5797"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:L","cwe":["CWE-1188"],"published":"2026-03-31"},{"id":"CVE-2026-41490","cve":"CVE-2026-41490","aliases":[],"title":"Dagster: Vulnerability in Dagster Core prior to 1.13.1","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Dagster","year":"2026","cvss_score":8.3,"severity":"high","kev":false,"impact":"Vulnerability in Dagster Core prior to 1.13.1","attack_vector":"Network user of the orchestrator","remediation":"Upgrade to 1.13.1+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-41490"],"status":"curated","published":"2026-05-07"},{"id":"CVE-2026-54100","cve":"CVE-2026-54100","aliases":[],"title":"Red Hat OpenShift Windows Machine Config Operator (unverified SSH host key): WMCO opens SSH to Windows worker nodes","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Red Hat OpenShift Windows Machine Config Operator (unverified SSH host key)","year":"2026","cvss_score":8.3,"severity":"high","kev":false,"impact":"WMCO opens SSH to Windows worker nodes without verifying the host key, so an adjacent-network attacker intercepting the session captures WICD and kubelet bootstrap credentials - which is cluster-node identity, with scope change.","attack_vector":"Adjacent network position able to intercept or redirect WMCO's SSH session to a Windows node.","remediation":"Apply RHSA-2026:47173. Operator update; afterwards rotate kubelet bootstrap credentials for Windows nodes, because captured bootstrap tokens remain valid until rotated.","references":["https://access.redhat.com/errata/RHSA-2026:47173"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-264","CWE-863"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2011-1898","cve":"CVE-2011-1898","aliases":["XSA-3"],"title":"Xen PCI passthrough on Intel VT-d chipsets without interrupt remapping: The founding GPU-passthrough escape. A guest","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen PCI passthrough on Intel VT-d chipsets without interrupt remapping","year":"2011","cvss_score":8.2,"severity":"high","kev":false,"impact":"The founding GPU-passthrough escape. A guest that owns a passed-through device programs it to DMA into the platform's interrupt-injection registers, synthesising MSIs it was never allowed to raise, and from there takes the hypervisor. The IOMMU alone does not stop this - DMA remapping constrains where a device may write data, but without interrupt remapping the MSI address range is still reachable, and an MSI is just a memory write. Directly applicable to any rented-GPU product: hand a tenant a GPU on a platform without IR (or with IR disabled) and you have handed them the host. This is the concrete instance of the Invisible Things Lab result that IOMMU-without-IR is not an isolation boundary.","attack_vector":"A guest administrator - i.e. the tenant renting the VM - who has been assigned any bus-mastering-capable PCI device. A GPU qualifies.","remediation":"Upgrade to Xen 4.1.1 / 4.0.2 or later, which refuses passthrough when interrupt remapping is unavailable, and confirm IR is actually enabled on every host: check the DMAR/IVRS tables and that the hypervisor did not silently fall back after a firmware quirk. The expensive part is not the hypervisor patch but the platform audit - some older boards report IR capability and have it errata-disabled, and the only safe response there is to stop selling passthrough on that SKU and retire it. Requires a host reboot per node with all tenants evacuated.","references":["https://xenbits.xen.org/xsa/advisory-3.html","https://invisiblethingslab.com/resources/2011/Software%20Attacks%20on%20Intel%20VT-d.pdf","https://nvd.nist.gov/vuln/detail/CVE-2011-1898"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-264","CWE-862"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2013-4329","cve":"CVE-2013-4329","aliases":["XSA-61"],"title":"Xen libxl (xenlight) PCI passthrough device setup: The toolstack hands a bus-mastering-capable PCI device to an HVM","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen libxl (xenlight) PCI passthrough device setup","year":"2013","cvss_score":8.2,"severity":"high","kev":false,"impact":"The toolstack hands a bus-mastering-capable PCI device to an HVM guest before IOMMU setup for that device has completed - and when the IOMMU is disabled outright, hands it over anyway. For the duration of that window the device DMAs against raw host physical addresses with nothing in the way, so a tenant whose guest driver is ready at attach time reads or writes anywhere in host memory and escalates. This is the textbook 'DMA before the IOMMU is configured' race, and it is exactly the class of flaw that makes GPU attach/detach on a busy multi-tenant node dangerous: the risky window opens on every VM start, not once at boot.","attack_vector":"A tenant's HVM guest with any passed-through bus-mastering device, exploiting the interval between device visibility and IOMMU programming.","remediation":"Apply the Xen 4.0-4.2 XSA-61 patches or move to a fixed release, and - independently of the patch - make it a hard platform invariant that passthrough is refused when the IOMMU is not active, rather than a configuration the operator can turn off. Fixing this needs a toolstack update plus a host restart of the affected domains; the durable control is verifying on every host boot that DMA remapping came up, and failing the node out of the scheduler if it did not, because a silently IOMMU-less host in a GPU fleet will otherwise keep accepting tenants.","references":["https://xenbits.xen.org/xsa/advisory-61.html","https://www.openwall.com/lists/oss-security/2013/09/10/4","https://nvd.nist.gov/vuln/detail/CVE-2013-4329"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-264"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2013-6375","cve":"CVE-2013-6375","aliases":["XSA-78"],"title":"Xen Intel VT-d IOMMU page-table handling for PCI passthrough: An inverted boolean means Xen clears a present IOMMU","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen Intel VT-d IOMMU page-table handling for PCI passthrough","year":"2013","cvss_score":8.2,"severity":"high","kev":false,"impact":"An inverted boolean means Xen clears a present IOMMU translation entry without flushing the IOMMU TLB. The device keeps translating through the stale cached mapping, so a passed-through GPU continues DMA-ing into host or previous-tenant memory after the mapping that authorised it has been revoked. This is the failure mode operators most underestimate: the IOMMU page tables look correct in memory while the hardware is still using an older view of them. In a cluster that recycles GPUs between tenants, stale IOMMU cache entries are exactly how tenant N reaches tenant N-1's pages.","attack_vector":"A guest administrator with a passed-through PCI device, arranging for mappings to be torn down while its device continues issuing DMA.","remediation":"Patch Xen 4.2.x/4.3.x per XSA-78 and reboot the hypervisor - IOMMU flush logic is not something that can be corrected at runtime. Pair the patch with the companion issue CVE-2013-6400 (XSA-80), which leaves the flush-suppression flag set on an error path and produces the same stale-mapping condition by a different route; patching one without the other leaves the hole open. Requires evacuating every tenant on the node.","references":["https://xenbits.xen.org/xsa/advisory-78.html","https://xenbits.xen.org/xsa/advisory-80.html","https://nvd.nist.gov/vuln/detail/CVE-2013-6375"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-312","CWE-522"],"fleet":{"pain_class":"firmware-flash"},"id":"CVE-2014-0860","cve":"CVE-2014-0860","aliases":[],"title":"IBM BladeCenter AMM (before 3.66E), IMM (before 1.43), IMM2: The in-band host-to-BMC pivot, with a CVE attached. The","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"IBM BladeCenter AMM (before 3.66E), IMM (before 1.43), IMM2","year":"2014","cvss_score":8.2,"severity":"high","kev":false,"impact":"The in-band host-to-BMC pivot, with a CVE attached. The management firmware stores IPMI credentials in cleartext and exposes them over two channels a tenant may already stand on - the chassis internal network, and the Ethernet-over-USB link that the host operating system sees as an ordinary network interface. A tenant who reaches root on the host OS of one blade reads those credentials off the in-band interface and then issues IPMI commands and opens remote-control sessions against blades belonging to other tenants. That is the whole out-of-band trust model inverted: the management plane was supposed to be the thing the host could not reach, and here the host reads its way in. Same primitive class as KCS host-to-BMC bridging, which has no CVE of its own because it is documented behaviour rather than a bug.","attack_vector":"Local root on the host OS of an affected blade (via the Ethernet-over-USB interface), or a position on the chassis internal network. Escalates from one tenant's node to chassis-wide control.","remediation":"Flash AMM to 3.66E+, IMM to 1.43+, IMM2 to 4.15+. Beyond the flash, the lasting control is to sever the in-band path: disable the Ethernet-over-USB / LAN-over-USB interface on hosts that do not need in-band management, and where in-band agents are required, treat the host-to-BMC channel as an untrusted boundary that needs its own authentication rather than an implicitly trusted sidechannel. Operators should assume any tenant with host root can talk to the BMC unless that interface has been explicitly turned off, and audit for it - it is enabled by default on a lot of hardware.","references":["https://exchange.xforce.ibmcloud.com/vulnerabilities/90880","https://www.cve.org/CVERecord?id=CVE-2014-0860","https://nvd.nist.gov/vuln/detail/CVE-2014-0860"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-863"],"fleet":{"pain_class":"node-drain"},"id":"CVE-2015-4106","cve":"CVE-2015-4106","aliases":["XSA-131"],"title":"QEMU xen_pt PCI passthrough config-space mediation (Xen 3.3.x-4.5.x): The device model failed to mediate guest writes","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"QEMU xen_pt PCI passthrough config-space mediation (Xen 3.3.x-4.5.x)","year":"2015","cvss_score":8.2,"severity":"high","kev":false,"impact":"The device model failed to mediate guest writes to PCI configuration space on passed-through devices, so a tenant reprograms registers the hypervisor believed it controlled - BARs, bus-mastering enables, capability structures. The advisory is explicit that privilege escalation, host crash and information leak all cannot be excluded. For a GPU rental business this is the fundamental mediation failure: config space is where the device's memory windows and DMA rights are declared, and letting a tenant edit it lets them redraw the map the IOMMU and the host were relying on.","attack_vector":"Guest administrator with an assigned PCI device, writing to its own device's PCI configuration space.","remediation":"Apply XSA-131 and restart the device models - a node drain, because guests with assigned devices cannot be live-migrated off. Apply alongside XSA-126/CVE-2015-2756 (command-register access) and XSA-128/CVE-2015-4103 (MSI message data), which are the same mediation failure reached through different registers; they were issued separately but an operator should treat them as one change window.","references":["https://xenbits.xen.org/xsa/advisory-131.html","https://xenbits.xen.org/xsa/advisory-126.html","https://nvd.nist.gov/vuln/detail/CVE-2015-4106"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2018-3643","cve":"CVE-2018-3643","aliases":["INTEL-SA-00131"],"title":"Power Management Controller (PMC) firmware in systems using Intel CSME 11.x/12.0 or Intel SPS 4.x: An administrative","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Power Management Controller (PMC) firmware in systems using Intel CSME 11.x/12.0 or Intel SPS 4.x","year":"2018","cvss_score":8.2,"severity":"high","kev":false,"impact":"An administrative attacker can reach the platform's Power Management Controller firmware. The PMC owns the platform power state machine and rails - so this is one of the few software-reachable paths with a direct physical outcome: forced or blocked power transitions on a node, and manipulation of the power/thermal control loop that the rest of the platform trusts. On GPU nodes drawing 6-10 kW, an attacker who can hold a node in the wrong power state or misreport its budget to Node Manager can trip breakers at the rack or hall level rather than merely killing one job. The PMC firmware also lives outside the host OS image, so a modification persists across reimage and tenant handoff.","attack_vector":"An attacker with administrative privileges on the platform - local root on the host reaching CSME/SPS via HECI, or an administrator on the management path.","remediation":"CSME/SPS firmware bundle including the PMC firmware, delivered by the OEM (Dell, HPE, Supermicro, Lenovo, Gigabyte, Quanta) as a BIOS/ME package. Host reboot and job drain required; HPE's bundle for this advisory came out well after Intel's September 2018 date, which is the normal pattern. There is no runtime mitigation - PMC firmware cannot be disabled. Operationally, pair the rollout with independent power telemetry at the PDU/branch level so you are not relying solely on platform-reported power for capacity protection.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-3643","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00131.html","https://support.hpe.com/hpsc/doc/public/display?docLocale=en_US&docId=emr_na-hpesbhf03873en_us"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2018-3682","cve":"CVE-2018-3682","aliases":["INTEL-SA-00130"],"title":"BMC firmware on Intel server boards, compute modules and systems - SMBus access control: An attacker","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"BMC firmware on Intel server boards, compute modules and systems - SMBus access control","year":"2018","cvss_score":8.2,"severity":"high","kev":false,"impact":"An attacker with administrative privileges on the BMC can issue unauthorized reads and writes on the platform SMBus. SMBus is the wire that reaches the power supplies (PMBus), the voltage regulators, the DIMM SPD EEPROMs and the temperature sensors. Write access there is a physical-consequence primitive: reprogram a VR or PSU setpoint, corrupt SPD so DIMMs no longer train, or falsify thermal telemetry so the platform does not throttle. On a dense GPU node this can mean a forced power-off, a bricked-until-RMA board, or a thermal event that the DCIM layer never sees coming. It is also persistence - SMBus-attached EEPROM contents survive any host reimage and therefore cross tenant handoff.","attack_vector":"Administrative access to the BMC. That is reached from the out-of-band management network, from any credential reuse across the IPMI/Redfish fleet, or from the host itself via the KCS/host interface if you have not disabled it - which means a tenant with root on a bare-metal node is one BMC bug away from the SMBus.","remediation":"BMC firmware update from Intel or the board OEM (Intel server boards, and the ODMs building on them - Quanta, Wiwynn, Supermicro). BMC flashes generally do not require a host reboot, which makes this one of the cheaper firmware rollouts, but the update must be staged per board family. The structural controls matter more: put the BMC on a network no tenant can reach, use unique per-node BMC credentials, and disable the host-to-BMC KCS/host interface on bare-metal SKUs so a tenant with root cannot talk to the BMC at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-3682","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00130.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2019-11248","cve":"CVE-2019-11248","aliases":[],"title":"Kubernetes (kubelet): /debug/pprof exposed on the unauthenticated kubelet healthz port","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kubelet)","year":"2019","cvss_score":8.2,"severity":"high","kev":false,"impact":"/debug/pprof exposed on the unauthenticated kubelet healthz port; leaks node and workload internals","attack_vector":"Any pod on the cluster network reaching the node's healthz port","remediation":"Rolling kubelet upgrade with node drain; firewall the healthz port to the control plane","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-11248"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2019-08-29"},{"id":"CVE-2020-10713","cve":"CVE-2020-10713","aliases":["BootHole"],"title":"GRUB2: Buffer overflow in `grub.cfg` parsing allowing Secure Boot bypass and arbitrary code execution inside GRUB","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2","year":"2020","cvss_score":8.2,"severity":"high","kev":false,"impact":"Buffer overflow in `grub.cfg` parsing allowing Secure Boot bypass and arbitrary code execution inside GRUB — a bootkit that persists across OS reinstall","attack_vector":"Local, or via a modified PXE-boot network","remediation":"dbx (Secure Boot revocation list) push plus a coordinated GRUB/shim/kernel update. The dbx push is the dangerous part: revoking the old shim before every node has the new bootloader leaves the node unbootable, and recovery is out-of-band console work per node","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-10713"],"status":"curated","fleet":{"ubiquity":"universal - GRUB2 + the Microsoft-signed shim is the boot path for nearly every Linux GPU node","remediation_pain":"node-reboot + firmware-flash-class pain: patching GRUB is easy, but the real fix is a **dbx revocation** update pushed into UEFI NVRAM on every node, and a botched dbx push makes the node unbootable - which is why operators delay it for years","pain_class":"firmware-flash","why_fleet_wide":"A config-file buffer overflow in GRUB2 lets attackers run pre-OS bootkits with Secure Boot enabled; because every old signed GRUB stays valid until revoked, the fleet remains exploitable until each node's firmware revocation list is updated."},"published":"2020-07-30"},{"id":"CVE-2020-15705","cve":"CVE-2020-15705","aliases":["BootHole family"],"title":"GRUB2 (direct kernel boot without shim): When GRUB is booted directly by UEFI rather than chained through shim, it does","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (direct kernel boot without shim)","year":"2020","cvss_score":8.2,"severity":"high","kev":false,"impact":"When GRUB is booted directly by UEFI rather than chained through shim, it does not verify the kernel signature at all. Any unsigned kernel boots with Secure Boot enabled and reporting healthy - so attestation and the operator's 'verified boot' control are simply false on those nodes. Confidential-computing claims built on measured boot become unverifiable.","attack_vector":"Applies to any node configured to load GRUB directly from the EFI System Partition. The attacker then only needs to drop a kernel, which any local root can do.","remediation":"grub2 package update + reboot, and audit the boot configuration on every node to confirm shim is actually in the chain - a surprising number of custom/netboot images skip it. This is one of the few in this family where a config check is a genuine part of the fix, not just a workaround.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-15705","https://ubuntu.com/security/CVE-2020-15705"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2020-07-29"},{"id":"CVE-2020-3165","cve":"CVE-2020-3165","aliases":[],"title":"Cisco NX-OS (BGP MD5 authentication): BGP MD5 authentication can be bypassed, so an attacker can bring up a BGP session","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco NX-OS (BGP MD5 authentication)","year":"2020","cvss_score":8.2,"severity":"high","kev":false,"impact":"BGP MD5 authentication can be bypassed, so an attacker can bring up a BGP session with the switch without the shared key. In an EVPN fabric that means injecting routes — including type-2 and type-5 EVPN routes — which is exactly how you steer one tenant's traffic to a machine you control. The authentication you configured to prevent unauthorized peering simply does not hold.","attack_vector":"Unauthenticated, remote — an attacker that can reach TCP/179 on the switch and is permitted by the peer-group/neighbor configuration's address range.","remediation":"NX-OS upgrade plus reload. Interim mitigation is control-plane policing plus tight neighbor prefix ACLs so only known peer addresses can open a session at all — live config changes. If you run EVPN, also verify no unexpected routes were learned before you patched.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-3165"],"status":"curated","tags":["tenant-isolation"],"published":"2020-02-26"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:L/A:N","cwe":["CWE-200"],"fleet":{"pain_class":"node-drain"},"id":"CVE-2020-4927","cve":"CVE-2020-4927","aliases":[],"title":"IBM Spectrum Scale / Storage Scale core daemon (cluster RPC transport): An attacker who can speak to the cluster's","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"IBM Spectrum Scale / Storage Scale core daemon (cluster RPC transport)","year":"2020","cvss_score":8.2,"severity":"high","kev":false,"impact":"An attacker who can speak to the cluster's internal RPC transport reads user data out of the filesystem and can inject data into the protocol stream, with no credentials. On a shared training cluster this is a direct read of another tenant's datasets and checkpoints, and a path to corrupting them.","attack_vector":"Network reachability to the Storage Scale core daemon ports on any cluster node. No account and no filesystem mount is required, so anything that lands on the storage VLAN - a compromised compute node, a mis-scoped tenant network, a jump host - is enough.","remediation":"Upgrade Storage Scale to 5.1.6.2 or later per IBM's bulletin, then confirm the cluster is running with the fixed daemon on every node. Until every node is upgraded, put the daemon ports behind an allowlist that only admits known cluster members and known client nodes.","references":["https://www.ibm.com/support/pages/node/6960571","https://nvd.nist.gov/vuln/detail/CVE-2020-4927"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2021-21378","cve":"CVE-2021-21378","aliases":[],"title":"Envoy: JWT with an issuer absent from the provider list bypasses JWT authentication","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy","year":"2021","cvss_score":8.2,"severity":"high","kev":false,"impact":"JWT with an issuer absent from the provider list bypasses JWT authentication","attack_vector":"Unauthenticated network","remediation":"Upgrade Envoy","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-21378"],"status":"curated","published":"2021-03-11"},{"id":"CVE-2021-21522","cve":"CVE-2021-21522","aliases":["CVE-2021-36285","Dell BIOS NVMe password bypass"],"title":"Dell client and server BIOS - NVMe drive password (SED credential) defeated by resetting the BIOS password","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell client and server BIOS - NVMe drive password (SED credential) defeated by resetting the BIOS password via the…","year":"2021","cvss_score":8.2,"severity":"high","kev":false,"impact":"The BIOS-managed NVMe drive password - the credential many operators rely on to keep a locked drive locked - can be defeated by resetting the BIOS password through the manageability interface, giving access to data on the NVMe device. The paired issue removes the limit on failed NVMe password attempts, so an administrator can brute-force the drive password instead. BREAKS TENANT HANDOFF wherever platform-level drive locking is your control: the lock lives in a BIOS that a local administrator can reset, so the drive credential inherits the security of the BIOS password rather than of the drive. This is the systems-integration failure mode of SEDs - the drive firmware may be perfectly sound while the platform that holds its credential hands it away.","attack_vector":"A local authenticated user with elevated privilege on the host - which on bare metal means the tenant you just rented the box to, if they have BIOS/manageability reach. The brute-force variant requires local administrator access.","remediation":"Apply the Dell BIOS updates named in Dell's advisory for the affected client and PowerEdge platforms; this is a host BIOS flash, so it needs a reboot and a maintenance window per node but not a drive teardown. Then fix the design, not just the bug: do not use BIOS-held NVMe passwords as your tenant-separation mechanism on bare metal, because the whole scheme assumes the tenant cannot reach platform firmware, and on a rented bare-metal node that assumption is false by definition. Lock down the manageability interface, set and monitor BIOS admin passwords, and move the actual data protection to LUKS/dm-crypt with a key your control plane holds and destroys at reclaim - a key the tenant's BIOS access cannot reach.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-21522","https://nvd.nist.gov/vuln/detail/CVE-2021-36285","https://www.dell.com/support/kbdoc/000191495"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2021-09-28"},{"id":"CVE-2021-33627","cve":"CVE-2021-33627","aliases":["INSYDE-SA-2022022","VU#796611"],"title":"Insyde InsydeH2O (FwBlockServiceSmm): Software SMI services reachable through EFI_SMM_COMMUNICATION_PROTOCOL never","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (FwBlockServiceSmm)","year":"2021","cvss_score":8.2,"severity":"high","kev":false,"impact":"Software SMI services reachable through EFI_SMM_COMMUNICATION_PROTOCOL never check whether the buffer address they were given points into SMRAM, MMIO or kernel memory. An OS-level attacker therefore gets SMM to write on their behalf - and FwBlockServiceSmm is the firmware-block service, so this sits directly on the path to the SPI flash. Result is a firmware implant that outlives every reimage and quietly breaks the root of trust the fleet's attestation depends on.","attack_vector":"Local admin/root on the host OS issuing a crafted SMM communication request. No physical access required.","remediation":"Fixed in InsydeH2O kernels 05.09.11 / 05.17.11 / 05.27.11 / 05.36.11 / 05.44.11 / 05.52.11 - delivered to you only as an OEM BIOS image, months downstream. Firmware flash, one reboot per node, drain the GPUs first. No configuration mitigates it. Enable and verify SPI write protection (BIOS Lock / protected range registers) as a partial hardening measure, but that does not close the SMM write primitive itself.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-33627","https://www.insyde.com/security-pledge/SA-2022022","https://kb.cert.org/vuls/id/796611"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-02-03"},{"id":"CVE-2021-4093","cve":"CVE-2021-4093","aliases":[],"title":"KVM (AMD SEV-ES): Out-of-bounds read/write in sev_es_string_io() - malicious SEV-ES guest corrupts host memory","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"KVM (AMD SEV-ES)","year":"2021","cvss_score":8.2,"severity":"high","kev":false,"impact":"Out-of-bounds read/write in sev_es_string_io() - malicious SEV-ES guest corrupts host memory","attack_vector":"Tenant VM guest (SEV-ES)","remediation":"Kernel/KVM patch + reboot. Directly relevant to any confidential-VM GPU offering","references":["https://access.redhat.com/security/cve/CVE-2021-4093"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2022-02-18"},{"id":"CVE-2022-28200","cve":"CVE-2022-28200","aliases":[],"title":"NVIDIA DGX A100 - SBIOS / SMM firmware: The BiosCfgTool reads and writes outside its bounds in SMRAM, handing","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX A100 - SBIOS / SMM firmware","year":"2022","cvss_score":8.2,"severity":"high","kev":false,"impact":"The BiosCfgTool reads and writes outside its bounds in SMRAM, handing a privileged local user code execution in System Management Mode with a changed scope. This is firmware-level persistence: it survives OS reinstall, image re-flash and tenant handoff, and it is invisible to anything running above it. On a bare-metal GPU rental business it is the difference between wiping a node between tenants and not actually being able to.","attack_vector":"Local and already privileged - host root, or code that has reached the platform firmware/SMM path. It is not a first foothold; it is what turns a one-time root compromise into something you cannot remediate by reimaging.","remediation":"Flash the fixed SBIOS from bulletin 5367. Cost: not live-patchable. Full node drain, host power cycle, and on DGX the SBIOS ships inside a firmware bundle alongside BMC and CPLD components, so budget 30-60 minutes of node downtime plus a post-flash health check. Firmware rollback protection means you cannot cleanly revert - stage on one node before the fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28200","https://github.com/NVIDIA/product-security/tree/main/2022/5367"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-119"],"fleet":{"pain_class":"firmware-flash"},"published":"2022-07-02"},{"id":"CVE-2022-28735","cve":"CVE-2022-28735","aliases":[],"title":"GRUB2 (shim_lock verifier): The shim_lock verifier let non-kernel files through, so an attacker could get unsigned","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (shim_lock verifier)","year":"2022","cvss_score":8.2,"severity":"high","kev":false,"impact":"The shim_lock verifier let non-kernel files through, so an attacker could get unsigned content loaded into the boot path while Secure Boot enforcement appeared intact. It defeats the exact control operators rely on to promise a clean handoff between bare-metal tenants.","attack_vector":"Local, with the ability to place a file GRUB will load.","remediation":"grub2 package update + reboot. This is one where the dbx revocation genuinely matters - without it the old signed GRUB remains a usable bypass tool that an attacker can simply drop onto a patched node.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28735","https://access.redhat.com/security/cve/CVE-2022-28735"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2023-07-20"},{"id":"CVE-2022-28737","cve":"CVE-2022-28737","aliases":[],"title":"shim (handle_image PE loader): Buffer overflow in shim's own image loader","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"shim (handle_image PE loader)","year":"2022","cvss_score":8.2,"severity":"high","kev":false,"impact":"Buffer overflow in shim's own image loader. Because shim is the Microsoft-signed component every Linux node chains through, a bug here is worse than a GRUB bug: it bypasses Secure Boot on any machine that trusts the Microsoft 3rd-party CA, regardless of which distro's GRUB sits behind it.","attack_vector":"Local, with the ability to present a crafted EFI image to shim.","remediation":"shim package update + reboot per node. Revoking the old shim means an SBAT generation bump pushed by Microsoft/vendor updates rather than a dbx entry - track SBAT levels, not just package versions, or you will believe you are fixed when you are not.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28737","https://access.redhat.com/security/cve/CVE-2022-28737"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2023-07-20"},{"id":"CVE-2022-29275","cve":"CVE-2022-29275","aliases":["INSYDE-SA-2022058"],"title":"Insyde InsydeH2O (UsbCoreDxe, untrusted pointer use): UsbCoreDxe uses pointers it was handed without establishing they","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (UsbCoreDxe, untrusted pointer use)","year":"2022","cvss_score":8.2,"severity":"high","kev":false,"impact":"UsbCoreDxe uses pointers it was handed without establishing they point outside SMRAM, so an OS-level caller gets SMM to tamper with either SMRAM or kernel memory. Ring -2 escalation from a driver that is resident on every node whether or not a USB device is plugged in. Not a DMA race - this one needs only host privilege, which makes it materially easier to exploit than the SA-2022042-057 set.","attack_vector":"Local admin/root on the host OS invoking the vulnerable software SMI with attacker-chosen pointers. On bare-metal GPU rental this is exactly the privilege the tenant already holds on their leased node.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in kernel 5.0 / 05.09.21, 5.1 / 05.17.21, 5.2 / 05.27.21, 5.3 / 05.36.21, 5.4 / 05.44.21, 5.5 / 05.52.21. Disabling USB legacy/emulation support in BIOS on headless nodes shrinks the reachable surface without a flash - confirm on your platform that it unloads the SMM module rather than only hiding the setup option. The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-29275","https://www.insyde.com/security-pledge/SA-2022058"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-11-15"},{"id":"CVE-2022-29276","cve":"CVE-2022-29276","aliases":["INSYDE-SA-2022059"],"title":"Insyde InsydeH2O (AhciBusDxe, untrusted SMI inputs): SMI functions in the AHCI/SATA driver consume untrusted inputs","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (AhciBusDxe, untrusted SMI inputs)","year":"2022","cvss_score":8.2,"severity":"high","kev":false,"impact":"SMI functions in the AHCI/SATA driver consume untrusted inputs and corrupt SMRAM. Straight privilege escalation to ring -2 for anyone with host root - firmware persistence that survives OS reinstall, plus control over the SATA path the node boots from.","attack_vector":"Local admin/root on the host OS invoking the vulnerable software SMI with attacker-chosen pointers. On bare-metal GPU rental this is exactly the privilege the tenant already holds on their leased node.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in kernel 5.0 / 05.09.18 through 5.5 / 05.52.18.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-29276","https://www.insyde.com/security-pledge/SA-2022059"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-11-15"},{"id":"CVE-2022-29278","cve":"CVE-2022-29278","aliases":["INSYDE-SA-2022061"],"title":"Insyde InsydeH2O (NvmExpressDxe, incorrect pointer checks): The NVMe driver's pointer validation is wrong, allowing","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (NvmExpressDxe, incorrect pointer checks)","year":"2022","cvss_score":8.2,"severity":"high","kev":false,"impact":"The NVMe driver's pointer validation is wrong, allowing tampering with both SMRAM and OS memory. The driver reaches the NVMe data path - datasets, checkpoints, weights on a GPU node - and the bug hands an OS-level attacker ring -2 on top of it. Distinct from the DMA race in SA-2022055 and separately fixed; a node can carry one and not the other.","attack_vector":"Local admin/root on the host OS invoking the vulnerable software SMI with attacker-chosen pointers. On bare-metal GPU rental this is exactly the privilege the tenant already holds on their leased node.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in kernel 5.1 / 05.17.23 through 5.5 / 05.52.23.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-29278","https://www.insyde.com/security-pledge/SA-2022061"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-11-15"},{"id":"CVE-2022-29279","cve":"CVE-2022-29279","aliases":["INSYDE-SA-2022062"],"title":"Insyde InsydeH2O (SdHostDriver and SdMmcDevice, untrusted pointer use): One advisory covering both SD layers: untrusted","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (SdHostDriver and SdMmcDevice, untrusted pointer use)","year":"2022","cvss_score":8.2,"severity":"high","kev":false,"impact":"One advisory covering both SD layers: untrusted pointers allow tampering with SMRAM and OS memory, giving ring -2 code execution. Easy to under-prioritise because SD/eMMC looks irrelevant on a GPU box, but the driver is compiled in and the SMI is reachable regardless of whether any SD media is present.","attack_vector":"Local admin/root on the host OS invoking the vulnerable software SMI with attacker-chosen pointers. On bare-metal GPU rental this is exactly the privilege the tenant already holds on their leased node.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in kernel 5.0 / 05.09.17 through 5.5 / 05.52.17.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-29279","https://www.insyde.com/security-pledge/SA-2022062"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-11-15"},{"id":"CVE-2022-30771","cve":"CVE-2022-30771","aliases":["INSYDE-SA-2022064"],"title":"Insyde InsydeH2O (PnpSmm initialization, SMRAM corruption via later PNP SMIs): An initialization-order defect: PnpSmm's","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (PnpSmm initialization, SMRAM corruption via later PNP SMIs)","year":"2022","cvss_score":8.2,"severity":"high","kev":false,"impact":"An initialization-order defect: PnpSmm's init function leaves state that causes SMRAM corruption when subsequent PNP SMI functions are called. The interesting operational property is that the damage is deferred - the node boots cleanly and the corruption only lands when something later exercises the PNP path, so this does not look like an attack when it fires.","attack_vector":"Local admin/root on the host OS invoking the vulnerable software SMI with attacker-chosen pointers. On bare-metal GPU rental this is exactly the privilege the tenant already holds on their leased node.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in kernel 5.1 / 05.17.25, 5.2 / 05.27.25, 5.3 / 05.36.25, 5.4 / 05.44.25, 5.5 / 05.52.25.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-30771","https://www.insyde.com/security-pledge/SA-2022064"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-11-15"},{"id":"CVE-2022-30772","cve":"CVE-2022-30772","aliases":["INSYDE-SA-2022065"],"title":"Insyde InsydeH2O (PnpSmm function 0x52, SMBIOS write address manipulation): PnpSmm function 0x52 takes an address","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (PnpSmm function 0x52, SMBIOS write address manipulation)","year":"2022","cvss_score":8.2,"severity":"high","kev":false,"impact":"PnpSmm function 0x52 takes an address and a size for data to write into the SMBIOS table and does not constrain where that address points. Malware supplies its own address and overwrites SMRAM or OS kernel memory - an arbitrary write primitive handed over by a documented firmware function, no race and no exotic hardware needed. The cleanest escalation in the batch.","attack_vector":"Local admin/root on the host OS invoking the vulnerable software SMI with attacker-chosen pointers. On bare-metal GPU rental this is exactly the privilege the tenant already holds on their leased node.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in kernel 5.0 / 05.09.41, 5.1 / 05.17.43, 5.2 / 05.27.30, 5.3 / 05.36.30, 5.4 / 05.44.30, 5.5 / 05.52.30.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-30772","https://www.insyde.com/security-pledge/SA-2022065"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-11-15"},{"id":"CVE-2022-31599","cve":"CVE-2022-31599","aliases":[],"title":"NVIDIA DGX A100 - SBIOS / SMM firmware: An uninitialised pointer in the Ofbd SMM handler gives a privileged local user","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX A100 - SBIOS / SMM firmware","year":"2022","cvss_score":8.2,"severity":"high","kev":false,"impact":"An uninitialised pointer in the Ofbd SMM handler gives a privileged local user SMM code execution and privilege escalation beyond the component. This is firmware-level persistence: it survives OS reinstall, image re-flash and tenant handoff, and it is invisible to anything running above it. On a bare-metal GPU rental business it is the difference between wiping a node between tenants and not actually being able to.","attack_vector":"Local and already privileged - host root, or code that has reached the platform firmware/SMM path. It is not a first foothold; it is what turns a one-time root compromise into something you cannot remediate by reimaging.","remediation":"Flash the fixed SBIOS from bulletin 5367. Cost: not live-patchable. Full node drain, host power cycle, and on DGX the SBIOS ships inside a firmware bundle alongside BMC and CPLD components, so budget 30-60 minutes of node downtime plus a post-flash health check. Firmware rollback protection means you cannot cleanly revert - stage on one node before the fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31599","https://github.com/NVIDIA/product-security/tree/main/2022/5367"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-824"],"fleet":{"pain_class":"firmware-flash"},"published":"2022-07-04"},{"id":"CVE-2022-31705","cve":"CVE-2022-31705","aliases":[],"title":"VMware ESXi / Workstation / Fusion: Heap out-of-bounds write in the USB 2.0 EHCI controller","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"VMware ESXi / Workstation / Fusion","year":"2022","cvss_score":8.2,"severity":"high","kev":false,"impact":"Heap out-of-bounds write in the USB 2.0 EHCI controller - VM escape to VMX process code execution (GeekPwn 2022)","attack_vector":"Tenant VM guest (local admin inside the VM)","remediation":"ESXi patch + host reboot with evacuation; remove USB controllers from tenant VM templates","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31705"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2022-12-14"},{"id":"CVE-2022-35408","cve":"CVE-2022-35408","aliases":["INSYDE-SA-2022031"],"title":"Insyde InsydeH2O (UsbLegacyControlSmm): A classic SMM callout: code running inside SMM calls out to a function pointer","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (UsbLegacyControlSmm)","year":"2022","cvss_score":8.2,"severity":"high","kev":false,"impact":"A classic SMM callout: code running inside SMM calls out to a function pointer that lives in memory the OS can write. An attacker plants their own pointer, triggers the SMI, and their code runs at ring -2. USB legacy support is enabled by default on most server BIOS images, so the attack surface is present on nodes that have no USB device attached at all.","attack_vector":"Local admin/root on the host OS, then a software SMI into the USB legacy handler.","remediation":"OEM BIOS update carrying the fixed Insyde kernel. Firmware flash, one reboot per node. Partial config workaround that is genuinely worth doing on servers: disable USB legacy support / USB emulation in BIOS setup - it is rarely needed on a headless GPU node and can be pushed via the OEM's remote BIOS-settings tooling without a flash. Verify on your platform that the setting actually unloads the driver rather than just hiding the option.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-35408","https://www.insyde.com/security-pledge/SA-2022031"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-09-22"},{"id":"CVE-2022-36337","cve":"CVE-2022-36337","aliases":["INSYDE-SA-2022039"],"title":"Insyde InsydeH2O (MebxConfiguration DXE driver): A UEFI variable that the OS can write is read back by BIOS code","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (MebxConfiguration DXE driver)","year":"2022","cvss_score":8.2,"severity":"high","kev":false,"impact":"A UEFI variable that the OS can write is read back by BIOS code into a fixed-size stack buffer without a length check. Set the variable from the OS, reboot, and your code runs during DXE - before Secure Boot has finished deciding what is allowed to run. The persistence mechanism is the variable store itself, which means the implant re-arms on every boot and survives disk replacement entirely.","attack_vector":"Local admin/root on the host OS with the ability to write UEFI variables (standard on Linux via efivarfs and on Windows via SetFirmwareEnvironmentVariable), then one reboot.","remediation":"OEM BIOS update built on the fixed Insyde kernel. Firmware flash, reboot per node. There is no config toggle. Detection is possible in the interim: monitor for unexpected writes to the relevant UEFI variables from the OS, and make efivarfs read-only where your workload does not need it. On a fleet, treat any node where firmware variables changed outside a maintenance window as suspect.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-36337","https://www.insyde.com/security-pledge/SA-2022039"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-11-23"},{"id":"CVE-2022-42291","cve":"CVE-2022-42291","aliases":[],"title":"GeForce Experience installer: Local privesc via untrusted search path","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GeForce Experience installer","year":"2022","cvss_score":8.2,"severity":"high","kev":false,"impact":"Local privesc via untrusted search path","attack_vector":"Local user","remediation":"Consumer-only; no DC action","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42291","https://github.com/NVIDIA/product-security/tree/main/2023/5384"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:C/C:N/I:H/A:H","cwe":["CWE-1386"],"published":"2023-02-07"},{"id":"CVE-2023-0209","cve":"CVE-2023-0209","aliases":[],"title":"NVIDIA DGX-1 - SBIOS / SMM firmware: The Uncore PEI module never authenticates the code executed by SSA, so","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX-1 - SBIOS / SMM firmware","year":"2023","cvss_score":8.2,"severity":"high","kev":false,"impact":"The Uncore PEI module never authenticates the code executed by SSA, so a privileged local attacker gets arbitrary firmware-phase code execution and a full Secure Boot bypass. This is the most serious DGX SBIOS issue in the set - the platform's chain of trust simply does not check. This is firmware-level persistence: it survives OS reinstall, image re-flash and tenant handoff, and it is invisible to anything running above it. On a bare-metal GPU rental business it is the difference between wiping a node between tenants and not actually being able to.","attack_vector":"Local and already privileged - host root, or code that has reached the platform firmware/SMM path. It is not a first foothold; it is what turns a one-time root compromise into something you cannot remediate by reimaging.","remediation":"Flash the fixed SBIOS from bulletin 5458. Cost: not live-patchable. Full node drain, host power cycle, and on DGX the SBIOS ships inside a firmware bundle alongside BMC and CPLD components, so budget 30-60 minutes of node downtime plus a post-flash health check. Firmware rollback protection means you cannot cleanly revert - stage on one node before the fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0209","https://github.com/NVIDIA/product-security/tree/main/2023/5458"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-287"],"fleet":{"pain_class":"firmware-flash"},"published":"2023-04-22"},{"id":"CVE-2023-2163","cve":"CVE-2023-2163","aliases":[],"title":"Linux kernel (eBPF verifier): Incorrect verifier pruning marks unsafe paths as safe","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (eBPF verifier)","year":"2023","cvss_score":8.2,"severity":"high","kev":false,"impact":"Incorrect verifier pruning marks unsafe paths as safe - arbitrary kernel read/write from BPF","attack_vector":"Any tenant process in a container where BPF is reachable","remediation":"Livepatchable; otherwise drain + reboot. `kernel.unprivileged_bpf_disabled=1`; audit any tenant-facing eBPF observability feature","references":["https://access.redhat.com/security/cve/CVE-2023-2163"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2023-09-20"},{"id":"CVE-2023-26484","cve":"CVE-2023-26484","aliases":[],"title":"KubeVirt: A compromised node's virt-handler service account can be abused cluster-wide","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"KubeVirt","year":"2023","cvss_score":8.2,"severity":"high","kev":false,"impact":"A compromised node's virt-handler service account can be abused cluster-wide","attack_vector":"An attacker who owns one node","remediation":"Upgrade KubeVirt; scope down the virt-handler service account","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-26484"],"status":"curated","published":"2023-03-15"},{"id":"CVE-2023-27487","cve":"CVE-2023-27487","aliases":[],"title":"Envoy: Client can forge the x-envoy-original-path header and bypass JWT checks","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy","year":"2023","cvss_score":8.2,"severity":"high","kev":false,"impact":"Client can forge the x-envoy-original-path header and bypass JWT checks","attack_vector":"Unauthenticated network","remediation":"Upgrade Envoy; strip x-envoy headers at the edge","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-27487"],"status":"curated","published":"2023-04-04"},{"id":"CVE-2023-31017","cve":"CVE-2023-31017","aliases":[],"title":"GPU Display Driver (Windows): Arbitrary write to privileged locations via reparse points","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver (Windows)","year":"2023","cvss_score":8.2,"severity":"high","kev":false,"impact":"Arbitrary write to privileged locations via reparse points","attack_vector":"Local low-priv user","remediation":"Upgrade Oct-2023 driver branch","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5491/5491.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-59"],"published":"2023-11-02"},{"id":"CVE-2023-31027","cve":"CVE-2023-31027","aliases":[],"title":"GPU Display Driver (Windows): Local privesc during driver update","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver (Windows)","year":"2023","cvss_score":8.2,"severity":"high","kev":false,"impact":"Local privesc during driver update","attack_vector":"Local low-priv user on the host","remediation":"Upgrade to Oct-2023 driver branch; Windows hosts only","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5491/5491.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:C/C:H/I:H/A:H","cwe":["CWE-427"],"published":"2023-11-02"},{"id":"CVE-2023-34330","cve":"CVE-2023-34330","aliases":[],"title":"AMI MegaRAC SPx (Dynamic Redfish Extension): Code injection executed via the Dynamic Redfish Extension interface","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx (Dynamic Redfish Extension)","year":"2023","cvss_score":8.2,"severity":"high","kev":false,"impact":"Code injection executed via the Dynamic Redfish Extension interface; BMC-level code execution","attack_vector":"Network / Redfish, authenticated-adjacent","remediation":"Same BMC flash cycle as CVE-2023-34329; ODM rebase required, cannot be mitigated in the host OS","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-34330"],"status":"curated","published":"2023-07-18"},{"id":"CVE-2023-35944","cve":"CVE-2023-35944","aliases":[],"title":"Envoy: Mixed-case HTTP/2 schemes defeat case-sensitive internal scheme checks","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy","year":"2023","cvss_score":8.2,"severity":"high","kev":false,"impact":"Mixed-case HTTP/2 schemes defeat case-sensitive internal scheme checks","attack_vector":"Unauthenticated network","remediation":"Upgrade Envoy","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-35944"],"status":"curated","published":"2023-07-25"},{"id":"CVE-2023-46805","cve":"CVE-2023-46805","aliases":[],"title":"Ivanti Connect Secure: Web-component authentication bypass reaching restricted resources","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ivanti Connect Secure","year":"2023","cvss_score":8.2,"severity":"high","kev":true,"impact":"Web-component authentication bypass reaching restricted resources; chained for unauth RCE","attack_vector":"Network (remote)","remediation":"Control-plane: patch and rebuild the appliance - the integrity checker is not sufficient","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-46805"],"status":"curated","published":"2024-01-12"},{"id":"CVE-2023-49938","cve":"CVE-2023-49938","aliases":[],"title":"Slurm: A user can modify their extended group list used by sbcast and open files with unauthorized permissions","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Slurm","year":"2023","cvss_score":8.2,"severity":"high","kev":false,"impact":"A user can modify their extended group list used by sbcast and open files with unauthorized permissions","attack_vector":"Any user who can submit a job","remediation":"Upgrade Slurm","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-49938"],"status":"curated","published":"2023-12-14"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:H/A:L","cwe":["CWE-664"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-52775","cve":"CVE-2023-52775","aliases":[],"title":"Linux kernel SMC-R (fallback path, DECLINE message leaking into the application stream): Silent data corruption, which","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel SMC-R (fallback path, DECLINE message leaking into the application stream)","year":"2023","cvss_score":8.2,"severity":"high","kev":false,"impact":"Silent data corruption, which is worse than a crash because nothing alerts. When SMC falls back to TCP, an in-flight SMC DECLINE control message could be delivered into the application's data stream - the reporters found Redis receiving raw SMC protocol bytes (0xE2 0xD4 0xC3 0xD9 ...) as if they were payload. On a training cluster the equivalent is protocol bytes landing inside a checkpoint shard or a gradient transfer: the job does not fail, it produces wrong results, and you find out weeks later if at all.","attack_vector":"Remote/network. A peer that triggers SMC fallback at the wrong moment - or ordinary network conditions that cause fallback - injects control bytes into the data stream. No authentication involved.","remediation":"Kernel update preventing the decline message from reaching the socket's data path. Until patched, treat any SMC-accelerated flow as unsuitable for data whose integrity you cannot verify end to end, and add application-level checksums to checkpoint and dataset transfers - which is good practice regardless of this bug.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=5ada292b5c504720a0acef8cae9acc62a694d19c","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2023/CVE-2023-52775.json"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2023-6549","cve":"CVE-2023-6549","aliases":[],"title":"Citrix NetScaler ADC/Gateway: Buffer overflow causing denial of service when configured as Gateway or AAA vserver","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Citrix NetScaler ADC/Gateway","year":"2023","cvss_score":8.2,"severity":"high","kev":true,"impact":"Buffer overflow causing denial of service when configured as Gateway or AAA vserver","attack_vector":"Network (remote)","remediation":"Control-plane: patch during the next maintenance window","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-6549"],"status":"curated","published":"2024-01-17"},{"id":"CVE-2024-0082","cve":"CVE-2024-0082","aliases":[],"title":"ChatRTX: Local privesc","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"ChatRTX","year":"2024","cvss_score":8.2,"severity":"high","kev":false,"impact":"Local privesc","attack_vector":"Local Windows user","remediation":"Consumer app; no DC action","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0082","https://github.com/NVIDIA/product-security/tree/main/2024/5532"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:C/C:H/I:H/A:H","cwe":["CWE-269"],"published":"2024-04-08"},{"id":"CVE-2024-0126","cve":"CVE-2024-0126","aliases":[],"title":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko): A privileged attacker escalates through","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko)","year":"2024","cvss_score":8.2,"severity":"high","kev":false,"impact":"A privileged attacker escalates through NVIDIA GPU software with a changed CVSS scope, reaching code execution, data corruption and information disclosure beyond the component. The Virtual GPU Manager is in the affected list, so on a vGPU host the scope change means the hypervisor. The attacker sits inside a tenant guest VM and the blast radius is the hypervisor host, which is running every other tenant's vGPU on the same physical GPU. This is precisely the boundary a vGPU-based multi-tenant offering is sold on. NVIDIA's description is unusually vague here; treat the scope-changed 8.2 on a vGPU host as the worst case until you have your own detail.","attack_vector":"A tenant inside their own guest VM, driving the paravirtualised vGPU control interface. Several of these need only an unprivileged process in the guest; the rest need guest root, which a tenant already has on a VM they rented. No host credentials are involved at any point.","remediation":"Patch the vGPU Manager on the hypervisor host per bulletin 5586. Cost: the highest of any class here. The host driver cannot be reloaded while vGPUs are attached, so every tenant VM on that hypervisor must be live-migrated or powered off - a full host drain. NVIDIA also enforces a supported host/guest driver skew, so budget a matching guest-driver campaign in the same window or tenants lose their vGPU on next boot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0126","https://github.com/NVIDIA/product-security/tree/main/2024/5586"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-20"],"fleet":{"pain_class":"node-drain"},"tags":["tenant-isolation"],"published":"2024-10-26"},{"id":"CVE-2024-0179","cve":"CVE-2024-0179","aliases":[],"title":"AmdCpmDisplayFeatureSMM - SMM callout (AMD-SB-7027): An SMM callout in the AmdCpmDisplayFeatureSMM driver lets ring-0","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AmdCpmDisplayFeatureSMM - SMM callout (AMD-SB-7027)","year":"2024","cvss_score":8.2,"severity":"high","kev":false,"impact":"An SMM callout in the AmdCpmDisplayFeatureSMM driver lets ring-0 code overwrite SMRAM. SMM callouts are the classic UEFI escalation pattern: SMM code calls outward into memory the OS controls, so the OS supplies the code SMM then runs. Reported by Quarkslab, scored 8.2 with changed scope - the attacker crosses from the OS into the platform's most privileged context.","attack_vector":"Local, ring-0 on the host.","remediation":"Fixed in AMD PI/AGESA firmware and delivered only as an OEM SBIOS package - AMD ships the PI drop to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before releasing BIOS. **Budget one to six months of OEM lag**, and note that several CVEs in this batch are marked 'no fix planned' on Naples (EPYC 7001) - for those the only remediation is retiring the hardware. Applying it means cordon, drain and a full power cycle per node; there is no driver reload, no live patch and no VBIOS step. Patch alongside the sibling AmdPspP2CmboxV2 issue in the same bulletin - they ship together and leaving one open leaves the class open.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0179","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-7027.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"tags":["tenant-isolation"],"published":"2025-02-11"},{"id":"CVE-2024-1220","cve":"CVE-2024-1220","aliases":["MPSA-238975"],"title":"Moxa NPort W2150A / W2250A wireless device server: A remote attacker can crash or potentially gain code execution","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Moxa NPort W2150A / W2250A wireless device server","year":"2024","cvss_score":8.2,"severity":"high","kev":false,"impact":"A remote attacker can crash or potentially gain code execution on the device server by sending a crafted payload to its web management service — the built-in web server has a stack-based buffer overflow.","attack_vector":"Remote, over the network — no authentication mentioned as a prerequisite in the vendor advisory; reachability to the web service is sufficient to trigger the overflow.","remediation":"Firmware upgrade to the version in Moxa's MPSA-238975 advisory. Flash and reboot each unit; serial sessions on that device drop briefly during the update.","references":["https://www.moxa.com/en/support/product-support/security-advisory/mpsa-238975-nport-w2150a-w2250a-series-web-server-stack-based-buffer-overflow-vulnerability"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-03-06"},{"id":"CVE-2024-21924","cve":"CVE-2024-21924","aliases":[],"title":"AmdPlatformRasSspSmm - SMM callout (AMD-SB-7028): An SMM callout in the platform RAS SMM driver lets ring-0 code modify","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AmdPlatformRasSspSmm - SMM callout (AMD-SB-7028)","year":"2024","cvss_score":8.2,"severity":"high","kev":false,"impact":"An SMM callout in the platform RAS SMM driver lets ring-0 code modify boot service handlers and execute at SMM. Reported by Eclypsium. Worth noting the irony for a GPU operator: the affected driver is the platform's *reliability and error-reporting* code, so the component you depend on to tell you a node is unhealthy is the one handing over the platform.","attack_vector":"Local, ring-0 on the host.","remediation":"Fixed in AMD PI/AGESA firmware and delivered only as an OEM SBIOS package - AMD ships the PI drop to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before releasing BIOS. **Budget one to six months of OEM lag**, and note that several CVEs in this batch are marked 'no fix planned' on Naples (EPYC 7001) - for those the only remediation is retiring the hardware. Applying it means cordon, drain and a full power cycle per node; there is no driver reload, no live patch and no VBIOS step.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21924","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-7028.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"tags":["tenant-isolation"],"published":"2025-02-11"},{"id":"CVE-2024-21925","cve":"CVE-2024-21925","aliases":[],"title":"AmdPspP2CmboxV2 - SMM input validation (AMD-SB-7027): Insufficient input validation in the AmdPspP2CmboxV2 SMM driver","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AmdPspP2CmboxV2 - SMM input validation (AMD-SB-7027)","year":"2024","cvss_score":8.2,"severity":"high","kev":false,"impact":"Insufficient input validation in the AmdPspP2CmboxV2 SMM driver - the PSP mailbox interface - lets ring-0 code overwrite SMRAM and execute at SMM. This one is notable for sitting on the PSP communication path, so a single bug hands the attacker both the SMM context and a channel to the secure processor.","attack_vector":"Local, ring-0 on the host.","remediation":"Fixed in AMD PI/AGESA firmware and delivered only as an OEM SBIOS package - AMD ships the PI drop to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before releasing BIOS. **Budget one to six months of OEM lag**, and note that several CVEs in this batch are marked 'no fix planned' on Naples (EPYC 7001) - for those the only remediation is retiring the hardware. Applying it means cordon, drain and a full power cycle per node; there is no driver reload, no live patch and no VBIOS step.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21925","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-7027.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"tags":["tenant-isolation"],"published":"2025-02-11"},{"id":"CVE-2024-3446","cve":"CVE-2024-3446","aliases":[],"title":"QEMU (virtio): DMA reentrancy leads to double free across virtio devices - guest-to-host code execution in QEMU","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"QEMU (virtio)","year":"2024","cvss_score":8.2,"severity":"high","kev":false,"impact":"DMA reentrancy leads to double free across virtio devices - guest-to-host code execution in QEMU","attack_vector":"Tenant VM guest","remediation":"QEMU update + VM restart or live-migration to patched hosts. GPU-passthrough VMs cannot be live-migrated, so this is a scheduled tenant-visible drain","references":["https://access.redhat.com/security/cve/CVE-2024-3446"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2024-04-09"},{"id":"CVE-2024-35199","cve":"CVE-2024-35199","aliases":[],"title":"TorchServe (gRPC 7070/7071): gRPC ports bound to all interfaces regardless of config","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"TorchServe (gRPC 7070/7071)","year":"2024","cvss_score":8.2,"severity":"high","kev":false,"impact":"gRPC ports bound to all interfaces regardless of config","attack_vector":"Unauthenticated network from a co-tenant or the internet","remediation":"Patch and enforce bind-address at the pod/network policy layer, not in app config","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-35199"],"status":"curated","published":"2024-07-19"},{"id":"CVE-2024-36129","cve":"CVE-2024-36129","aliases":[],"title":"OpenTelemetry Collector: Unsafe decompression","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"OpenTelemetry Collector","year":"2024","cvss_score":8.2,"severity":"high","kev":false,"impact":"Unsafe decompression -> unauthenticated attacker crashes the collector via excessive memory consumption","attack_vector":"Network (remote)","remediation":"Data-plane: upgrade both node agents and the gateway to 0.102.1+","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36129"],"status":"curated","published":"2024-06-05"},{"id":"CVE-2024-39720","cve":"CVE-2024-39720","aliases":[],"title":"Ollama (GGUF parser): Malformed 4-byte GGUF file crashes the server (two HTTP requests)","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ollama (GGUF parser)","year":"2024","cvss_score":8.2,"severity":"high","kev":false,"impact":"Malformed 4-byte GGUF file crashes the server (two HTTP requests)","attack_vector":"Unauthenticated network upload of a crafted GGUF","remediation":"Upgrade past 0.1.46; part of Oligo's six-issue Ollama disclosure","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-39720"],"status":"curated","published":"2024-10-31"},{"id":"CVE-2024-45067","cve":"CVE-2024-45067","aliases":[],"title":"Intel Gaudi software installer: The Gaudi software installer leaves files and directories with permissions that let","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Intel Gaudi software installer","year":"2024","cvss_score":8.2,"severity":"high","kev":false,"impact":"The Gaudi software installer leaves files and directories with permissions that let a non-root local user modify components that later run as root. That is a straight local root path on any node where the Gaudi stack was installed with the affected installer - and root on a Gaudi node means every tenant's job on that node.","attack_vector":"Any local authenticated user on a node that has the Gaudi stack installed. Node images built once and cloned across the fleet propagate the bad permissions everywhere.","remediation":"Upgrade the Gaudi software installer to 1.18 or later, and re-check permissions on nodes already provisioned - upgrading the package does not always repair permissions set by an earlier install. Rebuild the golden node image rather than patching in place. No firmware or BIOS component.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-45067","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01271.html"],"status":"curated","published":"2025-05-14"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:L/A:H","cwe":["CWE-20","CWE-843"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-58239","cve":"CVE-2024-58239","aliases":[],"title":"Linux kernel (net/tls): A non-DATA record already copied out of the pending list could be merged with a second record","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/tls)","year":"2024","cvss_score":8.2,"severity":"high","kev":false,"impact":"A non-DATA record already copied out of the pending list could be merged with a second record of the same type from the queue, so control-plane records (alerts, handshake/key-update messages) and application data get concatenated into one recv() result. The application's view of record boundaries and record types stops matching what the peer actually sent - a protocol-level record-type confusion the peer chooses.","attack_vector":"Remote: the peer decides record types and ordering, so it can arrange a non-DATA record on the rx_list followed by another of the same type. Any kTLS RX socket the peer can reach; no local privilege or device node required. TLS 1.3 makes it easier because the type is only known after decryption.","remediation":"Boot a kernel carrying the linked stable commits. Interim: terminate TLS in userspace for peer-facing services if the application relies on record-type boundaries (key update, alert handling).","references":["https://git.kernel.org/stable/c/f310143961e2d9a0479fca117ce869f8aaecc140","https://git.kernel.org/stable/c/31e10d6cb0c9532ff070cf50da1657c3acee9276","https://nvd.nist.gov/vuln/detail/CVE-2024-58239"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-7344","cve":"CVE-2024-7344","aliases":["Howyar Reloader"],"title":"Signed third-party UEFI application (Howyar Reloader and OEM rebrands): A Microsoft-signed UEFI recovery application","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Signed third-party UEFI application (Howyar Reloader and OEM rebrands)","year":"2024","cvss_score":8.2,"severity":"high","kev":false,"impact":"A Microsoft-signed UEFI recovery application loads an unsigned binary from a hardcoded path using its own loader instead of the firmware's verified LoadImage. Anyone holding a copy of that signed application can drop it on any Secure Boot machine that trusts the Microsoft third-party CA and boot arbitrary pre-OS code - the vulnerable system does not need the vendor's product installed. That is a universal, portable Secure Boot bypass usable to plant a bootkit on a rented GPU node.","attack_vector":"Write access to the EFI System Partition - local admin/root, prior bare-metal tenant, or BMC virtual media. No relationship to whether you use the affected recovery software.","remediation":"Apply the January 2025 UEFI revocation list (dbx) update that revokes the affected binaries - this is a firmware-level revocation, not a package update, so it lands via Windows Update, fwupd/LVFS, or an OEM BIOS update depending on the platform. Verify the revocation actually took on each node; dbx pushes silently no-op on some boards. Also audit the ESP for stray signed EFI applications you never installed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-7344","https://www.welivesecurity.com/en/eset-research/under-cloak-uefi-secure-boot-introducing-cve-2024-7344/","https://kb.cert.org/vuls/id/529659"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2025-01-14"},{"id":"CVE-2025-10451","cve":"CVE-2025-10451","aliases":["INSYDE-SA-2025009"],"title":"Insyde InsydeH2O (H19Int15CallbackSmm, combined DXE/SMM driver): An unchecked output buffer in a combined DXE/SMM","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (H19Int15CallbackSmm, combined DXE/SMM driver)","year":"2025","cvss_score":8.2,"severity":"high","kev":false,"impact":"An unchecked output buffer in a combined DXE/SMM driver lets an attacker write into SMRAM and reach arbitrary code execution in System Management Mode. The 2025 instalment of the same pattern Binarly and Insyde have been working through since 2021 - a driver that takes an address from the caller and writes to it without confirming the address is outside SMRAM. Ring -2 compromise: survives reinstall, defeats Secure Boot and attestation, invisible from the OS.","attack_vector":"Local admin/root on the host OS issuing the vulnerable SMI with a crafted output buffer address.","remediation":"OEM BIOS update carrying Insyde patch IB05690966. Affects Intel Ice Lake and Kaby Lake and AMD Picasso platforms; Insyde scopes the advisory by OEM feature version (HP feature version before 20C1) rather than by kernel version, so map it against your own OEM's BIOS release rather than against an Insyde kernel number. Firmware flash, reboot per node. No config workaround.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-10451","https://www.insyde.com/security-pledge/SA-2025009/"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2025-12-12"},{"id":"CVE-2025-20093","cve":"CVE-2025-20093","aliases":[],"title":"Intel ice driver (Ethernet 800 Series, Linux kernel mode): A missing check for an exceptional condition in the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel ice driver (Ethernet 800 Series, Linux kernel mode)","year":"2025","cvss_score":8.2,"severity":"high","kev":false,"impact":"A missing check for an exceptional condition in the 800-series Linux driver, scored high, reachable by an authenticated user. Same rollout unit as the rest of the ice 1.17.2 batch.","attack_vector":"Authenticated local user on the node.","remediation":"Fixed in the Intel out-of-tree ice driver (or the equivalent in-kernel version). Updating the driver requires unloading and reloading the module, which drops every link on that NIC - on a node whose RDMA fabric carries collective traffic, that is a job-killing event, so drain first. If you take it via a distro kernel update instead, it is a reboot. No firmware flash for the driver-side fixes. Target ice 1.17.2 or later.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-20093","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01296.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-08-12"},{"id":"CVE-2025-22225","cve":"CVE-2025-22225","aliases":[],"title":"VMware ESXi: Arbitrary kernel write from the VMX process - sandbox escape completing the zero-day chain","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"VMware ESXi","year":"2025","cvss_score":8.2,"severity":"high","kev":true,"impact":"Arbitrary kernel write from the VMX process - sandbox escape completing the zero-day chain [KEV]","attack_vector":"Tenant VM guest (chained after CVE-2025-22224)","remediation":"ESXi patch + host reboot with evacuation","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-22225"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-03-04"},{"id":"CVE-2025-22249","cve":"CVE-2025-22249","aliases":[],"title":"VMware Aria Automation (DOM-based XSS, token theft): A crafted URL steals the access token of a logged-in Aria","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"VMware Aria Automation (DOM-based XSS, token theft)","year":"2025","cvss_score":8.2,"severity":"high","kev":false,"impact":"A crafted URL steals the access token of a logged-in Aria Automation user, letting the attacker act as that user against the automation platform.","attack_vector":"Unauthenticated attacker who can get an operator to click a crafted link.","remediation":"Apply the Broadcom fix per advisory 25711.","references":["https://support.broadcom.com/web/ecx/support-content-notification/-/external/content/SecurityAdvisories/0/25711"],"status":"curated"},{"id":"CVE-2025-2241","cve":"CVE-2025-2241","aliases":[],"title":"OpenShift Hive / MCE / ACM (vCenter credential exposure): vCenter credentials are written into the ClusterProvision","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"OpenShift Hive / MCE / ACM (vCenter credential exposure)","year":"2025","cvss_score":8.2,"severity":"high","kev":false,"impact":"vCenter credentials are written into the ClusterProvision object after provisioning a vSphere cluster, so anyone with read access to those objects extracts them - a Kubernetes RBAC read grant becomes hypervisor admin.","attack_vector":"Any user or service account with read access to ClusterProvision objects in the management cluster.","remediation":"Apply the Red Hat fix, then rotate the exposed vCenter credentials and audit who holds read on ClusterProvision. Rotation is mandatory here - the credentials are already at rest in etcd and in any cluster backup.","references":["https://access.redhat.com/security/cve/CVE-2025-2241"],"status":"curated"},{"id":"CVE-2025-23309","cve":"CVE-2025-23309","aliases":[],"title":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys): An uncontrolled DLL loading path in the display","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys)","year":"2025","cvss_score":8.2,"severity":"high","kev":false,"impact":"An uncontrolled DLL loading path in the display driver lets a local attacker plant a DLL and get code execution at driver-install privilege - a classic path to SYSTEM on Windows GPU hosts. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5703. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23309","https://github.com/NVIDIA/product-security/tree/main/2025/5703"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:C/C:H/I:H/A:H","cwe":["CWE-427"],"fleet":{"pain_class":"node-drain"},"published":"2025-10-10"},{"id":"CVE-2025-23342","cve":"CVE-2025-23342","aliases":[],"title":"NVIDIA NVDebug tool: The NVDebug diagnostic collector lets an actor gain access to a privileged account, reaching code","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NVDebug tool","year":"2025","cvss_score":8.2,"severity":"high","kev":false,"impact":"The NVDebug diagnostic collector lets an actor gain access to a privileged account, reaching code execution and privilege escalation with a changed scope. NVDebug runs on DGX/HGX platform hosts and is exactly the tool an operator runs as root when something is already wrong.","attack_vector":"Local, low privileges, with user interaction - typically an operator running NVDebug on a node where an attacker already has an unprivileged foothold.","remediation":"Update NVDebug per bulletin 5696. Cost: it is a standalone tool, so updating costs nothing operationally. The real control is to stop running old NVDebug bundles as root on nodes you already suspect are compromised.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23342","https://github.com/NVIDIA/product-security/tree/main/2025/5696"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:C/C:H/I:H/A:H","cwe":["CWE-522"],"published":"2025-09-09"},{"id":"CVE-2025-25210","cve":"CVE-2025-25210","aliases":["INTEL-SA-01325","CVE-2025-22453","CVE-2025-35999","INTEL-SA-01412","CVE-2025-24918","INTEL-SA-01400"],"title":"Intel Server Firmware Update Utility (SysFwUpdt) and Server Configuration Utility before version 16.0.12: Improper","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Server Firmware Update Utility (SysFwUpdt) and Server Configuration Utility before version 16.0.12","year":"2025","cvss_score":8.2,"severity":"high","kev":false,"impact":"Improper input validation in the utility operators use to flash BIOS, BMC and ME firmware, letting a privileged local user escalate; the companion issues add an incorrect-permission assignment and a link-following flaw in the same tool family. The irony is the point: the tool you run to remediate firmware is itself the escalation path into firmware. A tenant or a compromised operator account on a node can subvert the update process so that the node ends up running attacker-chosen firmware while your fleet records show a successful patch. That gives below-the-OS persistence that survives reimage and crosses tenant handoff, and it corrupts the evidence you would use to detect it.","attack_vector":"A privileged local user on the host where the utility runs - which includes tenants on bare-metal nodes if the tooling is left in the host image, and any compromised operator or automation account that drives firmware rollouts.","remediation":"Update SysFwUpdt and the Server Configuration Utility to 16.0.12 or later before running any further firmware campaigns. Separately, treat firmware update tooling as privileged infrastructure: do not leave it installed in tenant-facing host images, run firmware updates from the BMC/Redfish out-of-band path rather than from the host OS wherever the platform supports it, and verify post-update firmware versions and measurements out of band instead of trusting the tool's own success report.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-25210","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01325.html","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01412.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-33045","cve":"CVE-2025-33045","aliases":["AMI-SA-2025007"],"title":"AMI AptioV UEFI BIOS (SMM): A write-what-where primitive plus an information leak in System Management Mode","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI AptioV UEFI BIOS (SMM)","year":"2025","cvss_score":8.2,"severity":"high","kev":false,"impact":"A write-what-where primitive plus an information leak in System Management Mode - the most privileged execution context on an x86 server, above the kernel and invisible to the hypervisor. An attacker who wins here can write anywhere in memory including SMRAM, disable firmware protections, and install a bootkit that persists across OS reinstall and survives every host-level detection you run. Scope is changed, so the compromise reaches beyond the BIOS into the running system. On a shared GPU host this defeats the boundary that VM isolation and confidential-computing attestation both rest on.","attack_vector":"Local, requires high privileges - root or kernel-level code on the host OS. So the precondition is that an attacker already owns the operating system on a node; this is what they use to convert a revocable OS compromise into permanent firmware residency. In a bare-metal GPU rental model, the tenant themselves have that privilege by design.","remediation":"BIOS update to AptioV_5.040 or later. That is a firmware flash plus a full host reboot per node, which on a GPU fleet means draining the node - and if it is part of a multi-node training job, draining the whole ring. Rollout is gated on your server vendor picking up the AMI BKC and publishing a rebased BIOS for your specific SKU, which typically lags AMI's advisory by months. There is no config-only mitigation for an SMM bug. If you cannot patch, the compensating control is to stop treating host root as a containable compromise: rebuild affected nodes with a verified BIOS reflash rather than an OS reimage.","references":["https://go.ami.com/hubfs/Security%20Advisories/2025/AMI-SA-2025007.pdf","https://nvd.nist.gov/vuln/detail/CVE-2025-33045"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2025-09-09"},{"id":"CVE-2025-53652","cve":"CVE-2025-53652","aliases":[],"title":"Jenkins (Git Parameter plugin): Git parameter value is not validated against the offered choices","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Jenkins (Git Parameter plugin)","year":"2025","cvss_score":8.2,"severity":"high","kev":false,"impact":"Git parameter value is not validated against the offered choices -> injection of arbitrary values into builds","attack_vector":"Network (remote)","remediation":"Control-plane: plugin upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-53652"],"status":"curated","published":"2025-07-09"},{"id":"CVE-2025-56547","cve":"CVE-2025-56547","aliases":["PT-2025-19"],"title":"Broadcom NetXtreme-E network adapter firmware: A high-severity flaw in the firmware of Broadcom NetXtreme-E adapters","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Broadcom NetXtreme-E network adapter firmware","year":"2025","cvss_score":8.2,"severity":"high","kev":false,"impact":"A high-severity flaw in the firmware of Broadcom NetXtreme-E adapters (found in firmware 231.1.162.1 and reported through Positive Technologies' responsible-disclosure process). NetXtreme-E is Broadcom's mainstream datacenter NIC line and the base for the Thor generation used in AI server designs. Adapter-firmware bugs matter more than their CVSS suggests: NIC firmware runs below the hypervisor and below the host OS, it persists across reinstall, and on many server designs the NIC also carries the NC-SI sideband to the BMC — so a compromised NIC is a candidate pivot into out-of-band management.","attack_vector":"Reachable through the adapter's firmware interfaces. Treat any party that can drive the NIC — a host-privileged tenant on bare metal, or network-side input depending on the affected path — as in scope until the vendor detail is public.","remediation":"Flash NetXtreme-E adapter firmware to the fixed release from Broadcom (or via your server OEM's firmware bundle). NIC firmware flash plus a cold power cycle, per node. On a GPU fleet, roll it into the same drain window you use for BMC and BIOS updates — doing NIC firmware as its own campaign is how it ends up never happening.","references":["https://global.ptsecurity.com/en/about/news/pt-expert-helped-patch-vulnerabilities-broadcom-network-adapter-firmware/","https://nvd.nist.gov/vuln/detail/CVE-2025-56547"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2026-24188","cve":"CVE-2026-24188","aliases":[],"title":"NVIDIA TensorRT: An out-of-bounds write reachable from the network reaches data tampering, scored 8.2 with no","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA TensorRT","year":"2026","cvss_score":8.2,"severity":"high","kev":false,"impact":"An out-of-bounds write reachable from the network reaches data tampering, scored 8.2 with no privileges required. TensorRT sits inside almost every optimised inference deployment, so the affected surface is wide even where TensorRT is not the thing you deployed by name.","attack_vector":"Network, unauthenticated. Reached through whatever service embeds the TensorRT runtime and passes it externally-influenced input.","remediation":"Update TensorRT per bulletin 5836 and rebuild every inference image that links it - including Triton images with the TensorRT backend. Cost: image rebuild plus rolling restart; inventory work is the expensive part because TensorRT is usually a transitive dependency.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24188","https://github.com/NVIDIA/product-security/tree/main/2026/5836"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:H/A:L","cwe":["CWE-787"],"fleet":{"pain_class":"daemon-restart"},"published":"2026-05-20"},{"id":"CVE-2026-24189","cve":"CVE-2026-24189","aliases":[],"title":"NVIDIA CUDA-Q: Info disclosure / code exec (OOB read in circuit compilation)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA CUDA-Q","year":"2026","cvss_score":8.2,"severity":"high","kev":false,"impact":"Info disclosure / code exec (OOB read in circuit compilation)","attack_vector":"Malicious user program","remediation":"Bump CUDA-Q; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24189","https://github.com/NVIDIA/product-security/tree/main/2026/5820"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:L/I:N/A:H","cwe":["CWE-125"],"published":"2026-04-21"},{"id":"CVE-2026-24253","cve":"CVE-2026-24253","aliases":[],"title":"NVIDIA Dynamo: RCE via buffer overflow in tensor shape validation","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Dynamo","year":"2026","cvss_score":8.2,"severity":"high","kev":false,"impact":"RCE via buffer overflow in tensor shape validation","attack_vector":"Any inference client","remediation":"Bump Dynamo; redeploy the serving stack","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24253","https://github.com/NVIDIA/product-security/tree/main/2026/5842"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:L/A:H","cwe":["CWE-787"],"published":"2026-08-04"},{"id":"CVE-2026-33748","cve":"CVE-2026-33748","aliases":[],"title":"BuildKit: Insufficient validation of git URL fragment subdir allows access to files outside the intended checkout","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"BuildKit","year":"2026","cvss_score":8.2,"severity":"high","kev":false,"impact":"Insufficient validation of git URL fragment subdir allows access to files outside the intended checkout","attack_vector":"Anyone who can submit a build","remediation":"Upgrade BuildKit to 0.28.1+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-33748"],"status":"curated","published":"2026-03-27"},{"id":"CVE-2026-41326","cve":"CVE-2026-41326","aliases":[],"title":"Kata Containers: Oversight in the CopyFile policy from v3.4.0 to v3.28.0","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kata Containers","year":"2026","cvss_score":8.2,"severity":"high","kev":false,"impact":"Oversight in the CopyFile policy from v3.4.0 to v3.28.0","attack_vector":"Any tenant workload under Kata","remediation":"Upgrade Kata","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-41326"],"status":"curated","published":"2026-04-24"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:L/A:H","cwe":["CWE-825"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-43466","cve":"CVE-2026-43466","aliases":[],"title":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core/en): After a transmit-queue error triggers driver recovery, the","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core/en)","year":"2026","cvss_score":8.2,"severity":"high","kev":false,"impact":"After a transmit-queue error triggers driver recovery, the software DMA FIFO's producer and consumer indices are left out of sync, so the driver unmaps DMA addresses recorded before the recovery. The node tears down IOMMU mappings that no longer correspond to the buffers being freed - stale-mapping teardown on a shared NIC, which means either a live mapping is revoked underneath in-flight DMA or a dead IOVA is left mapped, on top of repeated queue failures.","attack_vector":"Reachable over the network: the recovery path is entered on a transmit error completion, which a peer on the Ethernet/RDMA fabric can provoke through congestion, malformed or oversized traffic, and link-level abuse against the node's uplink. No tenant device node is required - the corrupted state is in the host's shared TX path, so any tenant sharing that NIC is exposed.","remediation":"Update to a kernel carrying the fix on your stream. There is no useful runtime workaround: the code runs whenever an error CQE occurs on a TX queue. Reduce exposure by rate-limiting or isolating untrusted senders on the fabric until nodes are rebooted onto a fixed kernel.","references":["https://git.kernel.org/stable/c/821f85d619f7f22cda7b9d7de89cf5eeb1d11544","https://git.kernel.org/stable/c/6eb68ecc5acc3b319986566c595990b8a7265b23","https://nvd.nist.gov/vuln/detail/CVE-2026-43466"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-47483","cve":"CVE-2026-47483","aliases":[],"title":"DCGM: DoS of the GPU telemetry/health daemon (resource exhaustion)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DCGM","year":"2026","cvss_score":8.2,"severity":"high","kev":false,"impact":"DoS of the GPU telemetry/health daemon (resource exhaustion)","attack_vector":"Any tenant able to reach the DCGM socket/endpoint on the node","remediation":"Bump DCGM / dcgm-exporter; upgrade GPU Operator chart; restart the daemonset, no tenant eviction needed","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47483","https://github.com/NVIDIA/product-security/tree/main/2026/5857"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:L/I:N/A:H","cwe":["CWE-770"],"published":"2026-07-28"},{"id":"CVE-2026-47623","cve":"CVE-2026-47623","aliases":[],"title":"NVIDIA Dynamo: Deserialization of untrusted data in Dynamo reaches denial of service and data tampering","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Dynamo","year":"2026","cvss_score":8.2,"severity":"high","kev":false,"impact":"Deserialization of untrusted data in Dynamo reaches denial of service and data tampering from an unauthenticated network caller, scored 8.2. Dynamo is NVIDIA's disaggregated serving framework, so this sits on the request path of a production inference tier.","attack_vector":"Network, unauthenticated, no user interaction. Any client that can submit to the Dynamo endpoint.","remediation":"Upgrade Dynamo per bulletin 5842 and roll the deployment. Cost: rolling restart of the serving tier; no driver or firmware change. Verify the endpoint is not exposed beyond the mesh.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47623","https://github.com/NVIDIA/product-security/tree/main/2026/5842"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:L/A:H","cwe":["CWE-502"],"fleet":{"pain_class":"daemon-restart"},"published":"2026-08-04"},{"id":"CVE-2026-53489","cve":"CVE-2026-53489","aliases":[],"title":"containerd: CRI restores container.log from a checkpoint image without validating symlinks","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2026","cvss_score":8.2,"severity":"high","kev":false,"impact":"CRI restores container.log from a checkpoint image without validating symlinks; arbitrary host file write","attack_vector":"Malicious checkpoint image","remediation":"Rolling containerd upgrade with node drain; disable CRI checkpoint/restore for tenants","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53489"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2026-07-01"},{"id":"CVE-2026-5817","cve":"CVE-2026-5817","aliases":[],"title":"Docker Model Runner (vllm-metal backend): `trust_remote_code=True` set unconditionally, no sandbox","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Docker Model Runner (vllm-metal backend)","year":"2026","cvss_score":8.2,"severity":"high","kev":false,"impact":"`trust_remote_code=True` set unconditionally, no sandbox → tokenizer code executes","attack_vector":"Customer-supplied model repo","remediation":"Rebuild/patch; no operator config can disable it","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-5817"],"status":"curated","published":"2026-05-22"},{"id":"CVE-2026-5944","cve":"CVE-2026-5944","aliases":[],"title":"Cisco Intersight Device Connector for Nutanix Prism Central: The device connector exposes an unauthenticated API","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Cisco Intersight Device Connector for Nutanix Prism Central","year":"2026","cvss_score":8.2,"severity":"high","kev":false,"impact":"The device connector exposes an unauthenticated API passthrough on TCP/7373 reachable within the deployment's network scope. An attacker with network access uses it to reach the Prism Central API without credentials - unauthenticated proxy into the virtualization control plane.","attack_vector":"Network access to TCP/7373 on the connector host. No authentication.","remediation":"Apply the Nutanix fix per Security Advisory 0046 and restrict TCP/7373 with host or network firewall rules. The port restriction is the immediate control and can be applied before the software update.","references":["https://download.nutanix.com/alerts/Security_Advisory_0046.pdf"],"status":"curated"},{"id":"CVE-2026-6484","cve":"CVE-2026-6484","aliases":["INSYDE-SA-2026003"],"title":"Insyde InsydeH2O (unverified firmware volume in the boot chain): Certain firmware volumes are executed without being","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (unverified firmware volume in the boot chain)","year":"2026","cvss_score":8.2,"severity":"high","kev":false,"impact":"Certain firmware volumes are executed without being verified, so the verified-boot chain has a hole in it: an attacker who can write to the unverified FV gets arbitrary code execution in firmware and the platform's own boot integrity check does not object. This is the root-of-trust failure rather than a memory-safety bug - the whole value of measured and verified boot on a GPU node is that unauthorised firmware cannot run, and here it can.","attack_vector":"An attacker able to modify the affected firmware volume - via SPI write access, a malicious capsule, or an earlier compromise with firmware-write privilege. Then any boot.","remediation":"OEM BIOS update. Insyde ships fixes per Intel platform: Arrow Lake H/U 05.56.17.0022, Arrow Lake S/HX 05.56.17.0037, Raptor Lake 05.47.24.0058 (mobile) and 05.47.24.0057 (server/embedded), Alder Lake 05.47.24.2057, Meteor Lake 05.56.07.0022, Elkhart Lake 05.48.17.0030 - so check your exact silicon generation, since several platforms are listed unaffected. Advisory dated 2026-08-12, meaning OEM images are only just starting to appear; expect the Dell/HPE/Lenovo/Supermicro rebase to trail by months. Firmware flash, reboot per node. No config workaround. Meanwhile, verify SPI flash write protection is actually enforced (BIOS Lock Enable, protected range registers) so the unverified FV is not writable in the first place.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-6484","https://www.insyde.com/security-pledge/SA-2026003/"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2026-08-12"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-665","CWE-787"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-74355","cve":"CVE-2026-74355","aliases":[],"title":"Linux kernel (drivers/iommu/intel): A device that does not support ATS never gets inserted into the VT-d device","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/intel)","year":"2026","cvss_score":8.2,"severity":"high","kev":false,"impact":"A device that does not support ATS never gets inserted into the VT-d device red-black tree, but a later probe failure still runs the removal, which treats the zeroed node as a tree root and corrupts the tree. That tree is what maps an incoming device request back to its device context, so corrupting it means faults and ATS lookups can resolve to the wrong device - and the corruption itself is an out-of-bounds write into kernel memory.","attack_vector":"Requires a probe failure on a device behind VT-d, so the trigger is host-side: driver binding during boot or during passthrough provisioning, on a device without ATS support where a later probe step fails. Not tenant-driven, but the damaged structure is shared by every device on the IOMMU, so the fallout lands on tenant devices.","remediation":"Update to a stable kernel carrying commits f5102e0f / d16923a4. Interim: watch for probe failures in the VT-d path during node bring-up and refuse to schedule tenants onto a node that logged one until it is rebooted.","references":["https://git.kernel.org/stable/c/f5102e0fc3c6dc8685549891a96e4589fdb3e211","https://git.kernel.org/stable/c/d16923a45d4d08367650fdc3451c89299ab6ac5a","https://nvd.nist.gov/vuln/detail/CVE-2026-74355"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-863","CWE-200"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-74516","cve":"CVE-2026-74516","aliases":[],"title":"Linux kernel (arch/x86/kvm/svm): If AVIC is inhibited while a nested guest is running, KVM leaves the x2APIC MSRs","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (arch/x86/kvm/svm)","year":"2026","cvss_score":8.2,"severity":"high","kev":false,"impact":"If AVIC is inhibited while a nested guest is running, KVM leaves the x2APIC MSRs unintercepted for the outer guest. That guest can then read most of the host's real APIC state, send arbitrary interrupts to host CPUs (including the posted-interrupt wakeup vector), change host task priority, and trivially take the node down. This is a guest reaching directly into host interrupt state.","attack_vector":"Driven from inside a guest on an AMD host: the guest starts a nested VM while AVIC is fully enabled, then triggers a VM-scoped AVIC inhibit, and afterwards issues raw x2APIC MSR reads/writes from L1. Requires AMD hardware with AVIC enabled and nested virtualization exposed to the tenant. Not reachable from a plain container tenant.","remediation":"Update to a kernel with the referenced stable commits. Interim: disable AVIC on affected AMD nodes (kvm_amd avic=0) and/or stop exposing nested virtualization to tenants (kvm_amd nested=0) until patched.","references":["https://git.kernel.org/stable/c/4ca05385b3ddbd463be17c6d69ec76fca657081d","https://git.kernel.org/stable/c/6664a5aea45318f4ec156a729949b474dd6e3159","https://nvd.nist.gov/vuln/detail/CVE-2026-74516"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:L/UI:N/S:C/C:H/I:H/A:N","cwe":["CWE-347"],"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2026-039-ceph-rgw-sigv4-signature-verifie","cve":null,"aliases":["GHSA-rmjq-ffrm-j6vj","CVE-2026-54330 (reserved)"],"title":"Ceph RGW (SigV4 signature verifier): Anyone handed a single presigned PUT URL gets more authority than whoever signed","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ceph RGW (SigV4 signature verifier)","year":"2026","cvss_score":8.2,"severity":"high","kev":false,"impact":"Anyone handed a single presigned PUT URL gets more authority than whoever signed it. AWS requires every x-amz-* header on a SigV4 request to be signed and rejects requests carrying extras; RGW only validates the headers named in X-Amz-SignedHeaders and silently honours any additional x-amz-* header the caller bolts on. So a presigned URL that was scoped to upload one object can be replayed with attacker-chosen x-amz-* directives attached, changing ACLs, storage class, object metadata and other server-side behaviour the signer never authorised. In a GPU cloud this is the classic 'here is a presigned link to drop your dataset' workflow turning into a write primitive against the bucket namespace, and presigned URLs are routinely pasted into notebooks, CI logs and Slack, so the holder set is much wider than the tenant who generated it.","attack_vector":"Network, from the public S3 endpoint. The attacker needs one presigned PUT URL, which is a low-privilege artifact normally treated as safe to share. No RGW account of their own is required beyond possession of that URL.","remediation":"Upgrade RGW to 20.2.4 or 19.2.6 and restart the radosgw daemons. Until then, shorten presigned URL lifetimes hard and stop treating a presigned URL as a capability safe to hand to a party you would not grant the underlying bucket permission to. Audit bucket ACLs and object metadata for changes made through presigned uploads.","references":["https://github.com/ceph/ceph/security/advisories/GHSA-rmjq-ffrm-j6vj","https://docs.ceph.com/en/latest/security/CVE-2026-54330"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:L/UI:N/S:C/C:H/I:L/A:L","cwe":["CWE-285"],"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2026-041-ceph-mon-config-key-store-mmonsu","cve":null,"aliases":["GHSA-rg9p-5xcp-wm8h","CVE-2026-50152 (reserved)"],"title":"Ceph MON (config-key store, MMonSubscribe handler): MULTI-TENANT ISOLATION AND HOST COMPROMISE: one crafted","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ceph MON (config-key store, MMonSubscribe handler)","year":"2026","cvss_score":8.2,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION AND HOST COMPROMISE: one crafted MMonSubscribe message from any account with 'mon allow r' dumps the entire monitor config-key store. That store is where Ceph keeps OSD LUKS passphrases and, on cephadm-managed clusters, the SSH private key cephadm uses to reach every host in the fleet. Under the default cephadm setup that key is root-equivalent on every storage node, so a read-only monitor cap converts directly into root on the whole storage tier — and the LUKS passphrases mean an attacker who can also touch the disks gets the data at rest. 'mon allow r' is a cap operators hand out casually to monitoring agents, dashboards and tenant-facing tooling because it reads like a harmless read grant; here it is the whole cluster.","attack_vector":"Adjacent network: the attacker must reach the Ceph monitors and hold any CephX identity with 'mon allow r' capabilities. That includes most metrics collectors, dashboards, and any service account provisioned for cluster read access. No user interaction.","remediation":"Upgrade to Ceph 20.2.4 or 19.2.6 and restart the monitors. Then rotate what the store held: regenerate the cephadm SSH key across the fleet and re-key OSD LUKS where feasible, because a patched monitor does not un-leak secrets already read. Audit which identities hold 'mon allow r' and cut it back to the ones that truly need it.","references":["https://github.com/ceph/ceph/security/advisories/GHSA-rg9p-5xcp-wm8h","https://docs.ceph.com/en/latest/security/CVE-2026-50152/"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-798","CWE-321"],"fleet":{"pain_class":"firmware-flash"},"id":"CVE-2013-3619","cve":"CVE-2013-3619","aliases":[],"title":"Supermicro IPMI BMC firmware: Every affected BMC shares one TLS private key and one SSH host key, baked into the","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro IPMI BMC firmware","year":"2013","cvss_score":8.1,"severity":"high","kev":false,"impact":"Every affected BMC shares one TLS private key and one SSH host key, baked into the firmware image. The consequence is that the encrypted management plane is not encrypted against anyone who has read the firmware: an attacker positioned on the management network decrypts recorded BMC sessions and harvests the administrator credentials inside them, or stands up a machine-in-the-middle that presents a certificate and SSH host key your tooling accepts without complaint. Host-key pinning and TLS verification - the two controls an operator would normally lean on - are both defeated, because the legitimate key and the attacker's key are the same key.","attack_vector":"Network, requires a position on the path to the BMC (or the ability to redirect traffic to it). Pre-auth: the attacker needs no BMC account, only the publicly extractable key material.","remediation":"Firmware flash to SMT_X9_317 / SMT X8 312 or later, and then verify the device actually regenerated unique keys rather than shipping a new shared pair - some BMC firmware regenerates on first boot after reset, some does not, so check the presented certificate fingerprint differs between two of your own nodes. Until that is done, assume all BMC traffic is readable and forgeable, and never carry a credential over it that is also valid elsewhere.","references":["https://www.rapid7.com/blog/post/2013/11/06/supermicro-ipmi-firmware-vulnerabilities/","https://www.kb.cert.org/vuls/id/648646","https://nvd.nist.gov/vuln/detail/CVE-2013-3619"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-330","CWE-384"],"fleet":{"pain_class":"firmware-flash"},"id":"CVE-2014-8272","cve":"CVE-2014-8272","aliases":["VU#843044"],"title":"Dell iDRAC6/iDRAC7 IPMI 1.5 session handling: IPMI 1.5 session IDs are handed out incrementally from a small pool, so","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC6/iDRAC7 IPMI 1.5 session handling","year":"2014","cvss_score":8.1,"severity":"high","kev":false,"impact":"IPMI 1.5 session IDs are handed out incrementally from a small pool, so an unauthenticated attacker guesses the session ID of an administrator's live session and injects IPMI commands into it. No credential is ever cracked - the attacker rides someone else's authentication. Because IPMI 1.5 has no per-message integrity and no encryption, there is nothing downstream to stop the injected command. What an operator gets out of this is power control, boot-device selection and account manipulation on servers they do not own. Dell's response is the tell: rather than fix session ID generation they deleted the IPMI 1.5 code path from the firmware, conceding the protocol version is not securable.","attack_vector":"Network, pre-auth, UDP/623. Needs an administrator session to be active or to be induced, then a short brute-force over the session ID space.","remediation":"Update iDRAC firmware to the versions that remove IPMI 1.5 (iDRAC6 modular 3.65, iDRAC6 monolithic 1.98, iDRAC7 1.57.57 or later). The flash itself is a routine iDRAC update that does not require host downtime, but you lose out-of-band access for a few minutes per node, so schedule it against nodes that are not mid-job. Independent of the flash, disable IPMI 1.5 wherever the BMC still offers it and require cipher suite 3 or better on IPMI 2.0 - a fleet that has patched the firmware but left IPMI 1.5 enabled on other vendors' BMCs has only moved the problem.","references":["https://www.kb.cert.org/vuls/id/843044","https://www.exploit-db.com/exploits/35770","https://nvd.nist.gov/vuln/detail/CVE-2014-8272"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.0/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:H/A:H","cwe":["CWE-285"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2018-10861","cve":"CVE-2018-10861","aliases":[],"title":"Ceph MON (ceph-mon): The monitor accepts pool create/delete and snapshot operations from any authenticated user that","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ceph MON (ceph-mon)","year":"2018","cvss_score":8.1,"severity":"high","kev":false,"impact":"The monitor accepts pool create/delete and snapshot operations from any authenticated user that only has read access. A read-only tenant key becomes a cluster-wide destructive capability - it can delete the pool holding another tenant's dataset or corrupt their RBD snapshots.","attack_vector":"Any authenticated Ceph user with read access that can reach the monitors, so any tenant node holding a client keyring.","remediation":"Upgrade ceph-mon to the fixed release and restart it. Then review pool-level and mon caps for every client key, and separate tenants into distinct pools with explicit per-pool caps rather than a broad read cap.","references":["https://access.redhat.com/security/cve/CVE-2018-10861","https://nvd.nist.gov/vuln/detail/CVE-2018-10861"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.0/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:N","cwe":["CWE-20"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2018-10926","cve":"CVE-2018-10926","aliases":[],"title":"GlusterFS (brick, gfs3_mknod_req): A crafted mknod RPC traverses out of the volume and writes a file anywhere the brick","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"GlusterFS (brick, gfs3_mknod_req)","year":"2018","cvss_score":8.1,"severity":"high","kev":false,"impact":"A crafted mknod RPC traverses out of the volume and writes a file anywhere the brick process can reach on the server node, which leads to code execution on the storage server. From a mounted volume, one tenant reaches the host filesystem underneath every tenant's data.","attack_vector":"Any authenticated gluster client that can mount a volume and issue RPCs to a brick - i.e. any tenant compute node with the share mounted.","remediation":"Upgrade glusterfs server to the fixed release (and apply CVE-2018-14651, the follow-up that completes this fix) and restart the brick processes. Run bricks as a non-root user where your deployment supports it and keep brick ports off tenant-routable networks.","references":["https://access.redhat.com/security/cve/CVE-2018-10926","https://nvd.nist.gov/vuln/detail/CVE-2018-10926"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2018-15372","cve":"CVE-2018-15372","aliases":[],"title":"Cisco IOS XE MACsec Key Agreement (MKA over EAP-TLS): A logic error in MKA over EAP-TLS lets an unauthenticated","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Cisco IOS XE MACsec Key Agreement (MKA over EAP-TLS)","year":"2018","cvss_score":8.1,"severity":"high","kev":false,"impact":"A logic error in MKA over EAP-TLS lets an unauthenticated adjacent attacker bypass authentication and pass traffic through a Layer 3 interface. MACsec here is doing double duty as link encryption and as port admission control, and both fail — an unauthenticated device on the wire gets its traffic forwarded as though it had authenticated.","attack_vector":"Unauthenticated attacker with adjacent (same-link) access to an interface configured for MKA with EAP-TLS.","remediation":"Software upgrade plus device reload. Do not treat MACsec/MKA as your only port-admission control — pair it with per-port VLAN pinning and MAC allowlisting, which are live config changes that keep an unauthenticated device from reaching anything useful even when the MKA check fails.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-15372"],"status":"curated","tags":["tenant-isolation"],"published":"2018-10-05"},{"id":"CVE-2019-11247","cve":"CVE-2019-11247","aliases":[],"title":"Kubernetes (kube-apiserver): Cluster-scoped custom resources reachable through namespaced requests, so namespace-scoped","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-apiserver)","year":"2019","cvss_score":8.1,"severity":"high","kev":false,"impact":"Cluster-scoped custom resources reachable through namespaced requests, so namespace-scoped RBAC grants cluster-wide CR access","attack_vector":"Cluster user with namespace access","remediation":"Rolling control-plane upgrade; no GPU drain. Audit CRD RBAC afterwards","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-11247"],"status":"curated","published":"2019-08-29"},{"id":"CVE-2020-12693","cve":"CVE-2020-12693","aliases":[],"title":"Slurm: Race condition in message aggregation allows launching a process as another user","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Slurm","year":"2020","cvss_score":8.1,"severity":"high","kev":false,"impact":"Race condition in message aggregation allows launching a process as another user","attack_vector":"Any user who can submit a job","remediation":"Upgrade Slurm","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-12693"],"status":"curated","published":"2020-05-21"},{"id":"CVE-2021-21540","cve":"CVE-2021-21540","aliases":[],"title":"Dell iDRAC9: Stack overflow overwriting iDRAC configuration via oversized payloads","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC9","year":"2021","cvss_score":8.1,"severity":"high","kev":false,"impact":"Stack overflow overwriting iDRAC configuration via oversized payloads","attack_vector":"Network, authenticated","remediation":"iDRAC firmware update","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-21540"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2021-04-30"},{"id":"CVE-2021-23214","cve":"CVE-2021-23214","aliases":[],"title":"PostgreSQL: With cert/trust+clientcert auth, a MITM can inject arbitrary SQL at connection setup","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"PostgreSQL","year":"2021","cvss_score":8.1,"severity":"high","kev":false,"impact":"With cert/trust+clientcert auth, a MITM can inject arbitrary SQL at connection setup","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; enforce full-verify TLS between control-plane services","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-23214"],"status":"curated","published":"2022-03-04"},{"id":"CVE-2021-29492","cve":"CVE-2021-29492","aliases":[],"title":"Envoy: Escaped slash sequences %2F and %5C not decoded","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy","year":"2021","cvss_score":8.1,"severity":"high","kev":false,"impact":"Escaped slash sequences %2F and %5C not decoded; path-based authorization bypass","attack_vector":"Unauthenticated network","remediation":"Upgrade Envoy; enable path normalization and reject encoded slashes","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-29492"],"status":"curated","published":"2021-05-28"},{"id":"CVE-2021-3139","cve":"CVE-2021-3139","aliases":[],"title":"tcmu-runner 1.3.x - 1.5.2 (userspace backstore handler for the Linux LIO target, used by Ceph iSCSI gateways and other","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"tcmu-runner 1.3.x - 1.5.2","year":"2021","cvss_score":8.1,"severity":"high","kev":false,"impact":"xcopy_locate_udev() does not enforce transport-layer restrictions, so an XCOPY (extended copy) request can name a source or destination by path traversal. An attacker who has been legitimately given one iSCSI LUN can therefore read or write files outside it - including other tenants' LUN backing files on the same target. This is the cleanest cross-tenant storage break in this set: no memory corruption, no crash, just a normal SCSI command that the target happily executes against the wrong tenant's data. It is the same mistake as CVE-2020-28374 in a different code path, so a target that patched only the earlier one is still exposed.","attack_vector":"An authenticated tenant with a single provisioned LUN on the affected target, issuing a crafted XCOPY over the normal iSCSI data path. No privilege escalation on the target and no access to the management network required.","remediation":"Upgrade tcmu-runner past 1.5.2 - distro package update plus a restart of tcmu-runner, which briefly stalls I/O on the LUNs it backs but does not require a kernel reboot. Check what actually ships tcmu-runner in your stack: Ceph iSCSI gateway deployments and several appliance images vendor it, so the fixed version may need to come from the appliance vendor rather than the distro. Verify the fix covers both this and CVE-2020-28374; patching one code path was the original mistake. Until patched, disable XCOPY/ODX support on the target if your backstore allows it.","references":["https://www.openwall.com/lists/oss-security/2021/01/12/12","https://bugzilla.suse.com/show_bug.cgi?id=1178372","https://nvd.nist.gov/vuln/detail/CVE-2021-3139"],"status":"curated","published":"2021-01-13"},{"id":"CVE-2021-39156","cve":"CVE-2021-39156","aliases":[],"title":"Istio: Host header with a port bypasses AuthorizationPolicy host matching","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Istio","year":"2021","cvss_score":8.1,"severity":"high","kev":false,"impact":"Host header with a port bypasses AuthorizationPolicy host matching","attack_vector":"Unauthenticated network","remediation":"Rolling istiod upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-39156"],"status":"curated","published":"2021-08-24"},{"id":"CVE-2021-42387","cve":"CVE-2021-42387","aliases":[],"title":"ClickHouse: Attacker-controlled offset in the LZ4 codec","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"ClickHouse","year":"2021","cvss_score":8.1,"severity":"high","kev":false,"impact":"Attacker-controlled offset in the LZ4 codec -> heap out-of-bounds read via a malicious query","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade the telemetry/usage-metering cluster","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-42387"],"status":"curated","published":"2022-03-14"},{"id":"CVE-2021-42388","cve":"CVE-2021-42388","aliases":[],"title":"ClickHouse: Second heap out-of-bounds read in LZ4::decompressImpl reachable from a client query","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"ClickHouse","year":"2021","cvss_score":8.1,"severity":"high","kev":false,"impact":"Second heap out-of-bounds read in LZ4::decompressImpl reachable from a client query","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; require auth on the native port","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-42388"],"status":"curated","published":"2022-03-14"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:N","cwe":["CWE-522"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2021-45101","cve":"CVE-2021-45101","aliases":["HTCONDOR-2021-0003"],"title":"HTCondor (condor_schedd, condor_collector): A user with nothing more than READ access to the schedd or collector can","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HTCondor (condor_schedd, condor_collector)","year":"2021","cvss_score":8.1,"severity":"high","kev":false,"impact":"A user with nothing more than READ access to the schedd or collector can pull out secrets - using stock command-line tools, no exploit development - that let them control other users' jobs and read their data. READ is the permission level sites hand out freely for monitoring, so the attacker population is broad.","attack_vector":"Any principal granted ALLOW_READ on the schedd or collector, using ordinary condor_q and condor_status style commands.","remediation":"Upgrade to HTCondor 8.8.15, 9.0.4 or 9.1.2 and restart the daemons. Review who actually holds READ on the collector - on many pools it is set to */* for convenience.","references":["https://htcondor.org/security/vulnerabilities/HTCONDOR-2021-0003.html","https://nvd.nist.gov/vuln/detail/CVE-2021-45101"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:N","cwe":["CWE-532"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2021-45103","cve":"CVE-2021-45103","aliases":["HTCONDOR-2022-0001"],"title":"HTCondor (S3 file transfer, daemon logs and job ClassAds): Pre-signed S3 URLs for a job's input and output are written","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HTCondor (S3 file transfer, daemon logs and job ClassAds)","year":"2021","cvss_score":8.1,"severity":"high","kev":false,"impact":"Pre-signed S3 URLs for a job's input and output are written into daemon logs and into the job ad. Anyone who can read the job queue or the logs gets working credentials to that tenant's private object storage - which on a GPU cluster is the training dataset and the checkpoints.","attack_vector":"Any user who can read job ClassAds (condor_q -l on another user's job in a default pool) or who has access to daemon log files on the access point.","remediation":"Upgrade to HTCondor 9.0.10 or 9.5.1 and restart the daemons. Then rotate the S3 credentials used for job transfers and expire any outstanding pre-signed URLs, and purge or restrict the old daemon logs - the leaked URLs stay valid in the log files after you patch.","references":["https://htcondor.org/security/vulnerabilities/HTCONDOR-2022-0001.html","https://nvd.nist.gov/vuln/detail/CVE-2021-45103"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2021-47131","cve":"CVE-2021-47131","aliases":[],"title":"Linux kernel (net/tls): When a NIC with active kTLS offload goes down, the offload teardown freed the TLS context while","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/tls)","year":"2021","cvss_score":8.1,"severity":"high","kev":false,"impact":"When a NIC with active kTLS offload goes down, the offload teardown freed the TLS context while sockets were still pointing at it. If the link comes back and the connection resumes after TCP retransmits, the kernel dereferences the freed context - a use-after-free driven by nothing more exotic than a link flap on a node carrying offloaded TLS connections. This is the offload-to-software fallback seam, and on a cluster fabric link flaps are routine rather than exceptional.","attack_vector":"No attacker privilege on the node is required. Any tenant or platform connection using NIC-offloaded kTLS is enough; the trigger is a netdev down/up transition (link flap, driver reset, ethtool reconfiguration, switch-side event) with live offloaded connections, followed by data resuming after retransmission. A peer on the fabric can help by keeping the connection alive across the flap. Conditional on TLS device offload actually being enabled on the NIC - check ethtool's tls-hw-tx-offload / tls-hw-rx-offload features.","remediation":"Boot a kernel carrying the fix commits below (no fixed stable version published; this is old enough that most maintained kernels already carry it - verify rather than assume). Interim control: disable kTLS NIC offload with ethtool (tls-hw-tx-offload off, tls-hw-rx-offload off) so kTLS runs in software and the offload teardown path is never taken.","references":["https://git.kernel.org/stable/c/f1d4184f128dede82a59a841658ed40d4e6d3aa2","https://git.kernel.org/stable/c/0f1e6fe66977a864fe850522316f713d7b926fd9","https://nvd.nist.gov/vuln/detail/CVE-2021-47131"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:H/A:H","cwe":["CWE-284"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-23241","cve":"CVE-2022-23241","aliases":[],"title":"NetApp ONTAP SnapLock on FlexGroup volumes: An authenticated remote user modifies or deletes WORM-locked data before","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"NetApp ONTAP SnapLock on FlexGroup volumes","year":"2022","cvss_score":8.1,"severity":"high","kev":false,"impact":"An authenticated remote user modifies or deletes WORM-locked data before its retention expires. The immutability guarantee that backups and compliance copies of training data rely on is simply not there.","attack_vector":"Any authenticated remote account on a Clustered Data ONTAP 9.11.1 through 9.11.1P2 system with SnapLock-configured FlexGroups.","remediation":"Upgrade to 9.11.1P3 or later. Then re-verify the retention state of every SnapLock FlexGroup - the fix stops new tampering but does not restore anything already deleted.","references":["https://security.netapp.com/advisory/ntap-20221017-0001/","https://nvd.nist.gov/vuln/detail/CVE-2022-23241"],"status":"curated"},{"id":"CVE-2022-28733","cve":"CVE-2022-28733","aliases":[],"title":"GRUB2 (net/ip IPv4 reassembly): Integer underflow in grub_net_recv_ip4_packets from a crafted IP packet","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (net/ip IPv4 reassembly)","year":"2022","cvss_score":8.1,"severity":"high","kev":false,"impact":"Integer underflow in grub_net_recv_ip4_packets from a crafted IP packet. This one matters far more than the filesystem bugs for a GPU cloud, because it is reachable over the network during PXE boot - an attacker who can answer on the provisioning VLAN owns the node before any OS, tenant, or agent exists.","attack_vector":"Anyone who can put packets on the provisioning/PXE network while a node is netbooting. No credentials, no prior access to the node.","remediation":"grub2 package update + reboot, and update the netboot GRUB image you actually serve - patching running nodes does nothing if the TFTP/HTTP-served binary is stale. Compensating control: put provisioning on an isolated L2 segment with DHCP snooping, and do not let tenant workloads share it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28733","https://access.redhat.com/security/cve/CVE-2022-28733"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2023-07-20"},{"id":"CVE-2022-42272","cve":"CVE-2022-42272","aliases":[],"title":"DGX servers BMC: RCE on BMC (buffer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX servers BMC","year":"2022","cvss_score":8.1,"severity":"high","kev":false,"impact":"RCE on BMC (buffer overflow)","attack_vector":"Network-adjacent","remediation":"Flash BMC 2.09.00+ out-of-band","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42272","https://github.com/NVIDIA/product-security/tree/main/2024/5513"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-120"],"published":"2023-01-12"},{"id":"CVE-2022-42273","cve":"CVE-2022-42273","aliases":[],"title":"DGX servers BMC: RCE on BMC (buffer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX servers BMC","year":"2022","cvss_score":8.1,"severity":"high","kev":false,"impact":"RCE on BMC (buffer overflow)","attack_vector":"Network-adjacent","remediation":"Flash BMC 2.09.00+ out-of-band","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42273","https://github.com/NVIDIA/product-security/tree/main/2024/5513"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:N","cwe":["CWE-120"],"published":"2023-01-12"},{"id":"CVE-2023-25409","cve":"CVE-2023-25409","aliases":[],"title":"ATEN PE8108 switched PDU: A restricted (non-admin) user account on the PDU's web interface can control outlets","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ATEN PE8108 switched PDU","year":"2023","cvss_score":8.1,"severity":"high","kev":false,"impact":"A restricted (non-admin) user account on the PDU's web interface can control outlets belonging to other users — meaning one tenant sharing this PDU with others can power-cycle or power-off outlets feeding another tenant's equipment, not just their own.","attack_vector":"Requires only a low-privileged, restricted user account on the PDU — no admin credentials needed to reach outlets outside the account's assigned scope.","remediation":"Firmware upgrade from ATEN to a version that enforces per-outlet authorization correctly. Flash each PDU; since this is the power path for the racks it feeds, coordinate the maintenance window with anyone whose equipment is on that PDU.","references":["https://www.pentagrid.ch/en/blog/multiple-vulnerabilities-in-aten-PE8108-power-distribution-unit"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2023-04-11"},{"id":"CVE-2023-25552","cve":"CVE-2023-25552","aliases":["SEVD-2023-101-04"],"title":"Schneider Electric StruxureWare Data Center Expert (V7.9.2 and prior) - Device File Transfer settings: Missing","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Schneider Electric StruxureWare Data Center Expert (V7.9.2 and prior) - Device File Transfer settings","year":"2023","cvss_score":8.1,"severity":"high","kev":false,"impact":"Missing authorisation on the Device File Transfer settings lets an attacker view, change or delete content and invoke functions they should not have. Device File Transfer is how DCE pushes firmware and config to managed power and cooling devices - so control of it is control of what firmware lands on your UPS and PDU fleet.","attack_vector":"Remote access to the DCE endpoints, without the authorisation the function should require.","remediation":"Upgrade past V7.9.2. Independently, gate firmware distribution: a DCIM system that can silently push firmware to power devices is a single point of catastrophic failure, and it deserves change control rather than trust.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25552"],"status":"curated","published":"2023-04-18"},{"id":"CVE-2023-27493","cve":"CVE-2023-27493","aliases":[],"title":"Envoy: Request properties are not escaped when generating request headers","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy","year":"2023","cvss_score":8.1,"severity":"high","kev":false,"impact":"Request properties are not escaped when generating request headers; header injection into upstreams","attack_vector":"Unauthenticated network","remediation":"Upgrade Envoy","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-27493"],"status":"curated","published":"2023-04-04"},{"id":"CVE-2023-31424","cve":"CVE-2023-31424","aliases":[],"title":"Brocade SANnav Management Portal web interface, before v2.3.0 and v2.2.2a: Remote unauthenticated users can bypass web","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Brocade SANnav Management Portal web interface, before v2.3.0 and v2.2.2a","year":"2023","cvss_score":8.1,"severity":"high","kev":false,"impact":"Remote unauthenticated users can bypass web authentication and authorization on the SANnav portal. That is the front door to the whole FC management estate - fabric inventory, zoning pushes, switch credentials, firmware distribution. Chained with the zone-management SQL injection above it turns a fully unauthenticated network position into control of tenant isolation across every managed fabric.","attack_vector":"Any host with network reachability to the SANnav web interface. No credentials at all.","remediation":"Upgrade SANnav to 2.3.0 or 2.2.2a. Management-plane upgrade only - no fabric or array disruption. Because the pre-fix window allowed unauthenticated access, also rotate stored switch credentials and diff the live zonesets against your intended configuration rather than assuming the upgrade closes the incident.","references":["https://support.broadcom.com/web/ecx/support-content-notification/-/external/content/SecurityAdvisories/0/22507","https://security.netapp.com/advisory/ntap-20240229-0004/","https://nvd.nist.gov/vuln/detail/CVE-2023-31424"],"status":"curated","published":"2023-08-31"},{"id":"CVE-2023-34336","cve":"CVE-2023-34336","aliases":["AMI-SA-2023005","NVIDIA OSR review"],"title":"AMI MegaRAC SPx (IPMI handler): Buffer overflow in the BMC's IPMI message handler leading to code execution","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx (IPMI handler)","year":"2023","cvss_score":8.1,"severity":"high","kev":false,"impact":"Buffer overflow in the BMC's IPMI message handler leading to code execution or privilege escalation inside the BMC. This is the pre-Redfish legacy protocol that almost every fleet still leaves enabled for ipmitool-based power control and sensor scraping, so the exposed surface is usually larger than operators assume. A win here means the attacker controls power, console and firmware update paths for the node.","attack_vector":"Network-reachable IPMI service, no credentials required per AMI's own CVSS vector, but high attack complexity. Any host that can send IPMI RMCP+ traffic to UDP/623 on the BMC is in range - which is every host on the management VLAN, and in badly-built estates any host that can route there.","remediation":"Firmware flash to SPx_12.7 / SPx_13.5, out-of-band per node, gated on ODM rebase. Unlike the network-stack bugs in this cluster there IS a meaningful config-only mitigation here: disable IPMI-over-LAN entirely and drive power/sensors through Redfish instead. That is a config change on the BMC, no reboot, no flash - but it breaks any ipmitool-based tooling in your provisioning and monitoring stack, so cost it as a tooling migration.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2023005.pdf","https://nvd.nist.gov/vuln/detail/CVE-2023-34336"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-06-12"},{"cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-362"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2023-41915","cve":"CVE-2023-41915","aliases":[],"title":"OpenPMIx (PMIx library used by Slurm and Open MPI for job launch): A race in PMIx library code that executes with UID 0","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"OpenPMIx (PMIx library used by Slurm and Open MPI for job launch)","year":"2023","cvss_score":8.1,"severity":"high","kev":false,"impact":"A race in PMIx library code that executes with UID 0 lets a local attacker take ownership of arbitrary files on the node. PMIx is the wire-up layer between the scheduler and MPI ranks, so this sits directly in the multi-node launch path of every distributed training job on the cluster.","attack_vector":"A local user on a node where a PMIx-using launcher runs with root privilege - which is the normal configuration for Slurm's PMIx MPI plugin and for Open MPI's runtime.","remediation":"Upgrade OpenPMIx to 4.2.6 or 5.0.1 across compute nodes and restart the launcher daemons. Note this is a separate package from Slurm - upgrading Slurm alone does not fix it, and sites frequently miss that because PMIx arrives as a distro dependency.","references":["https://github.com/openpmix/openpmix/releases/tag/v4.2.6","https://docs.openpmix.org/en/latest/security.html","https://nvd.nist.gov/vuln/detail/CVE-2023-41915"],"status":"curated"},{"id":"CVE-2023-4606","cve":"CVE-2023-4606","aliases":["LEN-140960"],"title":"Lenovo XClarity Controller (XCC) - user account API: A read-only XCC user can change any other user's password through","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lenovo XClarity Controller (XCC) - user account API","year":"2023","cvss_score":8.1,"severity":"high","kev":false,"impact":"A read-only XCC user can change any other user's password through a crafted API call. That is a direct path from the least-privileged BMC account you hand out to full administrative control of the service processor: change the admin's password, log in as admin, and you have power control, remote media, console and firmware on the node. It also locks out the legitimate administrator, which turns a quiet compromise into a visible outage. Affects ThinkSystem V2 and V3 servers - the generations that carry the SR670 V2 / SR675 V3 / SR685a GPU platforms. V1 servers are not affected.","attack_vector":"An authenticated XCC account holding only read-only permission - typically a monitoring collector, a DCIM integration, or an account issued to remote hands. Reachable over the out-of-band management VLAN.","remediation":"Flash XCC to the per-model version listed in Lenovo's advisory - out-of-band, per-node, no host reboot and no drain of running jobs. Model-specific version floors mean you cannot use one target build across a mixed fleet; pull the table from LEN-140960 and drive the campaign per SKU. Config-only mitigation in the meantime: audit and prune read-only XCC accounts, since 'read-only' provides no protection against this bug.","references":["https://support.lenovo.com/us/en/product_security/LEN-140960","https://nvd.nist.gov/vuln/detail/CVE-2023-4606"],"status":"curated","published":"2023-10-25"},{"id":"CVE-2023-6572","cve":"CVE-2023-6572","aliases":[],"title":"Gradio: Command injection","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Gradio","year":"2023","cvss_score":8.1,"severity":"high","kev":false,"impact":"Command injection","attack_vector":"Network user of the demo","remediation":"Upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-6572"],"status":"curated","published":"2023-12-14"},{"id":"CVE-2024-0114","cve":"CVE-2024-0114","aliases":[],"title":"Hopper HGX 8-GPU (HMC/firmware): Improper input validation in baseboard firmware","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Hopper HGX 8-GPU (HMC/firmware)","year":"2024","cvss_score":8.1,"severity":"high","kev":false,"impact":"Improper input validation in baseboard firmware -> code exec / tampering across the GPU baseboard","attack_vector":"Local privileged host access / mgmt path","remediation":"Flash HGX baseboard firmware out-of-band; full node drain and power cycle","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0114","https://github.com/NVIDIA/product-security/tree/main/2025/5561"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:C/C:L/I:H/A:H","cwe":["CWE-1244"],"fleet":{"pain_class":"node-reboot"},"published":"2025-03-05"},{"id":"CVE-2024-10220","cve":"CVE-2024-10220","aliases":[],"title":"Kubernetes (kubelet): Arbitrary command execution on the node via a gitRepo volume","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kubelet)","year":"2024","cvss_score":8.1,"severity":"high","kev":false,"impact":"Arbitrary command execution on the node via a gitRepo volume","attack_vector":"Cluster user able to create a pod with a gitRepo volume","remediation":"Rolling kubelet upgrade with node drain; block the deprecated gitRepo volume type by admission policy","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2024-11-22"},{"id":"CVE-2024-22273","cve":"CVE-2024-22273","aliases":[],"title":"VMware ESXi / Workstation / Fusion (storage controller out-of-bounds read/write): A malicious actor inside a VM","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"VMware ESXi / Workstation / Fusion (storage controller out-of-bounds read/write)","year":"2024","cvss_score":8.1,"severity":"high","kev":false,"impact":"A malicious actor inside a VM with storage controllers enabled triggers an out-of-bounds read/write and, chained with other issues, executes code on the hypervisor. That is a guest-to-host escape - the boundary your entire multi-tenant model rests on.","attack_vector":"A tenant with control of a VM on the host. Requires storage controllers enabled, which is the default.","remediation":"Patch ESXi and reboot the host. Hosts with GPUs in passthrough cannot be live-migrated the way ordinary VMs can, so this is a drain-and-reboot per host with real workload downtime - plan it as a rolling maintenance pass across the cluster. Prioritise above ordinary ESXi patches: escape-class bugs invalidate tenant isolation, and on GPU hosts the co-tenants are high-value.","references":["https://support.broadcom.com/web/ecx/support-content-notification/-/external/content/SecurityAdvisories/0/24308"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-28088","cve":"CVE-2024-28088","aliases":[],"title":"LangChain: Directory traversal via the template path parameter","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LangChain","year":"2024","cvss_score":8.1,"severity":"high","kev":false,"impact":"Directory traversal via the template path parameter","attack_vector":"Attacker-controlled final path segment","remediation":"Upgrade past 0.1.10","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-28088"],"status":"curated","published":"2024-03-04"},{"id":"CVE-2024-28233","cve":"CVE-2024-28233","aliases":[],"title":"JupyterHub: Malicious subdomain tricks a user","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"JupyterHub","year":"2024","cvss_score":8.1,"severity":"high","kev":false,"impact":"Malicious subdomain tricks a user → session takeover","attack_vector":"User visiting an attacker-controlled subdomain","remediation":"Upgrade; per-user subdomain isolation is the mitigation and the vuln","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-28233"],"status":"curated","published":"2024-03-27"},{"id":"CVE-2024-36623","cve":"CVE-2024-36623","aliases":[],"title":"Docker / moby: Race condition in the streamformatter package causing data corruption or daemon crash","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2024","cvss_score":8.1,"severity":"high","kev":false,"impact":"Race condition in the streamformatter package causing data corruption or daemon crash","attack_vector":"Any tenant workload driving concurrent daemon streams","remediation":"Upgrade moby; daemon restart","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36623"],"status":"curated","published":"2024-11-29"},{"cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-824"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-43878","cve":"CVE-2024-43878","aliases":[],"title":"Linux kernel (net/xfrm): The error path of xfrm_input leaves the secpath entry pointing at poisoned memory, and the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2024","cvss_score":8.1,"severity":"high","kev":false,"impact":"The error path of xfrm_input leaves the secpath entry pointing at poisoned memory, and the receive callback then dereferences it - KASAN reports a wild access at 0x6b6b6b6b in xfrmi_rcv_cb while handling an inbound ESP packet. An attacker who can make inbound state lookup fail gets the node reading attacker-influenced freed memory in the ESP receive path.","attack_vector":"Remote and pre-authentication: an ESP packet from anywhere on the fabric that hits a misconfigured or mismatched input state is enough - the report reproduces it with a plain ping over an xfrm interface. No local access, no device node, no valid SA required, since the trigger is precisely the failed-state path.","remediation":"Boot a kernel carrying the linked stable commits. Interim: filter ESP to known peer addresses at the node's ingress so unmatched inbound SPIs never reach xfrm_input.","references":["https://git.kernel.org/stable/c/a4c10813bc394ff2b5c61f913971be216f8f8834","https://git.kernel.org/stable/c/54fcc6189dfb822eea984fa2b3e477a02447279d","https://nvd.nist.gov/vuln/detail/CVE-2024-43878"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-4888","cve":"CVE-2024-4888","aliases":[],"title":"LiteLLM: Arbitrary file deletion via `/audio/transcriptions`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LiteLLM","year":"2024","cvss_score":8.1,"severity":"high","kev":false,"impact":"Arbitrary file deletion via `/audio/transcriptions`","attack_vector":"Network user of the proxy","remediation":"Upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-4888"],"status":"curated","published":"2024-06-06"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:N","cwe":["CWE-345"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2024-48916","cve":"CVE-2024-48916","aliases":[],"title":"Ceph RADOS Gateway (RGW): RGW accepts a JWT whose header declares alg \"none\" and never checks the signature, so anyone","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ceph RADOS Gateway (RGW)","year":"2024","cvss_score":8.1,"severity":"high","kev":false,"impact":"RGW accepts a JWT whose header declares alg \"none\" and never checks the signature, so anyone who can reach the gateway can mint a token claiming to be any OIDC identity. That is a full authentication bypass on the S3 endpoint - the attacker assumes another tenant's role and reads or writes their buckets.","attack_vector":"Any client that can open a TCP connection to the RGW S3/STS endpoint. Only affects clusters with OIDC/STS AssumeRoleWithWebIdentity configured, but there is no valid credential requirement at all.","remediation":"Upgrade RGW to a release past 19.2.3 that carries the fix and restart every radosgw daemon. Until then disable the OIDC/STS web-identity provider on the gateway, or front it with a proxy that rejects tokens whose alg header is not the one your IdP actually signs with.","references":["https://github.com/ceph/ceph/security/advisories/GHSA-5g9m-mmp6-93mq","https://nvd.nist.gov/vuln/detail/CVE-2024-48916"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-415"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-50215","cve":"CVE-2024-50215","aliases":[],"title":"Linux kernel NVMe target authentication (nvmet-auth DH group setup): Ctrl","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel NVMe target authentication (nvmet-auth DH group setup)","year":"2024","cvss_score":8.1,"severity":"high","kev":false,"impact":"Ctrl->dh_key is freed on the error path of nvmet_setup_dhgroup() but not nulled, and it survives across repeated calls for the same controller, so nvmet_destroy_auth() frees it a second time. The bug lives in the DH-HMAC-CHAP negotiation - the code that decides whether a connecting initiator is who it claims to be - and a remote initiator drives it by repeating a failing DH group negotiation. Corrupting the heap from inside the authentication handshake is the worst possible place for it, because it is reachable before the handshake grants anything.","attack_vector":"Remote, pre-authentication. An initiator repeatedly negotiates an invalid DH group against the target.","remediation":"Kernel update nulling dh_key after kfree_sensitive(). If in-band authentication is not required on a given subsystem, disabling DH-HMAC-CHAP removes the path; if it is required, patch the target nodes before widening initiator reachability.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=c60af16e1d6cc2237d58336546d6adfc067b6b8f","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2024/CVE-2024-50215.json"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-5154","cve":"CVE-2024-5154","aliases":[],"title":"CRI-O: Malicious container creates a symlink via directory traversal and gets arbitrary host read/write","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"CRI-O","year":"2024","cvss_score":8.1,"severity":"high","kev":false,"impact":"Malicious container creates a symlink via directory traversal and gets arbitrary host read/write","attack_vector":"Any tenant workload / malicious image","remediation":"Upgrade CRI-O on all nodes; drain and recreate pods","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-5154"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2024-06-12"},{"id":"CVE-2024-53673","cve":"CVE-2024-53673","aliases":[],"title":"HPE Insight Remote Support (Java deserialization): Java deserialization letting an unauthenticated attacker execute","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HPE Insight Remote Support (Java deserialization)","year":"2024","cvss_score":8.1,"severity":"high","kev":false,"impact":"Java deserialization letting an unauthenticated attacker execute code on the Insight RS server.","attack_vector":"Unauthenticated network access.","remediation":"Patch per HPESBGN04731 alongside the other IRS issues in the same bulletin.","references":["https://support.hpe.com/hpesc/public/docDisplay?docId=hpesbgn04731en_us&docLocale=en_US"],"status":"curated"},{"id":"CVE-2024-6387","cve":"CVE-2024-6387","aliases":[],"title":"OpenSSH (sshd): regreSSHion: signal-handler race in sshd giving unauthenticated remote root on glibc Linux","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"OpenSSH (sshd)","year":"2024","cvss_score":8.1,"severity":"high","kev":false,"impact":"regreSSHion: signal-handler race in sshd giving unauthenticated remote root on glibc Linux","attack_vector":"Unauthenticated network","remediation":"Package update + sshd restart; no reboot. Interim mitigation `LoginGraceTime 0` (costs DoS resilience). Highest-priority item for any tenant-reachable bastion or management SSH endpoint","references":["https://access.redhat.com/security/cve/CVE-2024-6387"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2024-07-01"},{"id":"CVE-2025-1094","cve":"CVE-2025-1094","aliases":[],"title":"PostgreSQL (libpq): Improper quoting in PQescape*","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"PostgreSQL (libpq)","year":"2025","cvss_score":8.1,"severity":"high","kev":false,"impact":"Improper quoting in PQescape* -> SQL injection, chainable to shell via psql \\!","attack_vector":"Network (remote)","remediation":"Control-plane: patch metadata/billing DB servers and all libpq clients","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-1094"],"status":"curated","published":"2025-02-13"},{"id":"CVE-2025-14279","cve":"CVE-2025-14279","aliases":[],"title":"MLflow (REST API): DNS rebinding — no Origin header validation","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"MLflow (REST API)","year":"2025","cvss_score":8.1,"severity":"high","kev":false,"impact":"DNS rebinding — no Origin header validation","attack_vector":"Operator's browser visiting a malicious page while on the cluster network","remediation":"Upgrade past 3.4.0","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-14279"],"status":"curated","published":"2026-01-12"},{"id":"CVE-2025-23318","cve":"CVE-2025-23318","aliases":[],"title":"NVIDIA Triton (Python backend): Out-of-bounds write in the Python backend","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"NVIDIA Triton (Python backend)","year":"2025","cvss_score":8.1,"severity":"high","kev":false,"impact":"Out-of-bounds write in the Python backend","attack_vector":"Unauthenticated network","remediation":"Patch","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23318"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-805"],"published":"2025-08-06"},{"id":"CVE-2025-23319","cve":"CVE-2025-23319","aliases":[],"title":"NVIDIA Triton (Python backend shared memory): Out-of-bounds write in the Python backend","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"NVIDIA Triton (Python backend shared memory)","year":"2025","cvss_score":8.1,"severity":"high","kev":false,"impact":"Out-of-bounds write in the Python backend","attack_vector":"Unauthenticated network","remediation":"Patch; the Python backend's shared-memory region is the pivot in the Wiz chain","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23319"],"status":"curated","fleet":{"ubiquity":"Very common - Triton is the default multi-model serving layer inside NIM and in many neocloud managed-inference products","remediation_pain":"`daemon-restart` (upgrade to 25.07, restart the serving process) - no host reboot, but it is a customer-visible endpoint outage","pain_class":"node-reboot","why_fleet_wide":"Unauthenticated remote chain: an oversized request leaks the internal shared-memory key, then an OOB write in the Python backend yields full RCE as the Triton process - i.e. theft of every model and every inference request on that server"},"cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-805"],"published":"2025-08-06"},{"id":"CVE-2025-27086","cve":"CVE-2025-27086","aliases":[],"title":"HPE Performance Cluster Manager (HPCM) GUI authentication bypass: Authentication bypass in the HPCM web GUI","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HPE Performance Cluster Manager (HPCM) GUI authentication bypass","year":"2025","cvss_score":8.1,"severity":"high","kev":false,"impact":"Authentication bypass in the HPCM web GUI. HPCM provisions and manages HPC/AI cluster nodes, so bypassing its authentication gives an attacker the ability to reimage and reconfigure compute nodes at will.","attack_vector":"Unauthenticated network access to the HPCM GUI. High attack complexity.","remediation":"Apply the HPCM update per HPESBCR04842. Management-server upgrade. Keep HPCM on an isolated provisioning network - it has PXE/imaging authority over the whole cluster.","references":["https://support.hpe.com/hpesc/public/docDisplay?docId=hpesbcr04842en_us&docLocale=en_US"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:N","cwe":["CWE-863"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2025-30093","cve":"CVE-2025-30093","aliases":["HTCONDOR-2025-0001"],"title":"HTCondor (IDToken authorization restrictions): The per-token authorization restrictions attached with","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HTCondor (IDToken authorization restrictions)","year":"2025","cvss_score":8.1,"severity":"high","kev":false,"impact":"The per-token authorization restrictions attached with condor_token_create -authz are not enforced. A token you deliberately minted as read-only, or scoped to one operation, performs everything the underlying identity is entitled to. Sites that use restricted tokens to hand limited access to a partner or a CI system have a boundary that does not exist.","attack_vector":"Any holder of an IDToken that the target daemon can validate. The attacker does not need to be permitted by the daemon's configured authorization policy for the restriction bypass itself.","remediation":"Upgrade to HTCondor 23.0.22, 23.10.22, 24.0.6 or 24.6.1 and restart the daemons. Then inventory every token issued with -authz, or approved via condor_token_request_approve, and reissue them - each one has been operating as an unrestricted token for its identity.","references":["https://htcondor.org/security/vulnerabilities/HTCONDOR-2025-0001.html","https://nvd.nist.gov/vuln/detail/CVE-2025-30093"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-3935","cve":"CVE-2025-3935","aliases":[],"title":"ConnectWise ScreenConnect: ViewState code injection","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"ConnectWise ScreenConnect","year":"2025","cvss_score":8.1,"severity":"high","kev":true,"impact":"ViewState code injection -> RCE once ASP.NET machine keys are obtained","attack_vector":"Network (remote)","remediation":"Control-plane: patch + rotate ASP.NET machine keys","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-3935"],"status":"curated","published":"2025-04-25"},{"cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-40168","cve":"CVE-2025-40168","aliases":[],"title":"Linux kernel (net/smc): The CLC prefix-match check on the listen path dereferences the destination cache entry's","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/smc)","year":"2025","cvss_score":8.1,"severity":"high","kev":false,"impact":"The CLC prefix-match check on the listen path dereferences the destination cache entry's net_device outside RCU and outside RTNL, so a device being torn down concurrently leaves the handshake reading freed memory. A peer connecting during any interface churn - VF teardown, bond member removal, container netns exit - gets a use-after-free on the server side of the SMC handshake.","attack_vector":"Remote-driven: smc_clc_prfx_match() runs inside smc_listen_work while processing an inbound CLC proposal, so an unauthenticated connecting peer supplies the timing. The race partner is ordinary netdev churn, which is constant on a multi-tenant node (per-container veths, SR-IOV VFs coming and going). Requires the smc module loaded, which any unprivileged socket(AF_SMC, ...) achieves.","remediation":"Boot a kernel carrying the fix commits (uses __sk_dst_get()/dst_dev_rcu() under RCU). Interim: blacklist the smc module on nodes not using SMC-R, and avoid exposing SMC listeners to untrusted peers.","references":["https://git.kernel.org/stable/c/d26e80f7fb62d77757b67a1b94e4ac756bc9c658","https://git.kernel.org/stable/c/235f81045c008169cc4e1955b4a64e118eebe61b","https://nvd.nist.gov/vuln/detail/CVE-2025-40168"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:H/A:H","cwe":["CWE-22","CWE-23"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2025-62156","cve":"CVE-2025-62156","aliases":["GHSA-p84v-gxvw-73pf"],"title":"Argo Workflows (workflow executor, artifact unpack path handling): Archive entries with traversal paths escape the","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo Workflows (workflow executor, artifact unpack path handling)","year":"2025","cvss_score":8.1,"severity":"high","kev":false,"impact":"Archive entries with traversal paths escape the temporary unpack directory during artifact extraction, letting a crafted archive overwrite files outside it - including paths under /etc in the executor. Because the executor carries the workflow's Kubernetes identity, a poisoned input artifact turns into control over a component that other tenants' workflows also run through.","attack_vector":"Any user able to submit a workflow that pulls an attacker-controlled input artifact, or able to write into a shared artifact repository path.","remediation":"Upgrade to 3.6.12 or 3.7.3, then immediately to 3.6.14 / 3.7.5 because the initial fix was bypassable via symlinks (CVE-2025-66626). Restart the controller so new workflow pods pick up the corrected executor.","references":["https://github.com/argoproj/argo-workflows/security/advisories/GHSA-p84v-gxvw-73pf","https://nvd.nist.gov/vuln/detail/CVE-2025-62156"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:N","cwe":["CWE-863"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2025-62506","cve":"CVE-2025-62506","aliases":[],"title":"MinIO (service accounts / STS session policies): The session policy attached to a service account or STS credential is","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"MinIO (service accounts / STS session policies)","year":"2025","cvss_score":8.1,"severity":"high","kev":false,"impact":"The session policy attached to a service account or STS credential is not enforced, so a credential that was deliberately scoped down to one prefix operates with the full rights of its parent identity. A key you handed a tenant job expecting it to see one bucket can read and write everything the parent can.","attack_vector":"Any holder of a MinIO service account or STS credential - typically every tenant workload, since scoped-down keys are the normal way to hand out access.","remediation":"Upgrade MinIO to the release in GHSA-jjjj-jwhf-8rgr and restart. Then re-issue every service account and STS credential that relied on a session policy for isolation, and back the boundary with distinct parent users per tenant rather than session policies alone.","references":["https://github.com/minio/minio/security/advisories/GHSA-jjjj-jwhf-8rgr","https://nvd.nist.gov/vuln/detail/CVE-2025-62506"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:H/A:H","cwe":["CWE-23","CWE-59","CWE-78"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2025-66626","cve":"CVE-2025-66626","aliases":["GHSA-xrqc-7xgx-c9vh"],"title":"Argo Workflows (workflow executor, tar extraction symlink handling): The patch for CVE-2025-62156 missed symlinks, so a","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo Workflows (workflow executor, tar extraction symlink handling)","year":"2025","cvss_score":8.1,"severity":"high","kev":false,"impact":"The patch for CVE-2025-62156 missed symlinks, so a crafted input artifact still escapes the extraction directory and writes anywhere the executor can reach. The executor container holds the workflow service account token and runs alongside the user's container, so an arbitrary write there converts a data-plane artifact into code execution inside Argo's own executor identity.","attack_vector":"Any user who can submit a workflow that consumes an attacker-supplied artifact, or who can plant an archive in a repository another tenant's workflow pulls from.","remediation":"Upgrade to 3.6.14 or 3.7.5 and restart the controller so new pods get the fixed executor image. Do not stop at 3.6.12/3.7.3 - that is the incomplete fix for CVE-2025-62156.","references":["https://github.com/argoproj/argo-workflows/security/advisories/GHSA-xrqc-7xgx-c9vh","https://nvd.nist.gov/vuln/detail/CVE-2025-66626"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-71340","cve":"CVE-2025-71340","aliases":[],"title":"picklescan: Misses `idlelib.pyshell.ModifiedInterpreter.runcode` gadget","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"picklescan","year":"2025","cvss_score":8.1,"severity":"high","kev":false,"impact":"Misses `idlelib.pyshell.ModifiedInterpreter.runcode` gadget","attack_vector":"Customer-supplied pickle","remediation":"Upgrade; the denylist approach is structurally incomplete","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-71340"],"status":"curated","published":"2026-06-25"},{"id":"CVE-2025-71342","cve":"CVE-2025-71342","aliases":[],"title":"picklescan: Misses `idlelib.run.Executive.runcode` gadget","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"picklescan","year":"2025","cvss_score":8.1,"severity":"high","kev":false,"impact":"Misses `idlelib.run.Executive.runcode` gadget","attack_vector":"Customer-supplied pickle","remediation":"Upgrade to 0.0.30+; treat gadget denylists as best-effort only","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-71342"],"status":"curated","published":"2026-07-04"},{"id":"CVE-2026-11816","cve":"CVE-2026-11816","aliases":[],"title":"Keras (archive extraction utils): Path traversal in `keras/src/utils/file_utils.py`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Keras (archive extraction utils)","year":"2026","cvss_score":8.1,"severity":"high","kev":false,"impact":"Path traversal in `keras/src/utils/file_utils.py`","attack_vector":"Customer-supplied archive","remediation":"Upgrade to 3.14.0+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-11816"],"status":"curated","published":"2026-06-11"},{"id":"CVE-2026-18577","cve":"CVE-2026-18577","aliases":[],"title":"N-able N-central: Incomplete patch for CVE-2026-18556","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"N-able N-central","year":"2026","cvss_score":8.1,"severity":"high","kev":true,"impact":"Incomplete patch for CVE-2026-18556 -> auth bypass and full account takeover","attack_vector":"Network (remote)","remediation":"Control-plane: apply the follow-up patch; audit N-central accounts created since","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-18577"],"status":"curated","published":"2026-08-02"},{"id":"CVE-2026-22719","cve":"CVE-2026-22719","aliases":[],"title":"VMware Aria Operations (command injection during assisted migration): An unauthenticated attacker injects commands","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"VMware Aria Operations (command injection during assisted migration)","year":"2026","cvss_score":8.1,"severity":"high","kev":true,"impact":"An unauthenticated attacker injects commands and reaches remote code execution on Aria Operations while a support-assisted product migration is running. The exposure window is operational rather than permanent - it opens exactly when you are mid-migration and least able to respond.","attack_vector":"Unauthenticated network access during a support-assisted migration window.","remediation":"Apply the patches in Broadcom advisory 36947 before undertaking any assisted migration. If a migration is already in flight, restrict network access to the Aria Operations appliance for its duration.","references":["https://support.broadcom.com/web/ecx/support-content-notification/-/external/content/SecurityAdvisories/0/36947"],"status":"curated"},{"id":"CVE-2026-24218","cve":"CVE-2026-24218","aliases":[],"title":"DGX Spark: Hardcoded credentials in config","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX Spark","year":"2026","cvss_score":8.1,"severity":"high","kev":false,"impact":"Hardcoded credentials in config -> unauthorized access","attack_vector":"Local or network attacker with config access","remediation":"Update DGX Spark software; rotate all affected credentials","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24218","https://github.com/NVIDIA/product-security/tree/main/2026/5835"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-321"],"published":"2026-05-20"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:N","cwe":["CWE-863"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2026-42296","cve":"CVE-2026-42296","aliases":["GHSA-3775-99mw-8rp4"],"title":"Argo Workflows (controller, hostNetwork / securityContext / serviceAccountName merge path): The first fix for","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo Workflows (controller, hostNetwork / securityContext / serviceAccountName merge path)","year":"2026","cvss_score":8.1,"severity":"high","kev":false,"impact":"The first fix for CVE-2026-31892 gated only on podSpecPatch, leaving hostNetwork, securityContext and serviceAccountName free to flow from the tenant's Workflow through the merge into the pod. Setting hostNetwork true puts a tenant pod on the node's network namespace, and overriding serviceAccountName lets it borrow a more privileged identity - both are direct escapes from the tenant's slice of a shared GPU node.","attack_vector":"Any user who can submit a Workflow that references a hardened template, on a controller running templateReferencing Strict or Secure.","remediation":"Upgrade the controller to 3.7.14 or 4.0.5 and restart, then continue to 3.7.15 / 4.0.6 to also cover CVE-2026-54526. Enforce hostNetwork and serviceAccountName restrictions with Pod Security Admission and an admission policy rather than relying on the controller's allow-list.","references":["https://github.com/argoproj/argo-workflows/security/advisories/GHSA-3775-99mw-8rp4","https://nvd.nist.gov/vuln/detail/CVE-2026-42296"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-43629","cve":"CVE-2026-43629","aliases":[],"title":"llama-server (KV cache state restore): Heap buffer overflow in `state_read_data`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"llama-server (KV cache state restore)","year":"2026","cvss_score":8.1,"severity":"high","kev":false,"impact":"Heap buffer overflow in `state_read_data`","attack_vector":"Attacker-supplied KV/session state file restored by the server","remediation":"Rebuild; session-state restore is an untrusted-input parser","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43629"],"status":"curated","published":"2026-08-06"},{"id":"CVE-2026-43632","cve":"CVE-2026-43632","aliases":[],"title":"llama-server (tokenization endpoints): Use-after-free across six tokenization endpoints","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"llama-server (tokenization endpoints)","year":"2026","cvss_score":8.1,"severity":"high","kev":false,"impact":"Use-after-free across six tokenization endpoints","attack_vector":"Unauthenticated network to llama-server","remediation":"Rebuild past b9060","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43632"],"status":"curated","published":"2026-08-06"},{"id":"CVE-2026-49121","cve":"CVE-2026-49121","aliases":["AITER pickle RCE"],"title":"AI Tensor Engine for ROCm (AITER) - MessageQueue.recv() in shm_broadcast.py: AITER's MessageQueue.recv() deserialises","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"AI Tensor Engine for ROCm (AITER) - MessageQueue.recv() in shm_broadcast.py","year":"2026","cvss_score":8.1,"severity":"high","kev":false,"impact":"AITER's MessageQueue.recv() deserialises whatever arrives on a ZeroMQ SUB socket with Python pickle, so anyone who can reach that socket gets **unauthenticated remote code execution** as the process running the ROCm inference or training job. This is the highest-value AMD-stack finding for an AI operator in the whole set: it needs no local access, no privilege and no GPU device handle - just network reach to a port that distributed ROCm jobs open between workers. If your tensor-parallel or pipeline-parallel workers talk over an unauthenticated ZMQ fabric on a shared cluster network, another tenant on that network owns your jobs.","attack_vector":"Network. Unauthenticated. The attacker needs only IP reachability to the ZMQ SUB socket that AITER opens for shared-memory broadcast between distributed workers. On a flat cluster network - which is the norm for RDMA/RoCE training fabrics - that means any other tenant on the fabric. Affects AITER through 0.1.14.","remediation":"Upgrade AITER past 0.1.14. Independently of the patch, fix the exposure: bind the ZMQ sockets to localhost or the job's private network namespace rather than 0.0.0.0, put distributed-training traffic on a per-job network segment, and enforce that with NetworkPolicy or equivalent so worker-to-worker ports are not reachable across tenants. No driver reload, no reboot, no firmware - this is an application and network-segmentation fix, which also means your firmware and kernel patch tooling will never surface it. Audit the rest of your stack for the same pattern: pickle-over-socket is endemic in distributed ML frameworks.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-49121","https://github.com/ROCm/aiter"],"status":"curated","tags":["tenant-isolation"],"published":"2026-06-01"},{"id":"CVE-2026-63925","cve":"CVE-2026-63925","aliases":[],"title":"Linux MACsec (replay protection at XPN lower-PN wrap): MACsec replay protection fails at the extended-packet-number","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux MACsec (replay protection at XPN lower-PN wrap)","year":"2026","cvss_score":8.1,"severity":"high","kev":false,"impact":"MACsec replay protection fails at the extended-packet-number lower-PN wrap. When the packet number is U32_MAX the increment overflows to zero and neither replay branch fires, so `next_pn_halves` is never advanced — an attacker who captured legitimate ciphertext can replay it and have it accepted. Replay protection is the property that stops a passive observer from becoming an active injector on an encrypted link; losing it turns a tap into a traffic-injection capability on links that carry multiple tenants.","attack_vector":"An attacker who can capture and re-transmit frames on a MACsec-protected link, timed to the PN wrap. Passive tap plus injection capability, no keys required.","remediation":"Kernel upgrade plus host reboot on any node terminating software MACsec (switch NOSes based on Linux included, where it arrives as a NOS image update plus reload). No config workaround — you cannot turn replay protection back on if the check itself is broken. Companion Linux MACsec issues in the same window: CVE-2026-72019, CVE-2022-48720.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63925","https://nvd.nist.gov/vuln/detail/CVE-2026-72019"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-07-19"},{"id":"NCVD-2021-004-infiniband-roce-memory-protectio","cve":null,"aliases":["ReDMArk","RDMA rkey brute-forcing","unauthorized RDMA memory access","memory window abuse"],"title":"InfiniBand / RoCE memory protection - memory region rkey/lkey namespace and protection domains: The only thing standing","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"InfiniBand / RoCE memory protection - memory region rkey/lkey namespace and protection domains","year":"2021","cvss_score":8.1,"severity":"high","kev":false,"impact":"The only thing standing between a remote peer and a registered memory region is the 32-bit rkey, and ReDMArk found that real RNIC firmware generates rkeys in a small, largely sequential, and therefore guessable space. Combined with the connection-injection weakness, an attacker can issue RDMA READ or WRITE against another tenant's memory region without ever having been granted access to it. Applications make this dramatically worse in practice by registering one huge region with IBV_ACCESS_REMOTE_WRITE covering far more than the buffers actually meant to be shared - a common shortcut in RDMA key-value stores, parameter servers, and disaggregated-memory layers. On a GPU node the registered region frequently includes GPU memory reachable through GPUDirect, so a guessed rkey reads model state directly out of HBM.","attack_vector":"From any node able to establish or inject into a queue pair on the target RNIC, the attacker sweeps rkey values in RDMA READ requests and watches for a completion instead of an error. Because the RNIC services the read entirely in hardware, the victim's CPU never runs and nothing is logged on the victim host - the sweep is silent and can run at line rate. Successful reads return the contents of whatever the victim registered; successful writes corrupt it. Memory windows (type 1/2) bind more narrowly but are seldom used, and applications that reuse a single protection domain across tenant contexts collapse the boundary entirely.","remediation":"No vendor patch. Application and config work: register the smallest possible regions, never grant REMOTE_WRITE where REMOTE_READ suffices, use a separate protection domain per tenant/connection, and prefer type-2 memory windows with short lifetimes so a guessed key expires. Where the RNIC firmware supports stronger key generation, apply the vendor firmware update (firmware flash, rolling per-node, ~5-15 min per host plus a reboot on most ConnectX generations). Fabric-level containment - one P_Key or VLAN per tenant - limits who can reach the RNIC to try keys at all and is a switch config change. Budget this as an application-rearchitecture item, not a maintenance window.","references":["https://www.usenix.org/conference/usenixsecurity21/presentation/rothenberger","https://netsec.ethz.ch/publications/papers/sec21summer-redmark.pdf","https://arxiv.org/abs/1903.09355"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"tags":["tenant-isolation"]},{"id":"NCVD-2021-010-infiniband-roce-memory-protectio","cve":null,"aliases":["ReDMArk","RDMA rkey brute-forcing","unauthorized RDMA memory access","memory window abuse"],"title":"InfiniBand / RoCE memory protection - memory region rkey/lkey namespace and protection domains: The only thing standing","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"InfiniBand / RoCE memory protection - memory region rkey/lkey namespace and protection domains","year":"2021","cvss_score":8.1,"severity":"high","kev":false,"impact":"The only thing standing between a remote peer and a registered memory region is the 32-bit rkey, and ReDMArk found that real RNIC firmware generates rkeys in a small, largely sequential, and therefore guessable space. Combined with the connection-injection weakness, an attacker can issue RDMA READ or WRITE against another tenant's memory region without ever having been granted access to it. Applications make this dramatically worse in practice by registering one huge region with IBV_ACCESS_REMOTE_WRITE covering far more than the buffers actually meant to be shared - a common shortcut in RDMA key-value stores, parameter servers, and disaggregated-memory layers. On a GPU node the registered region frequently includes GPU memory reachable through GPUDirect, so a guessed rkey reads model state directly out of HBM.","attack_vector":"From any node able to establish or inject into a queue pair on the target RNIC, the attacker sweeps rkey values in RDMA READ requests and watches for a completion instead of an error. Because the RNIC services the read entirely in hardware, the victim's CPU never runs and nothing is logged on the victim host - the sweep is silent and can run at line rate. Successful reads return the contents of whatever the victim registered; successful writes corrupt it. Memory windows (type 1/2) bind more narrowly but are seldom used, and applications that reuse a single protection domain across tenant contexts collapse the boundary entirely.","remediation":"No vendor patch. Application and config work: register the smallest possible regions, never grant REMOTE_WRITE where REMOTE_READ suffices, use a separate protection domain per tenant/connection, and prefer type-2 memory windows with short lifetimes so a guessed key expires. Where the RNIC firmware supports stronger key generation, apply the vendor firmware update (firmware flash, rolling per-node, ~5-15 min per host plus a reboot on most ConnectX generations). Fabric-level containment - one P_Key or VLAN per tenant - limits who can reach the RNIC to try keys at all and is a switch config change. Budget this as an application-rearchitecture item, not a maintenance window.","references":["https://www.usenix.org/conference/usenixsecurity21/presentation/rothenberger","https://netsec.ethz.ch/publications/papers/sec21summer-redmark.pdf","https://arxiv.org/abs/1903.09355"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.0/AV:A/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-287"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2018-1112","cve":"CVE-2018-1112","aliases":[],"title":"GlusterFS (glusterd, auth.allow): The auth.allow option does not actually restrict who may connect, so any","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"GlusterFS (glusterd, auth.allow)","year":"2018","cvss_score":8,"severity":"high","kev":false,"impact":"The auth.allow option does not actually restrict who may connect, so any unauthenticated gluster client on any network mounts the volume. Every dataset and checkpoint on that volume is readable and writable by anyone who can reach the bricks.","attack_vector":"Any host with network reach to glusterd/brick ports. No credential and no membership in the trusted pool required.","remediation":"Upgrade glusterfs server to 3.10.12 / 4.0.2 or later and restart glusterd and the brick processes. Do not rely on auth.allow as the boundary - enforce TLS with client certificates (transport.socket.ssl) and firewall brick ports to known client subnets.","references":["https://access.redhat.com/security/cve/CVE-2018-1112","https://nvd.nist.gov/vuln/detail/CVE-2018-1112"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-287"],"fleet":{"pain_class":"node-drain"},"id":"CVE-2019-15719","cve":"CVE-2019-15719","aliases":[],"title":"Altair PBS Professional / OpenPBS (pbs_mom): Pbs_mom, the daemon that executes jobs on every compute node, accepts","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Altair PBS Professional / OpenPBS (pbs_mom)","year":"2019","cvss_score":8,"severity":"high","kev":false,"impact":"Pbs_mom, the daemon that executes jobs on every compute node, accepts messages without authenticating them. Send it a message directly and you get code execution on that node with the daemon's privileges - which is how a tenant on one node reaches into the execution path of jobs belonging to everyone else on the cluster.","attack_vector":"Adjacent network - anything that can open a socket to pbs_mom on a compute node. On a flat cluster network that is every tenant with a running job.","remediation":"Upgrade past PBS Professional 19.1.2 / the corresponding OpenPBS release. Until then, firewall the pbs_mom port so that only the pbs_server host can reach it - the daemon has no business accepting connections from compute peers. Rolling the pbs_mom binary requires draining each node.","references":["https://www.hpcsec.com/2019/10/08/cve-2019-15719/","https://nvd.nist.gov/vuln/detail/CVE-2019-15719","http://packetstormsecurity.com/files/154782/PBS-Professional-19.2.3-Authentication-Bypass.html"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-285"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2020-10736","cve":"CVE-2020-10736","aliases":[],"title":"Ceph MON / MGR (ceph-mon, ceph-mgr): Ceph-mon and ceph-mgr fail to enforce the caps on an authenticated principal, so a","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ceph MON / MGR (ceph-mon, ceph-mgr)","year":"2020","cvss_score":8,"severity":"high","kev":false,"impact":"Ceph-mon and ceph-mgr fail to enforce the caps on an authenticated principal, so a low-privileged CephX user reaches admin-only commands. From there the attacker can read and change cluster configuration, create or delete pools and effectively take ownership of storage belonging to other tenants.","attack_vector":"Any authenticated CephX principal that can reach the mon or mgr on the Ceph public network, including tenant nodes that hold only a restricted client key.","remediation":"Upgrade to Ceph 15.2.2 or later and restart ceph-mon and ceph-mgr. Audit the caps on every CephX principal afterwards and revoke any that were widened. Never expose mon/mgr ports to tenant-routable networks.","references":["https://access.redhat.com/security/cve/CVE-2020-10736","https://nvd.nist.gov/vuln/detail/CVE-2020-10736"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2021-21505","cve":"CVE-2021-21505","aliases":["DSA-2021-020"],"title":"Dell EMC Integrated System for Microsoft Azure Stack Hub (undocumented iDRAC account): Dell shipped these integrated","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell EMC Integrated System for Microsoft Azure Stack Hub (undocumented iDRAC account)","year":"2021","cvss_score":8,"severity":"high","kev":false,"impact":"Dell shipped these integrated racks with an undocumented iDRAC account whose credentials are the same everywhere. Anyone who learns them owns the BMC on every node in the system - power control, Virtual Media boot of an attacker image, KVM into the console, and a persistent foothold under the hypervisor. A shared default credential is the worst shape of this class of bug because it does not need an exploit, scales across the whole install base at once, and is invisible to a vulnerability scan that only checks firmware versions. Affects builds 1906 through 2011.","attack_vector":"Anything routable to the iDRAC addresses on the out-of-band management VLAN, using credentials that are effectively public once disclosed. No exploit, no prior foothold, no privilege escalation step.","remediation":"Apply the Dell update package that brings the system to build 2102 or later. Because the root cause is an account rather than a code defect, also verify by hand afterward that the account is gone from every node's iDRAC user list, and rotate any other shared BMC credentials while you are in there. Firmware/update-package rollout is out-of-band; the account audit is config-only and should be done immediately regardless of the patch schedule.","references":["https://www.dell.com/support/kbdoc/en-us/000186008/dsa-2021-020-dell-emc-integrated-system-for-microsoft-azure-stack-hub-security-update-for-an-idrac-undocumented-account-vulnerability","https://nvd.nist.gov/vuln/detail/CVE-2021-21505"],"status":"curated","published":"2021-05-06"},{"id":"CVE-2021-23279","cve":"CVE-2021-23279","aliases":[],"title":"Eaton Intelligent Power Manager (IPM) prior to 1.69 - meta_driver_srv.js: Unauthenticated arbitrary file deletion","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Eaton Intelligent Power Manager (IPM) prior to 1.69 - meta_driver_srv.js","year":"2021","cvss_score":8,"severity":"high","kev":false,"impact":"Unauthenticated arbitrary file deletion on the IPM server. Less glamorous than the RCEs but operationally pointed: an attacker can delete the configuration and driver files that let IPM talk to your UPS estate, silently disabling the power-response layer without triggering anything that looks like an attack.","attack_vector":"Unauthenticated, remote, to the IPM server.","remediation":"Upgrade to IPM 1.69 or later. Also verify you have restorable backups of IPM configuration - the recovery path for this bug is restore, and most operators have never tested it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-23279"],"status":"curated","published":"2021-04-13"},{"id":"CVE-2021-43816","cve":"CVE-2021-43816","aliases":[],"title":"containerd: On SELinux hosts, an unprivileged pod with a hostPath volume can gain full read/write to the host filesystem","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2021","cvss_score":8,"severity":"high","kev":false,"impact":"On SELinux hosts, an unprivileged pod with a hostPath volume can gain full read/write to the host filesystem","attack_vector":"Cluster user able to create a pod with a hostPath volume","remediation":"Rolling containerd upgrade with node drain; block hostPath in tenant namespaces","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-43816"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2022-01-05"},{"id":"CVE-2022-21225","cve":"CVE-2022-21225","aliases":[],"title":"Intel Data Center Manager: Improper neutralisation (injection) in Data Center Manager lets an authenticated user","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel Data Center Manager","year":"2022","cvss_score":8,"severity":"high","kev":false,"impact":"Improper neutralisation (injection) in Data Center Manager lets an authenticated user with adjacent access escalate privilege on the management plane.","attack_vector":"Authenticated user with network adjacency to the DCM server.","remediation":"Upgrade the Intel Data Center Manager software. This is a management-plane application, so the update is an application upgrade and service restart - no node drain, no firmware, no reboot of managed hosts. The real work is deciding what DCM is allowed to reach: it holds credentials for platform power and telemetry across the fleet, so its network exposure matters more than its version. Upgrade to 4.1 or later.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-21225","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00662.html"],"status":"curated","fleet":{"pain_class":"daemon-restart"},"published":"2022-08-18"},{"id":"CVE-2022-32519","cve":"CVE-2022-32519","aliases":["SEVD-2023-010-06"],"title":"Schneider Electric Data Center Expert (versions prior to v7.9.0) - credential storage: DCE stores device passwords","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Schneider Electric Data Center Expert (versions prior to v7.9.0) - credential storage","year":"2022","cvss_score":8,"severity":"high","kev":false,"impact":"DCE stores device passwords in a recoverable format. Every UPS, PDU and cooling controller credential the DCIM system polls with can be recovered by an attacker who reaches the appliance. This is the amplifier that turns any of the other DCE bugs from 'one appliance' into 'the facility'.","attack_vector":"Network access to the DCE instance; recoverable-format storage means no cracking effort is needed once a foothold exists.","remediation":"Upgrade to v7.9.0 or later, then rotate every stored credential - the upgrade re-protects new secrets, it does not un-leak old ones. Budget the rotation properly: it means touching every managed device, and on a large site that is the expensive part, not the upgrade.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-32519"],"status":"curated","published":"2023-01-30"},{"id":"CVE-2023-25529","cve":"CVE-2023-25529","aliases":[],"title":"DGX H100 BMC (KVM daemon): Session token theft via timing side channel","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX H100 BMC (KVM daemon)","year":"2023","cvss_score":8,"severity":"high","kev":false,"impact":"Session token theft via timing side channel","attack_vector":"Network-adjacent unauthenticated","remediation":"Flash BMC 23.08.18; rotate BMC credentials","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5473/5473.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:H/PR:N/UI:N/S:C/C:H/I:H/A:N","cwe":["CWE-208"],"published":"2023-09-20"},{"id":"CVE-2023-25530","cve":"CVE-2023-25530","aliases":[],"title":"DGX H100 BMC (KVM): Code execution + privesc","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX H100 BMC (KVM)","year":"2023","cvss_score":8,"severity":"high","kev":false,"impact":"Code execution + privesc","attack_vector":"Network-adjacent BMC user","remediation":"Flash BMC 23.08.18 out-of-band","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5473/5473.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-20"],"published":"2023-09-20"},{"id":"CVE-2024-24590","cve":"CVE-2024-24590","aliases":[],"title":"ClearML client SDK: Deserialization of untrusted data","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"ClearML client SDK","year":"2024","cvss_score":8,"severity":"high","kev":false,"impact":"Deserialization of untrusted data — a malicious uploaded artifact executes code on the consuming host","attack_vector":"Customer-supplied artifact pulled by another user's ClearML job","remediation":"Upgrade past 1.14.2. Cross-tenant if a shared ClearML server serves multiple teams","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-24590"],"status":"curated","fleet":{"ubiquity":"Common - ClearML is a widely used MLOps/experiment platform beside GPU clusters; affects SDK 0.17.0-1.14.2","remediation_pain":"**Image rebuild** of every image containing the ClearML SDK, plus revocation of every artifact in the store","pain_class":"other","why_fleet_wide":"`Artifact.get()` pickle-deserializes a maliciously uploaded artifact, so one poisoned artifact runs code on every user or job that reads it - a supply-chain fan-out inside the platform, with 10+ public PoCs"},"published":"2024-02-06"},{"id":"CVE-2024-24591","cve":"CVE-2024-24591","aliases":[],"title":"ClearML client SDK: Path traversal — a malicious dataset writes arbitrary files on the consumer","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"ClearML client SDK","year":"2024","cvss_score":8,"severity":"high","kev":false,"impact":"Path traversal — a malicious dataset writes arbitrary files on the consumer","attack_vector":"Customer-supplied dataset","remediation":"Upgrade past 1.14.1","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-24591"],"status":"curated","published":"2024-02-06"},{"id":"CVE-2024-25951","cve":"CVE-2024-25951","aliases":[],"title":"Dell iDRAC8 (local RACADM): An authenticated user injects commands through local RACADM and takes control","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC8 (local RACADM)","year":"2024","cvss_score":8,"severity":"high","kev":false,"impact":"An authenticated user injects commands through local RACADM and takes control of the underlying BMC operating system - full out-of-band control of the server from an ordinary iDRAC account.","attack_vector":"Adjacent-network attacker holding any valid low-privilege iDRAC credential.","remediation":"Apply the iDRAC8 firmware update from DSA-2024-089. BMC firmware flash, no host reboot. Review iDRAC local accounts at the same time - the bug converts a low-privilege account into root on the BMC.","references":["https://www.dell.com/support/kbdoc/en-us/000222591/dsa-2024-089-security-update-for-dell-idrac8-local-racadm-vulnerability"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-28860","cve":"CVE-2024-28860","aliases":[],"title":"Cilium: IPsec transparent encryption is cryptographically ineffective","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2024","cvss_score":8,"severity":"high","kev":false,"impact":"IPsec transparent encryption is cryptographically ineffective; inter-node traffic can be decrypted or forged","attack_vector":"Anyone with access to the underlay network between nodes","remediation":"Upgrade Cilium and rotate IPsec keys; assume prior inter-node traffic was exposed","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-28860"],"status":"curated","published":"2024-03-27"},{"id":"CVE-2024-37774","cve":"CVE-2024-37774","aliases":[],"title":"Sunbird DCIM dcTrack v9.1.2: CSRF in admin screens lets an authenticated attacker escalate privileges by getting","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Sunbird DCIM dcTrack v9.1.2","year":"2024","cvss_score":8,"severity":"high","kev":false,"impact":"CSRF in admin screens lets an authenticated attacker escalate privileges by getting an administrator to load a crafted page. dcTrack is the system of record for where every asset, circuit and outlet lives - so administrator access is both a map of the facility and, through its integrations, a route into the devices themselves.","attack_vector":"Requires an authenticated dcTrack administrator to visit an attacker-controlled page while logged in.","remediation":"Upgrade dcTrack past 9.1.2. Software upgrade on one host. Also worth doing: check what credentials dcTrack holds for integrated PDU and power gear, since DCIM asset databases quietly accumulate device logins.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-37774"],"status":"curated","published":"2024-12-16"},{"id":"CVE-2024-39368","cve":"CVE-2024-39368","aliases":[],"title":"Intel Neural Compressor (SQL injection): SQL injection reachable by an authenticated user of Neural Compressor","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel Neural Compressor (SQL injection)","year":"2024","cvss_score":8,"severity":"high","kev":false,"impact":"SQL injection reachable by an authenticated user of Neural Compressor. Gets an attacker the backing database of the optimisation service - job metadata, model references, and whatever credentials the deployment stored there.","attack_vector":"Any authenticated user of the Neural Compressor service.","remediation":"Upgrade to Neural Compressor v3.0 or later. Application-level update, restart the service.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-39368","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01219.html"],"status":"curated","fleet":{"pain_class":"daemon-restart"},"published":"2024-11-13"},{"id":"CVE-2024-45766","cve":"CVE-2024-45766","aliases":[],"title":"Dell OpenManage Enterprise (code injection): A low-privileged remote user injects code into OME and executes","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Dell OpenManage Enterprise (code injection)","year":"2024","cvss_score":8,"severity":"high","kev":false,"impact":"A low-privileged remote user injects code into OME and executes it. OME manages iDRACs fleet-wide, so code execution there means credentialed access to every BMC it manages.","attack_vector":"Authenticated low-privilege user of the OME web console, with interaction.","remediation":"Upgrade OME past 4.1. Application upgrade with a service restart. Treat OME as tier-0: it holds fleet-wide BMC credentials, so compromise there is equivalent to compromising every server it manages.","references":["https://www.dell.com/support/kbdoc/en-us/000237300/dsa-2024-426-security-update-for-dell-openmanage-enterprise-vulnerabilities"],"status":"curated","fleet":{"pain_class":"daemon-restart"}},{"id":"CVE-2025-12638","cve":"CVE-2025-12638","aliases":[],"title":"Keras (`utils.get_file`): Path traversal in tar extraction in 3.11.3 (incomplete fix)","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Keras (`utils.get_file`)","year":"2025","cvss_score":8,"severity":"high","kev":false,"impact":"Path traversal in tar extraction in 3.11.3 (incomplete fix)","attack_vector":"Customer-supplied archive","remediation":"Upgrade past 3.11.3","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-12638"],"status":"curated","published":"2025-11-28"},{"id":"CVE-2025-23268","cve":"CVE-2025-23268","aliases":[],"title":"Triton Inference Server: RCE / privesc via input-validation failure","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2025","cvss_score":8,"severity":"high","kev":false,"impact":"RCE / privesc via input-validation failure","attack_vector":"Unauthenticated client of the inference endpoint","remediation":"Upgrade Triton; rebuild and redeploy all serving images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23268","https://github.com/NVIDIA/product-security/tree/main/2025/5691"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-20"],"published":"2025-09-17"},{"id":"CVE-2025-30165","cve":"CVE-2025-30165","aliases":[],"title":"vLLM (multi-node ZeroMQ): Secondary vLLM host trusts unauthenticated ZeroMQ messages","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (multi-node ZeroMQ)","year":"2025","cvss_score":8,"severity":"high","kev":false,"impact":"Secondary vLLM host trusts unauthenticated ZeroMQ messages","attack_vector":"Co-tenant or compromised worker on the multi-node deployment","remediation":"Upgrade; enforce mTLS or a private VRF for intra-job traffic","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-30165"],"status":"curated","published":"2025-05-06"},{"id":"CVE-2025-33179","cve":"CVE-2025-33179","aliases":[],"title":"Cumulus Linux / NVOS: Privesc to switch admin","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Cumulus Linux / NVOS","year":"2025","cvss_score":8,"severity":"high","kev":false,"impact":"Privesc to switch admin","attack_vector":"Authenticated switch user","remediation":"Upgrade Cumulus Linux / NVOS; rolling switch upgrade across the fabric","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33179","https://github.com/NVIDIA/product-security/tree/main/2026/5722"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-266"],"published":"2026-02-24"},{"id":"CVE-2025-33180","cve":"CVE-2025-33180","aliases":[],"title":"Cumulus Linux / NVOS: Command injection","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Cumulus Linux / NVOS","year":"2025","cvss_score":8,"severity":"high","kev":false,"impact":"Command injection -> switch takeover","attack_vector":"Authenticated switch user","remediation":"Upgrade Cumulus Linux / NVOS; rolling switch upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33180","https://github.com/NVIDIA/product-security/tree/main/2026/5722"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-77"],"published":"2026-02-24"},{"id":"CVE-2025-33188","cve":"CVE-2025-33188","aliases":[],"title":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware: An attacker tampers with hardware controls directly","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware","year":"2025","cvss_score":8,"severity":"high","kev":false,"impact":"An attacker tampers with hardware controls directly, reaching data tampering and denial of service at the platform level. These sit in the GB10 root-of-trust chain, so a successful exploit undermines the platform's own attestation and secure-boot story rather than just the OS above it.","attack_vector":"Local access to the DGX Spark. Several of the set need no privileges at all; the rest need host root. This is a desk-side developer box, so physical and local access assumptions are much weaker than for a racked DGX.","remediation":"Apply the DGX Spark firmware update from bulletin 5720. Cost: flash plus reboot, low drain cost given the form factor, but not live-patchable and root-of-trust firmware cannot be rolled back once applied.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33188","https://github.com/NVIDIA/product-security/tree/main/2025/5720"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:U/C:L/I:H/A:H","cwe":["CWE-269"],"fleet":{"pain_class":"firmware-flash"},"published":"2025-11-25"},{"id":"CVE-2025-33245","cve":"CVE-2025-33245","aliases":[],"title":"NeMo Framework: Remote RCE via insecure deserialization over the network","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NeMo Framework","year":"2025","cvss_score":8,"severity":"high","kev":false,"impact":"Remote RCE via insecure deserialization over the network","attack_vector":"Network attacker feeding a serialized payload","remediation":"Bump NeMo; rebuild and redeploy; restrict network exposure of NeMo services","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33245","https://github.com/NVIDIA/product-security/tree/main/2026/5762"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2026-02-18"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:R/S:U/C:H/I:H/A:H","fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-68803","cve":"CVE-2025-68803","aliases":[],"title":"Linux NFS server (nfsd, NFSv4 file creation ACL): When a client sets an ACL during NFSv4 file creation, nfsd silently","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux NFS server (nfsd, NFSv4 file creation ACL)","year":"2025","cvss_score":8,"severity":"high","kev":false,"impact":"When a client sets an ACL during NFSv4 file creation, nfsd silently drops it and falls back to an ACL derived from the mode bits. Files a tenant believed were restricted to a named principal are actually governed by looser mode-derived permissions, so other users on the shared export can read them.","attack_vector":"Any NFSv4 client creating files with an ACL naming a principal. The exposure is created by the server, not by an attacker action - the attacker just has to be another user on the export.","remediation":"Update the storage server kernel to one carrying the nfsd_create_setattr ACL fix and reboot. Then re-apply ACLs on files created during the exposure window, since the intended ACLs were never written to the inodes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-68803","https://git.kernel.org/pub/scm/linux/security/vulns.git"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-7766","cve":"CVE-2025-7766","aliases":[],"title":"Lantronix Provisioning Manager: Provisioning Manager reads configuration files supplied by the network devices","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Lantronix Provisioning Manager","year":"2025","cvss_score":8,"severity":"high","kev":false,"impact":"Provisioning Manager reads configuration files supplied by the network devices it manages. Because it doesn't lock down XML external entity resolution, a device (or something spoofing one) can hand it a poisoned config file that leads to unauthenticated remote code execution on the host running Provisioning Manager — which typically has admin reach into the whole fleet of Lantronix devices it provisions.","attack_vector":"Attacker needs to get a malicious XML config file processed by Provisioning Manager — either by compromising/spoofing a managed device on the network, or by feeding it a crafted import file if the workflow allows manual uploads.","remediation":"Software upgrade of Provisioning Manager to the patched release. This runs on a management workstation/server rather than the appliances themselves, so it's a single upgrade rather than a per-device fleet rollout — but treat it as high priority since it's the box with admin credentials to the whole console-server fleet.","references":["https://www.lantronix.com/"],"status":"curated","published":"2025-07-22"},{"id":"CVE-2026-24213","cve":"CVE-2026-24213","aliases":[],"title":"Triton Inference Server: Info disclosure / RCE (OOB read in tensor processing)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2026","cvss_score":8,"severity":"high","kev":false,"impact":"Info disclosure / RCE (OOB read in tensor processing)","attack_vector":"Network client sending crafted tensors","remediation":"Upgrade Triton; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24213","https://github.com/NVIDIA/product-security/tree/main/2026/5828"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-125"],"published":"2026-05-20"},{"id":"CVE-2026-24214","cve":"CVE-2026-24214","aliases":[],"title":"Triton Inference Server: RCE (integer overflow in model config parsing)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2026","cvss_score":8,"severity":"high","kev":false,"impact":"RCE (integer overflow in model config parsing)","attack_vector":"Malicious model config","remediation":"Upgrade Triton; redeploy; restrict model repo writes","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24214","https://github.com/NVIDIA/product-security/tree/main/2026/5828"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-190"],"published":"2026-05-20"},{"id":"CVE-2026-47237","cve":"CVE-2026-47237","aliases":[],"title":"Kubeflow Community Distribution: Insecure default in the platform install","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Kubeflow Community Distribution","year":"2026","cvss_score":8,"severity":"high","kev":false,"impact":"Insecure default in the platform install","attack_vector":"Tenant user of a Kubeflow-based platform","remediation":"Upgrade to 26.03-rc.1+; provider-owned if the provider ships managed Kubeflow","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47237"],"status":"curated","fleet":{"ubiquity":"Common - Kubeflow is a standard multi-tenant ML platform layer over GPU K8s; official manifests before 1.10 and most packaged distros affected","remediation_pain":"`daemon-restart` - Istio policy and manifest upgrade; the expensive part is invalidating every user token issued before the fix","pain_class":"daemon-restart","why_fleet_wide":"Overly permissive Istio permissions let any user with `kubeflow-edit` in *any* namespace create a notebook that steals other users' authorization tokens - a straight cross-tenant takeover on a shared AI platform"},"published":"2026-07-21"},{"cwe":["CWE-190","CWE-787"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-52969","cve":"CVE-2026-52969","aliases":[],"title":"Linux kernel (virt/kvm): The dirty-ring reset path bounds-checks an offset with unchecked 64-bit arithmetic, so a","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (virt/kvm)","year":"2026","cvss_score":8,"severity":"high","kev":false,"impact":"The dirty-ring reset path bounds-checks an offset with unchecked 64-bit arithmetic, so a crafted pair of ring entries wraps the sum past the check. KVM then indexes the memslot's rmap array with a near-U64_MAX GFN - an out-of-bounds load followed by a conditional bit clear through whatever pointer that load produced. That is an attacker-influenced write into host kernel memory and a straight path to root on the node.","attack_vector":"The commit states it plainly: reachable from any process holding /dev/kvm. The dirty ring is mapped MAP_SHARED into the process, so it rewrites the slot/offset payload of queued entries and then calls KVM_RESET_DIRTY_RINGS. Needs the legacy/shadow MMU path (shadow paging, any VM that allocated shadow roots, or a write-tracked slot). Relevant wherever tenants can open /dev/kvm - nested-virt-enabled VMs, or bare-metal tenants with KVM exposed.","remediation":"Update to a kernel carrying the referenced stable commits. Interim: do not expose /dev/kvm to tenant containers, and disable nested virtualization for tenant VMs so guests cannot run their own KVM.","references":["https://git.kernel.org/stable/c/74f1a22f7a80f03d28ad8551a2d25d563433addf","https://git.kernel.org/stable/c/0eb281eb95b2d4eea4db1da5fe91023aecc97095","https://nvd.nist.gov/vuln/detail/CVE-2026-52969"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-59690","cve":"CVE-2026-59690","aliases":[],"title":"Progress Kemp LoadMaster Multi Tenant: The Multi Tenant product line's REST API doesn't check whether a caller's","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Progress Kemp LoadMaster Multi Tenant","year":"2026","cvss_score":8,"severity":"high","kev":false,"impact":"The Multi Tenant product line's REST API doesn't check whether a caller's permission level actually allows the administrative operation they're requesting. A low-privileged tenant account can invoke privileged administrative operations meant only for the LoadMaster operator, reaching functionality that should be walled off from other tenants sharing the same appliance.","attack_vector":"Requires an authenticated low-privilege account on the Multi Tenant platform — no privilege escalation exploit needed, just calling the REST API endpoints the UI hides but the backend doesn't actually gate.","remediation":"Software upgrade to the fixed release per Progress's July 2026 LoadMaster Critical Security Bulletin (issued alongside four related CVEs). Prioritize this on any shared/multi-tenant LoadMaster deployment, since the whole point of the missing check is that tenant boundaries aren't being enforced.","references":["https://community.progress.com/s/article/LoadMaster-Critical-Security-Bulletin-July-2026-CVE-2026-59686-CVE-2026-59687-CVE-2026-59688-CVE-2026-59689-CVE-2026-59690"],"status":"curated","tags":["tenant-isolation"],"published":"2026-07-27"},{"cvss_vector":"CVSS:3.1/AV:A/AC:H/PR:N/UI:N/S:C/C:H/I:H/A:N","cwe":["CWE-350","CWE-940"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2012-4516","cve":"CVE-2012-4516","aliases":[],"title":"librdmacm 1.0.16 (userspace RDMA connection-manager library) - default fallback to ibacm port 6125: RDMA","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"librdmacm 1.0.16 (userspace RDMA connection-manager library) - default fallback to ibacm port 6125","year":"2012","cvss_score":7.9,"severity":"high","kev":false,"impact":"RDMA address-resolution poisoning, in the library every verbs application links against. When ibacm.port is unset, librdmacm silently falls back to connecting to port 6125 and trusts whatever ib_acm service answers there for its path records - GID, LID and path attributes. Whoever answers therefore decides where a victim application's queue pairs actually get established. This is the RDMA-layer equivalent of DNS or ARP poisoning, and it is worse than either because RDMA has no transport authentication to catch the redirect afterwards: the QP comes up, the RDMA operations succeed, and the tenant's gradients or checkpoints land at an endpoint of the attacker's choosing. It is also the one item in this collection that lives entirely in userspace, so no kernel patching regime catches it.","attack_vector":"Whoever can bind or answer on the ibacm port that the victim's librdmacm falls back to - a co-located process on the same node, or a rogue ib_acm reachable across the fabric. Chains directly from CVE-2012-4518, which lets any local user rewrite the ibacm.port file that decides where librdmacm looks.","remediation":"Update librdmacm past 1.0.16 (upstream commit 4b5c1aa734e0e734fc2ba3cd41d0ddf02170af6d) and explicitly pin ibacm.port rather than relying on the default. This is a userspace library, so the fix lands with a package update and a restart of the RDMA-using applications - no kernel reboot - but that is also the trap: it will not appear in any kernel CVE scan, and containerised AI workloads frequently ship their own vendored copy of the RDMA userspace stack inside the image, which means patching the host does nothing. Audit the librdmacm version inside your tenant base images, not just on the host.","references":["https://www.openwall.com/lists/oss-security/2012/10/11/6","https://bugzilla.redhat.com/show_bug.cgi?id=865483","https://access.redhat.com/security/cve/CVE-2012-4516"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2022-41739","cve":"CVE-2022-41739","aliases":[],"title":"IBM Spectrum Scale / Storage Scale Container Native Storage Access: Programs running inside a container can overcome","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"IBM Spectrum Scale / Storage Scale Container Native Storage Access","year":"2022","cvss_score":7.9,"severity":"high","kev":false,"impact":"Programs running inside a container can overcome the isolation mechanism of IBM Spectrum Scale Container Native Storage Access. Spectrum Scale (GPFS) is one of the two or three filesystems that actually keep up with large training clusters, and the container-native access layer is how Kubernetes-scheduled GPU jobs mount it. An isolation escape here means one tenant's pod reaching outside its intended storage boundary on shared cluster storage.","attack_vector":"A process inside a container that has Spectrum Scale container-native storage access — i.e. any tenant workload with a mounted volume.","remediation":"Upgrade Container Native Storage Access past 5.1.6.0. This is a rolling upgrade of the storage-access DaemonSet/operator; pods remount as it rolls, so drain latency-sensitive jobs. No filesystem downtime, but plan for I/O stalls during the roll.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-41739"],"status":"curated","fleet":{"pain_class":"node-drain"},"tags":["tenant-isolation"],"published":"2023-04-26"},{"id":"CVE-2023-45745","cve":"CVE-2023-45745","aliases":[],"title":"Intel TDX module: The TDX module is the software that stands between the host/VMM and every confidential VM on the box","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel TDX module","year":"2023","cvss_score":7.9,"severity":"high","kev":false,"impact":"The TDX module is the software that stands between the host/VMM and every confidential VM on the box; a privilege escalation inside it is a break of the boundary that separates a tenant's trust domain from the operator and from other TDs. Specific flaw: improper input validation reachable by a privileged host user.","attack_vector":"A privileged user on the host - which in the TDX threat model is the adversary the whole design exists to exclude, so 'requires host privilege' is not a mitigating factor here.","remediation":"Update the Intel TDX module. The TDX module is loaded by the SEAM loader at boot, so the practical rollout is: stage the new module, drain every trust domain off the node, and reboot. It is not a live-patchable component and running TDs cannot be migrated through it. After the update, every TD must re-attest because the TDX module SVN is part of the attestation report - so anything that pinned the old measurement will fail until you update your attestation policy too. No OEM BIOS release needed for the module itself, which makes this materially faster than a platform firmware update.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-45745","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01036.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2024-05-16"},{"id":"CVE-2024-0172","cve":"CVE-2024-0172","aliases":[],"title":"Dell PowerEdge Server BIOS / Precision Rack BIOS (improper privilege management): An unauthenticated local attacker","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell PowerEdge Server BIOS / Precision Rack BIOS (improper privilege management)","year":"2024","cvss_score":7.9,"severity":"high","kev":false,"impact":"An unauthenticated local attacker escalates privilege with scope change - a firmware-level escalation that survives OS reinstall and is invisible to host-based tooling.","attack_vector":"Local access to the server. On bare-metal GPU rental this is the tenant themselves.","remediation":"Flash the fixed PowerEdge BIOS. A BIOS update is a cold reboot per node and cannot be done live - on a GPU fleet that means draining jobs and taking the box out of the scheduler, so batch it with other firmware work rather than doing a standalone pass.","references":["https://www.dell.com/support/kbdoc/en-us/000223727/dsa-2024-035-security-update-for-dell-poweredge-server-bios-for-an-improper-privilege-management-security-vulnerability"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-21980","cve":"CVE-2024-21980","aliases":["SNP firmware write restriction bypass","UMC seed overwrite"],"title":"AMD SEV-SNP firmware (EPYC Milan, Genoa, Bergamo, Siena): SNP firmware fails to restrict where a hypervisor-driven","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV-SNP firmware (EPYC Milan, Genoa, Bergamo, Siena)","year":"2024","cvss_score":7.9,"severity":"high","kev":false,"impact":"SNP firmware fails to restrict where a hypervisor-driven write can land, letting a malicious host overwrite a confidential guest's memory or the UMC seed that derives that guest's memory encryption key. Overwriting guest memory defeats the integrity half of SNP outright; overwriting the UMC seed lets the operator choose the key material protecting a tenant. This is the highest-severity item in the AMD-SB-3011 batch and it is squarely an operator-versus-tenant break - the exact threat model a customer pays the SNP premium for.","attack_vector":"Malicious hypervisor / host root issuing SNP firmware commands against a running confidential guest. Local, high privilege.","remediation":"Two options per AMD-SB-3011: hot-loadable SEV firmware (1.37.14 hex on Milan, 1.37.24 on Genoa) which does NOT require a reboot, or a full Platform Initialization / BIOS flash (MilanPI 1.0.0.D, GenoaPI 1.0.0.C) which does. Take the SEV firmware path first to close the window without draining training jobs, then fold the PI update into your next maintenance cycle. Fixing it moves the SNP TCB version above 0x16 (Milan) / 0x15 (Genoa), which changes every attestation report the host issues - tenants' verifiers must be updated to accept the new TCB or their attestation will start failing.","references":["https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3011.html","https://nvd.nist.gov/vuln/detail/CVE-2024-21980"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-08-05"},{"id":"CVE-2024-23599","cve":"CVE-2024-23599","aliases":[],"title":"Intel reference platforms (Seamless Firmware Updates): A race condition in the seamless firmware update mechanism lets","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel reference platforms (Seamless Firmware Updates)","year":"2024","cvss_score":7.9,"severity":"high","kev":false,"impact":"A race condition in the seamless firmware update mechanism lets a privileged local user cause denial of service. Seamless update is the feature that is supposed to let you patch platform firmware without a reboot - a bug that turns it into an outage undermines the whole reason operators enable it.","attack_vector":"Privileged local access on the host.","remediation":"Platform firmware update from the OEM. Ironically this one does need a conventional drain and reboot to install, since the seamless path is what is broken.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-23599","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01071.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-09-16"},{"id":"CVE-2024-37307","cve":"CVE-2024-37307","aliases":[],"title":"Cilium: cilium-bugtool output contains sensitive data","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2024","cvss_score":7.9,"severity":"high","kev":false,"impact":"cilium-bugtool output contains sensitive data","attack_vector":"Anyone who receives a support bundle","remediation":"Upgrade Cilium; treat existing bugtool archives as secret-bearing","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-37307"],"status":"curated","published":"2024-06-13"},{"id":"CVE-2024-39585","cve":"CVE-2024-39585","aliases":[],"title":"Dell SmartFabric OS10 (hard-coded password): A hard-coded password in SmartFabric OS10 10.5.5.4-10.5.5.10 and 10.5.6.x","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell SmartFabric OS10 (hard-coded password)","year":"2024","cvss_score":7.9,"severity":"high","kev":false,"impact":"A hard-coded password in SmartFabric OS10 10.5.5.4-10.5.5.10 and 10.5.6.x, usable by a low-privileged attacker with remote access. A shipped credential in a switch NOS is identical on every unit in the fleet and on every other customer's fleet, so once it is known there is no per-device secrecy left. CVE-2024-48831 and CVE-2025-36609 are further hard-coded-password findings in the same product, which makes this a pattern rather than an incident.","attack_vector":"Low-privileged attacker with remote access to the switch. In practice any credential that gets you onto the box at all.","remediation":"OS10 upgrade plus switch reload. There is no config workaround — you cannot change a hard-coded credential. Until patched, the only real control is making the management interface unreachable from anything but a bastion. Rollout: one reload per switch across the leaf/spine.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-39585","https://nvd.nist.gov/vuln/detail/CVE-2024-48831","https://nvd.nist.gov/vuln/detail/CVE-2025-36609"],"status":"curated","published":"2024-09-06"},{"id":"CVE-2024-52880","cve":"CVE-2024-52880","aliases":["INSYDE-SA-2024016"],"title":"Insyde InsydeH2O (VariableRuntimeDxe, SecureBootHandler): The Secure Boot variable handler bounds-checks incoming data","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (VariableRuntimeDxe, SecureBootHandler)","year":"2024","cvss_score":7.9,"severity":"high","kev":false,"impact":"The Secure Boot variable handler bounds-checks incoming data using length fields that the caller supplies, so an attacker who lies about the sizes gets the handler to read and write outside the buffer. The affected code is the gatekeeper for the Secure Boot key databases (PK/KEK/db/dbx), which means the compromise targets the mechanism that decides what firmware and bootloaders are allowed to run. Highest-scored member of the four-CVE SA-2024016 VariableRuntimeDxe batch.","attack_vector":"Local admin/root on the host OS making crafted SetVariable / SMM variable-service calls.","remediation":"OEM BIOS update on Insyde kernel 5.2 / 05.29.50, 5.3 / 05.38.50, 5.4 / 05.46.50, 5.5 / 05.54.50, 5.6 / 05.61.50, 5.7 / 05.70.50 or later. Firmware flash, reboot per node. No config workaround. Patch the whole SA-2024016 set together - CVE-2024-52877, -52878 and -52879 are separate defects in the same driver and a partial fix leaves the driver reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-52880","https://www.insyde.com/security-pledge/sa-2024016/"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2025-05-15"},{"id":"CVE-2025-22889","cve":"CVE-2025-22889","aliases":[],"title":"Intel Xeon 6 with TDX (protected memory range handling): Improper handling of overlap between protected memory ranges","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Xeon 6 with TDX (protected memory range handling)","year":"2025","cvss_score":7.9,"severity":"high","kev":false,"impact":"Improper handling of overlap between protected memory ranges on Xeon 6 with TDX lets a privileged user escalate. Protected memory range enforcement is how TDX keeps one trust domain's pages away from the host and from other TDs, so overlap handling failing is the isolation primitive itself failing.","attack_vector":"Privileged host user on a Xeon 6 TDX platform.","remediation":"OEM platform firmware/BIOS update, not just a TDX module update - which means waiting on your server vendor, a per-node drain and a reboot. Re-attest all trust domains after.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-22889","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01311.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2025-08-12"},{"id":"CVE-2026-41520","cve":"CVE-2026-41520","aliases":[],"title":"Cilium: cilium-bugtool leaks sensitive data (recurrence of the 2024 issue)","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2026","cvss_score":7.9,"severity":"high","kev":false,"impact":"cilium-bugtool leaks sensitive data (recurrence of the 2024 issue)","attack_vector":"Anyone who receives a support bundle","remediation":"Upgrade Cilium; treat bugtool archives as secrets","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-41520"],"status":"curated","published":"2026-05-08"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:L/I:L/A:H","cwe":["CWE-843"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-43133","cve":"CVE-2026-43133","aliases":[],"title":"Linux kernel (arch/x86/kvm/svm): VMLOAD/VMSAVE executed by an L2 guest and not intercepted by L1 were emulated against","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (arch/x86/kvm/svm)","year":"2026","cvss_score":7.9,"severity":"high","kev":false,"impact":"VMLOAD/VMSAVE executed by an L2 guest and not intercepted by L1 were emulated against vmcb02 instead of vmcb01, so the nested guest reads and overwrites the wrong control block - segment bases, TR/LDTR, and the SYSCALL/SYSENTER MSR state belonging to the other level. A workload inside a nested guest can harvest or corrupt its hypervisor's saved state in host-managed pages.","attack_vector":"Guest-driven: requires nested SVM exposed to the tenant (AMD host, kvm_amd nested=1) and an L2 running with virtual VMLOAD/VMSAVE not intercepted by L1. The instruction is executed directly by guest code - no ioctl, no host privilege.","remediation":"Update to a stable kernel with the linked fix (no fixed release enumerated; take the branch carrying commit 3880e331b0b3). Interim control: withhold nested virtualization from tenant guests (kvm_amd.nested=0) or disable virtual VMLOAD/VMSAVE so the instructions always trap.","references":["https://git.kernel.org/stable/c/3880e331b0b31d0d5d3702b124f6c93539cd478a","https://git.kernel.org/stable/c/c3b7015000988ba35ecd5648f4b2283960f00543","https://nvd.nist.gov/vuln/detail/CVE-2026-43133"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-6540","cve":"CVE-2026-6540","aliases":[],"title":"Calico: Application Layer Policy (Dikastes) does not normalise URL paths, so path-traversal and encoded","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Calico","year":"2026","cvss_score":7.9,"severity":"high","kev":false,"impact":"Application Layer Policy (Dikastes) does not normalise URL paths, so path-traversal and encoded slashes bypass HTTP rules","attack_vector":"Unauthenticated network reaching a policy-protected service","remediation":"Rolling Calico upgrade; do not rely on ALP HTTP rules as the only authorization","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-6540"],"status":"curated","published":"2026-07-30"},{"id":"CVE-2026-72312","cve":"CVE-2026-72312","aliases":[],"title":"Linux octeontx2-af (VF rx-mode affecting PF promiscuous state): A VF setting its receive mode causes the *physical","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux octeontx2-af (VF rx-mode affecting PF promiscuous state)","year":"2026","cvss_score":7.9,"severity":"high","kev":false,"impact":"A VF setting its receive mode causes the *physical function's* promiscuous and all-multicast MCAM rules to be deleted, because the enable/disable APIs operate on the PF even when the request arrives over a VF's mailbox. One tenant's VF can therefore change what the host's own interface receives — either blinding the operator's PF, or, in the inverse direction, the coupling means VF-driven rx-mode changes have effects outside the VF's own scope. On a shared OCTEON adapter that is one tenant reaching across the SR-IOV boundary into the host's receive path.","attack_vector":"A tenant with an assigned OCTEON VF issuing a normal `nix_set_rx_mode` mailbox request — no exploit primitive needed, just the ordinary API.","remediation":"Kernel upgrade plus host reboot across OCTEON-equipped nodes. No config workaround; the coupling is in the mailbox handler. Rolling drain per node.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-72312"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-08-15"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-190"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2010-3865","cve":"CVE-2010-3865","aliases":[],"title":"Linux kernel RDS RDMA path net/rds/rdma.c - rds_rdma_pages: The page-count arithmetic for an RDS RDMA scatter-gather","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel RDS RDMA path net/rds/rdma.c - rds_rdma_pages","year":"2010","cvss_score":7.8,"severity":"high","kev":false,"impact":"The page-count arithmetic for an RDS RDMA scatter-gather request overflows on a crafted iovec, so the kernel allocates a short buffer and then fills it as if it were large - a heap overflow reachable by an unprivileged local user, with code execution not excluded. This is the RDMA-specific sibling of the RDS root bug above and lives in exactly the code path an operator would be exercising if they did adopt RDS over InfiniBand for a low-latency service.","attack_vector":"Local, unprivileged - an RDS socket plus a crafted iovec. Same autoload consideration as CVE-2010-3904.","remediation":"Kernel upgrade or vendor backport; rolling reboot across the fleet. As with CVE-2010-3904, blacklisting the rds module is the immediate, no-downtime control and is the right default on any node that does not deliberately run RDS - which is nearly all of them.","references":["https://access.redhat.com/security/cve/CVE-2010-3865","https://www.openwall.com/lists/oss-security/2010/11/19/4","https://nvd.nist.gov/vuln/detail/CVE-2010-3865"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-1284","CWE-264"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2010-3904","cve":"CVE-2010-3904","aliases":[],"title":"Linux kernel RDS (Reliable Datagram Sockets) net/rds/page.c - rds_page_copy_user: Straight local-to-root. RDS - the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel RDS (Reliable Datagram Sockets) net/rds/page.c - rds_page_copy_user","year":"2010","cvss_score":7.8,"severity":"high","kev":true,"impact":"Straight local-to-root. RDS - the datagram protocol built for InfiniBand cluster interconnects and still autoloaded by socket(PF_RDS,...) on stock distro kernels - copied to and from user-supplied addresses without validating them, so an unprivileged user writes to arbitrary kernel memory and takes the node. The operator-relevant twist is that nobody has to be using RDS: any tenant process that opens an RDS socket triggers the module autoload, so the attack surface exists on every node whose kernel shipped the module, RDMA fabric or not. CISA lists it as a known-exploited vulnerability; reliable public exploits have existed since 2010.","attack_vector":"Local, entirely unprivileged. One socket() call from inside any tenant container or VM that has not masked module autoloading.","remediation":"Kernel upgrade past 2.6.36 or a vendor backport (commit 799c10559d60f159ab2232203f222f18fa3c4a5f, de-pessimize rds_page_copy_user). Any kernel this old is well past support and the real remediation is a fleet-wide OS upgrade, which is a rolling drain-and-reboot. The zero-downtime mitigation that operators should apply regardless of patch state is to blacklist the module and stub the alias - install rds /bin/true in modprobe.d - which removes the surface immediately and costs nothing unless you are actually running RDS.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=799c10559d60f159ab2232203f222f18fa3c4a5f","https://access.redhat.com/security/cve/CVE-2010-3904","https://www.cisa.gov/known-exploited-vulnerabilities-catalog"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-190"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2010-4649","cve":"CVE-2010-4649","aliases":[],"title":"Linux kernel InfiniBand uverbs drivers/infiniband/core/uverbs_cmd.c - ib_uverbs_poll_cq: Integer overflow on the","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel InfiniBand uverbs drivers/infiniband/core/uverbs_cmd.c - ib_uverbs_poll_cq","year":"2010","cvss_score":7.8,"severity":"high","kev":false,"impact":"Integer overflow on the completion-queue poll path. A tenant passes an oversized entry count through the uverbs POLL_CQ command, the size computation wraps, and the kernel corrupts memory past the allocation - a heap-corruption primitive available to anyone holding the uverbs device node, with privilege escalation not excluded. Historically important because it is the first of the uverbs command-argument overflows and it establishes the pattern that CVE-2014-8159 and CVE-2016-8636 repeat: the verbs uAPI trusted tenant-supplied sizes.","attack_vector":"Local, unprivileged - read/write on /dev/infiniband/uverbsN, i.e. any RDMA-enabled tenant container.","remediation":"Kernel upgrade past 2.6.37 or a vendor backport; rolling reboot. Any fleet still exposed to this is running a kernel a decade and a half old, so the operator decision is really a platform upgrade, not a patch. Interim control is the same as for the other uverbs bugs: stop mapping /dev/infiniband/uverbs* into untrusted workloads.","references":["https://access.redhat.com/security/cve/CVE-2010-4649","https://bugzilla.redhat.com/show_bug.cgi?id=666556","https://nvd.nist.gov/vuln/detail/CVE-2010-4649"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-190","CWE-119"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2014-3601","cve":"CVE-2014-3601","aliases":[],"title":"Linux kernel KVM device assignment IOMMU path virt/kvm/iommu.c - kvm_iommu_map_pages: When an IOMMU mapping fails","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel KVM device assignment IOMMU path virt/kvm/iommu.c - kvm_iommu_map_pages","year":"2014","cvss_score":7.8,"severity":"high","kev":false,"impact":"When an IOMMU mapping fails partway through, KVM unwinds it with the wrong page count and corrupts host memory. A guest that supplies a large gfn drives the failure deliberately. This sits in the legacy KVM device-assignment code - the path used to give a VM a physical GPU - so the tenant's own act of attaching or reconfiguring its device is what triggers host memory corruption. It is also a good case study in patch discipline: the first fix was wrong and produced CVE-2014-8369, so operators who applied only the original are still exposed.","attack_vector":"Guest OS user on a host using KVM legacy device assignment with a passed-through PCI device.","remediation":"Kernel upgrade that includes both commit 350b8bdd689cd2ab2c67c8a86a0be86cfa0751a7 (the original fix) and 3d32e4dbe71374a6780eaf51d719d76f9a9bf22f (the correction, CVE-2014-8369). Rolling reboot. Also worth noting as a fleet-modernisation argument: the legacy KVM device-assignment path this lives in was removed from the kernel in favour of VFIO, so hosts still using it are running a code path upstream no longer maintains.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=350b8bdd689cd2ab2c67c8a86a0be86cfa0751a7","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=3d32e4dbe71374a6780eaf51d719d76f9a9bf22f","https://access.redhat.com/security/cve/CVE-2014-3601","https://bugzilla.redhat.com/show_bug.cgi?id=1131951"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2015-2291","cve":"CVE-2015-2291","aliases":["INTEL-SA-00051","iqvw64e.sys","Intel Ethernet Diagnostics Driver BYOVD"],"title":"Intel Ethernet diagnostics driver for Windows (iqvw64e.sys / iqvw32.sys), shipped with Intel network adapter tooling","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel Ethernet diagnostics driver for Windows (iqvw64e.sys / iqvw32.sys), shipped with Intel network adapter tooling","year":"2015","cvss_score":7.8,"severity":"high","kev":true,"impact":"A signed Intel kernel driver that exposes IOCTLs allowing arbitrary kernel memory read/write. This is the classic bring-your-own-vulnerable-driver primitive: an attacker who already has admin on a Windows host drops this legitimately signed Intel driver, loads it, and gets ring-0 - which they use to disable EDR, tamper with the boot chain, and install persistence. It is actively exploited in the wild (ransomware and intrusion crews). For a GPU operator running any Windows nodes, Windows-based fleet management, or Windows jump hosts on the management plane, this is a live escalation path from a compromised admin account to kernel and then to firmware tooling.","attack_vector":"Local administrator on a Windows host. The driver does not need to have been installed by you - the attacker brings the file, so the fleet's own driver inventory tells you nothing about exposure.","remediation":"There is no patch to deploy in the useful sense, because the attack supplies its own copy of the driver. The fix is blocklisting: enable the Microsoft vulnerable-driver blocklist (HVCI / Windows Defender Application Control), which lists iqvw64e.sys, and enforce driver-signature and WDAC policy on every Windows node and management jump host. Also remove Intel's diagnostic/adapter tooling from golden images where it is not operationally needed. No reboot-and-drain cost on the Linux GPU fleet, but real policy work on the Windows management surface.","references":["https://nvd.nist.gov/vuln/detail/CVE-2015-2291","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00051.html","https://www.cisa.gov/known-exploited-vulnerabilities-catalog"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-264"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2016-4565","cve":"CVE-2016-4565","aliases":[],"title":"Linux kernel InfiniBand/RDMA uAPI write() handlers (ib_uverbs, rdma_ucm, ib_ucm, ib_umad): The whole drivers/infiniband","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel InfiniBand/RDMA uAPI write() handlers (ib_uverbs, rdma_ucm, ib_ucm, ib_umad)","year":"2016","cvss_score":7.8,"severity":"high","kev":false,"impact":"The whole drivers/infiniband stack used write() as a bidirectional ioctl - the caller passes a response pointer inside the write buffer and the kernel writes the reply structure there. Reached through splice() or a writable pipe/aio path, that response pointer is resolved against kernel address space instead of user address space, giving a tenant an arbitrary kernel-memory write from any RDMA character device it can open. Same blast radius as the uverbs registration bug but a wider surface: rdma_ucm, ib_ucm and ib_umad are all affected, so even a tenant with only the connection-manager or MAD device exposed gets the primitive.","attack_vector":"Local, unprivileged - write access to any /dev/infiniband/* node (uverbsN, rdma_cm, ucmN, umadN). Standard for RDMA-enabled tenant containers.","remediation":"Kernel upgrade to 4.5.3+ or a vendor backport of commit e6bd18f57aad1a2d1ef40e646d03ed0f2515c9e3 (IB/security: restrict use of the write() interface), which detects and denies the suspicious write paths. Requires a rolling reboot; the fix touches the uAPI entry points so it cannot be hot-patched safely. Note the upstream commit is explicitly a stopgap - the durable fix was moving the RDMA uAPI to a structured ioctl() interface in later kernels, so a fleet stuck on a pre-4.5 vendor kernel is relying on a band-aid.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=e6bd18f57aad1a2d1ef40e646d03ed0f2515c9e3","https://access.redhat.com/security/cve/CVE-2016-4565","https://www.openwall.com/lists/oss-security/2016/04/27/5"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2016-5195","cve":"CVE-2016-5195","aliases":[],"title":"Linux kernel (mm, COW): Dirty COW: privilege escalation via MAP_PRIVATE COW breakage","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (mm, COW)","year":"2016","cvss_score":7.8,"severity":"high","kev":true,"impact":"Dirty COW: privilege escalation via MAP_PRIVATE COW breakage; classic container-escape-to-root primitive [KEV]","attack_vector":"Any tenant process in a container","remediation":"Livepatchable (kpatch/Ksplice/KernelCare shipped livepatches); otherwise drain + reboot. Any node still vulnerable is EOL and should be reimaged","references":["https://access.redhat.com/security/cve/CVE-2016-5195"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2016-11-10"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-190"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2016-8636","cve":"CVE-2016-8636","aliases":[],"title":"Linux kernel Soft-RoCE drivers/infiniband/sw/rxe/rxe_mr.c (mem_check_range): The bounds check that is supposed to","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel Soft-RoCE drivers/infiniband/sw/rxe/rxe_mr.c (mem_check_range)","year":"2016","cvss_score":7.8,"severity":"high","kev":false,"impact":"The bounds check that is supposed to confine an incoming RDMA READ/WRITE to the registered memory region overflows, so a request with a large offset passes validation and the software RoCE stack services it against memory outside the region. The result is out-of-bounds kernel read (leaking another tenant's buffers back over the wire) or write (corruption, escalation). This is the Soft-RoCE analogue of CVE-2014-8159 and matters disproportionately in AI clusters because rxe is what gets loaded when operators want RoCE semantics on nodes whose NICs lack hardware offload - test rigs, mixed-generation racks, and CPU-only staging nodes sitting on the same fabric as the GPU nodes.","attack_vector":"Local user with an rxe verbs device, and - because rxe services requests arriving over UDP from the network - a remote peer that already holds a valid QP/rkey pair for the target. On a flat fabric with no RDMA-level authentication, reaching that state is a low bar.","remediation":"Kernel upgrade to 4.9.10+ or a backport of commit 647bf3d8a8e5777319da92af672289b2a6c4dc66 (IB/rxe: fix mem_check_range integer overflow). Rolling reboot. Cheap interim control that most operators overlook: rxe is not needed on nodes with real RDMA hardware - blacklist the rdma_rxe module fleet-wide and the surface disappears without touching the kernel.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=647bf3d8a8e5777319da92af672289b2a6c4dc66","https://access.redhat.com/security/cve/CVE-2016-8636","https://bugzilla.redhat.com/show_bug.cgi?id=1390832"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-190","CWE-119"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2016-9083","cve":"CVE-2016-9083","aliases":[],"title":"Linux kernel VFIO drivers/vfio/pci/vfio_pci.c - VFIO_DEVICE_SET_IRQS ioctl: A state-machine confusion in","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel VFIO drivers/vfio/pci/vfio_pci.c - VFIO_DEVICE_SET_IRQS ioctl","year":"2016","cvss_score":7.8,"severity":"high","kev":false,"impact":"A state-machine confusion in VFIO_DEVICE_SET_IRQS lets a caller sidestep the integer-overflow checks on the interrupt-configuration path and write kernel memory outside the intended area. The holder of the VFIO device file descriptor is the process that owns a tenant's passed-through GPU - the QEMU instance, or in a container-based GPU cloud the tenant's own runtime. This is therefore the direct escape route from 'we gave the tenant a GPU' to 'the tenant is writing host kernel memory', in the code path that every VFIO passthrough setup executes.","attack_vector":"Local process holding an open VFIO PCI device fd - i.e. the per-tenant VMM, or any workload granted /dev/vfio access for userspace device drivers or DPDK-style datapaths.","remediation":"Kernel upgrade including commit 05692d7005a364add85c6e25a6c4447ce08f913a (which fixes this and CVE-2016-9084 together); RHEL shipped it in RHSA-2017:0386. Rolling reboot. Meaningful hardening independent of the patch: do not hand /dev/vfio/* into tenant containers. When passthrough is done through a VMM the fd stays in a process you control, and the exposure is bounded by that VMM's own hardening; handing the fd to tenant code removes that layer entirely.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=05692d7005a364add85c6e25a6c4447ce08f913a","https://access.redhat.com/security/cve/CVE-2016-9083","https://www.openwall.com/lists/oss-security/2016/10/26/11"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.0/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-190"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2016-9084","cve":"CVE-2016-9084","aliases":[],"title":"Linux kernel VFIO drivers/vfio/pci/vfio_pci_intrs.c - MSI/MSI-X allocation: Sibling of CVE-2016-9083 in the same","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel VFIO drivers/vfio/pci/vfio_pci_intrs.c - MSI/MSI-X allocation","year":"2016","cvss_score":7.8,"severity":"high","kev":false,"impact":"Sibling of CVE-2016-9083 in the same MSI/MSI-X setup path - the kzalloc size computation overflows, the kernel allocates a buffer far smaller than the interrupt array it is about to populate, and the heap is corrupted. Same reachability, same consequence: the process holding a tenant's passed-through GPU corrupts host kernel memory. Catalogued separately because the two were fixed in one commit and an operator checking only one CVE ID against a vendor changelog can convince themselves they are patched when they are not.","attack_vector":"Local process with access to a VFIO PCI device file - the tenant's VMM or a container granted /dev/vfio.","remediation":"Same fix as CVE-2016-9083 (commit 05692d7005a364add85c6e25a6c4447ce08f913a, RHSA-2017:0386); one kernel upgrade and rolling reboot closes both. Verify by kernel version rather than by CVE ID in a vendor advisory list, since backport changelogs frequently name only one of the pair.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=05692d7005a364add85c6e25a6c4447ce08f913a","https://access.redhat.com/security/cve/CVE-2016-9084","https://access.redhat.com/errata/RHSA-2017:0386"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.0/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-732"],"fleet":{"pain_class":"node-drain"},"id":"CVE-2017-0352","cve":"CVE-2017-0352","aliases":[],"title":"NVIDIA GPU Display Driver - GPU firmware access control over GPU control registers: Incorrect access control in the GPU","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - GPU firmware access control over GPU control registers","year":"2017","cvss_score":7.8,"severity":"high","kev":false,"impact":"Incorrect access control in the GPU firmware allows the CPU to reach sensitive GPU control registers it should not be able to touch, leading to privilege escalation. This is the register-level control plane of the accelerator - the surface that governs channel setup, memory-management-unit configuration and engine state, i.e. the machinery that is supposed to keep one tenant's GPU contexts from seeing another's. NVIDIA states all driver versions were affected. For an operator, this is the pre-2018 evidence that GPU isolation depends on firmware-enforced register permissions, not just on the driver, which is why 'we patched the driver' and 'the GPU is isolated' are different claims.","attack_vector":"Local - code running on the host CPU with GPU access, which on a shared node is any tenant workload holding a GPU.","remediation":"Update to a fixed NVIDIA driver branch; the correction shipped as part of the driver package because the driver loads the GPU firmware, so this is a driver update rather than a separate VBIOS flash. Same operational cost as any GPU driver bump - drain the node, stop every process holding the GPU, reload the kernel module. Confidence note for anyone reconstructing this: NVIDIA's original bulletin lived on the retired nvidia.custhelp.com host, so the primary vendor page is no longer resolvable and the record has to be read from the CVE and NVD entries.","references":["https://www.cve.org/CVERecord?id=CVE-2017-0352","https://nvd.nist.gov/vuln/detail/CVE-2017-0352","https://www.nvidia.com/en-us/security/"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2017-1000253","cve":"CVE-2017-1000253","aliases":[],"title":"Linux kernel (ELF loader): PIE stack buffer corruption, local root","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (ELF loader)","year":"2017","cvss_score":7.8,"severity":"high","kev":true,"impact":"PIE stack buffer corruption, local root [KEV]","attack_vector":"Local user / tenant process","remediation":"Livepatchable; otherwise drain + reboot. Only affects pre-2018 kernels - treat presence as a fleet-hygiene failure","references":["https://access.redhat.com/security/cve/CVE-2017-1000253"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2017-10-05"},{"id":"CVE-2017-5709","cve":"CVE-2017-5709","aliases":["INTEL-SA-00086","CVE-2017-5706","CVE-2017-5705"],"title":"Intel Server Platform Services (SPS) firmware 4.0 kernel","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Server Platform Services (SPS) firmware 4.0 kernel - the server-chipset variant of ME, Lewisburg PCH / Xeon…","year":"2017","cvss_score":7.8,"severity":"high","kev":false,"impact":"Multiple buffer overflows and privilege escalations in the SPS kernel let an unauthorized process on the host run code inside the SPS firmware itself. SPS is the management-engine flavour that ships on Xeon server chipsets - it owns Node Manager power capping, PECI thermal telemetry, and the sideband path to the BMC. Code execution there is below-the-OS persistence that survives a tenant reimage and is invisible to any host-based agent, and because SPS drives power and thermal management it carries direct physical consequence: an attacker in SPS can lie about power/thermal telemetry to the BMC and DCIM, or manipulate power limits on a node.","attack_vector":"A local unprivileged-to-privileged process on the host reaching the SPS firmware through the HECI/MEI interface. No network exposure required, no physical access. On a bare-metal GPU node this is reachable by any tenant who has root.","remediation":"SPS firmware flash, delivered only as an OEM BIOS/firmware bundle - Dell (iDRAC-driven DUP), HPE (SPP), Supermicro, Lenovo, Gigabyte, Quanta each ship their own, and for SA-00086 the OEM releases trailed Intel's November 2017 advisory by one to six months on server boards. Needs a full host reboot, so it drains every running training job on the node. There is no host-side mitigation: SPS cannot be disabled on a server chipset the way AMT can be unprovisioned. Verify the post-flash SPS version out of band via the BMC, not from the host.","references":["https://nvd.nist.gov/vuln/detail/CVE-2017-5709","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00086.html","https://cert-portal.siemens.com/productcert/pdf/ssa-892715.pdf"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"cvss_vector":"CVSS:3.0/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","fleet":{"pain_class":"node-drain"},"id":"CVE-2018-1431","cve":"CVE-2018-1431","aliases":[],"title":"IBM Spectrum Scale daemon (GSKit cryptographic library dependency): A local attacker takes control of the Spectrum","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"IBM Spectrum Scale daemon (GSKit cryptographic library dependency)","year":"2018","cvss_score":7.8,"severity":"high","kev":false,"impact":"A local attacker takes control of the Spectrum Scale daemon and from there reads and modifies any file in the shared filesystem, regardless of which tenant owns it. Training data, checkpoints and model artifacts are all in scope.","attack_vector":"Local account on a node running Spectrum Scale 4.1.1 through 5.0.0. The flaw is in the bundled GSKit crypto library that the daemon links, so it is reachable from ordinary node access rather than through a network service.","remediation":"Apply the Spectrum Scale efix that ships the fixed GSKit level. This one is worth treating as urgent rather than as routine third-party library hygiene, because the outcome is control of the daemon that enforces file ownership.","references":["https://www.ibm.com/support/docview.wss?uid=ssg1S1012049","https://nvd.nist.gov/vuln/detail/CVE-2018-1431"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2018-20669","cve":"CVE-2018-20669","aliases":[],"title":"Linux i915 GPU kernel driver (execbuffer2 ioctl): The execbuffer2 ioctl accepted a userspace-supplied address without","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux i915 GPU kernel driver (execbuffer2 ioctl)","year":"2018","cvss_score":7.8,"severity":"high","kev":false,"impact":"The execbuffer2 ioctl accepted a userspace-supplied address without an access_ok() check, so a local user submitting GPU work could get the kernel to touch an arbitrary address. Execbuffer is the hot path every GPU workload uses, which makes this trivially reachable from any GPU-enabled container.","attack_vector":"Any local user or container with a DRM render node - i.e. any tenant that was scheduled a GPU. No privileged capability needed.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-20669","http://git.kernel.org/cgit/linux/kernel/git/torvalds/linux.git/log/drivers/gpu/drm/i915/i915_gem_execbuffer.c"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2019-03-21"},{"id":"CVE-2018-6251","cve":"CVE-2018-6251","aliases":[],"title":"NVIDIA Windows GPU Display Driver, DirectX 10 user-mode driver: A crafted pixel shader writes into unallocated memory","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver, DirectX 10 user-mode driver","year":"2018","cvss_score":7.8,"severity":"high","kev":false,"impact":"A crafted pixel shader writes into unallocated memory in the user-mode driver. Because shaders are attacker-authored content in any remote-rendering, VDI or cloud-gaming path, this is a code-execution primitive reachable from whatever feeds you shaders.","attack_vector":"Anyone who can submit a shader to the host - a tenant VM in a vGPU/vSGA setup, a remote desktop session, or a local user.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://www.talosintelligence.com/vulnerability_reports/TALOS-2018-0514","https://nvd.nist.gov/vuln/detail/CVE-2018-6251"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2018-04-02"},{"id":"CVE-2019-0123","cve":"CVE-2019-0123","aliases":[],"title":"Intel processors supporting SGX (memory protection): Insufficient memory protection on SGX-capable processors gives","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors supporting SGX (memory protection)","year":"2019","cvss_score":7.8,"severity":"high","kev":false,"impact":"Insufficient memory protection on SGX-capable processors gives a privileged local user a privilege-escalation path. Practically, another entry on the list of reasons an SGX platform needs both microcode currency and enforced re-attestation.","attack_vector":"Privileged local access on the host.","remediation":"Microcode update and TCB recovery. Late-loadable microcode where the OS ships it; reboot required.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-0123","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00220.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"published":"2019-11-14"},{"id":"CVE-2019-0155","cve":"CVE-2019-0155","aliases":["iGPU Leak"],"title":"Intel processor graphics blitter command streamer: The graphics blitter accepted commands that could reference memory","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processor graphics blitter command streamer","year":"2019","cvss_score":7.8,"severity":"high","kev":false,"impact":"The graphics blitter accepted commands that could reference memory outside the submitting context, letting a local user with GPU access read or write memory belonging to another context. This is the archetype of the GPU isolation failure operators care about: one tenant's GPU commands reaching another tenant's GPU memory. Mitigated by kernel command parsing, which costs measurable submission throughput.","attack_vector":"Any local user or container able to submit GPU work through the DRM interface.","remediation":"Take the kernel fix (which re-enables blitter command parsing) plus the Intel graphics firmware/driver update, then reboot. Expect a performance regression on blitter-heavy workloads - that is the mitigation working, not a bug. Kernel + graphics driver, no BIOS.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-0155","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00242.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2019-11-14"},{"id":"CVE-2019-11085","cve":"CVE-2019-11085","aliases":[],"title":"Intel i915 graphics kernel-mode driver for Linux (< 5.0): Insufficient input validation in the i915 kernel-mode driver","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel i915 graphics kernel-mode driver for Linux (< 5.0)","year":"2019","cvss_score":7.8,"severity":"high","kev":false,"impact":"Insufficient input validation in the i915 kernel-mode driver gives a local authenticated user a path to privilege escalation. Old but still present on long-lived nodes running frozen enterprise kernels.","attack_vector":"Local user with access to the i915 DRM device - including any container granted /dev/dri.","remediation":"Move to a fixed kernel (5.0+ upstream or a distro backport) and reboot the node. Kernel-only fix, no firmware.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-11085","https://www.intel.com/content/www/us/en/security-center/advisory/INTEL-SA-00249.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2019-05-17"},{"id":"CVE-2019-13272","cve":"CVE-2019-13272","aliases":[],"title":"Linux kernel (ptrace): Broken permission and object lifetime handling for PTRACE_TRACEME, local root","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (ptrace)","year":"2019","cvss_score":7.8,"severity":"high","kev":true,"impact":"Broken permission and object lifetime handling for PTRACE_TRACEME, local root [KEV]","attack_vector":"Local user / tenant process","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2019-13272"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2019-07-17"},{"id":"CVE-2019-14565","cve":"CVE-2019-14565","aliases":[],"title":"Intel SGX SDK: Insufficient initialisation in the SGX SDK means enclaves built with the affected SDK can leak","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SGX SDK","year":"2019","cvss_score":7.8,"severity":"high","kev":false,"impact":"Insufficient initialisation in the SGX SDK means enclaves built with the affected SDK can leak uninitialised memory or be pushed into privilege escalation. The fix has to be applied by whoever builds the enclave, which for an operator means chasing your confidential-compute vendors rather than patching your own fleet.","attack_vector":"Local authenticated user interacting with an enclave built against a vulnerable SDK.","remediation":"Rebuild enclaves against SGX SDK 2.5 (Windows) / 2.7 (Linux) or later. Not an operator-side patch: it requires a new enclave binary from the software vendor, a re-signed enclave, and re-attestation. No node reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-14565","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00293.html"],"status":"curated","published":"2019-11-14"},{"id":"CVE-2019-14566","cve":"CVE-2019-14566","aliases":[],"title":"Intel SGX SDK: Insufficient input validation in the SGX SDK's generated edge routines, letting a local user","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SGX SDK","year":"2019","cvss_score":7.8,"severity":"high","kev":false,"impact":"Insufficient input validation in the SGX SDK's generated edge routines, letting a local user reach information disclosure, privilege escalation, or denial of service against an enclave built with the affected SDK.","attack_vector":"Local authenticated user calling into a vulnerable enclave.","remediation":"Rebuild and re-sign enclaves with a fixed SDK, then re-attest. Vendor-side fix; no operator reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-14566","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00293.html"],"status":"curated","published":"2019-11-14"},{"id":"CVE-2019-15752","cve":"CVE-2019-15752","aliases":[],"title":"Docker Desktop: Trojan docker-credential-wincred.exe in a world-writable path gives local privilege escalation","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker Desktop","year":"2019","cvss_score":7.8,"severity":"high","kev":true,"impact":"Trojan docker-credential-wincred.exe in a world-writable path gives local privilege escalation. [KEV]","attack_vector":"Local user on a developer/build Windows host","remediation":"Not a cluster-node issue; patch any Windows build/dev machines that run Docker Desktop","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-15752"],"status":"curated","published":"2019-08-28"},{"id":"CVE-2019-18181","cve":"CVE-2019-18181","aliases":[],"title":"Arista CloudVision Portal (Configlet Builder API): A read-only CloudVision user escapes their permissions through","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Arista CloudVision Portal (Configlet Builder API)","year":"2019","cvss_score":7.8,"severity":"high","kev":false,"impact":"A read-only CloudVision user escapes their permissions through Configlet Builder API calls and can execute restricted functionality. Read-only accounts are the ones handed out most freely — dashboards, auditors, tenant liaisons — so this is a large-blast-radius privilege escalation on the fabric controller.","attack_vector":"Any authenticated read-only CVP user with API access.","remediation":"CloudVision Portal upgrade. Controller-side, no switch impact. Review Configlet Builder execution history for anything run by a read-only principal.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-18181"],"status":"curated","published":"2019-12-19"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-59"],"fleet":{"pain_class":"hot-patch"},"id":"CVE-2019-3691","cve":"CVE-2019-3691","aliases":[],"title":"MUNGE (SUSE/openSUSE packaging): The munge package's install scripts follow symlinks, so a local attacker who controls","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"MUNGE (SUSE/openSUSE packaging)","year":"2019","cvss_score":7.8,"severity":"high","kev":false,"impact":"The munge package's install scripts follow symlinks, so a local attacker who controls the munge account can get root-owned files written to paths of their choosing. On a Slurm cluster the munge account is present on every node, which makes this a broad local escalation surface rather than a one-host issue.","attack_vector":"Local attacker with control of the munge user on a SUSE Linux Enterprise 15 or openSUSE host, exploited during package install or upgrade.","remediation":"Apply the SUSE munge package update on all nodes. Not applicable if you build MUNGE from source or run a non-SUSE distro.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-3691","https://bugzilla.suse.com/show_bug.cgi?id=1155075"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-74"],"fleet":{"pain_class":"node-drain"},"id":"CVE-2019-4558","cve":"CVE-2019-4558","aliases":[],"title":"IBM Spectrum Scale administrative command path: A local unprivileged user becomes root on a Storage Scale node by","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"IBM Spectrum Scale administrative command path","year":"2019","cvss_score":7.8,"severity":"high","kev":false,"impact":"A local unprivileged user becomes root on a Storage Scale node by injecting parameters into an administrative code path. Root on a storage node means the node's entire filesystem namespace, keys and cluster credentials are exposed.","attack_vector":"Local shell on any node running Storage Scale 4.2.0.0-4.2.3.17 or 5.0.0.0-5.0.3.2. No network position needed.","remediation":"Move to the fixed 4.2.3.18 / 5.0.3.3 level or later. Where an immediate upgrade is not possible, remove interactive shell access for non-admin users on storage and NSD server nodes.","references":["https://www.ibm.com/support/pages/node/1073732","https://nvd.nist.gov/vuln/detail/CVE-2019-4558"],"status":"curated"},{"id":"CVE-2019-5665","cve":"CVE-2019-5665","aliases":[],"title":"NVIDIA Windows GPU Display Driver, 3D Vision stereo service: A privileged NVIDIA service opens files without checking","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver, 3D Vision stereo service","year":"2019","cvss_score":7.8,"severity":"high","kev":false,"impact":"A privileged NVIDIA service opens files without checking for hard links, so an unprivileged user can point it at a file they should not be able to write and get arbitrary overwrite - the standard route to SYSTEM on Windows.","attack_vector":"Any local unprivileged user on a Windows host with the driver's stereo service running.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["http://support.lenovo.com/us/en/solutions/LEN-26250","https://nvd.nist.gov/vuln/detail/CVE-2019-5665"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2019-02-27"},{"id":"CVE-2019-5666","cve":"CVE-2019-5666","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): The context-creation DDI uses an untrusted array index","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2019","cvss_score":7.8,"severity":"high","kev":false,"impact":"The context-creation DDI uses an untrusted array index without validating it. Unprivileged local code gets an out-of-bounds kernel access, which NVIDIA rates as escalation-capable.","attack_vector":"Any local user with GPU device access on the host.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["http://support.lenovo.com/us/en/solutions/LEN-26250","https://nvd.nist.gov/vuln/detail/CVE-2019-5666"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2019-02-27"},{"id":"CVE-2019-5667","cve":"CVE-2019-5667","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): NULL dereference in the page-table DDI handler, reachable locally","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2019","cvss_score":7.8,"severity":"high","kev":false,"impact":"NULL dereference in the page-table DDI handler, reachable locally; crash the host, possibly worse.","attack_vector":"Any local user with GPU device access.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["http://support.lenovo.com/us/en/solutions/LEN-26250","https://nvd.nist.gov/vuln/detail/CVE-2019-5667"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2019-02-27"},{"id":"CVE-2019-5668","cve":"CVE-2019-5668","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): NULL dereference in the virtual command submission handler","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2019","cvss_score":7.8,"severity":"high","kev":false,"impact":"NULL dereference in the virtual command submission handler. Any local process that can submit work can bugcheck the node.","attack_vector":"Any local user with GPU device access.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["http://support.lenovo.com/us/en/solutions/LEN-26250","https://nvd.nist.gov/vuln/detail/CVE-2019-5668"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2019-02-27"},{"id":"CVE-2019-5669","cve":"CVE-2019-5669","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): Out-of-bounds kernel buffer access through the escape handler","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2019","cvss_score":7.8,"severity":"high","kev":false,"impact":"Out-of-bounds kernel buffer access through the escape handler from an unprivileged caller - denial of service, with escalation possible.","attack_vector":"Any local user with GPU device access.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["http://support.lenovo.com/us/en/solutions/LEN-26250","https://nvd.nist.gov/vuln/detail/CVE-2019-5669"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2019-02-27"},{"id":"CVE-2019-5670","cve":"CVE-2019-5670","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): Same escape-handler length bug but with information disclosure","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2019","cvss_score":7.8,"severity":"high","kev":false,"impact":"Same escape-handler length bug but with information disclosure and code execution called out: kernel memory contents can leak back to an unprivileged caller.","attack_vector":"Any local user with GPU device access.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["http://support.lenovo.com/us/en/solutions/LEN-26250","https://nvd.nist.gov/vuln/detail/CVE-2019-5670"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2019-02-27"},{"id":"CVE-2019-5675","cve":"CVE-2019-5675","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): Unsynchronized shared state (static variables across threads)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2019","cvss_score":7.8,"severity":"high","kev":false,"impact":"Unsynchronized shared state (static variables across threads) in the escape handler. Two threads racing produce undefined kernel behaviour - crash, escalation or information disclosure depending on what lands.","attack_vector":"Any local user able to issue concurrent escape calls, which is any process with GPU access.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-5675"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2019-05-10"},{"id":"CVE-2019-5683","cve":"CVE-2019-5683","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): Hard-link attack on the user-mode video driver's trace logger","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2019","cvss_score":7.8,"severity":"high","kev":false,"impact":"Hard-link attack on the user-mode video driver's trace logger: a privileged component writes where an unprivileged user points it. Arbitrary file overwrite as SYSTEM.","attack_vector":"Any local unprivileged user on the Windows host.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://support.lenovo.com/us/en/product_security/LEN-28096","https://nvd.nist.gov/vuln/detail/CVE-2019-5683"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2019-08-06"},{"id":"CVE-2019-5690","cve":"CVE-2019-5690","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): The escape handler does not validate input buffer size","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2019","cvss_score":7.8,"severity":"high","kev":false,"impact":"The escape handler does not validate input buffer size. Unprivileged local caller gets a kernel memory-safety bug, rated escalation-capable.","attack_vector":"Any local user with GPU device access.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-5690"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2019-11-09"},{"id":"CVE-2019-5691","cve":"CVE-2019-5691","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): NULL dereference in the escape handler","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2019","cvss_score":7.8,"severity":"high","kev":false,"impact":"NULL dereference in the escape handler - unprivileged local crash of the GPU node, escalation not excluded.","attack_vector":"Any local user with GPU device access.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-5691"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2019-11-09"},{"id":"CVE-2019-5692","cve":"CVE-2019-5692","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): Untrusted array index used in the escape handler: an unprivileged","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2019","cvss_score":7.8,"severity":"high","kev":false,"impact":"Untrusted array index used in the escape handler: an unprivileged process indexes kernel memory of its choosing.","attack_vector":"Any local user with GPU device access.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-5692"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2019-11-09"},{"id":"CVE-2020-0561","cve":"CVE-2020-0561","aliases":[],"title":"Intel SGX SDK (< 2.6.100.1): Improper initialisation in the SGX SDK gives an authenticated local user a privilege","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SGX SDK (< 2.6.100.1)","year":"2020","cvss_score":7.8,"severity":"high","kev":false,"impact":"Improper initialisation in the SGX SDK gives an authenticated local user a privilege escalation path against enclaves built with it.","attack_vector":"Local authenticated user against a vulnerable enclave.","remediation":"Rebuild enclaves with SGX SDK 2.6.100.1 or later and re-attest. Vendor-side.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-0561","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00336.html"],"status":"curated","published":"2020-02-13"},{"id":"CVE-2020-10699","cve":"CVE-2020-10699","aliases":["CVE-2020-13867","CVE-2020-14019"],"title":"targetcli-fb 2.1.50/2.1.51 and rtslib-fb through 2.1.72 (configuration tooling for the Linux LIO iSCSI/NVMe-oF target)","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"targetcli-fb 2.1.50/2.1.51 and rtslib-fb through 2.1.72 (configuration tooling for the Linux LIO iSCSI/NVMe-oF target)","year":"2020","cvss_score":7.8,"severity":"high","kev":false,"impact":"The targetclid socket is world-writable, so any local user on the storage node can drive the tool that configures LIO - creating backstores, changing LUN mappings and ACLs, and thereby escalating to root. The related defects leave /etc/target, its backup directory and saveconfig.json with weak permissions, exposing the entire target configuration including initiator ACLs and CHAP secrets to any local reader. Concretely: a non-root foothold on a storage node becomes 'map any tenant's LUN to an initiator I control', which is a cross-tenant data compromise achieved entirely through the target's own supported configuration path, leaving no memory-corruption artefacts to find afterwards.","attack_vector":"Any unprivileged local account on the LIO target host - a container escape, a compromised exporter or backup agent, a shared ops account. Requires that the targetclid socket unit is enabled, which several distributions do by default.","remediation":"Distro package update for targetcli-fb and rtslib-fb; no reboot and no I/O interruption. Independently, mask targetclid.socket unless you actually use the daemon mode - most operators drive targetcli interactively or from configuration management and never need it. Fix permissions on /etc/target, its backups and saveconfig.json, and rotate any CHAP secrets that were stored there, since the file was readable before you fixed it.","references":["https://github.com/open-iscsi/targetcli-fb/issues/162","https://bugzilla.redhat.com/show_bug.cgi?id=CVE-2020-10699","https://nvd.nist.gov/vuln/detail/CVE-2020-10699"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2020-04-15"},{"id":"CVE-2020-12929","cve":"CVE-2020-12929","aliases":[],"title":"AMD PSP trusted applications shipped in the AMD Graphics Driver: Trusted applications bundled with the AMD graphics","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD PSP trusted applications shipped in the AMD Graphics Driver","year":"2020","cvss_score":7.8,"severity":"high","kev":false,"impact":"Trusted applications bundled with the AMD graphics driver and running on the PSP do not validate their parameters, letting a local attacker bypass security restrictions and get arbitrary code execution in the secure processor. Notable because the entry point is the GPU driver stack rather than platform firmware - a GPU-adjacent compromise reaching the platform root of trust.","attack_vector":"Local, via the graphics driver's PSP interface.","remediation":"Fixed by updating the AMD graphics driver package (which carries the PSP trusted applications), then reloading the driver or rebooting the node. Unlike the pure-firmware ASP issues this does not wait on an OEM BIOS cycle - it moves at driver-release speed, so it is one of the cheaper ASP-class fixes to deploy.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-12929","https://www.amd.com/en/resources/product-security.html"],"status":"curated","tags":["tenant-isolation"],"published":"2021-11-15"},{"id":"CVE-2020-12930","cve":"CVE-2020-12930","aliases":[],"title":"AMD Secure Processor (ASP) drivers: Improper parameter handling in the ASP driver layer lets an already-privileged","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor (ASP) drivers","year":"2020","cvss_score":7.8,"severity":"high","kev":false,"impact":"Improper parameter handling in the ASP driver layer lets an already-privileged attacker escalate further and corrupt state the platform relies on for integrity. The realistic gain is turning root on the host into control over firmware-level components, i.e. persistence that survives reimaging.","attack_vector":"Local, privileged. Requires root or equivalent on the host first.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-12930","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-11-09"},{"id":"CVE-2020-12931","cve":"CVE-2020-12931","aliases":[],"title":"AMD Secure Processor (ASP) kernel: Improper parameter handling in the ASP's own kernel gives a privileged attacker","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor (ASP) kernel","year":"2020","cvss_score":7.8,"severity":"high","kev":false,"impact":"Improper parameter handling in the ASP's own kernel gives a privileged attacker a path to elevate inside the secure processor and damage platform integrity. Same practical outcome as the ASP driver flaw: root on the box becomes control of the firmware trust anchor.","attack_vector":"Local, privileged.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-12931","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-11-09"},{"id":"CVE-2020-12961","cve":"CVE-2020-12961","aliases":[],"title":"AMD PSP - System Management Network privileged register zeroing: An attacker can zero any privileged register on the","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD PSP - System Management Network privileged register zeroing","year":"2020","cvss_score":7.8,"severity":"high","kev":false,"impact":"An attacker can zero any privileged register on the System Management Network via the PSP. Zeroing SMN registers is a general-purpose way to disable platform protections - lock bits, access-control gates, security configuration - which turns into a bypass of whatever those registers were enforcing. It is a primitive rather than a single bug: whatever protection you were relying on at the SMN level can be switched off.","attack_vector":"Local, privileged, through the PSP interface.","remediation":"Fixed in AMD reference firmware (AGESA / PSP / SEV firmware) and delivered only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo and the ODMs each rebuild and requalify AMD's AGESA drop before shipping. **Expect one to six months of OEM lag**, and on end-of-support platforms expect nothing. Applying it is a drain plus full power cycle, not a driver reload. Verify by reading back the PSP/SMU firmware version afterwards rather than trusting the BIOS version string.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-12961","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2021-11-16"},{"id":"CVE-2020-12964","cve":"CVE-2020-12964","aliases":[],"title":"AMD Radeon Kernel Mode driver - Escape 0x2000c00 call handler: A low-privileged attacker can drive the Radeon","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"AMD Radeon Kernel Mode driver - Escape 0x2000c00 call handler","year":"2020","cvss_score":7.8,"severity":"high","kev":false,"impact":"A low-privileged attacker can drive the Radeon kernel-mode driver's Escape 0x2000c00 handler into privilege escalation or denial of service. Escape call handlers are the driver's catch-all ioctl surface and historically its weakest - a low-privilege caller reaching kernel-mode escalation is the pattern that makes GPU device nodes dangerous to hand out.","attack_vector":"Local, low privilege - reachable by an ordinary user with the GPU device open.","remediation":"Update the AMD graphics driver and reload or reboot. Note this is the Windows kernel-mode driver surface; Linux ROCm fleets are not affected by this specific handler, though the lesson about escape/ioctl surfaces carries over.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-12964","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2021-11-15"},{"id":"CVE-2020-36787","cve":"CVE-2020-36787","aliases":[],"title":"ASPEED video engine driver clock/reset sequencing (drivers/media/platform/aspeed): The driver brings the video engine","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ASPEED video engine driver clock/reset sequencing (drivers/media/platform/aspeed)","year":"2020","cvss_score":7.8,"severity":"high","kev":false,"impact":"The driver brings the video engine out of reset in the wrong order relative to its two clocks, and the hardware responds by issuing DMA writes to effectively random BMC memory. This is a DMA engine scribbling on the management processor's RAM with no software mediating the target address - the failure mode is silent BMC state corruption and unexplained BMC hangs or reboots. Operators usually log these as flaky hardware and RMA the board; the actual cost is a management processor whose memory integrity you cannot reason about, on every node running an affected image with video capture enabled.","attack_vector":"Triggers on video engine initialization, i.e. whenever iKVM/video capture is started or restarted on an affected BMC image. No attacker required for the corruption itself, but a host-side tenant who can force display-mode changes can force the reinit repeatedly.","remediation":"Fixed in the kernel driver and backported to stable branches; delivery is a BMC firmware flash, per node, out-of-band, ODM-gated. Config-only mitigation is to leave the video capture service stopped on nodes that do not use graphical console. Worth pairing with a check of your BMC crash/reboot telemetry - if you have unexplained BMC resets on ASPEED nodes with iKVM enabled, this is a candidate cause rather than bad silicon.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-36787","https://git.kernel.org/stable/c/1dc1d30ac101bb8335d9852de2107af60c2580e7"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-02-28"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-732"],"fleet":{"pain_class":"hot-patch"},"id":"CVE-2020-4278","cve":"CVE-2020-4278","aliases":["IBM X-Force 176137"],"title":"IBM Platform LSF / Spectrum LSF Suite (debug configuration file permissions): With specific debug settings enabled, LSF","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"IBM Platform LSF / Spectrum LSF Suite (debug configuration file permissions)","year":"2020","cvss_score":7.8,"severity":"high","kev":false,"impact":"With specific debug settings enabled, LSF writes files with permissions loose enough for a local user to escalate to the privileges of the LSF daemons. On the master host that is effectively control of the scheduler, so it is a route from one tenant's shell to deciding what runs where.","attack_vector":"A local user on an LSF host where the debug configuration is active. Affects Platform LSF 9.1 and 10.1, Spectrum LSF Suite 10.2 and Spectrum Suite for HPA 10.2.","remediation":"Apply IBM's fix, and turn off the LSF debug settings in production - they are a troubleshooting aid, and leaving them on permanently is what makes this exploitable at all.","references":["https://www.ibm.com/support/pages/node/3357549","https://nvd.nist.gov/vuln/detail/CVE-2020-4278"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-287","CWE-798"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2020-4983","cve":"CVE-2020-4983","aliases":["IBM X-Force 192586"],"title":"IBM Spectrum LSF / LSF Suite (authentication, hard-coded credentials): A user who is merely allowed to submit LSF jobs","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"IBM Spectrum LSF / LSF Suite (authentication, hard-coded credentials)","year":"2020","cvss_score":7.8,"severity":"high","kev":false,"impact":"A user who is merely allowed to submit LSF jobs can execute arbitrary commands, via an authentication weakness backed by hard-coded credentials. Job-submission rights are the lowest privilege a cluster hands out, so this promotes every tenant to command execution in the scheduler's context.","attack_vector":"A user on the local network holding ordinary LSF job-submission privileges. Affects Spectrum LSF 10.1 and LSF Suite 10.2.","remediation":"Apply IBM's fix for Spectrum LSF. Hard-coded credentials mean the secret is public once the binary is - patching is the only real mitigation, network restriction only narrows who can try.","references":["https://www.ibm.com/support/pages/node/6395478","https://www.hpcsec.com/2021/01/14/cve-2020-4983/","https://nvd.nist.gov/vuln/detail/CVE-2020-4983"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2020-5963","cve":"CVE-2020-5963","aliases":[],"title":"NVIDIA GPU Display Driver, inter-process communication APIs: Improper access control on the driver's IPC surface gives","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver, inter-process communication APIs","year":"2020","cvss_score":7.8,"severity":"high","kev":false,"impact":"Improper access control on the driver's IPC surface gives a local attacker code execution, a crash, or a read of data crossing that IPC. This one also shipped as an Ubuntu security update, so Linux compute nodes are in scope, not just Windows workstations.","attack_vector":"Any local user or container on the host that can reach the driver's IPC endpoints.","remediation":"Install the fixed GPU Display Driver branch on both Windows and Linux nodes. The kernel component (nvlddmkm.sys / nvidia.ko) cannot be hot-swapped under load, so this is a node drain and reboot per host; restart the container runtime afterwards so mounted driver libraries match the kernel module. No VBIOS or BMC flash.","references":["https://usn.ubuntu.com/4404-1/","https://usn.ubuntu.com/4404-2/","https://nvd.nist.gov/vuln/detail/CVE-2020-5963"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2020-06-25"},{"id":"CVE-2020-5964","cve":"CVE-2020-5964","aliases":[],"title":"NVIDIA GPU Display Driver service host component: The service host can skip its integrity check on application","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver service host component","year":"2020","cvss_score":7.8,"severity":"high","kev":false,"impact":"The service host can skip its integrity check on application resources, so a local attacker who swaps a resource gets code execution in a privileged NVIDIA service.","attack_vector":"Local user able to write the resources the service loads.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5964"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2020-06-25"},{"id":"CVE-2020-5966","cve":"CVE-2020-5966","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): NULL dereference in the escape handler","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2020","cvss_score":7.8,"severity":"high","kev":false,"impact":"NULL dereference in the escape handler; unprivileged local caller crashes the node, escalation not excluded.","attack_vector":"Any local user with GPU device access.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5966"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2020-06-25"},{"id":"CVE-2020-5968","cve":"CVE-2020-5968","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): The vGPU plugin fails to bound an indexed or pointer-based access, and NVIDIA calls","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2020","cvss_score":7.8,"severity":"high","kev":false,"impact":"The vGPU plugin fails to bound an indexed or pointer-based access, and NVIDIA calls out code execution, escalation and information disclosure. This is the guest-to-host escape shape - a tenant VM running code in the hypervisor's vGPU plugin owns every other tenant on that GPU. Affects vGPU 8.x before 8.4, 9.x before 9.4, 10.x before 10.3.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU assigned.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5968"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2020-06-30"},{"id":"CVE-2020-5971","cve":"CVE-2020-5971","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): Out-of-bounds read in the vGPU plugin that NVIDIA rates as code-execution capable. A","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2020","cvss_score":7.8,"severity":"high","kev":false,"impact":"Out-of-bounds read in the vGPU plugin that NVIDIA rates as code-execution capable. A tenant VM reads past a host buffer - the direct route to leaking another tenant's data or to a host escape. vGPU 8.x before 8.4, 9.x before 9.4, 10.x before 10.3.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5971"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2020-06-30"},{"id":"CVE-2020-5980","cve":"CVE-2020-5980","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): A securely loaded system DLL then loads its own dependencies","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2020","cvss_score":7.8,"severity":"high","kev":false,"impact":"A securely loaded system DLL then loads its own dependencies insecurely - the fix for one preloading bug leaving a second-order one behind. Local user gets code execution in a privileged driver component.","attack_vector":"Local user with write access on the dependency search path.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5980"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2020-10-02"},{"id":"CVE-2020-5981","cve":"CVE-2020-5981","aliases":[],"title":"NVIDIA Windows GPU Display Driver, DirectX 11 user-mode driver (nvwgf2um.dll): Crafted shader triggers an out-of-bounds","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver, DirectX 11 user-mode driver (nvwgf2um.dll)","year":"2020","cvss_score":7.8,"severity":"high","kev":false,"impact":"Crafted shader triggers an out-of-bounds access in the DX11 user-mode driver, this time rated code-execution capable. Anywhere untrusted shaders reach the GPU - VDI, cloud gaming, remote rendering - treat it as RCE at the session's privilege level.","attack_vector":"Anyone who can submit shaders to the host, including from a guest VM or remote session.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5981"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2020-10-02"},{"id":"CVE-2020-5984","cve":"CVE-2020-5984","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): Use-after-free while releasing resources in the vGPU plugin, rated code-execution","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2020","cvss_score":7.8,"severity":"high","kev":false,"impact":"Use-after-free while releasing resources in the vGPU plugin, rated code-execution capable. Guest-triggered UAF in the host plugin is the classic guest-to-host escape. vGPU 8.x before 8.5, 10.x before 10.4, and 11.0.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5984"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2020-10-02"},{"id":"CVE-2020-5987","cve":"CVE-2020-5987","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): Parameters stay writable by the guest after the plugin has validated them - a","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2020","cvss_score":7.8,"severity":"high","kev":false,"impact":"Parameters stay writable by the guest after the plugin has validated them - a textbook double-fetch. The guest passes validation with benign values then swaps them, feeding invalid parameters into host handlers. Escalation on the hypervisor host. vGPU 8.x before 8.5, 10.x before 10.4, and 11.0.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU who can race the host's validation window.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5987"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2020-10-02"},{"id":"CVE-2020-5991","cve":"CVE-2020-5991","aliases":[],"title":"NVIDIA CUDA Toolkit, nvJPEG library: Out-of-bounds read/write in nvJPEG while decoding an image, rated code-execution","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA CUDA Toolkit, nvJPEG library","year":"2020","cvss_score":7.8,"severity":"high","kev":false,"impact":"Out-of-bounds read/write in nvJPEG while decoding an image, rated code-execution capable. nvJPEG sits in exactly the place you feed untrusted data: GPU-accelerated image decode in training and inference pipelines, DALI preprocessing, and media services. If your ingest path decodes user-uploaded JPEGs on the GPU, this is remote code execution reachable by whoever can upload an image.","attack_vector":"Anyone who can get a malformed JPEG into a pipeline that decodes with nvJPEG - typically an unauthenticated uploader, several hops upstream of the GPU.","remediation":"Upgrade the CUDA Toolkit to 11.1.1 or later. The real cost is not the host install: every container image, wheel, conda package and vendored build that statically carries the affected library has to be rebuilt and re-pushed, and running jobs restarted onto the new image. No driver reload, no node drain, no firmware flash - but a full image-fleet rebuild.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5991"],"status":"curated","published":"2020-10-30"},{"id":"CVE-2020-7053","cve":"CVE-2020-7053","aliases":[],"title":"Linux i915 GPU kernel driver: A use-after-free in the i915 GPU kernel driver. The general shape is that a GPU object is","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux i915 GPU kernel driver","year":"2020","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the i915 GPU kernel driver. The general shape is that a GPU object is freed on one path while another path still holds a reference to it, so a local user with GPU access can get the kernel to read or write freed memory. Exploitability varies by heap layout, but on a GPU node every such bug is reachable from inside a container that was granted /dev/dri - the same boundary that is supposed to separate tenants. Specific trigger: the per-process GTT close path freeing structures the VM still references.","attack_vector":"Any local user or container with a DRM render node - i.e. any tenant that was scheduled a GPU. No privileged capability needed.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-7053","https://git.kernel.org/cgit/linux/kernel/git/torvalds/linux.git/commit/?id=7dc40713618c884bf07c030d1ab1f47a9dc1f310"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2020-01-14"},{"id":"CVE-2021-0084","cve":"CVE-2021-0084","aliases":[],"title":"Intel RDMA driver for Ethernet X722 and 800 series (Linux): Improper input validation in the Intel RDMA Linux driver","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel RDMA driver for Ethernet X722 and 800 series (Linux)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"Improper input validation in the Intel RDMA Linux driver gives an authenticated user privilege escalation. Earlier member of the same family as the irdma access-control issue and the same reason to care: RDMA queue pairs are mapped into userspace by design.","attack_vector":"Authenticated local user with access to the RDMA verbs interface.","remediation":"Update the Intel RDMA driver to 1.3.19 or later. Module reload drops RDMA links - drain first.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-0084","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00515.html"],"status":"curated","fleet":{"pain_class":"node-drain"},"tags":["tenant-isolation"],"published":"2021-08-11"},{"id":"CVE-2021-1052","cve":"CVE-2021-1052","aliases":[],"title":"NVIDIA GPU Display Driver (Windows nvlddmkm.sys + Linux nvidia.ko): User-mode clients can reach legacy privileged APIs","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver (Windows nvlddmkm.sys + Linux nvidia.ko)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"User-mode clients can reach legacy privileged APIs in the kernel-mode layer through DxgkDdiEscape or ioctl. This is a whole class of privileged driver functionality left exposed to unprivileged callers - NVIDIA rates it escalation and disclosure capable on both Windows and Linux. On a Linux GPU node, any container with /dev/nvidiactl can reach it.","attack_vector":"Any local user or GPU container on the host. Container isolation does not help: the device node is the attack surface and the standard container toolkit maps it in.","remediation":"Install the fixed GPU Display Driver branch on both Windows and Linux nodes. The kernel component (nvlddmkm.sys / nvidia.ko) cannot be hot-swapped under load, so this is a node drain and reboot per host; restart the container runtime afterwards so mounted driver libraries match the kernel module. No VBIOS or BMC flash.","references":["https://security.gentoo.org/glsa/202310-02","https://nvd.nist.gov/vuln/detail/CVE-2021-1052"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-01-08"},{"id":"CVE-2021-1057","cve":"CVE-2021-1057","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): The vGPU plugin lets guests allocate resources they are not authorised to hold","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"The vGPU plugin lets guests allocate resources they are not authorised to hold, which NVIDIA describes as integrity and confidentiality loss. A tenant VM grabbing unauthorised host GPU resources is the isolation boundary failing outright. vGPU 8.x before 8.6, 11.0 before 11.3.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU on the affected host.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1057"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2021-01-08"},{"id":"CVE-2021-1059","cve":"CVE-2021-1059","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): Unvalidated guest index leads to integer overflow in the host plugin, then tampering","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"Unvalidated guest index leads to integer overflow in the host plugin, then tampering or disclosure. vGPU 8.x before 8.6, 11.0 before 11.3.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1059"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2021-01-08"},{"id":"CVE-2021-1063","cve":"CVE-2021-1063","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): Unvalidated offset produces a buffer overread in the host plugin. A tenant reads","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"Unvalidated offset produces a buffer overread in the host plugin. A tenant reads host memory past the intended buffer - the direct path to another tenant's data. vGPU 8.x before 8.6, 11.0 before 11.3.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1063"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2021-01-08"},{"id":"CVE-2021-1080","cve":"CVE-2021-1080","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): Unvalidated guest input in the vGPU Manager plugin giving host information","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"Unvalidated guest input in the vGPU Manager plugin giving host information disclosure, data tampering or a shared-GPU outage. Affects vGPU 12.x before 12.2, 11.x before 11.4, 8.x before 8.7.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1080"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2021-04-29"},{"id":"CVE-2021-1081","cve":"CVE-2021-1081","aliases":[],"title":"NVIDIA vGPU software (guest kernel-mode driver + vGPU plugin): Unvalidated length across the guest kernel-mode driver","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software (guest kernel-mode driver + vGPU plugin)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"Unvalidated length across the guest kernel-mode driver and vGPU plugin boundary. Tenant-to-host disclosure or tampering. vGPU 12.x before 12.2, 11.x before 11.4, 8.x before 8.7.","attack_vector":"Any user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the host and the vGPU guest driver inside each tenant VM to the fixed release. Host side is a node drain plus reboot; guest side is a per-VM driver install and reboot. Because the guest driver is inside tenant-controlled VMs, in a multi-tenant estate you cannot fully remediate the guest half yourself - the host-side upgrade is the control you own. No VBIOS flash.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1081"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2021-04-29"},{"id":"CVE-2021-1082","cve":"CVE-2021-1082","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): Unvalidated input length in the vGPU Manager plugin, yielding host disclosure","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"Unvalidated input length in the vGPU Manager plugin, yielding host disclosure, tampering or denial of service on the shared GPU. vGPU 12.x before 12.2, 11.x before 11.4, 8.x before 8.7.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1082"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2021-04-29"},{"id":"CVE-2021-1083","cve":"CVE-2021-1083","aliases":[],"title":"NVIDIA vGPU software (guest kernel-mode driver + vGPU plugin): Unvalidated length across the guest driver and vGPU","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software (guest kernel-mode driver + vGPU plugin)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"Unvalidated length across the guest driver and vGPU Manager boundary. vGPU 12.x before 12.2, 11.x before 11.4.","attack_vector":"Any user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the host and the vGPU guest driver inside each tenant VM to the fixed release. Host side is a node drain plus reboot; guest side is a per-VM driver install and reboot. Because the guest driver is inside tenant-controlled VMs, in a multi-tenant estate you cannot fully remediate the guest half yourself - the host-side upgrade is the control you own. No VBIOS flash.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1083"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2021-04-29"},{"id":"CVE-2021-1084","cve":"CVE-2021-1084","aliases":[],"title":"NVIDIA vGPU software (guest kernel-mode driver + vGPU plugin): Another unvalidated-length path between guest driver and","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software (guest kernel-mode driver + vGPU plugin)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"Another unvalidated-length path between guest driver and vGPU Manager, with the same tenant-to-host disclosure and tampering outcome. vGPU 12.x before 12.2, 11.x before 11.4.","attack_vector":"Any user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the host and the vGPU guest driver inside each tenant VM to the fixed release. Host side is a node drain plus reboot; guest side is a per-VM driver install and reboot. Because the guest driver is inside tenant-controlled VMs, in a multi-tenant estate you cannot fully remediate the guest half yourself - the host-side upgrade is the control you own. No VBIOS flash.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1084"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2021-04-29"},{"id":"CVE-2021-1089","cve":"CVE-2021-1089","aliases":[],"title":"NVIDIA GPU Display Driver for Windows, nvidia-smi: nvidia-smi loads DLLs from an uncontrolled path","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver for Windows, nvidia-smi","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"nvidia-smi loads DLLs from an uncontrolled path. nvidia-smi is run by monitoring agents, schedulers and health checks, usually as a privileged service and often on a timer - so an unprivileged user who can plant a DLL gets scheduled code execution as that service on every GPU node running the same monitoring stack.","attack_vector":"Any local user who can write to a directory on nvidia-smi's DLL search path on a Windows GPU host.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1089"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-07-22"},{"id":"CVE-2021-1097","cve":"CVE-2021-1097","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): The vGPU Manager trusts a length field in a guest request that does not match the","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"The vGPU Manager trusts a length field in a guest request that does not match the actual input. A malicious tenant lies about the size and gets host-side disclosure, tampering or a shared-GPU outage. vGPU 12.x before 12.3, 11.x before 11.5, 8.x before 8.8.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1097"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2021-07-21"},{"id":"CVE-2021-1098","cve":"CVE-2021-1098","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): The vGPU Manager fails to release resources on guest driver unload, and the guest","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"The vGPU Manager fails to release resources on guest driver unload, and the guest can then reuse them. A tenant that unloads and reloads its driver operates on stale host resources - potentially ones now belonging to another tenant. vGPU 12.x before 12.3, 11.x before 11.5, 8.x before 8.8.","attack_vector":"Any user inside a guest VM who can unload and reload the vGPU guest driver, which requires only admin inside their own VM.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1098"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2021-07-21"},{"id":"CVE-2021-1118","cve":"CVE-2021-1118","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): The guest OS can execute privileged operations through the vGPU Manager, which","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"The guest OS can execute privileged operations through the vGPU Manager, which NVIDIA rates as escalation, tampering and disclosure. Privileged host operations driven from inside a tenant VM is the escape scenario, not a degraded-service scenario.","attack_vector":"Any user inside a guest VM with a vGPU on the affected host.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1118"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2021-10-29"},{"id":"CVE-2021-22555","cve":"CVE-2021-22555","aliases":[],"title":"Linux kernel (netfilter x_tables): Heap out-of-bounds write in xt_compat_target_from_user()","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (netfilter x_tables)","year":"2021","cvss_score":7.8,"severity":"high","kev":true,"impact":"Heap out-of-bounds write in xt_compat_target_from_user(); reliable container escape via unprivileged userns + CAP_NET_ADMIN [KEV]","attack_vector":"Any tenant process in a container with a user namespace","remediation":"Livepatchable; otherwise drain + reboot. Compensating control: disable unprivileged user namespaces (`kernel.unprivileged_userns_clone=0`) - breaks rootless Podman/Apptainer, which many HPC tenants rely on","references":["https://access.redhat.com/security/cve/CVE-2021-22555"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-07-07"},{"id":"CVE-2021-25124","cve":"CVE-2021-25124","aliases":[],"title":"BMC firmware on the HPE Cloudline whitebox line: An attacker directs the BMC's video-deletion routine at arbitrary","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"BMC firmware on the HPE Cloudline whitebox line","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"An attacker directs the BMC's video-deletion routine at arbitrary filesystem paths, deleting files inside the controller. Destroying the right files on a BMC means destroying its configuration, its credential store, or its ability to boot - a denial of service against the out-of-band plane on nodes you may not be able to reach any other way. It also destroys the BMC-side record of what happened, which makes it a useful anti-forensics step after a more serious compromise. The broader operator lesson is inventory: Cloudline nodes look like HPE in your asset database but share a BMC codebase with the ODM whiteboxes, and this is one of sixteen CVEs published against that BMC's REST service in a single disclosure. CL5800 Gen9, CL5200 Gen9, CL4100 Gen10, CL3100 Gen10 and CL5800 Gen10. Path traversal in the deletevideo_func handler of spx_restservice. Cloudline is HPE's ODM-manufactured whitebox range and runs a MegaRAC-derived BMC, not iLO, so iLO advisories and iLO tooling do not cover it.","attack_vector":"Access to the BMC's spx_restservice REST interface on the management network. The advisory characterises the path as local to the BMC, meaning it needs a session against the controller rather than pure network anonymity.","remediation":"BMC firmware update per HPE's Cloudline advisory - HPE does publish a readable advisory document for this line, which puts it ahead of most whitebox vendors. Flash per node, out of band. The inventory action matters as much as the flash: audit your fleet for Cloudline nodes specifically and confirm they are being tracked against Cloudline BMC advisories rather than iLO ones, because tooling that assumes 'HPE server means iLO' will silently report them as covered when they are not.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-25124","https://support.hpe.com/hpsc/doc/public/display?docLocale=en_US&docId=emr_na-hpesbhf04073en_us"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-25749","cve":"CVE-2021-25749","aliases":[],"title":"Kubernetes (kubelet): Windows workloads run as ContainerAdministrator despite runAsNonRoot","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kubelet)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"Windows workloads run as ContainerAdministrator despite runAsNonRoot","attack_vector":"Any tenant workload on a Windows node","remediation":"Rolling kubelet upgrade with Windows node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-25749"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2023-05-24"},{"id":"CVE-2021-26315","cve":"CVE-2021-26315","aliases":[],"title":"AMD PSP boot ROM - integrity of decrypted firmware image: The PSP boot ROM authenticates and decrypts firmware but does","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD PSP boot ROM - integrity of decrypted firmware image","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"The PSP boot ROM authenticates and decrypts firmware but does not sufficiently verify the integrity of the *decrypted* image before using it. An attacker who can influence the encrypted blob can therefore get the boot ROM to execute content it never really validated - code execution in the earliest, most privileged stage of the platform, in mask ROM territory where no patch can reach the flawed check itself.","attack_vector":"Local, requires the ability to modify the firmware image in SPI ROM - root plus flash write, a compromised BMC, or supply-chain access.","remediation":"Fixed in AMD reference firmware (AGESA / PSP / SEV firmware) and delivered only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo and the ODMs each rebuild and requalify AMD's AGESA drop before shipping. **Expect one to six months of OEM lag**, and on end-of-support platforms expect nothing. Applying it is a drain plus full power cycle, not a driver reload. Verify by reading back the PSP/SMU firmware version afterwards rather than trusting the BIOS version string. Since the flawed logic is in boot ROM, the mitigation is in the firmware AMD ships around it rather than a fix to the ROM. Practical compensating controls: enforce SPI write protection, require signed BIOS update packages, and restrict which management paths can drive host flash.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26315","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2021-11-16"},{"id":"CVE-2021-26324","cve":"CVE-2021-26324","aliases":[],"title":"AMD SEV-ES Trusted Memory Region - SNP guest memory integrity: A bug in the SEV-ES Trusted Memory Region handling costs","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV-ES Trusted Memory Region - SNP guest memory integrity","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"A bug in the SEV-ES Trusted Memory Region handling costs memory integrity for SNP-active VMs. The TMR is the region the SEV firmware itself works in; a defect there means the component enforcing confidential-VM isolation can have its own working memory disturbed, and SNP guests lose the integrity guarantee they were sold.","attack_vector":"Requires host/hypervisor privilege on a machine running SNP guests.","remediation":"Fixed in AMD reference firmware (AGESA / PSP / SEV firmware) and delivered only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo and the ODMs each rebuild and requalify AMD's AGESA drop before shipping. **Expect one to six months of OEM lag**, and on end-of-support platforms expect nothing. Applying it is a drain plus full power cycle, not a driver reload. Verify by reading back the PSP/SMU firmware version afterwards rather than trusting the BIOS version string. This sits inside the SEV-SNP trust boundary, so the update moves the platform's reported TCB version: refresh VCEK certificates from AMD's KDS and update any attestation policy your tenants pin, or confidential guest launches will start failing right after the BIOS lands.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26324","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2022-05-10"},{"id":"CVE-2021-26331","cve":"CVE-2021-26331","aliases":["SMU mailbox manipulation"],"title":"AMD System Management Unit (SMU) mailbox interface: A malicious user can manipulate SMU mailbox entries and reach","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD System Management Unit (SMU) mailbox interface","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"A malicious user can manipulate SMU mailbox entries and reach arbitrary code execution in the System Management Unit. The SMU is the microcontroller that owns voltage, clock and thermal control for the package. Code execution there is not just another privilege boundary - it is PHYSICAL control of the part: undervolting to induce computational faults (the technique behind fault-injection attacks on secure enclaves), forcing thermal or power states that throttle or hard-shut-down a node, and doing it in a way the OS reports as a normal thermal event. On a dense GPU rack that is a denial-of-service lever against neighbours and potentially a hardware-damage lever. Sibling issues CVE-2021-26329 and CVE-2021-26330 are the overflow variants in the same interface.","attack_vector":"Local attacker able to reach the SMU mailbox - typically ring 0 on the host, or a bare-metal tenant.","remediation":"AGESA / BIOS update per AMD-SB-1021 from the OEM, flash plus reboot. Independently, restrict tenant access to power and thermal management interfaces (no raw MSR access, no vendor overclocking or power-tuning drivers in tenant images) - that mitigation is under your control and does not wait on a BIOS drop. Add out-of-band power and thermal telemetry so an SMU-driven event is distinguishable from a genuine cooling fault.","references":["https://www.amd.com/en/corporate/product-security/bulletin/amd-sb-1021","https://nvd.nist.gov/vuln/detail/CVE-2021-26331"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2021-11-16"},{"id":"CVE-2021-26335","cve":"CVE-2021-26335","aliases":[],"title":"AMD Secure Processor (ASP) bootloader - image header parsing: The ASP bootloader reads and acts on fields from a","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor (ASP) bootloader - image header parsing","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"The ASP bootloader reads and acts on fields from a firmware image header *before* it verifies that image's signature. Attacker-controlled values therefore reach range checks and pointer arithmetic in the pre-verification window, giving code execution in the secure processor's own bootloader. This is upstream of every signature check on the platform: an attacker who lands here owns the root of trust and can persist beneath any OS reinstall or disk wipe.","attack_vector":"Local. Requires the ability to place a crafted image where the ASP bootloader will parse it - in practice SPI ROM write access or a compromised firmware update path, so root plus flash access, or a supply-chain/refurbishment scenario.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. Because this is bootloader code inside the ASP, there is no software workaround and no kernel-side mitigation - you either get the OEM BIOS or you do not. In the meantime, the compensating control is guarding SPI write access: enable the platform's SPI ROM protection and BIOS write-protect, and treat any node that has been through third-party hands as untrusted.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26335","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2021-11-16"},{"id":"CVE-2021-26360","cve":"CVE-2021-26360","aliases":[],"title":"AMD Secure Processor - SoC security-configuration registers: A local attacker can make unauthorised changes to the","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor - SoC security-configuration registers","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"A local attacker can make unauthorised changes to the SoC's security configuration registers and from there corrupt the AMD Secure Processor's encrypted memory. Corrupting ASP memory is not just a crash - it is a route to influencing what the secure processor computes, which is the same engine that gates memory encryption and attestation for every confidential guest on the node.","attack_vector":"Local, with access to the SoC register interface - root on the host.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. Because this sits inside the SEV-SNP trust boundary, the update also moves the platform's reported TCB version: after patching you must refresh VCEK certificates from AMD's KDS and update whatever attestation policy your tenants (or your own confidential-VM control plane) pin against, or every guest launch will start failing validation.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26360","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2022-11-09"},{"id":"CVE-2021-26409","cve":"CVE-2021-26409","aliases":[],"title":"AMD SEV-ES - bounds checking on Reverse Map table memory: Insufficient bounds checking in SEV-ES lets an attacker","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV-ES - bounds checking on Reverse Map table memory","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"Insufficient bounds checking in SEV-ES lets an attacker corrupt Reverse Map table memory, breaking SEV-SNP memory integrity. The RMP is the single data structure that decides which physical page belongs to which guest; corrupting it is the most direct possible attack on confidential-VM isolation, because after that the hardware itself believes the wrong owner.","attack_vector":"Host/hypervisor-privileged attacker.","remediation":"Fixed in AMD reference firmware (AGESA / PSP / SEV firmware) and delivered only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo and the ODMs each rebuild and requalify AMD's AGESA drop before shipping. **Expect one to six months of OEM lag**, and on end-of-support platforms expect nothing. Applying it is a drain plus full power cycle, not a driver reload. Verify by reading back the PSP/SMU firmware version afterwards rather than trusting the BIOS version string. This sits inside the SEV-SNP trust boundary, so the update moves the platform's reported TCB version: refresh VCEK certificates from AMD's KDS and update any attestation policy your tenants pin, or confidential guest launches will start failing right after the BIOS lands. Treat RMP-corruption issues as the top tier of your SEV patch queue - everything else in SNP rests on the RMP being correct.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26409","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2023-01-11"},{"id":"CVE-2021-27365","cve":"CVE-2021-27365","aliases":[],"title":"Linux iSCSI: iSCSI netlink structures lack length checks","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux iSCSI","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"iSCSI netlink structures lack length checks -> heap overflow, unprivileged local root","attack_vector":"Local","remediation":"Data-plane: kernel patch + reboot on every node using iSCSI LUNs","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-27365"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-03-07"},{"id":"CVE-2021-28500","cve":"CVE-2021-28500","aliases":[],"title":"Arista EOS (AAA API): Incorrect AAA API usage enables unrestricted local device access","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (AAA API)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"Incorrect AAA API usage enables unrestricted local device access — an operator with limited role gets full switch control","attack_vector":"Local","remediation":"EOS upgrade fabric-wide","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-28500"],"status":"curated","published":"2022-01-14"},{"id":"CVE-2021-28501","cve":"CVE-2021-28501","aliases":[],"title":"Arista EOS (TerminAttr AAA): TerminAttr streaming-telemetry agent bypasses AAA, giving unauthorized local device access","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (TerminAttr AAA)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"TerminAttr streaming-telemetry agent bypasses AAA, giving unauthorized local device access — TerminAttr is the CloudVision telemetry agent running on essentially every Arista switch in a managed fabric","attack_vector":"Local","remediation":"EOS + TerminAttr upgrade across the fabric","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-28501"],"status":"curated","published":"2022-01-14"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-134"],"fleet":{"pain_class":"node-drain"},"id":"CVE-2021-29740","cve":"CVE-2021-29740","aliases":[],"title":"IBM Spectrum Scale core component (format string handling): A user with a shell on any node that runs Storage Scale","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"IBM Spectrum Scale core component (format string handling)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"A user with a shell on any node that runs Storage Scale gets arbitrary code execution inside the core storage process, which runs privileged. From there the whole node's view of the filesystem is under attacker control.","attack_vector":"Local, low-privileged account on a node running the Storage Scale core component - in practice any compute node with the client mounted, including a tenant's own job container if it can reach the host.","remediation":"Apply the Storage Scale efix listed in IBM's bulletin for the 5.0.x / 5.1.0.x line and restart the daemon on each node in a rolling fashion. Audit which non-admin accounts have local shells on nodes that run the core component.","references":["https://www.ibm.com/support/pages/node/6457629","https://nvd.nist.gov/vuln/detail/CVE-2021-29740"],"status":"curated"},{"id":"CVE-2021-3156","cve":"CVE-2021-3156","aliases":[],"title":"sudo: Baron Samedit: heap overflow in sudo argument parsing, root from any local account","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"sudo","year":"2021","cvss_score":7.8,"severity":"high","kev":true,"impact":"Baron Samedit: heap overflow in sudo argument parsing, root from any local account [KEV]","attack_vector":"Local user","remediation":"Package update only; no reboot","references":["https://access.redhat.com/security/cve/CVE-2021-3156"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2021-01-26"},{"id":"CVE-2021-33123","cve":"CVE-2021-33123","aliases":["INTEL-SA-00601","CVE-2021-0159","CVE-2021-33124","CVE-2021-33103"],"title":"BIOS Authenticated Code Module (ACM) for a broad set of Intel processors, including Xeon Scalable: Improper access","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"BIOS Authenticated Code Module (ACM) for a broad set of Intel processors, including Xeon Scalable","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"Improper access control (plus, in the sibling CVEs, an out-of-bounds write and an input-validation flaw) in the BIOS ACM. ACMs are Intel-signed code that runs in the authenticated-code execution mode used to establish the root of trust for Boot Guard and TXT - it is more privileged than SMM and more privileged than the hypervisor. Code execution or state corruption at ACM level lets an attacker subvert the measurement chain from underneath, which means a platform can present a valid measured-boot report while running attacker-controlled firmware. Persistence at this level is below-the-OS, survives reimage, and defeats attestation-based tenant-handoff checks.","attack_vector":"A privileged local user - local root or SMM-capable code on the host. On bare-metal GPU nodes handed to tenants with root, that is the tenant.","remediation":"The ACM ships inside the BIOS image, so the fix is a BIOS/platform-firmware update from the OEM (Dell, HPE, Supermicro, Lenovo, Gigabyte, Quanta, Wiwynn) - reboot and job drain, and for a May 2022 Intel advisory the OEM server BIOS releases spread across the rest of 2022. There is no configuration workaround: you cannot disable the ACM. If your fleet's trust story depends on Boot Guard or TXT measurements, treat un-updated nodes as not attestable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-33123","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00601.html","https://security.netapp.com/advisory/ntap-20220818-0003/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-33909","cve":"CVE-2021-33909","aliases":[],"title":"Linux kernel (seq_file / fs layer): Sequoia: size_t-to-int conversion in the filesystem layer, local root on default","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (seq_file / fs layer)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"Sequoia: size_t-to-int conversion in the filesystem layer, local root on default configs","attack_vector":"Local user / tenant process able to mount or traverse deep paths","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2021-33909"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-07-20"},{"id":"CVE-2021-34398","cve":"CVE-2021-34398","aliases":[],"title":"NVIDIA DCGM (nv-hostengine, DIAG module): Any local user can inject a shared library into the DCGM server","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DCGM (nv-hostengine, DIAG module)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"Any local user can inject a shared library into the DCGM server, which normally runs as root - so DCGM turns every unprivileged account on a GPU node into root. This is the single most operationally relevant NVIDIA bug of 2021 for a cluster operator, because DCGM is not optional infrastructure: it is what feeds dcgm-exporter, Prometheus, node health checks and scheduler telemetry, so it is installed and running as root on essentially every GPU node in a modern fleet. Combined with a container that has host access to the DCGM socket, it is a container-to-host root escape. Versions before 2.2.9.","attack_vector":"Any local user on a node running nv-hostengine, including anything running inside a container that can reach the DCGM socket or port - which is the normal configuration for dcgm-exporter deployments.","remediation":"Upgrade DCGM to 2.2.9 or later on every node running nv-hostengine, then restart the service (systemctl restart nvidia-dcgm) and restart dcgm-exporter containers so they link the new library. No GPU drain, no reboot, no firmware flash. Until you patch, stop running nv-hostengine as root or block local access to its socket/port - the DIAG path is reachable by any local user.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-34398"],"status":"curated","fleet":{"pain_class":"daemon-restart"},"published":"2021-08-13"},{"id":"CVE-2021-3490","cve":"CVE-2021-3490","aliases":[],"title":"Linux kernel (eBPF verifier): eBPF ALU32 bitwise-op bounds tracking flaw","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (eBPF verifier)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"eBPF ALU32 bitwise-op bounds tracking flaw - arbitrary kernel read/write from unprivileged BPF","attack_vector":"Any tenant process in a container where unprivileged BPF is enabled","remediation":"Livepatchable; otherwise drain + reboot. Durable control: `kernel.unprivileged_bpf_disabled=1` fleet-wide","references":["https://access.redhat.com/security/cve/CVE-2021-3490"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-06-04"},{"id":"CVE-2021-3560","cve":"CVE-2021-3560","aliases":[],"title":"polkit: Local privilege escalation via polkit_system_bus_name_get_creds_sync() race","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"polkit","year":"2021","cvss_score":7.8,"severity":"high","kev":true,"impact":"Local privilege escalation via polkit_system_bus_name_get_creds_sync() race [KEV]","attack_vector":"Local user","remediation":"Package update + restart polkitd; no reboot","references":["https://access.redhat.com/security/cve/CVE-2021-3560"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2022-02-16"},{"id":"CVE-2021-4034","cve":"CVE-2021-4034","aliases":[],"title":"polkit (pkexec): PwnKit: local privilege escalation to root via argv handling","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"polkit (pkexec)","year":"2021","cvss_score":7.8,"severity":"high","kev":true,"impact":"PwnKit: local privilege escalation to root via argv handling; no exploit prerequisites [KEV]","attack_vector":"Local user, incl. any shell inside a privileged/host-namespace container","remediation":"Package update only; no reboot. Interim mitigation: `chmod 0755 /usr/bin/pkexec` (strip setuid)","references":["https://access.redhat.com/security/cve/CVE-2021-4034"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-20"],"fleet":{"pain_class":"hot-patch"},"published":"2022-01-28"},{"id":"CVE-2021-41103","cve":"CVE-2021-41103","aliases":[],"title":"containerd: Container root dirs and plugin dirs created world-traversable","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"Container root dirs and plugin dirs created world-traversable; unprivileged host user can read/modify container state","attack_vector":"Any process on the node, including a container that has partial host filesystem access","remediation":"Rolling containerd upgrade plus permission fix; drain node","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-41103"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2021-10-04"},{"id":"CVE-2021-42252","cve":"CVE-2021-42252","aliases":[],"title":"ASPEED LPC control driver (drivers/soc/aspeed/aspeed-lpc-ctrl.c) in the OpenBMC kernel: A process on the BMC that can","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ASPEED LPC control driver (drivers/soc/aspeed/aspeed-lpc-ctrl.c) in the OpenBMC kernel","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"A process on the BMC that can open the Aspeed LPC control device gets to mmap past the region it is supposed to own and write into BMC kernel memory. The size check compares the wrong quantities, so it is not a bounds check at all. The payoff is BMC root/kernel from a lower-privileged BMC daemon - which matters because OpenBMC's whole isolation story is that bmcweb, ipmid and the host-mailbox daemons run as separate constrained users. Chain it behind any of the bmcweb memory-corruption bugs and you go from a crashed web server to full control of the management processor, which is the position you need to write flash and persist.","attack_vector":"Local on the BMC itself. Requires code execution as a user with access to the LPC control character device - i.e. an attacker who already landed on the BMC via a network daemon bug, or a malicious/compromised OpenBMC package.","remediation":"Kernel fix, landed in Linux 5.14.6 and backported. For an operator this is not a package update - the BMC kernel is baked into the firmware image, so it means a full BMC firmware flash per node, out-of-band, gated on your ODM rebasing their OpenBMC tree. Many ODM images sit years behind upstream. No config-only mitigation; the device node has to exist for host-BMC mailbox and flash-sharing features to work.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-42252","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=b49a0e69a7b1a68c8d3f64097d06dabb770fec96","https://security.netapp.com/advisory/ntap-20211112-0006/"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2021-10-11"},{"id":"CVE-2021-46771","cve":"CVE-2021-46771","aliases":[],"title":"AMD Secure Processor (ASP) firmware system-call interface: The ASP firmware does not validate addresses passed across","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor (ASP) firmware system-call interface","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"The ASP firmware does not validate addresses passed across its system-call boundary, so a compromised user application that can issue ASP syscalls can steer the secure processor into reading or writing memory of the caller's choosing, ending in code execution inside the ASP. That converts a userspace compromise into control of the platform's security engine.","attack_vector":"Local. Requires an already-compromised application with ASP syscall access - typically a privileged agent or a trusted application, not a plain tenant container.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-46771","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2022-05-10"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2021-46925","cve":"CVE-2021-46925","aliases":[],"title":"Linux kernel (net/smc): The CDC send-completion handler takes a lock inside an smc_sock that close() has already freed","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/smc)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"The CDC send-completion handler takes a lock inside an smc_sock that close() has already freed, so a tenant that closes an SMC connection while a control-data-connection message is still in flight gets a use-after-free write in tasklet context. The published crash is a fatal page fault in _raw_spin_lock from the RDMA completion path - a hard node panic, with the freed-object write usable as a corruption primitive.","attack_vector":"Local and unprivileged: open an SMC-R connection, send, then close while a CDC message is outstanding - the race is between smc_release() and the WR completion tasklet, so it is a timing loop any tenant can run. Requires SMC-R actually negotiating over an RDMA device on the node; the smc module itself autoloads from an unprivileged socket(AF_SMC, ...) call.","remediation":"Boot a kernel carrying the fix commits (adds a refcount so CDC completions cannot outlive the socket). Interim: blacklist smc, or deny socket family 43 to tenants, on nodes where SMC-R is not deliberately in use.","references":["https://git.kernel.org/stable/c/e8a5988a85c719ce7205cb00dcf0716dcf611332","https://git.kernel.org/stable/c/349d43127dac00c15231e8ffbcaabd70f7b0e544","https://nvd.nist.gov/vuln/detail/CVE-2021-46925"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2021-47012","cve":"CVE-2021-47012","aliases":[],"title":"Linux kernel (drivers/infiniband/sw/siw): Soft-iWARP memory-region allocation stores the memory object into the MR and","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/sw/siw)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"Soft-iWARP memory-region allocation stores the memory object into the MR and then frees it if ID allocation fails, leaving the MR pointing at freed memory that the error path immediately dereferences. A use-after-free on the memory-registration path, reachable without any RDMA hardware.","attack_vector":"An unprivileged process on a node with siw loaded registers memory regions until the ID allocator fails - a container can drive that with its own resource limits. No HCA, no fabric peer, no host root.","remediation":"No fixed version is listed in the record - take the stable kernel carrying 30b9e92d0b5e (or 608a4b90ece0 / 3e22b88e02c1) and reboot. Interim: unload and blacklist siw where soft-iWARP is not required.","references":["https://git.kernel.org/stable/c/30b9e92d0b5e5d5dc1101ab856c17009537cbca4","https://git.kernel.org/stable/c/608a4b90ece039940e9425ee2b39c8beff27e00c","https://nvd.nist.gov/vuln/detail/CVE-2021-47012"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2021-47046","cve":"CVE-2021-47046","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Fix off by one in hdmi_14_process_transaction()","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-47046","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-02-28"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-908"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2021-47078","cve":"CVE-2021-47078","aliases":[],"title":"Linux kernel (drivers/infiniband/sw/rxe): When soft-RoCE queue-pair initialisation fails, the QP structure is left full","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/sw/rxe)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"When soft-RoCE queue-pair initialisation fails, the QP structure is left full of uninitialised garbage and the cleanup path treats that garbage as valid pointers - refcount underflow and use-after-free on freed or bogus addresses. Found by syzkaller from unprivileged userspace, so it is a proven reachable memory-corruption primitive.","attack_vector":"An unprivileged process on any node with the rdma_rxe module loaded issues QP creation calls that fail partway (bad attributes or resource pressure). No hardware HCA, no fabric peer, no host root.","remediation":"No fixed version is listed in the record - take the stable kernel carrying c65391dd9f0a (or 6a8086a42dfb / f3783c415bf6) and reboot. Interim: unload and blacklist rdma_rxe where soft-RoCE is not required.","references":["https://git.kernel.org/stable/c/c65391dd9f0a47617e96e38bd27e277cbe1c40b0","https://git.kernel.org/stable/c/6a8086a42dfbf548a42bf2ae4faa291645c72c66","https://nvd.nist.gov/vuln/detail/CVE-2021-47078"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2021-47081","cve":"CVE-2021-47081","aliases":[],"title":"habanalabs kernel driver (gaudi_memset_device_memory): Use-after-free in the Gaudi device-memory memset path","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"habanalabs kernel driver (gaudi_memset_device_memory)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"Use-after-free in the Gaudi device-memory memset path: the command buffer is released on the error path and then dereferenced again. Gives a local accelerator user a kernel UAF - crash at minimum, and the usual UAF privilege-escalation potential with enough heap grooming.","attack_vector":"Local user holding the habanalabs device node, reached by driving the memset ioctl down an error path.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-47081","https://git.kernel.org/stable/c/115726c5d312b462c9d9931ea42becdfa838a076"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-03-01"},{"id":"CVE-2021-47142","cve":"CVE-2021-47142","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdgpu): A use-after-free in the amdkfd (KFD compute driver","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdgpu)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the amdkfd (KFD compute driver, /dev/kfd). Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amdgpu: Fix a use-after-free","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-47142","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-03-25"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2021-47196","cve":"CVE-2021-47196","aliases":[],"title":"Linux kernel (drivers/infiniband/core): The core set the send and receive completion-queue pointers on a queue pair","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/core)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"The core set the send and receive completion-queue pointers on a queue pair only after handing it to the driver, so when driver-side creation fails the destroy path walks unset pointers and writes into freed memory. The upstream report is a KASAN use-after-free reached from an ordinary userspace program on mlx5 - the exact NIC under most GPU clusters.","attack_vector":"A tenant container holding /dev/infiniband/uverbs* issues an ibv_create_qp that the hardware driver rejects. Failing QP creation is trivially arranged with bad attributes or resource exhaustion, so this is a one-syscall reach from inside a container - no fabric peer, no host root.","remediation":"No fixed version is listed in the record - take the stable kernel carrying b70e072feffa (or 6cd7397d01c4) and reboot. Interim: remove /dev/infiniband/* from untrusted containers.","references":["https://git.kernel.org/stable/c/b70e072feffa0ba5c41a99b9524b9878dee7748e","https://git.kernel.org/stable/c/6cd7397d01c4a3e09757840299e4f114f0aa5fa0","https://nvd.nist.gov/vuln/detail/CVE-2021-47196"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2021-47247","cve":"CVE-2021-47247","aliases":[],"title":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core/en/rep): The neighbour-update worker takes a reference on an","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core/en/rep)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"The neighbour-update worker takes a reference on an encapsulation entry that a concurrent TC filter update is already freeing. KASAN reports a use-after-free while the worker is deciding which hardware encap rule to reprogram, so freed memory influences the steering state of the shared eswitch and the node crashes.","attack_vector":"Two concurrent drivers are needed and both are cheap to supply: neighbour (ARP/ND) churn from a host on the fabric - including a tenant VM with an address on the same segment - and concurrent TC filter add/delete from the host's offload agent. Requires switchdev/eswitch mode with tunnel-encap TC offload, the standard configuration for OVS hardware offload on a multi-tenant node.","remediation":"Update to a patched kernel on your stream. Interim: disable hw-tc-offload on the mlx5 uplink or stop offloading tunnel-encap rules, and keep untrusted tenants off the L2 segment that feeds neighbour updates.","references":["https://git.kernel.org/stable/c/0d1e7a7964ce6abb28883a3906bbc20fe0009f03","https://git.kernel.org/stable/c/b6447b72aca571632e71bb73a797118d5ce46a93","https://nvd.nist.gov/vuln/detail/CVE-2021-47247"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787","CWE-125"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2021-47261","cve":"CVE-2021-47261","aliases":[],"title":"NVIDIA/Mellanox ConnectX driver (mlx5_ib completion-queue resize, init_cq_frag_buf): CQ resize initialised the wrong","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA/Mellanox ConnectX driver (mlx5_ib completion-queue resize, init_cq_frag_buf)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"CQ resize initialised the wrong buffer. Because get_cqe() always returns entries from the current cq->buf, enlarging a completion queue made the driver write initialisation patterns past the end of the smaller live buffer instead of into the new one - an out-of-bounds write whose length the tenant controls by choosing the resize delta, ending in a kernel panic in the reported case.","attack_vector":"Local, unprivileged. A tenant calls resize_cq with a larger size on an mlx5 device.","remediation":"Kernel update making init_cq_frag_buf() address the buffer actually being initialised. No configuration workaround - CQ resize is part of the standard verbs API.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=1ec2dcd680c71d0d36fa25638b327a468babd5c9","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2021/CVE-2021-47261.json"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2021-47391","cve":"CVE-2021-47391","aliases":[],"title":"Linux kernel (drivers/infiniband/core): The RDMA connection-manager state machine can be driven in a circle so two","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/core)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"The RDMA connection-manager state machine can be driven in a circle so two address-resolution requests are outstanding for the same connection ID. Cancellation only removes the first, so the ID is freed while the second request is still queued, and the address handler then runs against freed memory - a use-after-free in the shared connection-setup path.","attack_vector":"A tenant container holding /dev/infiniband/rdma_cm drives rdma_resolve_addr through the address-error/rebind flow twice and then destroys the ID. Timing is helped by whatever the fabric does to route resolution, but no fabric peer cooperation is required and no host root is needed.","remediation":"No fixed version is listed in the record - take the stable kernel carrying 9a085fa9b7d6 (or 03d884671572 / 305d568b72f1) and reboot. Interim: remove /dev/infiniband/rdma_cm from containers that do not need connection management (many RDMA workloads only need uverbs).","references":["https://git.kernel.org/stable/c/9a085fa9b7d644a234465091e038c1911e1a4f2a","https://git.kernel.org/stable/c/03d884671572af8bcfbc9e63944c1021efce7589","https://nvd.nist.gov/vuln/detail/CVE-2021-47391"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2021-47421","cve":"CVE-2021-47421","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): A race condition or locking defect in the amdgpu kernel driver","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"A race condition or locking defect in the amdgpu kernel driver core. Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdgpu: handle the case of pci_channel_io_frozen only in amdgpu_pci_resume","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-47421","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-21"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-190","CWE-787"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2021-47485","cve":"CVE-2021-47485","aliases":[],"title":"Linux kernel InfiniBand qib driver (user SDMA path, qib_user_sdma_pkt): The user SDMA descriptor path did arithmetic on","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel InfiniBand qib driver (user SDMA path, qib_user_sdma_pkt)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"The user SDMA descriptor path did arithmetic on user-supplied buffer sizes without overflow checks, so a tenant could overflow addrlimit or bytes_togo and drive a kernel heap buffer overflow directly from the send path. The qib user SDMA interface is a zero-copy DMA submission channel - the whole point is that userspace hands the adapter descriptors - which makes an integer overflow here a straight route to controlled kernel memory corruption on a shared node.","attack_vector":"Local, unprivileged, via the qib character device on nodes carrying QLogic/Intel TrueScale InfiniBand HCAs. Still relevant because older IB fabrics get recycled into cheap secondary GPU clusters long after the frontier fleet has moved to ConnectX.","remediation":"Kernel update adding the overflow checks across the user-controlled arithmetic. If the fleet has no qib hardware, blacklist the ib_qib module - that is a cheap, reboot-free way to remove the surface entirely, and worth doing on any node where the driver is present but the hardware is not.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=0d4395477741608d123dad51def9fe50b7ebe952","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2021/CVE-2021-47485.json"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2021-47489","cve":"CVE-2021-47489","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amdgpu): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amdgpu)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: Fix even more out of bound writes from debugfs","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-47489","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-22"},{"id":"CVE-2021-47551","cve":"CVE-2021-47551","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amd/amdkfd): A correctness defect in the amdkfd (KFD compute","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amd/amdkfd)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdkfd (KFD compute driver, /dev/kfd) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/amdkfd: Fix kernel panic when reset failed and been triggered again","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-47551","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-24"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2021-47614","cve":"CVE-2021-47614","aliases":[],"title":"Linux kernel (drivers/infiniband/hw/irdma): In the physical-buffer-list allocator, a chunk is freed while still linked","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/hw/irdma)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"In the physical-buffer-list allocator, a chunk is freed while still linked on the PBLE info list, so the driver keeps walking and using freed memory. This is a use-after-free on the memory-registration path - a heap-grooming primitive for a container that wants host kernel code execution.","attack_vector":"A tenant container holding /dev/infiniband/uverbs* on an irdma node registers memory regions until the host-memory-cache segment allocation fails. A container can create that pressure deliberately with its own memory cgroup, so the failure branch is attacker-reachable rather than incidental.","remediation":"No fixed version is listed in the record - take the stable kernel carrying 11eebcf63e98 (or 1e11a39a82e9) and reboot. Interim: drop /dev/infiniband/* from untrusted containers on irdma nodes.","references":["https://git.kernel.org/stable/c/11eebcf63e98fcf047a876a51d76afdabc3b8b9b","https://git.kernel.org/stable/c/1e11a39a82e95ce86f849f40dda0d9c0498cebd9","https://nvd.nist.gov/vuln/detail/CVE-2021-47614"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-415"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2021-47616","cve":"CVE-2021-47616","aliases":[],"title":"Linux kernel (drivers/infiniband/sw/rxe): The soft-RoCE queue-pair init error path frees the send-queue ring and leaves","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/sw/rxe)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"The soft-RoCE queue-pair init error path frees the send-queue ring and leaves the pointer dangling, so the final reference drop frees it a second time. A double free of an attacker-sized ring buffer is a strong heap-corruption primitive for escaping a container.","attack_vector":"An unprivileged process on a node with rdma_rxe loaded creates queue pairs whose initialisation fails after the send queue is allocated. No HCA, no fabric peer, no host root.","remediation":"No fixed version is listed in the record - take the stable kernel carrying acb53e47db1f (or 84b01721e804) and reboot. Interim: unload and blacklist rdma_rxe on nodes that do not need soft-RoCE.","references":["https://git.kernel.org/stable/c/acb53e47db1fbc7cd37ab10b46388f045a76e383","https://git.kernel.org/stable/c/84b01721e8042cdd1e8ffeb648844a09cd4213e0","https://nvd.nist.gov/vuln/detail/CVE-2021-47616"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2021-47639","cve":"CVE-2021-47639","aliases":[],"title":"Linux kernel (arch/x86/kvm/mmu): The TDP MMU skipped invalid roots when unmapping a GFN range, so KVM could still hold","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (arch/x86/kvm/mmu)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"The TDP MMU skipped invalid roots when unmapping a GFN range, so KVM could still hold references to host pages after an MMU-notifier callback had returned - the exact guarantee that keeps a guest from touching host memory the kernel has already reclaimed. Completing the zap afterwards means KVM writes to and dirties pages that no longer belong to the VM.","attack_vector":"No special guest action needed: the window opens whenever a root is invalidated (memslot updates, VM teardown, nx_huge_pages toggling) while the host reclaims guest memory through the MMU notifier - swap, KSM, THP collapse, ballooning. On a memory-oversubscribed GPU node with several tenants, that is routine operation rather than an exotic race.","remediation":"Update to a kernel with the referenced stable commits (no fixed release string in the record - match by commit). Interim: avoid memory overcommit / aggressive reclaim on unpatched hypervisor nodes and do not toggle kvm.nx_huge_pages at runtime on a node with live guests.","references":["https://git.kernel.org/stable/c/af47248407c0c5ae52a752af1ab5ce5b0db91502","https://git.kernel.org/stable/c/0c8a8da182d4333d9bbb9131d765145568c847b2","https://nvd.nist.gov/vuln/detail/CVE-2021-47639"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2022-0185","cve":"CVE-2022-0185","aliases":[],"title":"Linux kernel (fs_context): Heap overflow in legacy filesystem parameter handling","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (fs_context)","year":"2022","cvss_score":7.8,"severity":"high","kev":true,"impact":"Heap overflow in legacy filesystem parameter handling; escapes unprivileged containers to host root [KEV]","attack_vector":"Any tenant process in a container with a user namespace","remediation":"Livepatchable; otherwise drain + reboot. Mitigate by disabling unprivileged user namespaces","references":["https://access.redhat.com/security/cve/CVE-2022-0185"],"status":"curated","fleet":{"ubiquity":"Universal - kernel 5.1 through 5.16.1; exploitable wherever unprivileged user namespaces are on, which is the default on Ubuntu GPU images","remediation_pain":"`node-reboot` - kernel upgrade; the only no-reboot mitigation is disabling unprivileged user namespaces, which breaks rootless/Podman-style tenant workflows","pain_class":"node-reboot","why_fleet_wide":"Heap overflow in `fs_context` gives a container-confined attacker full host root, demonstrated as a Kubernetes container escape on GKE/EKS/AKS-class engines - one tenant image compromises the whole node and its co-tenants"},"published":"2022-02-11"},{"id":"CVE-2022-0330","cve":"CVE-2022-0330","aliases":[],"title":"Linux i915 GPU kernel driver (GTT TLB handling): Stale GPU TLB entries let the GPU keep reading physical pages after","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux i915 GPU kernel driver (GTT TLB handling)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"Stale GPU TLB entries let the GPU keep reading physical pages after they were unmapped and handed to somebody else. A tenant running crafted GPU code reads whatever the host recycled those pages into - other tenants' data, or kernel memory. This is a true cross-tenant memory disclosure on Intel GPU nodes, not a crash bug.","attack_vector":"Any local user or container with a DRM render node - i.e. any tenant that was scheduled a GPU. No privileged capability needed.","remediation":"Update the kernel and reboot; the fix forces a full TLB flush on unbind, which costs GPU unbind throughput on memory-churning workloads. Drain the node - the driver cannot be swapped under live GPU jobs. Kernel-only, no firmware or microcode.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-0330","https://access.redhat.com/security/cve/CVE-2022-0330"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2022-03-25"},{"id":"CVE-2022-0847","cve":"CVE-2022-0847","aliases":[],"title":"Linux kernel (pipe): Dirty Pipe: uninitialised pipe_buffer flags allow overwriting read-only files, incl","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (pipe)","year":"2022","cvss_score":7.8,"severity":"high","kev":true,"impact":"Dirty Pipe: uninitialised pipe_buffer flags allow overwriting read-only files, incl. host binaries from inside a container [KEV]","attack_vector":"Any tenant process in a container","remediation":"Livepatchable (all major vendors shipped livepatches); otherwise drain + reboot. Highest-priority historical container-escape primitive","references":["https://access.redhat.com/security/cve/CVE-2022-0847"],"status":"curated","fleet":{"ubiquity":"Universal - every kernel 5.8+ before 5.16.11/5.15.25/5.10.102, which was the mainstream range on GPU hosts at the time","remediation_pain":"`node-reboot` - a kernel fix means a reboot of every GPU host unless the operator runs livepatch/kpatch; either way jobs must be drained first","pain_class":"node-reboot","why_fleet_wide":"An unprivileged process in a container overwrites read-only files, and the modification lands on the *host* file (page cache is shared), so a tenant overwrites host SUID binaries and takes the node"},"published":"2022-03-10"},{"id":"CVE-2022-21821","cve":"CVE-2022-21821","aliases":[],"title":"NVIDIA CUDA Toolkit - cuobjdump: An integer overflow reached by disassembling a corrupted fatbin gives remote code","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA CUDA Toolkit - cuobjdump","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"An integer overflow reached by disassembling a corrupted fatbin gives remote code execution in the context of the user running the tool. The realistic exposure is your build and profiling pipeline, not your runtime fleet: anything that automatically disassembles third-party fatbins, vendor kernels or model artifacts is running this parser on attacker-influenced input.","attack_vector":"Local, and requires a user or an automated job to run cuobjdump over an attacker-supplied file. CI jobs that inspect third-party CUDA binaries are the usual path.","remediation":"Update the CUDA Toolkit package (bulletin 5334). Cost: effectively zero - userspace SDK only, no driver reload, no node drain, no running-job impact. Rebuild build/CI images and move on.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-21821","https://github.com/NVIDIA/product-security/tree/main/2022/5334"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-190"],"published":"2022-03-29"},{"id":"CVE-2022-2588","cve":"CVE-2022-2588","aliases":[],"title":"Linux kernel (net/sched cls_route): Use-after-free in the cls_route filter","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/sched cls_route)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"Use-after-free in the cls_route filter - local privilege escalation, publicly exploited in container escapes","attack_vector":"Any tenant process in a container with CAP_NET_ADMIN in a userns","remediation":"Livepatchable; otherwise drain + reboot. Blacklist `cls_route` module as a stopgap","references":["https://access.redhat.com/security/cve/CVE-2022-2588"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-01-08"},{"id":"CVE-2022-26358","cve":"CVE-2022-26358","aliases":[],"title":"Xen on AMD-Vi - unity map handling on device reassignment: AMD-Vi unity mappings are not correctly torn down or","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen on AMD-Vi - unity map handling on device reassignment","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"AMD-Vi unity mappings are not correctly torn down or re-established when a device moves between guests, so mappings from a previous owner can persist. On a GPU cloud this is the device-handoff problem in its purest form: the accelerator you just reassigned to a new tenant may still carry DMA reach into the previous tenant's memory.","attack_vector":"Requires device reassignment between guests - i.e. exactly what happens when you recycle a passed-through GPU from one customer to the next.","remediation":"Fixed in Xen via XSA-400. Update the hypervisor and reboot. Operationally, treat GPU reassignment as a security transition: reset the device, verify IOMMU mappings are torn down, and prefer a host reboot between tenants where your margins allow it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-26358","https://xenbits.xen.org/xsa/advisory-400.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2022-04-05"},{"id":"CVE-2022-26359","cve":"CVE-2022-26359","aliases":[],"title":"Xen on AMD-Vi - unity map handling: Second XSA-400 AMD-Vi unity-map issue. Stale or incorrect IOMMU mappings across","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen on AMD-Vi - unity map handling","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"Second XSA-400 AMD-Vi unity-map issue. Stale or incorrect IOMMU mappings across device assignment leave DMA windows open into memory the current device owner should not reach.","attack_vector":"Guest with an assigned device, particularly across reassignment.","remediation":"Fixed in Xen (XSA-400). Hypervisor update plus reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-26359","https://xenbits.xen.org/xsa/advisory-400.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2022-04-05"},{"id":"CVE-2022-26360","cve":"CVE-2022-26360","aliases":[],"title":"Xen on AMD-Vi - unity map handling: Third XSA-400 AMD-Vi issue. Same class - IOMMU mappings that outlive their","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen on AMD-Vi - unity map handling","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"Third XSA-400 AMD-Vi issue. Same class - IOMMU mappings that outlive their justification give an assigned device cross-guest DMA reach.","attack_vector":"Guest with an assigned device.","remediation":"Fixed in Xen (XSA-400). Hypervisor update plus reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-26360","https://xenbits.xen.org/xsa/advisory-400.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2022-04-05"},{"id":"CVE-2022-26361","cve":"CVE-2022-26361","aliases":[],"title":"Xen on AMD-Vi - unity map handling: Fourth XSA-400 AMD-Vi issue. Patch the set together","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen on AMD-Vi - unity map handling","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"Fourth XSA-400 AMD-Vi issue. Patch the set together; individually they are variants of the same broken teardown of device DMA permissions.","attack_vector":"Guest with an assigned device.","remediation":"Fixed in Xen (XSA-400). Hypervisor update plus reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-26361","https://xenbits.xen.org/xsa/advisory-400.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2022-04-05"},{"id":"CVE-2022-27666","cve":"CVE-2022-27666","aliases":[],"title":"Linux kernel (IPsec ESP): Buffer overflow in the IPsec ESP transformation code - local root","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (IPsec ESP)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"Buffer overflow in the IPsec ESP transformation code - local root","attack_vector":"Any tenant process in a container with CAP_NET_ADMIN","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2022-27666"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2022-03-23"},{"id":"CVE-2022-29919","cve":"CVE-2022-29919","aliases":["INTEL-SA-00692","CVE-2022-45112","INTEL-SA-00846","CVE-2023-31271","INTEL-SA-00953","CVE-2024-23489"],"title":"Intel Virtual RAID on CPU (VROC) software before 7.7.6.1003, with follow-on issues through 8.6.0.1191: Use-after-free","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel Virtual RAID on CPU (VROC) software before 7.7.6.1003, with follow-on issues through 8.6.0.1191","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"Use-after-free in the VROC software giving an authenticated local user privilege escalation, followed by a run of access-control, default-permission and path-traversal escalations in later VROC releases. VROC is the software RAID layer sitting directly on NVMe on Xeon platforms - it is in the storage path for boot volumes and local scratch on many server designs. A local escalation here is a tenant-to-root path on any node where VROC is installed, and root on the node is the gateway to the whole firmware stack below it. The recurring pattern across four advisories is the useful signal: this component has a weak security history and should not be left installed where it is not needed.","attack_vector":"Authenticated local user on the host with the VROC software installed.","remediation":"Update VROC to 8.6.0.1191 or later (or the latest available for your platform) via the OEM's storage software package - Dell, HPE, Lenovo and Supermicro all redistribute it. Software update, so no firmware flash, but a reboot is typical. The better move for most GPU fleets: if you are not actually using VROC RAID, uninstall it rather than patch it - it is frequently present in golden images purely because it came with the platform driver bundle, and each release has brought a new local escalation.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-29919","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00692.html","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00953.html"],"status":"curated"},{"id":"CVE-2022-31606","cve":"CVE-2022-31606","aliases":[],"title":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys): Missing data validation lets a basic user cause","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing data validation lets a basic user cause an out-of-bounds access in kernel mode through DxgkDdiEscape, reaching privilege escalation and kernel information disclosure. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5383. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31606","https://github.com/NVIDIA/product-security/tree/main/2022/5383"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-drain"},"published":"2022-11-19"},{"id":"CVE-2022-31607","cve":"CVE-2022-31607","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): Missing input validation in nvidia.ko lets a basic local","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing input validation in nvidia.ko lets a basic local user escalate privileges, tamper with data, or leak a limited amount of kernel information - a genuine local root path from inside a GPU container. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5383. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31607","https://github.com/NVIDIA/product-security/tree/main/2022/5383"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-20"],"fleet":{"pain_class":"node-drain"},"published":"2022-11-19"},{"id":"CVE-2022-31608","cve":"CVE-2022-31608","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): An optional D-Bus configuration file shipped","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"An optional D-Bus configuration file shipped with the Linux driver leaves protected D-Bus endpoints reachable by a basic local user, which chains to code execution and privilege escalation on the host. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5383. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31608","https://github.com/NVIDIA/product-security/tree/main/2022/5383"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-281"],"fleet":{"pain_class":"node-drain"},"published":"2022-11-19"},{"id":"CVE-2022-31609","cve":"CVE-2022-31609","aliases":[],"title":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko): The vGPU plugin lets a guest VM","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"The vGPU plugin lets a guest VM allocate resources it is not authorised to hold, breaking the resource partition between tenants and reaching information disclosure and data tampering. The attacker sits inside a tenant guest VM and the blast radius is the hypervisor host, which is running every other tenant's vGPU on the same physical GPU. This is precisely the boundary a vGPU-based multi-tenant offering is sold on. This is a direct authorisation failure on the guest/host boundary, not a memory-safety accident - the guest simply asks for something it should not get and the plugin agrees.","attack_vector":"A tenant inside their own guest VM, driving the paravirtualised vGPU control interface. Several of these need only an unprivileged process in the guest; the rest need guest root, which a tenant already has on a VM they rented. No host credentials are involved at any point.","remediation":"Patch the vGPU Manager on the hypervisor host per bulletin 5383. Cost: the highest of any class here. The host driver cannot be reloaded while vGPUs are attached, so every tenant VM on that hypervisor must be live-migrated or powered off - a full host drain. NVIDIA also enforces a supported host/guest driver skew, so budget a matching guest-driver campaign in the same window or tenants lose their vGPU on next boot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31609","https://github.com/NVIDIA/product-security/tree/main/2022/5383"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-285"],"fleet":{"pain_class":"node-drain"},"tags":["tenant-isolation"],"published":"2022-08-05"},{"id":"CVE-2022-31610","cve":"CVE-2022-31610","aliases":[],"title":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys): An out-of-bounds write in the kernel mode layer","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds write in the kernel mode layer gives a basic local user code execution in kernel context - full compromise of the Windows GPU node. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5383. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31610","https://github.com/NVIDIA/product-security/tree/main/2022/5383"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-drain"},"published":"2022-11-19"},{"id":"CVE-2022-31617","cve":"CVE-2022-31617","aliases":[],"title":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys): An out-of-bounds read in the kernel mode layer","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds read in the kernel mode layer chains to code execution and privilege escalation on the node. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5383. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31617","https://github.com/NVIDIA/product-security/tree/main/2022/5383"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-125"],"fleet":{"pain_class":"node-drain"},"published":"2022-11-19"},{"id":"CVE-2022-32250","cve":"CVE-2022-32250","aliases":[],"title":"Linux kernel (netfilter): Use-after-free write in the netfilter subsystem - privilege escalation to root","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (netfilter)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"Use-after-free write in the netfilter subsystem - privilege escalation to root","attack_vector":"Any tenant process in a container with CAP_NET_ADMIN in a userns","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2022-32250"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2022-06-02"},{"id":"CVE-2022-34325","cve":"CVE-2022-34325","aliases":["INSYDE-SA-2022057"],"title":"Insyde InsydeH2O (StorageSecurityCommandDxe SMI input buffer, DMA TOCTOU): Highest-scored DMA entry in the 2022 batch","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (StorageSecurityCommandDxe SMI input buffer, DMA TOCTOU)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"Highest-scored DMA entry in the 2022 batch. StorageSecurityCommandDxe issues TCG/Opal security commands to self-encrypting drives, so this driver reaches SED authentication material. An attacker winning the race gets SMRAM corruption plus a position inside the code path that unlocks encrypted drives - which is precisely the control a GPU cloud relies on to claim tenant data is protected at rest between leases.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in the kernel releases named in the advisory (Insyde does not enumerate per-kernel versions for this one).  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34325","https://www.insyde.com/security-pledge/SA-2022057"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-11-14"},{"id":"CVE-2022-34670","cve":"CVE-2022-34670","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): Truncation when casting to a smaller primitive loses data","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"Truncation when casting to a smaller primitive loses data inside the kernel handler, giving an unprivileged user a denial of service or an information leak out of kernel memory. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5415. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34670","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-197"],"fleet":{"pain_class":"node-drain"},"published":"2022-12-30"},{"id":"CVE-2022-34918","cve":"CVE-2022-34918","aliases":[],"title":"Linux kernel (nf_tables): Heap overflow in nft_set_elem_init() - local root, weaponised inside containers","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (nf_tables)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"Heap overflow in nft_set_elem_init() - local root, weaponised inside containers","attack_vector":"Any tenant process in a container with CAP_NET_ADMIN in a userns","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2022-34918"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2022-07-04"},{"id":"CVE-2022-3650","cve":"CVE-2022-3650","aliases":[],"title":"Ceph: ceph-crash.service local privilege escalation to root plus privileged crash-dump disclosure","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ceph","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"ceph-crash.service local privilege escalation to root plus privileged crash-dump disclosure","attack_vector":"Local","remediation":"Data-plane: package update on every OSD/MON host","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-3650"],"status":"curated","published":"2023-01-17"},{"id":"CVE-2022-36763","cve":"CVE-2022-36763","aliases":["GHSA-xvv8-66cq-prwr"],"title":"EDK II SecurityPkg (Tcg2Dxe, Tcg2MeasureGptTable): A crafted GPT partition table overflows the heap inside the very","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"EDK II SecurityPkg (Tcg2Dxe, Tcg2MeasureGptTable)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"A crafted GPT partition table overflows the heap inside the very code that is supposed to measure the disk layout into the TPM. Two consequences that matter for a fleet: the attacker gets code execution during boot, and the compromise happens inside the measured-boot machinery itself, so the PCR values that downstream attestation trusts are produced by code the attacker already controls. Remote attestation of the node becomes meaningless while telling you everything is fine.","attack_vector":"Anyone who can present a disk with an attacker-controlled GPT to the node - a tenant who had the box before you and wrote to a local drive, a removable device, or an iSCSI/SAN LUN whose contents the attacker influences. Requires the node to boot with that disk attached.","remediation":"OEM BIOS update; the upstream edk2 fix predates public disclosure by over a year, so most current server BIOS lines already carry it - confirm against the OEM release notes for your exact platform generation rather than assuming. Flash + reboot per node. No config workaround inside firmware; operationally, wiping and re-partitioning tenant disks between leases reduces exposure but does not close the bug.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-36763","https://github.com/tianocore/edk2/security/advisories/GHSA-xvv8-66cq-prwr"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-01-09"},{"id":"CVE-2022-36765","cve":"CVE-2022-36765","aliases":["GHSA-ch4w-v7m3-g8wx"],"title":"EDK II MdePkg (CreateHob, HOB list construction): An integer overflow in the routine that allocates Hand-Off Blocks","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"EDK II MdePkg (CreateHob, HOB list construction)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"An integer overflow in the routine that allocates Hand-Off Blocks lets an attacker read and write outside the HOB list. HOBs are the structure PEI uses to describe memory, security state and platform configuration to DXE, so corrupting them is a way to lie to the rest of the boot about what is trusted - including the memory ranges Secure Boot and SMM protections are supposed to cover. Result is early-boot code execution with the firmware's own privileges.","attack_vector":"Requires influence over PEI-phase input on the node - typically a local attacker with the ability to feed the early boot path (attacker-controlled firmware volume content, a malicious capsule, or an earlier compromise that persists into PEI). Not remotely reachable.","remediation":"OEM BIOS update - the fix is a core MdePkg change, so every IBV downstream of edk2 had to rebase it, and coverage in shipped server images varies by platform generation. Flash + reboot per node. No configuration mitigates it; the only operational lever is keeping firmware-write paths (capsule update, SPI programming) locked down so an attacker cannot get to PEI in the first place.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-36765","https://github.com/tianocore/edk2/security/advisories/GHSA-ch4w-v7m3-g8wx"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-01-09"},{"id":"CVE-2022-37459","cve":"CVE-2022-37459","aliases":["AMP-SB-0004","Retbleed on Arm","Arm KA005138"],"title":"Ampere Altra before 1.08g and Altra Max before 2.05a - return address prediction: An attacker can control","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ampere Altra before 1.08g and Altra Max before 2.05a - return address prediction","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"An attacker can control return-address predictions and steer speculative execution into a chosen gadget, then read the result out of the cache. Same practical outcome as Spectre-v2 on these parts: cross-privilege and, where cores are shared, cross-tenant reads of memory the attacker has no rights to. Relevant to anyone running Altra as a CPU head node in front of GPU workers, since that node typically holds cluster credentials and customer data in flight.","attack_vector":"Unprivileged local code on an affected Altra / Altra Max part, targeting a victim on the same core or the same branch predictor structures. Local only.","remediation":"Update Altra firmware to 1.08g / Altra Max 2.05a or later from the board OEM, plus the corresponding kernel mitigations. Flash + reboot + drain. Speculation mitigations on Arm cost real throughput on syscall- and context-switch-heavy paths, so benchmark your actual serving stack rather than accepting the vendor's number. As with every predictor-sharing bug, not co-scheduling untrusted tenants on the same physical core is the mitigation that does not degrade over time.","references":["https://amperecomputing.com/products/security-bulletins/retbleed.html","https://developer.arm.com/documentation/ka005138/1-0/","https://nvd.nist.gov/vuln/detail/CVE-2022-37459"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2022-08-17"},{"id":"CVE-2022-4139","cve":"CVE-2022-4139","aliases":[],"title":"Linux i915 GPU kernel driver (TLB invalidation): An incorrect TLB flush in i915 leaves the GPU able to reach memory it","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux i915 GPU kernel driver (TLB invalidation)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"An incorrect TLB flush in i915 leaves the GPU able to reach memory it should no longer see, producing random memory corruption or leakage across contexts. Same family as the earlier GTT TLB bug and with the same consequence for a shared GPU node: one tenant's GPU reads or corrupts pages that now belong to someone else.","attack_vector":"Any local user or container with a DRM render node - i.e. any tenant that was scheduled a GPU. No privileged capability needed.","remediation":"Kernel update and reboot. Drain the node first. No firmware or microcode component; the fix is entirely in the driver's invalidation logic.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-4139","https://access.redhat.com/security/cve/CVE-2022-4139"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2023-01-27"},{"id":"CVE-2022-42260","cve":"CVE-2022-42260","aliases":[],"title":"NVIDIA vGPU software - guest driver (inside tenant VM): A D-Bus configuration file shipped with the Linux vGPU guest","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software - guest driver (inside tenant VM)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"A D-Bus configuration file shipped with the Linux vGPU guest driver leaves protected endpoints reachable, so an unauthorised user inside the guest VM reaches code execution and privilege escalation within that VM. Scope is confined to the guest VM, so this is a tenant-internal privilege problem rather than a break of your isolation boundary - but it is the first half of a chain if a vGPU Manager bug is also unpatched.","attack_vector":"An unprivileged user inside a guest VM that has a vGPU attached. You may not control these VMs at all if tenants bring their own images.","remediation":"Ship the fixed guest driver (bulletin 5415) to tenant VMs. Cost: low on your side, but you often cannot force it - if tenants own their guest images, the realistic control is a supported-driver-version policy plus refusing host attach below a floor version. No host drain required.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42260","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-281"],"published":"2022-12-30"},{"id":"CVE-2022-42261","cve":"CVE-2022-42261","aliases":[],"title":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko): An unvalidated input index in the vGPU","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"An unvalidated input index in the vGPU plugin produces a host-side buffer overrun, reaching data tampering, information disclosure or denial of service from inside a guest. The attacker sits inside a tenant guest VM and the blast radius is the hypervisor host, which is running every other tenant's vGPU on the same physical GPU. This is precisely the boundary a vGPU-based multi-tenant offering is sold on. A buffer overrun in host context driven by a guest-controlled index is the shape of a hypervisor escape; treat it as high priority even though NVIDIA scores it local.","attack_vector":"A tenant inside their own guest VM, driving the paravirtualised vGPU control interface. Several of these need only an unprivileged process in the guest; the rest need guest root, which a tenant already has on a VM they rented. No host credentials are involved at any point.","remediation":"Patch the vGPU Manager on the hypervisor host per bulletin 5415. Cost: the highest of any class here. The host driver cannot be reloaded while vGPUs are attached, so every tenant VM on that hypervisor must be live-migrated or powered off - a full host drain. NVIDIA also enforces a supported host/guest driver skew, so budget a matching guest-driver campaign in the same window or tenants lose their vGPU on next boot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42261","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-120"],"fleet":{"pain_class":"node-drain"},"tags":["tenant-isolation"],"published":"2022-12-30"},{"id":"CVE-2022-42267","cve":"CVE-2022-42267","aliases":[],"title":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys): An out-of-bounds read in the Windows driver","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds read in the Windows driver escalates to code execution and full node compromise from an ordinary user account. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5415. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42267","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-345"],"fleet":{"pain_class":"node-drain"},"published":"2022-12-30"},{"id":"CVE-2022-42268","cve":"CVE-2022-42268","aliases":[],"title":"Isaac Sim / Omniverse: Local privesc (improper input validation in config)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Isaac Sim / Omniverse","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc (improper input validation in config)","attack_vector":"Local user on the workstation/node","remediation":"Upgrade Omniverse packages; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42268","https://github.com/NVIDIA/product-security/tree/main/2023/5418"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-74"],"published":"2023-01-13"},{"id":"CVE-2022-42270","cve":"CVE-2022-42270","aliases":[],"title":"Jetson AGX Xavier / Orin bootloader: Local privesc (stack overflow in bootloader)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Jetson AGX Xavier / Orin bootloader","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc (stack overflow in bootloader)","attack_vector":"Local operator / physical","remediation":"Flash bootloader firmware out-of-band","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42270","https://github.com/NVIDIA/product-security/tree/main/2023/5442"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-121"],"published":"2022-12-30"},{"id":"CVE-2022-42274","cve":"CVE-2022-42274","aliases":[],"title":"DGX-2 BMC: RCE on BMC (buffer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX-2 BMC","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE on BMC (buffer overflow)","attack_vector":"Network-adjacent mgmt-LAN attacker","remediation":"Flash DGX-2 BMC firmware out-of-band","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42274","https://github.com/NVIDIA/product-security/tree/main/2024/5513"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-120"],"fleet":{"pain_class":"firmware-flash"},"published":"2023-01-13"},{"id":"CVE-2022-42332","cve":"CVE-2022-42332","aliases":["XSA-427"],"title":"Xen (shadow paging): x86 shadow plus log-dirty mode use-after-free - guest to host","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen (shadow paging)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"x86 shadow plus log-dirty mode use-after-free - guest to host","attack_vector":"Tenant VM guest","remediation":"Hypervisor patch + reboot/evacuation. Shadow paging is used during live migration, so this is reachable in normal operations","references":["https://xenbits.xen.org/xsa/advisory-427.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2023-03-21"},{"id":"CVE-2022-42973","cve":"CVE-2022-42973","aliases":["SEVD-2022-256-01"],"title":"APC Easy UPS Online Monitoring Software - embedded database credentials: Hardcoded credentials let any local user","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"APC Easy UPS Online Monitoring Software - embedded database credentials","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"Hardcoded credentials let any local user connect to the software's database and escalate. The database holds the site's UPS inventory and the credentials used to command shutdowns, so this converts a low-privilege foothold on one Windows host into control of the power-shutdown path.","attack_vector":"Local access to the machine running the monitoring software.","remediation":"Software upgrade. Hardcoded credentials mean the value is public once the advisory ships, so also confirm the database is not listening beyond localhost. Cheap to fix, and worth doing in the same window as the two RCEs above.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42973"],"status":"curated","published":"2023-02-01"},{"id":"CVE-2022-4318","cve":"CVE-2022-4318","aliases":[],"title":"CRI-O: Crafted environment variable injects arbitrary lines into /etc/passwd","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"CRI-O","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"Crafted environment variable injects arbitrary lines into /etc/passwd","attack_vector":"Malicious image or tenant-controlled pod env","remediation":"Upgrade CRI-O; drain node","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-4318"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2023-09-25"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-269"],"fleet":{"pain_class":"node-drain"},"id":"CVE-2022-43831","cve":"CVE-2022-43831","aliases":[],"title":"IBM Storage Scale Container Native Storage Access (pod security context): A local user in a CNSA-served container","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"IBM Storage Scale Container Native Storage Access (pod security context)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"A local user in a CNSA-served container escalates to privileged access on the host node when security contexts are not set as intended. Owning the GPU node means owning every other tenant's container on it.","attack_vector":"Execution inside a container on a node running Storage Scale CNSA 5.1.2.1 through 5.1.6.1 where the security context is left at its permissive default.","remediation":"Upgrade CNSA to the fixed level, and independently enforce restrictive pod security standards so the driver's containers cannot request host privileges. Verify by inspecting the running security context, not the chart defaults.","references":["https://www.ibm.com/support/pages/node/7015067","https://nvd.nist.gov/vuln/detail/CVE-2022-43831"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-78"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2022-43867","cve":"CVE-2022-43867","aliases":[],"title":"IBM Spectrum Scale container image (command execution): A local attacker runs arbitrary commands inside the Spectrum","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"IBM Spectrum Scale container image (command execution)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"A local attacker runs arbitrary commands inside the Spectrum Scale container, which is the process that brokers filesystem access for everything scheduled on that node.","attack_vector":"Local, low-privileged access to a node running Spectrum Scale 5.1.0.1 through 5.1.4.1 in container form.","remediation":"Pull the fixed Storage Scale container image per IBM's bulletin and redeploy the DaemonSet. Confirm the running image digest afterwards rather than trusting the tag.","references":["https://www.ibm.com/support/pages/node/6844771","https://nvd.nist.gov/vuln/detail/CVE-2022-43867"],"status":"curated"},{"id":"CVE-2022-48632","cve":"CVE-2022-48632","aliases":[],"title":"Linux kernel i2c-mlxbf (BlueField DPU I2C/SMBus controller): memcpy() is called in a loop with no upper bound","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel i2c-mlxbf (BlueField DPU I2C/SMBus controller)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"memcpy() is called in a loop with no upper bound on the operation length while the index keeps incrementing - a kernel stack overflow on the BlueField DPU itself. The DPU Arm cores run your infrastructure services (telemetry, storage emulation, security agents) alongside tenant-facing datapaths, so stack corruption there is compromise of the control point you deployed the DPU to be.","attack_vector":"Local on the DPU Arm side - a process with access to the SMBus/I2C interface. Relevant when tenants or third-party agents get any foothold on the DPU OS.","remediation":"Upgrade the DPU Arm-side kernel: 6.0 or a stable backport (5.10.146, 5.15.71, 5.19.12). In practice this arrives as a BFB bundle re-image or a DOCA/BFOS package upgrade on the DPU, followed by a DPU reset - which drops the host's network and storage while it reboots. Plan it as a per-node maintenance window, not a live update.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-48632","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2022/CVE-2022-48632.json"],"status":"curated","published":"2024-04-28"},{"id":"CVE-2022-48662","cve":"CVE-2022-48662","aliases":[],"title":"Linux i915 GPU kernel driver: A use-after-free in the i915 GPU kernel driver. The general shape is that a GPU object is","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux i915 GPU kernel driver","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the i915 GPU kernel driver. The general shape is that a GPU object is freed on one path while another path still holds a reference to it, so a local user with GPU access can get the kernel to read or write freed memory. Exploitability varies by heap layout, but on a GPU node every such bug is reachable from inside a container that was granted /dev/dri - the same boundary that is supposed to separate tenants. Specific trigger: the GEM context link being manipulated outside reference protection, which i915_perf then follows into freed memory.","attack_vector":"Any local user or container with a DRM render node - i.e. any tenant that was scheduled a GPU. No privileged capability needed.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-48662","https://git.kernel.org/stable/c/713fa3e4591f65f804bdc88e8648e219fabc9ee1"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-04-28"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-48721","cve":"CVE-2022-48721","aliases":[],"title":"Linux kernel (net/smc): An unprivileged tenant that opens an AF_SMC socket, registers it with epoll, and lets the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/smc)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"An unprivileged tenant that opens an AF_SMC socket, registers it with epoll, and lets the connection fall back to plain TCP leaves poisoned waitqueue entries behind; the TCP receive path then walks a corrupted list from softirq and takes a general protection fault. Any tenant can panic a shared node with a few dozen lines of ordinary socket code.","attack_vector":"Local, fully unprivileged, and on the common path: SMC falls back to TCP whenever the peer does not speak SMC, which is the default outcome for almost all traffic. socket(AF_SMC, ...) autoloads the smc module via the net-pf-43 alias with no capability check, so a plain tenant container - no /dev/infiniband, no RDMA device, no privileges - reaches it. The fault lands in NAPI/softirq context, so it takes the whole node down, not just the calling task.","remediation":"Update to 5.15.22 or later on the 5.15 branch, or any kernel carrying the fix commits. Interim: blacklist the smc module (`install smc /bin/false`) or deny socket family 43 in the tenant seccomp profile - there is no in-kernel toggle for this path.","references":["https://git.kernel.org/stable/c/0ef6049f664941bc0f75828b3a61877635048b27","https://git.kernel.org/stable/c/341adeec9adad0874f29a0a1af35638207352a39","https://nvd.nist.gov/vuln/detail/CVE-2022-48721"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-48726","cve":"CVE-2022-48726","aliases":[],"title":"Linux kernel (drivers/infiniband/core): A heap use-after-free in the userspace RDMA connection-manager interface.","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/core)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"A heap use-after-free in the userspace RDMA connection-manager interface. Multicast state is scanned without a lock while another thread is freeing it, so a tenant reads and follows a dangling pointer in kernel memory - the classic starting point for local privilege escalation out of a container.","attack_vector":"Any process holding /dev/infiniband/rdma_cm (ucma) - no capabilities needed. One thread leaves a multicast group while another destroys the CM context; syzkaller reaches it through plain write() calls on the device. Provider-independent: it works over soft-RoCE as well as real HCAs, so a node with rxe loaded is exposed even without RDMA hardware.","remediation":"No fixed release is published in this record - apply the listed stable fix commits or run a current stable kernel. Interim: remove /dev/infiniband/rdma_cm from tenant containers (most workloads that only do verbs do not need the CM character device) and unload rdma_rxe if unused.","references":["https://git.kernel.org/stable/c/75c610212b9f1756b9384911d3a2c347eee8031c","https://git.kernel.org/stable/c/2923948ffe0835f7114e948b35bcc42bc9b3baa1","https://nvd.nist.gov/vuln/detail/CVE-2022-48726"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-48771","cve":"CVE-2022-48771","aliases":[],"title":"Linux kernel (drivers/gpu/drm/vmwgfx): When the copy of the fence reply back to userspace fails, the driver installs a","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/vmwgfx)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"When the copy of the fence reply back to userspace fails, the driver installs a file descriptor it then tears down without releasing the fd table slot, leaving the tenant with a valid descriptor pointing at a freed file object. The tenant chooses when the copy fails, so this is a reliable, on-demand use-after-free on a core kernel object - the classic path to full kernel code execution from an unprivileged process.","attack_vector":"A tenant process holding /dev/dri/renderD* (or card*) on a vmwgfx device issues the execbuf or fence-event ioctl with a fence-reply pointer it has arranged to fault (unmapped or read-only page), then keeps using the leftover descriptor. Requires vmwgfx to be the DRM driver, i.e. workloads running inside VMware VMs. No privilege beyond the device node.","remediation":"Boot a kernel carrying the vmwgfx fd_install ordering fix below. Interim: drop /dev/dri/* from untrusted containers running on VMware-backed nodes, or blacklist vmwgfx on nodes where the guest does not need 3D.","references":["https://git.kernel.org/stable/c/e8d092a62449dcfc73517ca43963d2b8f44d0516","https://git.kernel.org/stable/c/0008a0c78fc33a84e2212a7c04e6b21a36ca6f4d","https://nvd.nist.gov/vuln/detail/CVE-2022-48771"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787","CWE-129"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-48883","cve":"CVE-2022-48883","aliases":[],"title":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core/ipoib): Creating an IPoIB PKEY child interface with fewer RX","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core/ipoib)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"Creating an IPoIB PKEY child interface with fewer RX queues than its parent lets the parent's RX channel index run off the end of the child's per-channel stats array on every packet, giving an out-of-bounds access driven by receive traffic. The parent IPoIB device is shared, so the corrupted memory and the resulting crash hit whoever else is using that interface.","attack_vector":"Reachable by anyone who can create an IPoIB child PKEY interface over netlink, i.e. CAP_NET_ADMIN in the namespace that owns the IPoIB parent. In InfiniBand GPU clusters where IPoIB netdevs are handed into tenant containers or pods that also hold NET_ADMIN, this is directly tenant-reachable; otherwise it is operator-only. Requires the mlx5 IPoIB path (ib_ipoib) in use - irrelevant on pure RoCE/Ethernet nodes.","remediation":"Boot a kernel with the mlx5e IPoIB validation fix (no fixed-version list published; match the stable commits below to your distro backport). Interim controls: do not grant NET_ADMIN to containers that hold an IPoIB parent netdev, and create PKEY child interfaces yourself with the parent's queue count rather than letting tenants do it.","references":["https://git.kernel.org/stable/c/5844a46f09f768da866d6b0ffbf1a9073266bf24","https://git.kernel.org/stable/c/31c70bfe58ef09fe36327ddcced9143a16e9e83d","https://nvd.nist.gov/vuln/detail/CVE-2022-48883"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-362","CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-48887","cve":"CVE-2022-48887","aliases":[],"title":"Linux kernel (drivers/gpu/drm/vmwgfx): User-resource lookup during command submission used a broken RCU fast path, so","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/vmwgfx)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"User-resource lookup during command submission used a broken RCU fast path, so two threads submitting command buffers against shared resources corrupt the lookup state and crash the kernel. The same race is a resource-lifetime bug, not just a crash: a tenant can drive the lookup to operate on a resource that is being torn down.","attack_vector":"A tenant process holding /dev/dri/renderD* on a vmwgfx device submits command buffers from two threads referencing shared surfaces - the upstream reproducer is the stock IGT vmwgfx execbuf stress test, so weaponizing it requires no novel research. Conditional on vmwgfx being the DRM driver (VMware-backed nodes).","remediation":"Boot a kernel carrying the vmwgfx resource-locking fix below. Interim: withhold /dev/dri/* from untrusted tenants on VMware-backed nodes, or blacklist vmwgfx where guests do not need 3D or shared surfaces.","references":["https://git.kernel.org/stable/c/7ac9578e45b20e3f3c0c8eb71f5417a499a7226a","https://git.kernel.org/stable/c/a309c7194e8a2f8bd4539b9449917913f6c2cd50","https://nvd.nist.gov/vuln/detail/CVE-2022-48887"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-48899","cve":"CVE-2022-48899","aliases":[],"title":"Linux kernel (drivers/gpu/drm/virtio): GEM handle values are guessable, and the driver dereferences the buffer object","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/virtio)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"GEM handle values are guessable, and the driver dereferences the buffer object after dropping the handle's reference, so a tenant that races object creation against a close of the guessed handle gets a use-after-free on a kernel object. That is an exploitable heap primitive from an unprivileged process inside the guest, i.e. container-to-guest-kernel privilege escalation.","attack_vector":"Any process inside a VM tenant holding /dev/dri/renderD* on a virtio-gpu device: two threads, one creating GEM objects, one closing the guessed handle number in a loop. Requires only the render node - no display access, no master, no host root. Reachable for any tenant given a virtio-gpu device (paravirtual GPU / vGPU-backed VMs).","remediation":"Boot a kernel carrying the virtgpu handle-creation fix below. Interim: remove /dev/dri/renderD* from workloads that do not actually need virtio-gpu acceleration, and prefer passing a real GPU device rather than virtio-gpu where a hostile guest process is in the threat model.","references":["https://git.kernel.org/stable/c/19ec87d06acfab2313ee82b2a689bf0c154e57ea","https://git.kernel.org/stable/c/d01d6d2b06c0d8390adf8f3ba08aa60b5642ef73","https://nvd.nist.gov/vuln/detail/CVE-2022-48899"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-48925","cve":"CVE-2022-48925","aliases":[],"title":"Linux kernel (drivers/infiniband/core): An unprivileged tenant corrupts RDMA connection-manager state and lands a","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/core)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"An unprivileged tenant corrupts RDMA connection-manager state and lands a use-after-free. An address-resolve call that should have been rejected outright still overwrites the source address of a listening connection id, which later makes cancellation walk a list it does not own - freed memory is read and linked into live kernel lists.","attack_vector":"Any process with /dev/infiniband/rdma_cm (ucma) inside a tenant container: put an id into LISTEN, then issue a resolve on the same id, then destroy it. Plain write() calls on the device, no capabilities and no cooperating peer required. Provider-independent, so soft-RoCE nodes are exposed too.","remediation":"No fixed release is published in this record - apply the listed stable fix commits or run a current stable kernel. Interim: drop /dev/infiniband/rdma_cm from tenant containers that do not need connection-manager access.","references":["https://git.kernel.org/stable/c/5b1cef5798b4fd6e4fd5522e7b8a26248beeacaa","https://git.kernel.org/stable/c/00265efbd3e5705038c9492a434fda8cf960c8a2","https://nvd.nist.gov/vuln/detail/CVE-2022-48925"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2022-48932","cve":"CVE-2022-48932","aliases":["net/mlx5 DR slab-out-of-bounds in mlx5_cmd_dr_create_fte"],"title":"Linux kernel mlx5_core software steering (fs_dr): Adding a flow rule with 32 destinations overflows an undersized","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_core software steering (fs_dr)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"Adding a flow rule with 32 destinations overflows an undersized action buffer in the software steering path - there was no check on the action count. Multi-destination rules are how you build mirroring and multicast replication in the NIC, so an operator or automation system building a large fan-out rule corrupts kernel slab memory.","attack_vector":"Local with network-configuration privilege (CAP_NET_ADMIN), including inside a user namespace on many setups. Triggered by installing a steering rule with a large destination list.","remediation":"Upgrade the host kernel to 5.17 or the 5.16.12 stable backport. Rolling reboot. Interim: cap the number of destinations your SDN controller or CNI is allowed to program per rule.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-48932","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2022/CVE-2022-48932.json"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-08-22"},{"id":"CVE-2022-48979","cve":"CVE-2022-48979","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: fix array index out of bound error in DCN32 DML","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-48979","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-48981","cve":"CVE-2022-48981","aliases":[],"title":"Linux kernel (drivers/gpu/drm): The shared shmem GEM mmap helper dropped a reference it never owned, so the buffer","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"The shared shmem GEM mmap helper dropped a reference it never owned, so the buffer object is freed while still mapped and still referenced - a use-after-free on a GEM object that the tenant continues to hold a mapping to. This is a strong exploitation primitive: the tenant keeps a live mapping over freed kernel-managed pages.","attack_vector":"A tenant process holding /dev/dri/renderD* on any driver built on the shmem GEM helpers (virtio-gpu and the shmem-backed DRM drivers) triggers it by mmap'ing a GEM object and forcing the helper's error path. Unprivileged, render-node only.","remediation":"Boot a kernel carrying the drm_gem_shmem_helper fix below. Interim: remove /dev/dri/renderD* from untrusted containers on nodes whose DRM driver uses the shmem helpers.","references":["https://git.kernel.org/stable/c/585a07b820059462e0c93b76c7de2cd946b26b40","https://git.kernel.org/stable/c/6a4da05acd062ae7774b6b19cef2b7d922902d36","https://nvd.nist.gov/vuln/detail/CVE-2022-48981"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2022-48990","cve":"CVE-2022-48990","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): A use-after-free in the amdgpu RAS / GPU reset and","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the amdgpu RAS / GPU reset and recovery path. Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amdgpu: fix use-after-free during gpu recovery","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-48990","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-10-21"},{"id":"CVE-2022-49025","cve":"CVE-2022-49025","aliases":["net/mlx5e use-after-free reverting termination table"],"title":"Linux kernel mlx5_core eswitch offloads (termination tables): Adding a multi-destination eswitch rule that partially","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel mlx5_core eswitch offloads (termination tables)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"Adding a multi-destination eswitch rule that partially fails leaves a stale termination-table pointer, and releasing the rule triggers a use-after-free in the eswitch - the exact subsystem that enforces which VF sees which traffic. Corruption here is a plausible route to host kernel control from a workload that only has network-namespace privilege.","attack_vector":"A local user who can add and delete tc flower rules with CAP_NET_ADMIN - obtainable in an unprivileged user namespace (unshare -Urn), so reachable from inside many container runtimes, not just from host root.","remediation":"Upgrade the host kernel to 6.1 or a stable backport (5.4.226, 5.10.158, 5.15.82, 6.0.12). Rolling reboot of the fleet. Interim: disable unprivileged user namespaces (kernel.unprivileged_userns_clone=0 / user.max_user_namespaces=0) where your container runtime does not need them - a sysctl config change, no reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49025","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2022/CVE-2022-49025.json"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-10-21"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-49076","cve":"CVE-2022-49076","aliases":[],"title":"Linux kernel (drivers/infiniband/hw/hfi1): The driver drops the last reference on a process's memory-descriptor","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/hw/hfi1)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"The driver drops the last reference on a process's memory-descriptor structure and then keeps using it while unpinning pages. A newly started process can be handed that same structure, so one tenant's cleanup corrupts another running task's address-space state - the reported symptoms include a mangled mmap lock leading to a system-wide hang.","attack_vector":"A tenant with the hfi1 character device in its container, doing pinned-memory RDMA (typical MPI workload), that exits or aborts abruptly while page unpinning is still in flight - an MPI_Abort is the documented trigger, so the tenant controls the timing. Requires hfi1 hardware (Intel Omni-Path).","remediation":"Update to 5.10 or later per the record, or apply the listed stable commits on your branch. Interim: keep the hfi1 device node out of untrusted containers; there is no runtime toggle that closes the race.","references":["https://git.kernel.org/stable/c/5f54364ff6cfcd14cddf5441c4a490bb28dd69f7","https://git.kernel.org/stable/c/9ca11bd8222a612de0d2f54d050bfcf61ae2883f","https://nvd.nist.gov/vuln/detail/CVE-2022-49076"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2022-49203","cve":"CVE-2022-49203","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A double free in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"A double free in the amdgpu display core (DC/DM). The same allocation is released twice, corrupting the slab allocator's freelist. This is a classic heap-corruption primitive: with slab grooming it becomes arbitrary kernel memory write and therefore host compromise from an unprivileged GPU workload. The cheap outcome is a node panic. Upstream fix: drm/amd/display: Fix double free during GPU reset on DC streams","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49203","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-02-26"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-125","CWE-787"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-49261","cve":"CVE-2022-49261","aliases":[],"title":"Linux kernel (drivers/gpu/drm/i915/gem): The access handler for a memory-mapped GEM object never bounds-checks the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/i915/gem)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"The access handler for a memory-mapped GEM object never bounds-checks the requested length before the memcpy, so a tenant can read and write kernel memory adjacent to the object's mapping. Out-of-bounds read gives another tenant's leftover buffer contents; out-of-bounds write gives kernel heap corruption from an unprivileged process.","attack_vector":"A tenant process holding /dev/dri/renderD* on an i915 GPU maps a GEM object and then drives the vm_access path against it (process_vm_readv/writev or ptrace-style access to its own mapping, which is what the public PoC does). No privilege beyond the render node, no display access needed.","remediation":"Boot a kernel carrying the i915 vm_access bounds-check fix below. Interim: drop /dev/dri/renderD* from containers that do not need i915 acceleration; there is no runtime toggle that closes this path while the device node is exposed.","references":["https://git.kernel.org/stable/c/89ddcc81914ab58cc203acc844f27d55ada8ec0e","https://git.kernel.org/stable/c/312d3d4f49e12f97260bcf972c848c3562126a18","https://nvd.nist.gov/vuln/detail/CVE-2022-49261"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-49426","cve":"CVE-2022-49426","aliases":[],"title":"Linux kernel (drivers/iommu/arm/arm-smmu-v3): The SMMUv3 SVA path released the pinned ASID without holding a reference","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/arm/arm-smmu-v3)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"The SMMUv3 SVA path released the pinned ASID without holding a reference on the mm, so the mm can be freed while the IOMMU still believes it owns that address-space ID. That is a use-after-free on the structure that decides which process's page tables a device is allowed to walk - the classic route from a local unprivileged process to kernel memory corruption, and on Arm hosts to an accelerator pointed at recycled page tables.","attack_vector":"Local, on Arm64 hosts using SMMUv3 SVA (this is the IOMMU on Grace/Grace-Hopper class systems). Any process that can bind an SVA/PASID context through an accelerator character device - a compute accelerator, an SVA-capable NIC queue, or an in-kernel SVA consumer - and then exit while the ASID is still pinned reaches it. No host root required; requires CONFIG_ARM_SMMU_V3_SVA and a device driver that offers SVA binding to userspace.","remediation":"No fixed release is listed in this record; apply the linked stable commits or run a current stable/LTS kernel on Arm nodes. Interim: disable SVA for accelerator drivers that expose it to tenants, or do not expose SVA-capable device nodes into tenant containers.","references":["https://git.kernel.org/stable/c/fc90f13ea0dcd960e5002d204fa55cec4e0db2fa","https://git.kernel.org/stable/c/e3cbbdbff8a4db5d053c53fd71be62ccccdb52b0","https://nvd.nist.gov/vuln/detail/CVE-2022-49426"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2022-49530","cve":"CVE-2022-49530","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): A use-after-free in the amdgpu power management","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the amdgpu power management (SMU/powerplay). Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amd/pm: fix double free in si_parse_power_table()","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49530","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-02-26"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-49562","cve":"CVE-2022-49562","aliases":[],"title":"Linux kernel (arch/x86/kvm/mmu): When guest memory is backed by a VM_PFNMAP mapping, KVM derived the target page frame","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (arch/x86/kvm/mmu)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"When guest memory is backed by a VM_PFNMAP mapping, KVM derived the target page frame from vm_pgoff, which is a file offset and has nothing to do with the mapped pfn. Updating guest page-table accessed/dirty bits therefore wrote into effectively arbitrary host physical pages - a tenant's ordinary page-table walks corrupt host memory outside its own VM.","attack_vector":"Guest-driven with no special guest privilege: every vCPU memory access can cause KVM to set A/D bits during a shadow page-table walk. Reachable whenever any part of the guest's address space is backed by a VM_PFNMAP VMA - device memory, /dev/mem backing, or a passthrough BAR mapped into the guest, which is the normal shape of a GPU-passthrough VM.","remediation":"Update to a stable kernel containing the linked fix (no fixed release is enumerated in the record; take the branch carrying commit f122dfe44768). Interim control: back tenant VMs with ordinary anonymous/hugetlb memory only and avoid VM_PFNMAP-backed memslots, including /dev/mem-backed regions, on unpatched hosts.","references":["https://git.kernel.org/stable/c/f122dfe4476890d60b8c679128cd2259ec96a24c","https://git.kernel.org/stable/c/8089e5e1d18402fb8152d6b6815450a36fffa9b0","https://nvd.nist.gov/vuln/detail/CVE-2022-49562"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-190","CWE-787"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-49963","cve":"CVE-2022-49963","aliases":[],"title":"Linux kernel (drivers/gpu/drm/i915/gt): The GPU migration copy path used plain ints for sizes that a tenant controls","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/i915/gt)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"The GPU migration copy path used plain ints for sizes that a tenant controls and fed the whole object size into a fixed-size mapping window, so any buffer object larger than the 8MB chunk size is copied incorrectly and objects past 64MB blow through the compression-block limit. Migration is how buffers move between VRAM and system memory, so the consequence is corrupted or partially-copied tenant data and out-of-range block counts on discrete Intel GPUs.","attack_vector":"A tenant container holding /dev/dri/renderD* on a discrete Intel GPU (DG2 class) only has to allocate large buffer objects and let them be evicted or swapped - the upstream reproducers are the stock GPU memory-swapping IGT tests with slightly larger object sizes. No privilege beyond the render node.","remediation":"Boot a kernel carrying the i915 migration/CCS fix below. Interim: cap per-tenant buffer-object and VRAM sizes to keep large-object migration out of play, and avoid overcommitting VRAM so eviction is rare.","references":["https://git.kernel.org/stable/c/97434cb55bd884bd268626ec41489f79b261b2d4","https://git.kernel.org/stable/c/8d905254162965c8e6be697d82c7dbf5d08f574d","https://nvd.nist.gov/vuln/detail/CVE-2022-49963"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2022-49969","cve":"CVE-2022-49969","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: clear optc underflow before turn off odm clock","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49969","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-06-18"},{"id":"CVE-2022-50035","cve":"CVE-2022-50035","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A use-after-free in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the amdgpu GEM/VM/command-submission ioctl surface. Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amdgpu: Fix use-after-free on amdgpu_bo_list mutex","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-50035","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-06-18"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-50137","cve":"CVE-2022-50137","aliases":[],"title":"Linux kernel (drivers/infiniband/hw/irdma): Use-after-free on completion-queue teardown. The driver frees the CQ","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/hw/irdma)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"Use-after-free on completion-queue teardown. The driver frees the CQ backing resources before it has stopped the interrupt path from processing completions against them, so an in-flight CQE lands in freed memory - a heap write primitive from an unprivileged tenant, not merely a crash.","attack_vector":"A tenant container holding /dev/infiniband/uverbs* on an Intel E810 / irdma node destroys a CQ while traffic is still arriving on the associated QP. The race window is widened by inbound fabric traffic, so a cooperating peer on the RDMA fabric makes it far easier to win. Requires the irdma module and Intel RDMA-capable NICs.","remediation":"No fixed release is published in this record - apply the listed stable fix commits or move to a current stable kernel. Interim: if tenants do not need RDMA, drop /dev/infiniband/* from their containers; if they do, treat irdma nodes as untrusted-tenant boundary hosts and prioritize the kernel update.","references":["https://git.kernel.org/stable/c/92520864ef9f912f38b403d172a0ded020683d55","https://git.kernel.org/stable/c/0abf2eef80295923b819ce89ff9edc1fe61be17c","https://nvd.nist.gov/vuln/detail/CVE-2022-50137"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2022-50303","cve":"CVE-2022-50303","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A double free in the amdkfd (KFD compute driver","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"A double free in the amdkfd (KFD compute driver, /dev/kfd). The same allocation is released twice, corrupting the slab allocator's freelist. This is a classic heap-corruption primitive: with slab grooming it becomes arbitrary kernel memory write and therefore host compromise from an unprivileged GPU workload. The cheap outcome is a node panic. Upstream fix: drm/amdkfd: Fix double release compute pasid","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-50303","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-09-15"},{"id":"CVE-2022-50354","cve":"CVE-2022-50354","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A NULL pointer dereference in the amdkfd (KFD compute","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"A NULL pointer dereference in the amdkfd (KFD compute driver, /dev/kfd). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdkfd: Fix kfd_process_device_init_vm error handling","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-50354","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-09-17"},{"id":"CVE-2022-50393","cve":"CVE-2022-50393","aliases":[],"title":"Linux kernel DRM scheduler / TTM / dma-buf shared layer used by amdgpu (drm/amdgpu): Missing or insufficient validation","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel DRM scheduler / TTM / dma-buf shared layer used by amdgpu (drm/amdgpu)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the DRM scheduler / TTM / dma-buf shared layer used by amdgpu. A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdgpu: SDMA update use unlocked iterator","attack_vector":"Local. Reachable by any process that can submit GPU work or import/export a dma-buf - i.e. any ROCm or graphics tenant on the node. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-50393","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-09-18"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787","CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-50440","cve":"CVE-2022-50440","aliases":[],"title":"Linux kernel (drivers/gpu/drm/vmwgfx): The dimensions of a DMA surface-copy box submitted in the command stream were","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/vmwgfx)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"The dimensions of a DMA surface-copy box submitted in the command stream were never validated against the size of the destination image buffer, so a tenant-chosen copy box overflows the kernel-side memcpy. That is an out-of-bounds kernel heap write with attacker-controlled length and content - memory corruption, not just a crash.","attack_vector":"A tenant process holding /dev/dri/renderD* on a vmwgfx device submits a surface DMA command with an oversized copy box through the normal command-submission ioctl; the overflow happens while the kernel validates and snoops that command, before any display involvement. Conditional on vmwgfx being the DRM driver (VMware-backed nodes).","remediation":"Boot a kernel carrying the vmwgfx copybox validation fix below. Interim: remove /dev/dri/* from untrusted containers on VMware-backed nodes, or blacklist vmwgfx where 3D is not required.","references":["https://git.kernel.org/stable/c/ee8d31836cbe7c26e207bfa0a4a726f0a25cfcf6","https://git.kernel.org/stable/c/50d177f90b63ea4138560e500d92be5e4c928186","https://nvd.nist.gov/vuln/detail/CVE-2022-50440"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-50454","cve":"CVE-2022-50454","aliases":[],"title":"Linux kernel (drivers/gpu/drm/nouveau): Importing a dma-buf whose backing buffer object fails to initialize leaves the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/nouveau)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"Importing a dma-buf whose backing buffer object fails to initialize leaves the driver taking a reference on memory TTM has already freed - a use-after-free driven entirely from the PRIME import path. A tenant that can make buffer-object init fail (memory pressure, oversized or badly aligned imports) gets a freed-object reference count bumped under its control.","attack_vector":"A tenant process holding /dev/dri/renderD* on a nouveau GPU calls the PRIME/dma-buf import ioctl on a descriptor it crafted, repeatedly, under memory pressure it creates itself. Requires nouveau to be the driver in use for the NVIDIA card (not the proprietary stack), which is the case for hosts running the open upstream driver.","remediation":"Boot a kernel carrying the nouveau prime import fix below. Interim: if the node runs nouveau, drop /dev/dri/renderD* from untrusted containers or block dma-buf import from tenant workloads that do not need cross-device buffer sharing.","references":["https://git.kernel.org/stable/c/56ee9577915dc06f55309901012a9ef68dbdb5a8","https://git.kernel.org/stable/c/5d6093c49c098d86c7b136aba9922df44aeb6944","https://nvd.nist.gov/vuln/detail/CVE-2022-50454"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2022-50528","cve":"CVE-2022-50528","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A memory or reference-count leak in the amdkfd (KFD","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"A memory or reference-count leak in the amdkfd (KFD compute driver, /dev/kfd). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amdkfd: Fix memory leakage","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-50528","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-10-07"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-415"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-50543","cve":"CVE-2022-50543","aliases":[],"title":"Linux kernel (drivers/infiniband/sw/rxe): A tenant gets a double free in the kernel heap through a failed memory","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/sw/rxe)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"A tenant gets a double free in the kernel heap through a failed memory registration. When user MR setup fails, both the error path and the object cleanup release the same page-map allocation, giving the tenant a controllable allocator-corruption primitive rather than a simple crash.","attack_vector":"Tenant container holding /dev/infiniband/uverbs* on a node with soft-RoCE (rdma_rxe) loaded: call reg_mr with an address/length that makes the user-memory pinning fail, which the tenant chooses freely. Unprivileged and deterministic. Hardware HCAs do not use this code.","remediation":"Update to a kernel carrying the fix (record cites 5.20-era mainline) or apply the listed stable commits. Interim: blacklist/unload rdma_rxe where soft-RoCE is not required, and remove /dev/infiniband/* from containers that do not use verbs.","references":["https://git.kernel.org/stable/c/6ce577f09013206e36e674cd27da3707b2278268","https://git.kernel.org/stable/c/06f73568f553b5be6ba7f6fe274d333ea29fc46d","https://nvd.nist.gov/vuln/detail/CVE-2022-50543"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-50726","cve":"CVE-2022-50726","aliases":[],"title":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core): The async firmware-command context can be freed while a","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"The async firmware-command context can be freed while a completion callback is still running, so the completion handler writes into a freed slab object from interrupt context. That is a kernel heap corruption primitive on a node shared with other tenants, and at minimum a panic that takes every workload on the box down with it.","attack_vector":"Needs code executing on the node that opens mlx5 async firmware-command channels and then tears them down while commands are in flight. The realistic holder is a tenant container with /dev/infiniband/uverbs* and DEVX enabled, or a VF driver instance inside a tenant VM; host-root paths (devlink reload, eswitch mode changes) reach the same race. Not reachable from the RDMA/IP fabric.","remediation":"Boot a kernel carrying the mlx5_core fix (the kernel CNA published no fixed-version list for this ID - match the stable commits below against your distro's mlx5_core backport). Interim control: remove /dev/infiniband/* from untrusted containers and do not grant DEVX to tenant workloads.","references":["https://git.kernel.org/stable/c/69dd3ad406c49aa69ce4852c15231ac56af8caf9","https://git.kernel.org/stable/c/ab3de780c176bb91995c6166a576b370d9726e17","https://nvd.nist.gov/vuln/detail/CVE-2022-50726"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-50756","cve":"CVE-2022-50756","aliases":[],"title":"Linux kernel (drivers/nvme/host): The PRP list mempool is sized in the wrong units, so a large I/O that needs two PRP","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/nvme/host)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"The PRP list mempool is sized in the wrong units, so a large I/O that needs two PRP lists writes past the end of the pool allocation. This is a genuine heap overflow on the local NVMe data path - kfence caught it in practice - so one tenant's I/O can corrupt kernel memory belonging to the rest of the node.","attack_vector":"Reachable by any unprivileged process on the node that can issue I/O to a local PCIe NVMe device (a container with a scratch volume is enough) - no /dev/nvme passthrough required. The window is narrow: it needs roughly a 4MB transfer split into 127 physical segments on a submission queue whose controller does not support SGLs. Most enterprise NVMe negotiates SGLs and is therefore not on this path, so check controller capability before rating this urgent on a given SKU.","remediation":"No fixed version is listed on this record - boot a kernel carrying the linked stable commits. Interim: confirm whether the node's NVMe controllers advertise SGL support (SGL-capable queues do not take this path), and cap the maximum I/O size handed to tenants on any controller that does not.","references":["https://git.kernel.org/stable/c/dfb6d54893d544151e7f480bc44cfe7823f5ad23","https://git.kernel.org/stable/c/9141144b37f30e3e7fa024bcfa0a13011e546ba9","https://nvd.nist.gov/vuln/detail/CVE-2022-50756"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2023-0179","cve":"CVE-2023-0179","aliases":[],"title":"Linux kernel (netfilter): Integer overflow in nft_payload_copy_vlan - stack leak plus local privilege escalation","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (netfilter)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Integer overflow in nft_payload_copy_vlan - stack leak plus local privilege escalation","attack_vector":"Any tenant process in a container with CAP_NET_ADMIN in a userns","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2023-0179"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2023-03-27"},{"id":"CVE-2023-0182","cve":"CVE-2023-0182","aliases":[],"title":"GPU Display Driver: Local privesc to host root (kernel buffer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc to host root (kernel buffer overflow)","attack_vector":"Any tenant with a container holding /dev/nvidia*","remediation":"Driver upgrade; drain + reboot node, evict tenant workloads","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0182","https://github.com/NVIDIA/product-security/tree/main/2023/5452"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-reboot"},"published":"2023-04-01"},{"id":"CVE-2023-0266","cve":"CVE-2023-0266","aliases":[],"title":"Linux kernel (ALSA): Use-after-free in snd_ctl_elem_read - local privilege escalation","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (ALSA)","year":"2023","cvss_score":7.8,"severity":"high","kev":true,"impact":"Use-after-free in snd_ctl_elem_read - local privilege escalation [KEV]","attack_vector":"Local user with access to /dev/snd","remediation":"Livepatchable; otherwise drain + reboot. Low exposure on headless GPU nodes - blacklist sound modules","references":["https://access.redhat.com/security/cve/CVE-2023-0266"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2023-01-30"},{"id":"CVE-2023-1017","cve":"CVE-2023-1017","aliases":[],"title":"TPM 2.0 reference implementation: Out-of-bounds write in `CryptParameterDecryption`","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"TPM 2.0 reference implementation","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Out-of-bounds write in `CryptParameterDecryption` — code execution inside the TPM or a bricked TPM. If the TPM is the root of trust for attestation, a compromised TPM invalidates every measured-boot claim the cloud makes to its tenants","attack_vector":"Local, low privilege","remediation":"TPM firmware update from the TPM/platform vendor; TPM firmware updates frequently clear the TPM, which invalidates sealed keys and any disk encryption bound to PCRs — this is why fleets skip it","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-1017"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-02-28"},{"id":"CVE-2023-20548","cve":"CVE-2023-20548","aliases":[],"title":"AMD Secure Processor - TOCTOU race: A time-of-check-to-time-of-use race in the ASP lets an attacker swap a value","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor - TOCTOU race","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"A time-of-check-to-time-of-use race in the ASP lets an attacker swap a value between validation and use, corrupting secure-processor memory. Winning the race costs integrity, confidentiality or availability of the ASP depending on what gets corrupted - and the ASP is the component vouching for the node's confidential-computing posture.","attack_vector":"Local. Requires the attacker to run concurrently with the ASP operation and to be able to modify the memory being checked, i.e. host-privileged code.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. Because this sits inside the SEV-SNP trust boundary, the update also moves the platform's reported TCB version: after patching you must refresh VCEK certificates from AMD's KDS and update whatever attestation policy your tenants (or your own confidential-VM control plane) pin against, or every guest launch will start failing validation.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-20548","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2026-02-11"},{"id":"CVE-2023-20555","cve":"CVE-2023-20555","aliases":[],"title":"AMD SMM - memory corruption (AMD-SB-4003): Memory corruption reachable in System Management Mode. Same class as the","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD SMM - memory corruption (AMD-SB-4003)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Memory corruption reachable in System Management Mode. Same class as the rest of the SMM cluster - a ring-0 attacker escalates into the one execution context that no hypervisor, kernel or EDR can observe, and the foothold survives OS reinstallation.","attack_vector":"Local, ring-0 privilege required.","remediation":"Fixed in AMD PI/AGESA firmware and delivered only as an OEM SBIOS package - AMD ships the PI drop to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before releasing BIOS. **Budget one to six months of OEM lag**, and note that several CVEs in this batch are marked 'no fix planned' on Naples (EPYC 7001) - for those the only remediation is retiring the hardware. Applying it means cordon, drain and a full power cycle per node; there is no driver reload, no live patch and no VBIOS step.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-20555","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"tags":["tenant-isolation"],"published":"2023-08-08"},{"id":"CVE-2023-20598","cve":"CVE-2023-20598","aliases":[],"title":"AMD Radeon Graphics driver - IOCTL granting arbitrary I/O port and physical memory access: Improper privilege","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD Radeon Graphics driver - IOCTL granting arbitrary I/O port and physical memory access","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Improper privilege management in the AMD graphics driver lets an authenticated attacker craft an IOCTL that gives I/O control over arbitrary hardware ports or physical memory. This is about as broad as a driver bug gets: arbitrary physical memory access from a user-issued IOCTL is a complete bypass of kernel memory protection, so any workload holding the GPU device node can read every other tenant's memory and take the host. AMD's PSP driver component (AMDPSP) was the affected surface.","attack_vector":"Local, authenticated, via a GPU driver IOCTL - i.e. reachable by anything that has the GPU device handed to it, which on a GPU cloud is every tenant container.","remediation":"Update the AMD graphics driver package and reload the driver or reboot the node. Driver-speed fix - no BIOS, no VBIOS - so there is no excuse for leaving it. Treat as top priority on any node where untrusted workloads get a GPU device node.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-20598","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2023-10-17"},{"id":"CVE-2023-21987","cve":"CVE-2023-21987","aliases":[],"title":"Oracle VirtualBox: Core component flaw allowing a low-privileged guest user to take over the host VirtualBox","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Oracle VirtualBox","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Core component flaw allowing a low-privileged guest user to take over the host VirtualBox installation","attack_vector":"Tenant VM guest","remediation":"VirtualBox update + VM restart. VirtualBox is not a production hypervisor - its presence on a neocloud node is itself the finding","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-21987"],"status":"curated","published":"2023-04-18"},{"id":"CVE-2023-25505","cve":"CVE-2023-25505","aliases":[],"title":"NVIDIA DGX BMC (IPMI handler): Buffer overflow in the IPMI handler of the NVIDIA DGX BMC","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVIDIA DGX BMC (IPMI handler)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Buffer overflow in the IPMI handler of the NVIDIA DGX BMC — code execution on the BMC of a GPU node","attack_vector":"Local / IPMI","remediation":"DGX BMC firmware update through NVIDIA's own bundle; DGX firmware bundles are monolithic, so this pulls in unrelated component updates and a longer maintenance window","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25505"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-120"],"fleet":{"pain_class":"firmware-flash"},"published":"2023-04-22"},{"id":"CVE-2023-25515","cve":"CVE-2023-25515","aliases":[],"title":"NVIDIA GPU Display Driver - kernel mode layer (Windows nvlddmkm.sys and Linux nvidia.ko): The driver parses unexpected","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - kernel mode layer (Windows nvlddmkm.sys and Linux nvidia.ko)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"The driver parses unexpected untrusted data, reaching code execution, privilege escalation and information disclosure from an unprivileged local account on either OS. Both the Windows and Linux datacenter drivers are affected, so a mixed fleet needs two separate rollouts.","attack_vector":"Local and unprivileged on either OS. On Linux it is reachable from any GPU container via /dev/nvidia*; on Windows from any session holding a GPU handle.","remediation":"Upgrade both the Linux and the Windows datacenter driver branches listed in bulletin 5468. Cost: Linux needs a drain and nvidia.ko reload per node; Windows needs a reboot per node. Two change windows unless your fleet is homogeneous.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25515","https://github.com/NVIDIA/product-security/tree/main/2023/5468"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-822"],"fleet":{"pain_class":"node-reboot"},"published":"2023-06-23"},{"id":"CVE-2023-25519","cve":"CVE-2023-25519","aliases":[],"title":"NVIDIA BlueField: Privilege escalation from incorrect user management on the DPU","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA BlueField","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Privilege escalation from incorrect user management on the DPU","attack_vector":"Local","remediation":"BlueField DOCA/BFB update; requires draining the node because the DPU carries the tenant's network path","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25519"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-286"],"published":"2023-09-12"},{"id":"CVE-2023-25527","cve":"CVE-2023-25527","aliases":[],"title":"DGX H100 BMC: Kernel memory corruption","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX H100 BMC","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Kernel memory corruption -> arbitrary code exec on BMC","attack_vector":"Authenticated local BMC access","remediation":"Flash BMC 23.08.18 out-of-band","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5473/5473.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-119"],"published":"2023-09-20"},{"id":"CVE-2023-2598","cve":"CVE-2023-2598","aliases":[],"title":"Linux kernel (io_uring): io_uring fixed-buffer registration gives out-of-bounds access to physical memory","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (io_uring)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"io_uring fixed-buffer registration gives out-of-bounds access to physical memory - full host compromise from an unprivileged process","attack_vector":"Any tenant process in a container with io_uring enabled","remediation":"Livepatchable; otherwise drain + reboot. Strategic answer for a neocloud is to disable io_uring in the default container seccomp profile (`kernel.io_uring_disabled=2` on 6.6+)","references":["https://access.redhat.com/security/cve/CVE-2023-2598"],"status":"curated","fleet":{"ubiquity":"Very common - io_uring is on by default in modern kernels and is heavily used by high-throughput data loaders on GPU nodes","remediation_pain":"`node-reboot` - kernel 6.4-rc1+; the practical stopgap is disabling io_uring via sysctl, which degrades storage throughput for training jobs","pain_class":"node-reboot","why_fleet_wide":"Out-of-bounds physical-memory access from `IORING_REGISTER_BUFFERS` gives a low-privilege local user full root; from inside a container with io_uring allowed, that is a host takeover on a shared GPU node"},"published":"2023-06-01"},{"id":"CVE-2023-31019","cve":"CVE-2023-31019","aliases":[],"title":"GPU Display Driver (Windows): Local privesc via named-pipe server access","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver (Windows)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc via named-pipe server access","attack_vector":"Local low-priv user","remediation":"Upgrade Oct-2023 driver branch","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5491/5491.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-284"],"published":"2023-11-02"},{"id":"CVE-2023-31248","cve":"CVE-2023-31248","aliases":[],"title":"Linux kernel (nf_tables): Use-after-free in nft_chain_lookup_byid() - local root (Pwn2Own Vancouver chain)","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (nf_tables)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Use-after-free in nft_chain_lookup_byid() - local root (Pwn2Own Vancouver chain)","attack_vector":"Any tenant process in a container with CAP_NET_ADMIN in a userns","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2023-31248"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2023-07-05"},{"id":"CVE-2023-31324","cve":"CVE-2023-31324","aliases":[],"title":"AMD Secure Processor - XGMI Trusted Agent (TOCTOU): A TOCTOU race in the ASP's XGMI Trusted Agent lets an attacker","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor - XGMI Trusted Agent (TOCTOU)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"A TOCTOU race in the ASP's XGMI Trusted Agent lets an attacker modify Infinity-Fabric/XGMI link commands after they are validated but before they execute. XGMI is the coherent interconnect binding MI-series GPUs together in a node, so tampering with its trusted-agent commands is tampering with the fabric that carries other tenants' model traffic between accelerators.","attack_vector":"Local, host-privileged, and specific to multi-GPU platforms that actually use XGMI - which is every MI200/MI250/MI300 node in an 8-GPU configuration.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. This one is squarely an Instinct-node issue, not a generic EPYC one: prioritise it on MI250/MI300 hosts where XGMI is carrying real inter-GPU traffic.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31324","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2026-02-11"},{"id":"CVE-2023-32233","cve":"CVE-2023-32233","aliases":[],"title":"Linux kernel (nf_tables): Use-after-free in nf_tables anonymous-set batch processing","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (nf_tables)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Use-after-free in nf_tables anonymous-set batch processing - unprivileged local user to root, public exploit","attack_vector":"Any tenant process in a container with CAP_NET_ADMIN in a userns","remediation":"Livepatchable; otherwise drain + reboot. Disable unprivileged user namespaces to blunt it","references":["https://access.redhat.com/security/cve/CVE-2023-32233"],"status":"curated","fleet":{"ubiquity":"Universal - `nf_tables` is enabled by default in most distributions and every kernel through 6.3.1 is affected","remediation_pain":"`node-reboot` - kernel upgrade to 6.3.2+; mitigation is blocking `CAP_NET_ADMIN`/user namespaces, which many tenant workloads legitimately need","pain_class":"node-reboot","why_fleet_wide":"Use-after-free in nf_tables batch processing gives arbitrary kernel read/write and root from an unprivileged local user - in a container with a user namespace this is a straight escape to the GPU host"},"published":"2023-05-08"},{"id":"CVE-2023-3269","cve":"CVE-2023-3269","aliases":[],"title":"Linux kernel (mm VMA): StackRot: privilege escalation via non-RCU-protected VMA traversal","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (mm VMA)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"StackRot: privilege escalation via non-RCU-protected VMA traversal; affects 6.1-6.4","attack_vector":"Local user / tenant process","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2023-3269"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2023-07-11"},{"id":"CVE-2023-3390","cve":"CVE-2023-3390","aliases":[],"title":"Linux kernel (nf_tables): UAF in nft_set_lookup_global after mixed named/anonymous set batches - local root","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (nf_tables)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"UAF in nft_set_lookup_global after mixed named/anonymous set batches - local root","attack_vector":"Any tenant process in a container with CAP_NET_ADMIN in a userns","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2023-3390"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2023-06-28"},{"id":"CVE-2023-34195","cve":"CVE-2023-34195","aliases":["INSYDE-SA-2023052"],"title":"Insyde InsydeH2O (SystemFirmwareManagementRuntimeDxe, GetImage method): The firmware reads a runtime UEFI variable","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (SystemFirmwareManagementRuntimeDxe, GetImage method)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"The firmware reads a runtime UEFI variable called GetImageProgress and then calls it as a function pointer. An attacker sets that variable from the OS to point at code they control and the firmware jumps to it during the DXE phase. This is a firmware-update-service driver, so the attacker ends up executing inside the machinery responsible for validating the next BIOS image - a direct route to a persistent, self-reinstalling firmware implant on a GPU node.","attack_vector":"Local admin/root on the host OS with UEFI variable write access, then a reboot or a firmware-management call that reaches GetImage.","remediation":"OEM BIOS update carrying the fixed Insyde kernel (5.0-5.5 affected). Firmware flash, one reboot per node. No configuration mitigates it. Interim hardening: restrict OS-side UEFI variable writes, and where the platform supports it verify that capsule updates require a signed payload so an implant cannot re-flash itself through the same service.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-34195","https://www.insyde.com/security-pledge/SA-2023052"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-09-18"},{"id":"CVE-2023-34319","cve":"CVE-2023-34319","aliases":["XSA-432"],"title":"Linux (Xen netback): Buffer overrun in netback due to an unusual packet - guest attacks dom0","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux (Xen netback)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Buffer overrun in netback due to an unusual packet - guest attacks dom0","attack_vector":"Tenant VM guest","remediation":"dom0 kernel patch; livepatchable, otherwise dom0 reboot (evacuates every guest on the host)","references":["https://xenbits.xen.org/xsa/advisory-432.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2023-09-22"},{"id":"CVE-2023-34322","cve":"CVE-2023-34322","aliases":["XSA-438"],"title":"Xen (64-bit PV): Top-level shadow reference dropped too early for 64-bit PV guests - privilege escalation to host","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen (64-bit PV)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Top-level shadow reference dropped too early for 64-bit PV guests - privilege escalation to host","attack_vector":"Tenant VM guest (PV)","remediation":"Hypervisor patch + reboot/evacuation","references":["https://xenbits.xen.org/xsa/advisory-438.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-01-05"},{"id":"CVE-2023-34325","cve":"CVE-2023-34325","aliases":["XSA-443"],"title":"Xen (libfsimage/pygrub): Multiple vulnerabilities in libfsimage disk handling","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen (libfsimage/pygrub)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Multiple vulnerabilities in libfsimage disk handling - a hostile guest disk image compromises the toolstack in dom0","attack_vector":"Tenant-supplied disk image","remediation":"Patch libfsimage and run pygrub de-privileged (see also XSA-508); toolstack restart, no full host reboot","references":["https://xenbits.xen.org/xsa/advisory-443.html"],"status":"curated","published":"2024-01-05"},{"id":"CVE-2023-34332","cve":"CVE-2023-34332","aliases":["AMI-SA-2023010"],"title":"AMI MegaRAC SPx 12 / SPx 13 (BMC): Untrusted pointer dereference in the BMC that a low-privileged actor can turn","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx 12 / SPx 13 (BMC)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Untrusted pointer dereference in the BMC that a low-privileged actor can turn into code execution or a controller crash. As a second-stage bug it is how an attacker who already got a foothold - a read-only monitoring account, a low-privilege Redfish user, or a partial exploit of one of the network bugs - upgrades to full BMC control and firmware persistence.","attack_vector":"AMI's CVSS vector scores this as local access with low privileges required, while AMI's own prose calls it reachable from the local network; treat it as reachable by anyone who already holds a low-privilege position on or adjacent to the BMC. In a fleet, the realistic precondition is a leaked low-tier BMC credential - which is common, because BMC passwords are frequently shared across a whole rack or SKU by the provisioning system.","remediation":"Firmware flash to SPx_12.7 / SPx_13.6, out-of-band per node, ODM-gated. Alongside the flash, the cheap wins are config-only: give every BMC a unique password (kill any shared default from the deployment template), delete unused BMC accounts, and drop any monitoring account down to the minimum Redfish role.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2023010.pdf","https://nvd.nist.gov/vuln/detail/CVE-2023-34332"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-01-09"},{"id":"CVE-2023-34333","cve":"CVE-2023-34333","aliases":[],"title":"AMI MegaRAC SPx (untrusted pointer dereference): Untrusted pointer dereference in the BMC allowing a local-network","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx (untrusted pointer dereference)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Untrusted pointer dereference in the BMC allowing a local-network attacker to compromise confidentiality, integrity and availability of the management processor.","attack_vector":"Local-network access to the BMC with low privilege.","remediation":"Obtain and flash updated BMC firmware from your board OEM.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2023010.pdf"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-34853","cve":"CVE-2023-34853","aliases":[],"title":"Supermicro X12DPG-QR BIOS 1.4b: Control-flow hijack inside platform firmware, driven by an NVRAM variable","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro X12DPG-QR BIOS 1.4b","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Control-flow hijack inside platform firmware, driven by an NVRAM variable that host-side privileged code can set. The attacker moves from OS root into firmware-privileged execution and can establish a UEFI-resident implant. On a X12DPG-QR - a dual-socket GPU platform board - the payoff is a persistent presence on an accelerator node that reimaging does not touch and that the operator has no host-side way to detect. A buffer overflow reachable by manipulating the SmcSecurityEraseSetupVar UEFI variable, i.e. an NVRAM variable that firmware trusts without validating its contents.","attack_vector":"Local, host-side privileged code that can write UEFI variables. On Linux that means root with access to efivarfs. This is the standard escalation available to anyone who has rented the metal or otherwise obtained host root.","remediation":"BIOS flash to a fixed image per Supermicro's August 2023 BIOS advisory, delivered as the X12DPG-QR BIOS package. Requires a host reboot. A partial hardening step that costs nothing: restrict or remove write access to efivarfs in tenant-facing bare-metal images, which raises the bar for any NVRAM-variable attack, not just this one. It does not substitute for the flash, because a tenant with root can usually reach the variable store another way.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-34853","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2023/34xxx/CVE-2023-34853.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-35001","cve":"CVE-2023-35001","aliases":[],"title":"Linux kernel (nf_tables): Stack out-of-bounds read/write in nft_byteorder_eval() - local root (Pwn2Own)","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (nf_tables)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Stack out-of-bounds read/write in nft_byteorder_eval() - local root (Pwn2Own)","attack_vector":"Any tenant process in a container with CAP_NET_ADMIN in a userns","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2023-35001"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2023-07-05"},{"id":"CVE-2023-4004","cve":"CVE-2023-4004","aliases":[],"title":"Linux kernel (netfilter pipapo): UAF from improper element removal in nft_pipapo_remove() - local root","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (netfilter pipapo)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"UAF from improper element removal in nft_pipapo_remove() - local root","attack_vector":"Any tenant process in a container with CAP_NET_ADMIN in a userns","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2023-4004"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2023-07-31"},{"id":"CVE-2023-4147","cve":"CVE-2023-4147","aliases":[],"title":"Linux kernel (nf_tables): UAF adding a rule with NFTA_RULE_CHAIN_ID - local root","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (nf_tables)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"UAF adding a rule with NFTA_RULE_CHAIN_ID - local root","attack_vector":"Any tenant process in a container with CAP_NET_ADMIN in a userns","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2023-4147"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2023-08-07"},{"id":"CVE-2023-43068","cve":"CVE-2023-43068","aliases":[],"title":"Dell SmartFabric Storage Software (restricted shell in SSH): OS command injection escaping the restricted shell of the","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell SmartFabric Storage Software (restricted shell in SSH)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"OS command injection escaping the restricted shell of the NVMe-oF fabric controller, from an authenticated remote user. Escaping the restricted shell gives full control of the appliance that enforces which host can attach to which NVMe namespace — the storage equivalent of owning the fabric's ACLs. Related CLI and path-traversal issues: CVE-2023-43069, CVE-2023-43070, CVE-2023-4401.","attack_vector":"Authenticated remote user with SSH access to SmartFabric Storage Software v1.4 or earlier.","remediation":"Upgrade SmartFabric Storage Software past v1.4 — appliance software upgrade plus restart. Restrict SSH to a management bastion, and audit the NVMe-oF zoning/namespace-masking configuration afterwards rather than assuming the patch is sufficient.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-43068","https://nvd.nist.gov/vuln/detail/CVE-2023-43069","https://nvd.nist.gov/vuln/detail/CVE-2023-4401"],"status":"curated","tags":["tenant-isolation"],"published":"2023-10-05"},{"id":"CVE-2023-4623","cve":"CVE-2023-4623","aliases":[],"title":"Linux kernel (net/sched hfsc): Use-after-free in sch_hfsc qdisc - local root","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/sched hfsc)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Use-after-free in sch_hfsc qdisc - local root","attack_vector":"Any tenant process in a container with CAP_NET_ADMIN in a userns","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2023-4623"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2023-09-06"},{"id":"CVE-2023-4692","cve":"CVE-2023-4692","aliases":[],"title":"GRUB2 (NTFS filesystem parser): Out-of-bounds write parsing a crafted NTFS volume","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (NTFS filesystem parser)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Out-of-bounds write parsing a crafted NTFS volume. Relevant to any node that dual-boots, mounts a Windows-formatted staging volume, or is handed a raw disk between tenants - the NTFS parser runs before anything verifies the disk's provenance.","attack_vector":"An attacker-supplied NTFS volume attached to the node, including via BMC virtual media.","remediation":"grub2 package update + reboot per node. Where you never need NTFS, building GRUB without the module is a permanent fix rather than a patch treadmill.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-4692","https://access.redhat.com/security/cve/CVE-2023-4692"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2023-10-25"},{"id":"CVE-2023-4911","cve":"CVE-2023-4911","aliases":[],"title":"glibc (ld.so): Looney Tunables: buffer overflow in the ld.so GLIBC_TUNABLES parser - local root on default installs","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"glibc (ld.so)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Looney Tunables: buffer overflow in the ld.so GLIBC_TUNABLES parser - local root on default installs; used by Kinsing cloud crypto-mining crews","attack_vector":"Local user, incl. any shell inside a container that shares the host glibc","remediation":"Package update; every long-running process must be restarted to pick up the new loader, which in practice means a rolling node restart or reboot for full coverage","references":["https://access.redhat.com/security/cve/CVE-2023-4911"],"status":"curated","fleet":{"ubiquity":"Universal - glibc is in essentially every base image and on every host; Qualys got root on default Fedora, Ubuntu 22.04/23.04 and Debian 12/13","remediation_pain":"`node-drain` for the host glibc (SUID binaries must be re-executed) plus a **rebuild of every container image** in the fleet - the pain is the image fan-out, not the host","pain_class":"node-drain","why_fleet_wide":"Buffer overflow in `GLIBC_TUNABLES` parsing gives local root from any SUID binary, so it converts every low-privilege foothold - in a container or on the host - into root, across every image and host simultaneously"},"published":"2023-10-03"},{"id":"CVE-2023-50274","cve":"CVE-2023-50274","aliases":[],"title":"HPE OneView (command injection with local privilege escalation): A low-privileged local user on the OneView appliance","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HPE OneView (command injection with local privilege escalation)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"A low-privileged local user on the OneView appliance injects commands and escalates - full control of the fleet management appliance.","attack_vector":"Local low-privilege access to the OneView appliance.","remediation":"Apply the OneView update per HPESBGN04586. Appliance update with restart.","references":["https://support.hpe.com/hpesc/public/docDisplay?docId=hpesbgn04586en_us&docLocale=en_US"],"status":"curated"},{"id":"CVE-2023-5058","cve":"CVE-2023-5058","aliases":["VU#811862"],"title":"Phoenix SecureCore Technology 4 (boot splash screen image parsing): The firmware parses a user-supplied boot logo image","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Phoenix SecureCore Technology 4 (boot splash screen image parsing)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"The firmware parses a user-supplied boot logo image without validating it, giving denial of service or arbitrary code execution in the DXE phase. This is the Phoenix instance of the wider image-parser problem in UEFI firmware: the logo is attacker-replaceable data sitting inside the firmware volume, parsed by privileged code long before Secure Boot has any say. For an operator the uncomfortable part is that a custom boot logo is a supported, documented OEM feature, so the write path exists by design.","attack_vector":"An attacker who can replace the boot logo image in the firmware volume - requiring firmware-write access from the OS (root plus a writable ESP or an unlocked SPI region), then a reboot.","remediation":"OEM BIOS update on the fixed SecureCore Technology 4 build. Firmware flash, reboot per node. Config-side hardening that helps immediately: ensure the SPI flash and the logo storage region are write-protected at the platform level, and do not deploy custom OEM boot logos on nodes where that means leaving the region writable. Track this alongside the LogoFAIL family from other IBVs - the same image-parser class was found across Insyde, AMI and Phoenix, so a fleet with mixed OEMs needs all three checked.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-5058","https://www.kb.cert.org/vuls/id/811862"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-12-07"},{"id":"CVE-2023-51042","cve":"CVE-2023-51042","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface: A use-after-free in the amdgpu GEM/VM/command-submission","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the amdgpu GEM/VM/command-submission ioctl surface. Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: In the Linux kernel before 6.4.12, amdgpu_cs_wait_all_fences in drivers/gpu/drm/amd/amdgpu/amdgpu_cs.c has a fence use-after-free.","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-51042","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-01-23"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-125","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-52474","cve":"CVE-2023-52474","aliases":[],"title":"Linux kernel (drivers/infiniband/hw/hfi1): User SDMA requests with multiple payload buffers are read past the declared","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/hw/hfi1)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"User SDMA requests with multiple payload buffers are read past the declared length of each buffer and the wrong pages are put on the wire, so data the sender never asked to transmit leaves the node. The same fix closes pin-cache races that produce duplicate page pinnings and a window where an entry is removed, the lock dropped, and new pages pinned - stale pinned-page state a tenant can steer.","attack_vector":"Local and unprivileged: a tenant with access to the hfi1 user device submits an SDMA request whose non-tail iovec does not end on a page boundary. No fabric peer or root is needed. Conditional on hfi1 (Omni-Path) hardware and the user SDMA path being exposed to tenants.","remediation":"No fixed version is recorded in this entry; boot a stable kernel carrying the multi-iovec SDMA and mmu_rb fixes (commits 9c4c6512d733 / a2bd706ab635). Interim: remove hfi1 user device nodes from tenant containers, or blacklist hfi1 on nodes where Omni-Path is not the production fabric.","references":["https://git.kernel.org/stable/c/9c4c6512d7330b743c4ffd18bd999a86ca26db0d","https://git.kernel.org/stable/c/a2bd706ab63509793b5cd5065e685b7ef5cba678","https://nvd.nist.gov/vuln/detail/CVE-2023-52474"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2023-52624","cve":"CVE-2023-52624","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Wake DMCUB before executing GPINT commands","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52624","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-03-26"},{"id":"CVE-2023-52678","cve":"CVE-2023-52678","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): Missing or insufficient validation of user-supplied","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdkfd (KFD compute driver, /dev/kfd). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdkfd: Confirm list is non-empty before utilizing list_first_entry in kfd_topology.c","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52678","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-05-17"},{"id":"CVE-2023-52691","cve":"CVE-2023-52691","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): A double free in the amdgpu power management","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"A double free in the amdgpu power management (SMU/powerplay). The same allocation is released twice, corrupting the slab allocator's freelist. This is a classic heap-corruption primitive: with slab grooming it becomes arbitrary kernel memory write and therefore host compromise from an unprivileged GPU workload. The cheap outcome is a node panic. Upstream fix: drm/amd/pm: fix a double-free in si_dpm_init","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52691","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-17"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-52745","cve":"CVE-2023-52745","aliases":[],"title":"Linux kernel (drivers/infiniband/ulp/ipoib): A PKEY child interface created over netlink comes up with multiple TX/RX","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/ulp/ipoib)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"A PKEY child interface created over netlink comes up with multiple TX/RX queues even when the parent device supports only one, and the first packet sent on it dereferences a NULL pointer and panics the node. PKEY partitions are exactly how an InfiniBand fabric separates tenants, so the isolation mechanism itself becomes the crash trigger.","attack_vector":"Two-stage: the partition child interface is created by the operator or orchestration over netlink (needs CAP_NET_ADMIN), after which any unprivileged workload sending traffic on that interface panics the node. Conditional on legacy IPoIB with PKEY child interfaces on a device that supports a single queue.","remediation":"Update to 5.4.232 / 5.10.168 / 5.15.94 / 6.1.12 or later. Interim: create PKEY child interfaces through the legacy sysfs path rather than netlink, or avoid PKEY child interfaces on single-queue IPoIB devices until patched.","references":["https://git.kernel.org/stable/c/4a779187db39b2f32d048a752573e56e4e77807f","https://git.kernel.org/stable/c/b1afb666c32931667c15ad1b58e7203f0119dcaf","https://nvd.nist.gov/vuln/detail/CVE-2023-52745"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2023-52812","cve":"CVE-2023-52812","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amd): An out-of-bounds access in the amdgpu kernel driver core - a length","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amd)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu kernel driver core - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd: check num of link levels when update pcie param","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52812","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-05-21"},{"id":"CVE-2023-52816","cve":"CVE-2023-52816","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): An out-of-bounds access in the amdkfd (KFD compute","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdkfd (KFD compute driver, /dev/kfd) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdkfd: Fix shift out-of-bounds issue","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52816","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-05-21"},{"id":"CVE-2023-52818","cve":"CVE-2023-52818","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd): An out-of-bounds access in the amdgpu power management","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu power management (SMU/powerplay) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd: Fix UBSAN array-index-out-of-bounds for SMU7","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52818","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-21"},{"id":"CVE-2023-52825","cve":"CVE-2023-52825","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A use-after-free in the amdkfd (KFD compute driver","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the amdkfd (KFD compute driver, /dev/kfd). Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amdkfd: Fix a race condition of vram buffer unref in svm code","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52825","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-05-21"},{"id":"CVE-2023-52913","cve":"CVE-2023-52913","aliases":[],"title":"Linux i915 GPU kernel driver: A use-after-free in the i915 GPU kernel driver. The general shape is that a GPU object is","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux i915 GPU kernel driver","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the i915 GPU kernel driver. The general shape is that a GPU object is freed on one path while another path still holds a reference to it, so a local user with GPU access can get the kernel to read or write freed memory. Exploitability varies by heap layout, but on a GPU node every such bug is reachable from inside a container that was granted /dev/dri - the same boundary that is supposed to separate tenants. Specific trigger: GEM context registration making a context visible to userspace before it is fully owned, so a second thread can free it mid-setup.","attack_vector":"Any local user or container with a DRM render node - i.e. any tenant that was scheduled a GPU. No privileged capability needed.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52913","https://git.kernel.org/stable/c/ae278887193110dfeb857ea63e243a3851fbb0bc"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-08-21"},{"id":"CVE-2023-52916","cve":"CVE-2023-52916","aliases":[],"title":"ASPEED video engine capture driver (drivers/media/platform/aspeed) - iKVM path: The video engine writes past","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ASPEED video engine capture driver (drivers/media/platform/aspeed) - iKVM path","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"The video engine writes past its capture buffer when the host is driving a 1600x900 mode, corrupting whatever BMC kernel memory sits after it. The reported repro is mundane: run iKVM with virtual media mounted while the BMC is under memory pressure. Since the resolution is chosen by the host, a tenant with control of the server's display output can steer the BMC into the corrupting mode on demand. Realistically this is a BMC crash - which on a GPU node means losing remote power control and console right when you need it - but out-of-bounds writes driven by attacker-chosen geometry are the raw material for something worse.","attack_vector":"The host side picks the video mode, so any tenant with root on the bare-metal node can select it. Triggering the corruption additionally needs the BMC's iKVM/video capture to be running, which is the normal state on fleets that leave remote console available.","remediation":"Kernel fix backported into stable; reaching your fleet means a BMC firmware flash per node, out-of-band, waiting on the ODM rebase. Config-only stopgap: disable the iKVM/video capture service on nodes that do not need graphical remote console. On a GPU fleet that is usually acceptable - operators run serial-over-LAN and Redfish, not KVM - and it removes both this bug and the video-engine DMA issue below without touching firmware.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52916","https://git.kernel.org/stable/c/4c823e4027dd1d6e88c31028dec13dd19bc7b02d","https://lists.debian.org/debian-lts-announce/2025/03/msg00001.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-09-06"},{"id":"CVE-2023-52921","cve":"CVE-2023-52921","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A use-after-free in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the amdgpu GEM/VM/command-submission ioctl surface. Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amdgpu: fix possible UAF in amdgpu_cs_pass1()","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52921","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-11-19"},{"id":"CVE-2023-52930","cve":"CVE-2023-52930","aliases":[],"title":"Linux i915 GPU kernel driver (GEM tiling): A double-free reachable by racing I915_GEM_SET_TILING from multiple threads.","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux i915 GPU kernel driver (GEM tiling)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"A double-free reachable by racing I915_GEM_SET_TILING from multiple threads. Double-free in the kernel slab allocator is the most directly weaponisable class here - it gives an attacker with GPU access a well-understood route to arbitrary kernel write and therefore to the host and every co-tenant on the node.","attack_vector":"Any local user or container with a DRM render node - i.e. any tenant that was scheduled a GPU. No privileged capability needed.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52930","https://git.kernel.org/stable/c/0769f997a7b6d5cb8336db0b4ec3d2d311b8097c"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-03-27"},{"id":"CVE-2023-53009","cve":"CVE-2023-53009","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A correctness defect in the amdkfd (KFD compute","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdkfd (KFD compute driver, /dev/kfd) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdkfd: Add sync after creating vram bo","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53009","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-03-27"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-190","CWE-835"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-53026","cve":"CVE-2023-53026","aliases":[],"title":"Linux kernel (drivers/infiniband/core): A 32-bit advance counter in the core RDMA block iterator wraps when a single","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/core)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"A 32-bit advance counter in the core RDMA block iterator wraps when a single scatter-gather entry needs more than 4GB of page-aligned coverage, and the iterator never terminates. Registering one very large memory region spins a CPU inside the kernel forever - a whole-node hang that every co-tenant on the box pays for, and it lives in core code shared by all RDMA drivers.","attack_vector":"Local: any tenant holding /dev/infiniband/uverbs* can register a large, misaligned memory region (the upstream trace is an EFA dmabuf registration with a 3GB SG entry and a 2GB page size) and hang the registering CPU. No fabric peer needed; the bug is in drivers/infiniband/core/verbs.c, so it applies across mlx5, efa, irdma and the rest.","remediation":"No fixed version is recorded in this entry; boot a stable kernel carrying the block-iterator overflow fix (commits 902063a9fea5 / d66c1d4178c2). Interim: cap tenant memory-registration size / locked-memory limits, and drop /dev/infiniband from containers that do not need native verbs.","references":["https://git.kernel.org/stable/c/902063a9fea5f8252df392ade746bc9cfd07a5ae","https://git.kernel.org/stable/c/d66c1d4178c219b6e7d7a6f714e3e3656faccc36","https://nvd.nist.gov/vuln/detail/CVE-2023-53026"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2023-53077","cve":"CVE-2023-53077","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: fix shift-out-of-bounds in CalculateVMAndRowBytes","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53077","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-05-02"},{"id":"CVE-2023-53087","cve":"CVE-2023-53087","aliases":[],"title":"Linux i915 GPU kernel driver (active barrier tracking): Non-idle barriers were misused as fence trackers, corrupting","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux i915 GPU kernel driver (active barrier tracking)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Non-idle barriers were misused as fence trackers, corrupting kernel lists when i915 perf is in use. Users hit this as an oops on list corruption - node down, jobs lost.","attack_vector":"Any local user or container with a DRM render node - i.e. any tenant that was scheduled a GPU. No privileged capability needed.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53087","https://git.kernel.org/stable/c/5c7591b8574c52c56b3994c2fbef1a3a311b5715"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-05-02"},{"id":"CVE-2023-53090","cve":"CVE-2023-53090","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): Missing or insufficient validation of user-supplied","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdkfd (KFD compute driver, /dev/kfd). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdkfd: Fix an illegal memory access","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53090","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-05-02"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787","CWE-682"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-53236","cve":"CVE-2023-53236","aliases":[],"title":"Linux kernel (drivers/iommu/iommufd): The pfn batch carries the wrong page-frame number forward when a mapping spans a","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/iommufd)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"The pfn batch carries the wrong page-frame number forward when a mapping spans a batch boundary, so page-pin accounting is applied to pages the tenant never mapped. Host page metadata gets corrupted - pages pinned and unpinned out from under whoever actually owns them.","attack_vector":"A tenant or VMM holding /dev/iommu doing an ordinary IOMMU_IOAS_MAP over a region large enough to cross a pfn batch boundary. No race required, no host root, no special hardware beyond iommufd being the passthrough path.","remediation":"The record lists no fixed release; boot a kernel carrying the stable fix commits below. Interim control: keep /dev/iommu out of tenant containers.","references":["https://git.kernel.org/stable/c/6ed5784526ddc0fb58b1798af36ec0c3139a8dca","https://git.kernel.org/stable/c/13a0d1ae7ee6b438f5537711a8c60cba00554943","https://nvd.nist.gov/vuln/detail/CVE-2023-53236"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-129","CWE-787"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-53340","cve":"CVE-2023-53340","aliases":[],"title":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core): A userspace DEVX consumer can issue a firmware command opcode","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"A userspace DEVX consumer can issue a firmware command opcode the driver does not track, and when that command fails the driver indexes its per-opcode failure-statistics array with the untracked opcode. The result is an out-of-bounds array access in shared driver state driven by a value the caller chose - memory corruption reachable from a tenant's RDMA device node, not just a crash.","attack_vector":"A tenant container holding /dev/infiniband/uverbs* with DEVX enabled (rdma-core mlx5dv DEVX general command) can send arbitrary firmware command opcodes and make them fail on purpose. No host root needed; the only precondition is that the tenant is allowed to open the mlx5 RDMA device and use DEVX, which is exactly what a bare RDMA/GPUDirect tenant is given.","remediation":"Update to a patched kernel on your stream. Interim: drop /dev/infiniband/* from tenant containers that do not need verbs, or block DEVX for tenants (do not grant the RDMA device to untrusted workloads) until the node is rebooted onto a fixed kernel.","references":["https://git.kernel.org/stable/c/411e4d6caa7f7169192b8dacc8421ac4fd64a354","https://git.kernel.org/stable/c/d8b6f175235d7327b4e1b13216859e89496dfbd5","https://nvd.nist.gov/vuln/detail/CVE-2023-53340"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2023-53378","cve":"CVE-2023-53378","aliases":[],"title":"Linux i915 GPU kernel driver (display page table objects): The buffer object backing a display page table","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux i915 GPU kernel driver (display page table objects)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"The buffer object backing a display page table was not treated as a framebuffer, letting it be moved or reused while the display engine still pointed at it. Practical outcome on a headless GPU node is a driver crash rather than a tenant boundary break, but on nodes that do run display output it can surface other memory on screen.","attack_vector":"Any local user or container with a DRM render node - i.e. any tenant that was scheduled a GPU. No privileged capability needed.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53378","https://git.kernel.org/stable/c/3413881e1ecc3cba722a2e87ec099692eed5be28"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-09-18"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-362","CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-53394","cve":"CVE-2023-53394","aliases":[],"title":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core): When a regular receive queue is reactivated after an AF_XDP","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"When a regular receive queue is reactivated after an AF_XDP socket closes, it consumes completion entries left over from the previous incarnation of the queue. Those stale descriptors point at buffers that no longer belong to the queue, which corrupts the RQ: the interface silently stops receiving for every workload on the node and then crashes on the next queue close.","attack_vector":"A local principal that can open and close AF_XDP zero-copy sockets on an mlx5 interface while traffic is flowing - a tenant container with a passed-through VF netdev and the capability to run XDP, or any host dataplane agent - can drive this by repeatedly starting and killing an XSK application. Conditional on AF_XDP zero-copy being permitted on mlx5 interfaces.","remediation":"Update to a patched kernel on your stream. Interim: deny AF_XDP/XDP program attachment on mlx5 interfaces for tenant workloads (drop CAP_NET_ADMIN and CAP_BPF from those containers) until the node is rebooted onto a fixed kernel.","references":["https://git.kernel.org/stable/c/02a84eb2af6bea7871cd34264fb27f141f005fd9","https://git.kernel.org/stable/c/39646d9bcd1a65d2396328026626859a1dab59d7","https://nvd.nist.gov/vuln/detail/CVE-2023-53394"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2023-53471","cve":"CVE-2023-53471","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu/gfx): A race condition or locking defect in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu/gfx)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"A race condition or locking defect in the amdgpu power management (SMU/powerplay). Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdgpu/gfx: disable gfx9 cp_ecc_error_irq only when enabling legacy gfx ras","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53471","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-10-01"},{"id":"CVE-2023-53545","cve":"CVE-2023-53545","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A NULL pointer dereference in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"A NULL pointer dereference in the amdgpu GEM/VM/command-submission ioctl surface. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: unmap and remove csa_va properly","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53545","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-10-04"},{"id":"CVE-2023-53552","cve":"CVE-2023-53552","aliases":[],"title":"Linux i915 GPU kernel driver: A use-after-free in the i915 GPU kernel driver. The general shape is that a GPU object is","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux i915 GPU kernel driver","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the i915 GPU kernel driver. The general shape is that a GPU object is freed on one path while another path still holds a reference to it, so a local user with GPU access can get the kernel to read or write freed memory. Exploitability varies by heap layout, but on a GPU node every such bug is reachable from inside a container that was granted /dev/dri - the same boundary that is supposed to separate tenants. Specific trigger: requests belonging to GuC virtual engines outliving the engine, so userspace-held request references dangle.","attack_vector":"Any local user or container with a DRM render node - i.e. any tenant that was scheduled a GPU. No privileged capability needed.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53552","https://git.kernel.org/stable/c/5eefc5307c983b59344a4cb89009819f580c84fa"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-10-04"},{"id":"CVE-2023-53707","cve":"CVE-2023-53707","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): An out-of-bounds access in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu GEM/VM/command-submission ioctl surface - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: Fix integer overflow in amdgpu_cs_pass1","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53707","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-10-22"},{"id":"CVE-2023-53753","cve":"CVE-2023-53753","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: fix mapping to non-allocated address","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53753","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-12-08"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-53781","cve":"CVE-2023-53781","aliases":[],"title":"Linux kernel (net/smc): Closing an SMC socket can leave the internal TCP kernel socket with its timers still armed and","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/smc)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Closing an SMC socket can leave the internal TCP kernel socket with its timers still armed and its netns reference already dropped, so a TCP retransmit timer later fires against freed memory. KASAN confirms a slab-use-after-free read in tcp_write_timer_handler - reached from timer softirq, so it destabilises the whole node rather than the calling tenant.","attack_vector":"Local and unprivileged: create and close AF_SMC sockets in a loop (the syzbot reproducer does exactly this). socket(AF_SMC, ...) needs no capability and autoloads the smc module through the net-pf-43 alias, so any tenant container reaches it - no RDMA device or /dev/infiniband access required, because the leaked object is the clcsock TCP socket.","remediation":"Boot a kernel carrying the fix commits (takes a netns reference for SMC's internal kernel sockets). Interim: blacklist smc or deny socket family 43 in tenant seccomp profiles.","references":["https://git.kernel.org/stable/c/1cc41c8acfc1ee30b4868559058db97fa44b0137","https://git.kernel.org/stable/c/9744d2bf19762703704ecba885b7ac282c02eacf","https://nvd.nist.gov/vuln/detail/CVE-2023-53781"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-362","CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-53795","cve":"CVE-2023-53795","aliases":[],"title":"Linux kernel (drivers/iommu/iommufd): The destroy ioctl takes a temporary reference on an iommufd object without the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/iommufd)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"The destroy ioctl takes a temporary reference on an iommufd object without the lock that every other temporary reference is required to hold. Two racing destroys can therefore drop the last reference on an object still in use, breaking the lifetime rule for the fd that owns a tenant's entire IOMMU address space and page tables.","attack_vector":"A holder of /dev/iommu racing two IOMMUFD_DESTROY ioctls against each other, or a destroy against close(). syzkaller-reachable from plain userspace ioctls; no host root, no hardware precondition beyond iommufd being the passthrough path.","remediation":"The record lists no fixed release; boot a kernel carrying the stable fix commits below. Interim control: keep /dev/iommu out of tenant containers.","references":["https://git.kernel.org/stable/c/495b327435b0298e9b3b434f5834d459a93673ce","https://git.kernel.org/stable/c/99f98a7c0d6985d5507c8130a981972e4b7b3bdc","https://nvd.nist.gov/vuln/detail/CVE-2023-53795"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2023-53806","cve":"CVE-2023-53806","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: populate subvp cmd info only for the top pipe","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53806","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-12-09"},{"id":"CVE-2023-53816","cve":"CVE-2023-53816","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A use-after-free in the amdkfd (KFD compute driver","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the amdkfd (KFD compute driver, /dev/kfd). Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amdkfd: fix potential kgd_mem UAFs","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53816","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-12-09"},{"id":"CVE-2023-53819","cve":"CVE-2023-53819","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (amdgpu): An out-of-bounds access in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (amdgpu)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu GEM/VM/command-submission ioctl surface - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: amdgpu: validate offset_in_bo of drm_amdgpu_gem_va","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53819","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-12-09"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-54048","cve":"CVE-2023-54048","aliases":[],"title":"Linux kernel (drivers/infiniband/hw/bnxt_re): The driver keeps scheduling completion handlers for a queue pair after","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/hw/bnxt_re)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"The driver keeps scheduling completion handlers for a queue pair after that QP is destroyed, so a poll can run against a completion queue that has already been freed. Upstream captured the resulting panic in bnxt_re_poll_cq. A tenant driving normal destroy-QP / destroy-CQ sequences turns this into a use-after-free on a shared node.","attack_vector":"Local: a tenant holding /dev/infiniband/uverbs* on a Broadcom bnxt_re adapter destroys a QP while its completion queue still has work scheduled, then frees the CQ - the race is reachable entirely through ordinary verbs lifecycle calls. Conditional on bnxt_re hardware.","remediation":"No fixed version is recorded in this entry; boot a stable kernel carrying the completion-flush-before-destroy fix (commits b79a0e71d6e8 / b8500538b8f5). Interim: remove /dev/infiniband device nodes from untrusted tenant containers on bnxt_re nodes.","references":["https://git.kernel.org/stable/c/b79a0e71d6e8692e0b6da05f8aaa7d69191cf7e7","https://git.kernel.org/stable/c/b8500538b8f5b2cd86b02754c8de83eaa7a2d6ba","https://nvd.nist.gov/vuln/detail/CVE-2023-54048"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2023-54168","cve":"CVE-2023-54168","aliases":["RDMA/mlx4 prevent shift wrapping in set_user_sq_size()"],"title":"Linux kernel mlx4_ib (legacy ConnectX-3 RDMA): Same class of bug on the older mlx4 stack: the user-supplied","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel mlx4_ib (legacy ConnectX-3 RDMA)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Same class of bug on the older mlx4 stack: the user-supplied log_sq_bb_count shift wraps in set_user_sq_size(). Any tenant with RDMA access on a ConnectX-3-era host gets a user-controlled shift into kernel sizing logic. Matters for operators still running legacy mlx4 nodes alongside a modern fleet.","attack_vector":"Local, low-privileged process creating an RDMA queue pair on an mlx4 device.","remediation":"Upgrade the host kernel to 6.4 or a stable backport (4.19.283, 5.4.243, 5.10.180, 5.15.111, 6.1.28, 6.2.15, 6.3.2). Host reboot. Strategically, this is a prompt to retire remaining ConnectX-3/mlx4 hardware rather than keep patching it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-54168","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2023/CVE-2023-54168.json"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-12-30"},{"id":"CVE-2023-54202","cve":"CVE-2023-54202","aliases":[],"title":"Linux i915 GPU kernel driver: A use-after-free in the i915 GPU kernel driver. The general shape is that a GPU object is","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux i915 GPU kernel driver","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the i915 GPU kernel driver. The general shape is that a GPU object is freed on one path while another path still holds a reference to it, so a local user with GPU access can get the kernel to read or write freed memory. Exploitability varies by heap layout, but on a GPU node every such bug is reachable from inside a container that was granted /dev/dri - the same boundary that is supposed to separate tenants. Specific trigger: a race in the perf config ioctl where a guessable object id lets two threads free the same OA config.","attack_vector":"Any local user or container with a DRM render node - i.e. any tenant that was scheduled a GPU. No privileged capability needed.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-54202","https://git.kernel.org/stable/c/240b1502708858b5e3f10b6dc5ca3f148a322fef"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-12-30"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-54216","cve":"CVE-2023-54216","aliases":[],"title":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core): Adding a TC flower rule while the device is in NIC mode makes","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Adding a TC flower rule while the device is in NIC mode makes the driver draw from an eswitch object-mapping pool that was never initialised, and the freed/uninitialised object is then fed to mlx5_add_flow_rules. KASAN reports a slab use-after-free inside the flow-steering rule-insertion path - the shared table that decides where every tenant's packets go - with a crash of the node as the visible outcome.","attack_vector":"Triggered by a `tc filter add` on an mlx5 netdev when the eswitch is not enabled. Any principal that can program TC offload on an mlx5 interface reaches it: the host network agent, or a tenant container/VM that was given a VF netdev plus CAP_NET_ADMIN in its own netns. It does not require a tenant RDMA or VFIO device node.","remediation":"Update to a patched kernel on your stream. Interim: do not grant CAP_NET_ADMIN over an mlx5 netdev to tenant workloads, and disable hw-tc-offload on mlx5 interfaces running in NIC (non-switchdev) mode.","references":["https://git.kernel.org/stable/c/4150441c010dec36abc389828e2e4758bd8ad4b3","https://git.kernel.org/stable/c/dfa1e46d6093831b9d49f0f350227a1d13644a2f","https://nvd.nist.gov/vuln/detail/CVE-2023-54216"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2023-6817","cve":"CVE-2023-6817","aliases":[],"title":"Linux kernel (netfilter pipapo): Inactive elements mishandled in nft_pipapo_walk - use-after-free, local root","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (netfilter pipapo)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Inactive elements mishandled in nft_pipapo_walk - use-after-free, local root","attack_vector":"Any tenant process in a container with CAP_NET_ADMIN in a userns","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2023-6817"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2023-12-18"},{"id":"CVE-2024-0071","cve":"CVE-2024-0071","aliases":[],"title":"GPU Display Driver: Local privesc (kernel buffer over-read)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc (kernel buffer over-read)","attack_vector":"Any tenant with a container; vGPU guest","remediation":"Driver upgrade (551.61 / 550.54.14 branch); drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0071","https://github.com/NVIDIA/product-security/tree/main/2024/5520"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-125"],"fleet":{"pain_class":"node-reboot"},"published":"2024-03-27"},{"id":"CVE-2024-0073","cve":"CVE-2024-0073","aliases":[],"title":"GPU Display Driver: Local privesc (improper privilege management)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc (improper privilege management)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0073","https://github.com/NVIDIA/product-security/tree/main/2024/5520"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-250"],"fleet":{"pain_class":"node-reboot"},"published":"2024-03-27"},{"id":"CVE-2024-0077","cve":"CVE-2024-0077","aliases":[],"title":"vGPU Manager: Guest-to-host privesc (improper privilege management)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU Manager","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Guest-to-host privesc (improper privilege management)","attack_vector":"Tenant VM guest","remediation":"Upgrade vGPU Manager on hypervisor; evacuate guest VMs, reboot host","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0077","https://github.com/NVIDIA/product-security/tree/main/2024/5520"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-285"],"fleet":{"pain_class":"node-reboot"},"published":"2024-03-27"},{"id":"CVE-2024-0084","cve":"CVE-2024-0084","aliases":[],"title":"vGPU Manager: Guest-to-host privesc","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU Manager","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Guest-to-host privesc","attack_vector":"Tenant VM guest","remediation":"vGPU Manager upgrade; evacuate VMs, reboot host","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0084","https://github.com/NVIDIA/product-security/tree/main/2024/5551"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-250"],"fleet":{"pain_class":"node-reboot"},"published":"2024-06-13"},{"id":"CVE-2024-0089","cve":"CVE-2024-0089","aliases":[],"title":"GPU Display Driver: Local privesc (improper resource validation / uninit memory)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc (improper resource validation / uninit memory)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0089","https://github.com/NVIDIA/product-security/tree/main/2024/5551"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-665"],"fleet":{"pain_class":"node-reboot"},"published":"2024-06-13"},{"id":"CVE-2024-0090","cve":"CVE-2024-0090","aliases":[],"title":"GPU Display Driver: Local privesc to host root (OOB write)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc to host root (OOB write)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0090","https://github.com/NVIDIA/product-security/tree/main/2024/5551"],"status":"curated","fleet":{"ubiquity":"Universal - Windows and Linux driver, all datacenter SKUs","remediation_pain":"`node-reboot` (driver replacement)","pain_class":"node-reboot","why_fleet_wide":"Out-of-bounds write reachable from an unprivileged local user through the driver API: any tenant process on a GPU node can attempt host code execution"},"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787"],"published":"2024-06-13"},{"id":"CVE-2024-0091","cve":"CVE-2024-0091","aliases":[],"title":"GPU Display Driver: Local privesc to host root (use-after-free)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc to host root (use-after-free)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0091","https://github.com/NVIDIA/product-security/tree/main/2024/5551"],"status":"curated","fleet":{"ubiquity":"Universal - same driver","remediation_pain":"`node-reboot`","pain_class":"node-reboot","why_fleet_wide":"Untrusted-pointer dereference via a driver API call from a low-privileged tenant process, giving DoS / info disclosure / tampering on a shared host"},"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-822"],"published":"2024-06-13"},{"id":"CVE-2024-0099","cve":"CVE-2024-0099","aliases":[],"title":"vGPU Manager: Guest-to-host escape (buffer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU Manager","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Guest-to-host escape (buffer overflow)","attack_vector":"Tenant VM guest","remediation":"vGPU Manager upgrade; evacuate all guest VMs, reboot host","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0099","https://github.com/NVIDIA/product-security/tree/main/2024/5551"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-120"],"fleet":{"pain_class":"node-reboot"},"published":"2024-06-13"},{"id":"CVE-2024-0107","cve":"CVE-2024-0107","aliases":[],"title":"GPU Display Driver: Local privesc (buffer over-read in driver)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc (buffer over-read in driver)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0107","https://github.com/NVIDIA/product-security/tree/main/2024/5557"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-125"],"fleet":{"pain_class":"node-reboot"},"published":"2024-08-08"},{"id":"CVE-2024-0117","cve":"CVE-2024-0117","aliases":[],"title":"NVIDIA GPU Display Driver - Windows user mode layer: An out-of-bounds read in the Windows user mode driver layer chains","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows user mode layer","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds read in the Windows user mode driver layer chains to code execution and privilege escalation. Note the CVSS assumes user interaction, which on a render/VDI host is trivially satisfied. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5586. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0117","https://github.com/NVIDIA/product-security/tree/main/2024/5586"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-125"],"fleet":{"pain_class":"node-drain"},"published":"2024-10-26"},{"id":"CVE-2024-0118","cve":"CVE-2024-0118","aliases":[],"title":"NVIDIA GPU Display Driver - Windows user mode layer: A second out-of-bounds read in the Windows user mode driver layer","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows user mode layer","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A second out-of-bounds read in the Windows user mode driver layer with the same code-execution outcome. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5586. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0118","https://github.com/NVIDIA/product-security/tree/main/2024/5586"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-125"],"fleet":{"pain_class":"node-drain"},"published":"2024-10-26"},{"id":"CVE-2024-0119","cve":"CVE-2024-0119","aliases":[],"title":"NVIDIA GPU Display Driver - Windows user mode layer: A third out-of-bounds read in the Windows user mode driver layer","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows user mode layer","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A third out-of-bounds read in the Windows user mode driver layer reaching code execution and privilege escalation. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5586. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0119","https://github.com/NVIDIA/product-security/tree/main/2024/5586"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-125"],"fleet":{"pain_class":"node-drain"},"published":"2024-10-26"},{"id":"CVE-2024-0120","cve":"CVE-2024-0120","aliases":[],"title":"NVIDIA GPU Display Driver - Windows user mode layer: A fourth out-of-bounds read in the Windows user mode driver layer","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows user mode layer","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A fourth out-of-bounds read in the Windows user mode driver layer with the same impact. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5586. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0120","https://github.com/NVIDIA/product-security/tree/main/2024/5586"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-125"],"fleet":{"pain_class":"node-drain"},"published":"2024-10-26"},{"id":"CVE-2024-0121","cve":"CVE-2024-0121","aliases":[],"title":"NVIDIA GPU Display Driver - Windows user mode layer: A fifth out-of-bounds read in the Windows user mode driver layer","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows user mode layer","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A fifth out-of-bounds read in the Windows user mode driver layer, fixed in the same bulletin as the other four. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5586. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0121","https://github.com/NVIDIA/product-security/tree/main/2024/5586"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-125"],"fleet":{"pain_class":"node-drain"},"published":"2024-10-26"},{"id":"CVE-2024-0127","cve":"CVE-2024-0127","aliases":[],"title":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko): A tenant who has compromised their own","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A tenant who has compromised their own guest kernel feeds bad input to the host GPU kernel driver and reaches code execution and privilege escalation on the host. The attacker sits inside a tenant guest VM and the blast radius is the hypervisor host, which is running every other tenant's vGPU on the same physical GPU. This is precisely the boundary a vGPU-based multi-tenant offering is sold on. Guest root is the assumed starting point, which is exactly what every vGPU tenant has. This is the cleanest guest-to-host escape candidate in the 2024 set.","attack_vector":"A tenant inside their own guest VM, driving the paravirtualised vGPU control interface. Several of these need only an unprivileged process in the guest; the rest need guest root, which a tenant already has on a VM they rented. No host credentials are involved at any point.","remediation":"Patch the vGPU Manager on the hypervisor host per bulletin 5586. Cost: the highest of any class here. The host driver cannot be reloaded while vGPUs are attached, so every tenant VM on that hypervisor must be live-migrated or powered off - a full host drain. NVIDIA also enforces a supported host/guest driver skew, so budget a matching guest-driver campaign in the same window or tenants lose their vGPU on next boot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0127","https://github.com/NVIDIA/product-security/tree/main/2024/5586"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-20"],"fleet":{"pain_class":"node-drain"},"tags":["tenant-isolation"],"published":"2024-10-26"},{"id":"CVE-2024-0146","cve":"CVE-2024-0146","aliases":[],"title":"vGPU Manager: Guest-to-host escape via GPU firmware buffer overflow","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU Manager","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Guest-to-host escape via GPU firmware buffer overflow","attack_vector":"Tenant VM guest","remediation":"Upgrade vGPU Manager + GPU firmware; evacuate all guest VMs, reboot host","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0146","https://github.com/NVIDIA/product-security/tree/main/2025/5614"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-120"],"fleet":{"pain_class":"node-reboot"},"published":"2025-01-28"},{"id":"CVE-2024-0582","cve":"CVE-2024-0582","aliases":[],"title":"Linux kernel (io_uring): Page use-after-free via io_uring buffer-ring mmap - unprivileged local user to root","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (io_uring)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Page use-after-free via io_uring buffer-ring mmap - unprivileged local user to root","attack_vector":"Any tenant process in a container with io_uring enabled","remediation":"Livepatchable; otherwise drain + reboot. Or disable io_uring for tenant containers","references":["https://access.redhat.com/security/cve/CVE-2024-0582"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-01-16"},{"id":"CVE-2024-1086","cve":"CVE-2024-1086","aliases":[],"title":"Linux kernel (nf_tables): Use-after-free in nft_verdict_init() - double-free to local root","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (nf_tables)","year":"2024","cvss_score":7.8,"severity":"high","kev":true,"impact":"Use-after-free in nft_verdict_init() - double-free to local root; weaponised public exploit with a very high success rate [KEV]","attack_vector":"Any tenant process in a container with CAP_NET_ADMIN in a userns","remediation":"Livepatchable on supported kernels (Canonical/TuxCare/Ksplice all shipped it); otherwise drain + reboot. The single highest-priority container-escape CVE of the 2024 set - treat unpatched nodes as compromised-by-default in a shared-tenant fleet","references":["https://access.redhat.com/security/cve/CVE-2024-1086"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-01-31"},{"id":"CVE-2024-26586","cve":"CVE-2024-26586","aliases":["mlxsw spectrum_acl_tcam fix stack corruption"],"title":"Linux kernel mlxsw (Spectrum switch ASIC ACL TCAM): On Spectrum-2 and newer, firmware reports more than 16 ACLs per","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel mlxsw (Spectrum switch ASIC ACL TCAM)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"On Spectrum-2 and newer, firmware reports more than 16 ACLs per group but the driver's register layout was never widened, so putting more than 16 ACLs in a group corrupts the kernel stack and panics the switch. Triggered by adding tc filters with decreasing priority in alternating order - a shape that ordinary policy automation produces. Relevant to anyone running Linux-based switch control on Spectrum silicon.","attack_vector":"Local on the switch with network-configuration privilege - an operator or automation system installing tc filters. Not remotely reachable.","remediation":"Upgrade the switch's kernel to 6.8 or a stable backport (5.10.209, 5.15.148, 6.1.79, 6.6.14, 6.7.2). On a Cumulus/NVOS switch that means an OS image upgrade and a switch reload - a fabric rolling window. Interim: constrain your ACL automation so a single group never exceeds 16 ACLs.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26586","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2024/CVE-2024-26586.json"],"status":"curated","published":"2024-02-22"},{"id":"CVE-2024-26592","cve":"CVE-2024-26592","aliases":[],"title":"Linux kernel (ksmbd): Use-after-free in ksmbd_tcp_new_connection() - in-kernel SMB server","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (ksmbd)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Use-after-free in ksmbd_tcp_new_connection() - in-kernel SMB server","attack_vector":"Unauthenticated network (if ksmbd is exposed)","remediation":"Livepatchable; otherwise drain + reboot. Correct answer for a neocloud is to not ship ksmbd at all - blacklist the module fleet-wide","references":["https://access.redhat.com/security/cve/CVE-2024-26592"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-02-22"},{"id":"CVE-2024-26656","cve":"CVE-2024-26656","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A use-after-free in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the amdgpu GEM/VM/command-submission ioctl surface. Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amdgpu: fix use-after-free bug","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26656","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-04-02"},{"id":"CVE-2024-26699","cve":"CVE-2024-26699","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Fix array-index-out-of-bounds in dcn35_clkmgr","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26699","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-04-03"},{"id":"CVE-2024-26728","cve":"CVE-2024-26728","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: fix null-pointer dereference on edid reading","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26728","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-04-03"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-193","CWE-787"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-26766","cve":"CVE-2024-26766","aliases":[],"title":"Linux kernel (drivers/infiniband/hw/hfi1): An off-by-one in the SDMA descriptor accounting lets the descriptor array in","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/hw/hfi1)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"An off-by-one in the SDMA descriptor accounting lets the descriptor array in a transmit request overflow, overwriting the neighbouring fields of the request structure - upstream shows a corrupted header pointer and a general protection fault. A tenant sending ordinary traffic thus corrupts kernel memory and can take the node down or steer a kernel pointer.","attack_vector":"Local and unprivileged: reproducible from the `sendmsg` syscall over an IPoIB interface backed by hfi1, per the upstream report. Any tenant that can send on an Omni-Path / hfi1 IPoIB interface reaches it; no fabric peer or privileged device node is required. Conditional on hfi1 hardware being present and in use for the data path.","remediation":"Update to 4.19.308 / 5.4.270 / 5.10.211 / 5.15.150 / 6.1.80 / 6.3 or later. Interim: if hfi1 is not the production fabric on a node, blacklist the hfi1 module; otherwise drain untrusted tenants from Omni-Path nodes until patched.","references":["https://git.kernel.org/stable/c/115b7f3bc1dce590a6851a2dcf23dc1100c49790","https://git.kernel.org/stable/c/5833024a9856f454a964a198c63a57e59e07baf5","https://nvd.nist.gov/vuln/detail/CVE-2024-26766"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-26786","cve":"CVE-2024-26786","aliases":[],"title":"Linux kernel (drivers/iommu/iommufd): On a partially-failed access attach, iommufd overwrites the xarray id that tracks","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/iommufd)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"On a partially-failed access attach, iommufd overwrites the xarray id that tracks a live access object, so a later destroy erases and frees the wrong entry. Object bookkeeping and object lifetime diverge, giving a use-after-free against the structure that owns a tenant's IOVA mappings.","attack_vector":"A holder of /dev/iommu racing an IOAS change against destroy/close on the same access object - reachable with ordinary ioctls plus close(), which is how syzkaller found it. Needs the alignment-recalculation step to fail after the id allocation succeeds, which the caller can arrange by choosing the IOAS. No host root, no hardware precondition beyond iommufd being in use.","remediation":"The record lists no fixed release; boot a kernel carrying the stable fix commits below. Interim control: keep /dev/iommu out of tenant containers.","references":["https://git.kernel.org/stable/c/f1fb745ee0a6fe43f1d84ec369c7e6af2310fda9","https://git.kernel.org/stable/c/9526a46cc0c378d381560279bea9aa34c84298a0","https://nvd.nist.gov/vuln/detail/CVE-2024-26786"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-26797","cve":"CVE-2024-26797","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Prevent potential buffer overflow in map_hw_resources","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26797","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-04-04"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-362","CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-26810","cve":"CVE-2024-26810","aliases":[],"title":"Linux kernel (drivers/vfio/pci): A tenant races a DisINTx write to emulated config space against a SET_IRQS ioctl, so","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/vfio/pci)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A tenant races a DisINTx write to emulated config space against a SET_IRQS ioctl, so the interrupt configuration changes underneath the code that fires the eventfd. The interrupt path then uses a stale or NULL trigger, corrupting host kernel state or crashing the node from inside a tenant's own passthrough device.","attack_vector":"A tenant holding the vfio-pci device fd. Both halves of the race are plain tenant-issued operations: a config-space write clearing DisINTx in the PCI command register, and VFIO_DEVICE_SET_IRQS changing the INTx configuration. Conditional on a passthrough device using legacy INTx. No host root required.","remediation":"The record lists no fixed release; boot a kernel carrying the stable fix commits below (this fix and CVE-2024-26812 are the two halves of the same INTx hardening series - take both). Interim controls: pass through MSI/MSI-X-capable devices only, or drop /dev/vfio from the container.","references":["https://git.kernel.org/stable/c/1e71b6449d55179170efc8dee8664510bb813b42","https://git.kernel.org/stable/c/3dd9be6cb55e0f47544e7cdda486413f7134e3b3","https://nvd.nist.gov/vuln/detail/CVE-2024-26810"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-476","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-26812","cve":"CVE-2024-26812","aliases":[],"title":"Linux kernel (drivers/vfio/pci): A tenant holding a passthrough PCI device can make the kernel signal an interrupt","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/vfio/pci)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A tenant holding a passthrough PCI device can make the kernel signal an interrupt eventfd whose context has already been torn down, dereferencing a NULL/stale pointer from interrupt context. The host kernel corrupts or dies, and every other tenant sharing the node goes down with it.","attack_vector":"A container or VM holding /dev/vfio/<group> plus the device fd. The tenant deconfigures the INTx eventfd (VFIO_DEVICE_SET_IRQS with fd -1) while a device interrupt is pending, then drives the loopback trigger through SET_IRQS or the unmask irqfd - the irqfd path runs asynchronously to the ioctl mutex, so it is not serialized. Conditional on vfio-pci and a passthrough device that uses legacy INTx rather than MSI/MSI-X. No host root required.","remediation":"The record lists no fixed release; boot a kernel carrying the stable fix commits below (backported across the 6.1/6.6/6.7 stable lines). Interim controls: only pass through devices using MSI/MSI-X, or remove the /dev/vfio device nodes from tenant containers.","references":["https://git.kernel.org/stable/c/b18fa894d615c8527e15d96b76c7448800e13899","https://git.kernel.org/stable/c/27d40bf72dd9a6600b76ad05859176ea9a1b4897","https://nvd.nist.gov/vuln/detail/CVE-2024-26812"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-393","CWE-787"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-26911","cve":"CVE-2024-26911","aliases":[],"title":"Linux kernel (drivers/gpu/drm): The shared VRAM buddy allocator reports success for a ranged allocation it never","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"The shared VRAM buddy allocator reports success for a ranged allocation it never actually satisfied. The caller then treats memory it does not own as its own buffer, so one tenant's buffer object can overlap VRAM the allocator subsequently hands to another tenant - corruption in both directions and a read path into someone else's frame data.","attack_vector":"Any tenant holding /dev/dri/renderD* on a driver built on drm_buddy (amdgpu, i915, xe) can drive it: fragment or exhaust VRAM, then request ranged allocations until one hits the corner case. Purely local, no capabilities, no display access needed.","remediation":"Boot a kernel with the drm_buddy fix (stable commits below; no fixed_in published). There is no meaningful interim control short of not sharing a GPU between tenants - the allocator is on every VRAM allocation path.","references":["https://git.kernel.org/stable/c/4b59c3fada06e5e8010ef7700689c71986e667a2","https://git.kernel.org/stable/c/8746c6c9dfa31d269c65dd52ab42fde0720b7d91","https://nvd.nist.gov/vuln/detail/CVE-2024-26911"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-26913","cve":"CVE-2024-26913","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Fix dcn35 8k30 Underflow/Corruption Issue","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26913","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-04-17"},{"id":"CVE-2024-26914","cve":"CVE-2024-26914","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: fix incorrect mpc_combine array size","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26914","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-04-17"},{"id":"CVE-2024-26922","cve":"CVE-2024-26922","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): Missing or insufficient validation of","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu GEM/VM/command-submission ioctl surface. A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdgpu: validate the parameters of bo mapping operations more clearly","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26922","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-04-23"},{"id":"CVE-2024-26939","cve":"CVE-2024-26939","aliases":[],"title":"Linux i915 GPU kernel driver: A use-after-free in the i915 GPU kernel driver. The general shape is that a GPU object is","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux i915 GPU kernel driver","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the i915 GPU kernel driver. The general shape is that a GPU object is freed on one path while another path still holds a reference to it, so a local user with GPU access can get the kernel to read or write freed memory. Exploitability varies by heap layout, but on a GPU node every such bug is reachable from inside a container that was granted /dev/dri - the same boundary that is supposed to separate tenants. Specific trigger: the VMA destroy path racing against retire, freeing a virtual-memory area object still in use.","attack_vector":"Any local user or container with a DRM render node - i.e. any tenant that was scheduled a GPU. No privileged capability needed.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26939","https://git.kernel.org/stable/c/0e45882ca829b26b915162e8e86dbb1095768e9e"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-05-01"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-362","CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-27062","cve":"CVE-2024-27062","aliases":[],"title":"Linux kernel (drivers/gpu/drm/nouveau/nvkm/core): Nouveau's per-client object tree had no locking at all, so concurrent","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/nouveau/nvkm/core)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Nouveau's per-client object tree had no locking at all, so concurrent object creation and destruction - most visibly VRAM BAR mappings - corrupts the tree. The reported crash is a general protection fault on a wild pointer during object lookup, which is the signature of an attacker-influenceable dangling pointer in the object namespace that backs GPU memory mappings.","attack_vector":"An unprivileged multi-threaded process in a container holding /dev/dri/renderD* on a nouveau-driven GPU races object allocation against object free; the crash was reproduced by a plain Vulkan conformance run, so no exotic ioctl sequence is needed. Conditional on the open nouveau driver being the one bound to the GPU.","remediation":"Update to a kernel with the fix commits below, which adds locking around the client object tree. Interim: blacklist nouveau where the proprietary NVIDIA driver is used, and otherwise keep /dev/dri away from untrusted tenants.","references":["https://git.kernel.org/stable/c/6887314f5356389fc219b8152e951ac084a10ef7","https://git.kernel.org/stable/c/96c8751844171af4b3898fee3857ee180586f589","https://nvd.nist.gov/vuln/detail/CVE-2024-27062"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-27400","cve":"CVE-2024-27400","aliases":[],"title":"Linux kernel DRM scheduler / TTM / dma-buf shared layer used by amdgpu (drm/amdgpu): A NULL pointer dereference","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel DRM scheduler / TTM / dma-buf shared layer used by amdgpu (drm/amdgpu)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A NULL pointer dereference in the DRM scheduler / TTM / dma-buf shared layer used by amdgpu. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: once more fix the call oder in amdgpu_ttm_move() v2","attack_vector":"Local. Reachable by any process that can submit GPU work or import/export a dma-buf - i.e. any ROCm or graphics tenant on the node. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-27400","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-14"},{"id":"CVE-2024-27459","cve":"CVE-2024-27459","aliases":[],"title":"OpenVPN: Stack overflow in the interactive service","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"OpenVPN","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Stack overflow in the interactive service -> local privilege escalation on the client host","attack_vector":"Local","remediation":"Control-plane: jump-host and operator workstation update","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-27459"],"status":"curated","published":"2024-07-08"},{"id":"CVE-2024-31583","cve":"CVE-2024-31583","aliases":[],"title":"PyTorch (mobile interpreter): Use-after-free in `torch/csrc/jit/mobile/interpreter.cpp`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"PyTorch (mobile interpreter)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Use-after-free in `torch/csrc/jit/mobile/interpreter.cpp`","attack_vector":"Customer-supplied mobile/lite model file","remediation":"Ship torch >= 2.2.0 in base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-31583"],"status":"curated","published":"2024-04-17"},{"id":"CVE-2024-31858","cve":"CVE-2024-31858","aliases":["INTEL-SA-01124","CVE-2025-33000","INTEL-SA-01373","CVE-2022-21804","CVE-2020-12333"],"title":"Intel QuickAssist Technology (QAT) software and drivers","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel QuickAssist Technology (QAT) software and drivers - QAT software before 2.2.0, with a 2025 batch through 2.6.0","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Out-of-bounds write in the QAT software stack giving an authenticated local user privilege escalation, with a further improper-input-validation escalation (CVSS 8.8) in the 2025 batch and an earlier credential-exposure issue in the Linux QAT package. QAT is the crypto and compression offload engine on Xeon platforms - it terminates TLS and does bulk compression for storage paths, so it handles key material by design, and it is a DMA-capable PCIe device. A local escalation through the QAT driver is a container-to-root path on nodes where QAT is enabled, and QAT's position in the TLS path makes credential exposure in the same stack materially worse than a generic driver bug.","attack_vector":"Authenticated local user on the host with access to the QAT device interfaces. Where QAT is exposed into containers or VMs for offload, that is the tenant.","remediation":"Update the QAT driver and software package to 2.2.0 or later (2.6.0+ for the 2025 batch) - a software/driver update from Intel, not a firmware flash, so it can go out with a service restart or reboot rather than a full firmware maintenance window. If QAT is not actually in use on a node, unbind and blacklist the driver rather than leaving an unused DMA-capable offload path exposed to tenants. Where you do expose QAT to tenants, review whether the crypto offload path is carrying keys that a tenant-side escalation would reach.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-31858","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01124.html","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01373.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-31891","cve":"CVE-2024-31891","aliases":[],"title":"IBM Storage Scale GUI (local privilege escalation): A local privilege escalation in the Storage Scale GUI available","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"IBM Storage Scale GUI (local privilege escalation)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A local privilege escalation in the Storage Scale GUI available to an actor with command-line access to the GUI service account. Escalating to root on a GUI node is escalating on a node that holds cluster-wide storage credentials, which is why this scores higher operationally than 'it's just the web UI' suggests.","attack_vector":"Local, requires command-line access as the GUI service user on Storage Scale GUI 5.1.9.0-5.1.9.6 or 5.2.0.0-5.2.1.1.","remediation":"Upgrade the Storage Scale GUI. Service-level upgrade with a restart; filesystem I/O is unaffected. Companion CSV-handling issue CVE-2024-31892 is fixed in the same range.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-31891","https://nvd.nist.gov/vuln/detail/CVE-2024-31892"],"status":"curated","published":"2024-12-14"},{"id":"CVE-2024-33656","cve":"CVE-2024-33656","aliases":["AMI-SA-2024003"],"title":"AMI AptioV UEFI BIOS (SmmComputrace DXE module): The SmmComputrace DXE module leaks stack and global memory to a local","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI AptioV UEFI BIOS (SmmComputrace DXE module)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"The SmmComputrace DXE module leaks stack and global memory to a local attacker, giving up the addresses and secrets needed to defeat firmware memory protections and escalate to arbitrary code execution and OS security bypass. Computrace is the anti-theft persistence module - it exists specifically to survive OS reinstall, so a bug in it lands in code designed for durability. Most datacenter operators do not use Computrace at all and do not realise the module is compiled into their BIOS anyway.","attack_vector":"Local, low privileges, no interaction. Any code on the host OS can start reading. On a bare-metal GPU rental the tenant qualifies without doing anything unusual.","remediation":"BIOS update from your server vendor with the fixed AptioV build - firmware flash plus host reboot per node, vendor-gated, and AMI names only 'AptioV' as the fix version so you have to confirm the specific BIOS release with your OEM. The useful config-only step here is subtractive: check whether Computrace is enabled in BIOS setup on your fleet and disable it, since datacenter operators almost never need it and it removes the attack surface with a setup change plus one reboot rather than a firmware flash campaign.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/2024/AMI-SA-2024003.pdf","https://nvd.nist.gov/vuln/detail/CVE-2024-33656"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-08-21"},{"id":"CVE-2024-33657","cve":"CVE-2024-33657","aliases":["AMI-SA-2024003"],"title":"AMI AptioV UEFI BIOS (SMM modules): An SMM vulnerability letting a privileged local attacker execute arbitrary code","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI AptioV UEFI BIOS (SMM modules)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"An SMM vulnerability letting a privileged local attacker execute arbitrary code in System Management Mode, manipulate SMM stack memory, and leak SMRAM contents into kernel space. SMRAM is supposed to be opaque to the OS; leaking it hands the attacker firmware secrets and the layout information needed to build a reliable bootkit. Once code runs in SMM the attacker is above the hypervisor and can survive OS reinstall, so a node that was compromised once should be considered compromised until its firmware is reflashed and verified, not just reimaged.","attack_vector":"Local, low privileges required per AMI's CVSS vector, no user interaction. Needs code on the host - which on any node running untrusted tenant workloads or on any node after an initial OS compromise is a given.","remediation":"BIOS update from your server vendor carrying the fixed AptioV build - firmware flash plus a full host reboot, per node. AMI's advisory names only 'AptioV' as the fix version rather than a specific BKC, so you must confirm with your OEM which BIOS release for your SKU actually contains it; do not assume 'latest' covers it. No config-only mitigation for an SMM bug.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/2024/AMI-SA-2024003.pdf","https://nvd.nist.gov/vuln/detail/CVE-2024-33657"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-08-21"},{"id":"CVE-2024-33658","cve":"CVE-2024-33658","aliases":[],"title":"AMI AptioV BIOS (memory buffer restriction failure): Local privilege escalation and potentially arbitrary code","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI AptioV BIOS (memory buffer restriction failure)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privilege escalation and potentially arbitrary code execution in firmware.","attack_vector":"Local low-privilege access.","remediation":"AMI ships the fix to OEMs, not to you - obtain the updated BIOS from your board/server vendor (Supermicro, Gigabyte, ASRock Rack, Quanta, Tyan etc.) and flash it. Expect a lag of weeks to months between the AMI advisory and an OEM image for your exact SKU, and expect some SKUs never to get one. Cold reboot per node.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/2024/AMI-SA-2024004.pdf"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-35817","cve":"CVE-2024-35817","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A correctness defect in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu GEM/VM/command-submission ioctl surface reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu: amdgpu_ttm_gart_bind set gtt bound flag","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-35817","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-17"},{"id":"CVE-2024-35931","cve":"CVE-2024-35931","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A race condition or locking defect","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A race condition or locking defect in the amdgpu GEM/VM/command-submission ioctl surface. Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdgpu: Skip do PCI error slot reset during RAS recovery","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-35931","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-19"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787","CWE-131"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-36018","cve":"CVE-2024-36018","aliases":[],"title":"Linux kernel (drivers/gpu/drm/nouveau): Nouveau's VM_BIND remap path miscalculates the address and range of the unmap","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/nouveau)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Nouveau's VM_BIND remap path miscalculates the address and range of the unmap that accompanies a partial rebind, so it tears down GPU page-table entries well outside the region the tenant actually asked to remap. The reported effect is corrupted page tables and a kernel oops - meaning a tenant's bind operations can scribble on GPU mappings that are not theirs and take the node down.","attack_vector":"Driven from userspace by a container holding /dev/dri/renderD* on a nouveau-driven NVIDIA GPU: it was found by a standard Vulkan sparse-resources conformance test, i.e. ordinary sparse-binding workloads hit it without trying. Conditional on the open nouveau driver with the uvmm/VM_BIND interface in use.","remediation":"Update to a kernel with the fix commits below. Interim: blacklist nouveau on nodes that run the proprietary NVIDIA driver, or block sparse-binding tenants from the affected hosts.","references":["https://git.kernel.org/stable/c/692a51bebf4552bdf0a79ccd68d291182a26a569","https://git.kernel.org/stable/c/0c16020d2b69a602c8ae6a1dd2aac9a3023249d6","https://nvd.nist.gov/vuln/detail/CVE-2024-36018"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-36914","cve":"CVE-2024-36914","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Skip on writeback when it's not applicable","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36914","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-30"},{"id":"CVE-2024-36971","cve":"CVE-2024-36971","aliases":[],"title":"Linux kernel (net routing): Use-after-free in network route management (__dst_negative_advice) - actively exploited","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net routing)","year":"2024","cvss_score":7.8,"severity":"high","kev":true,"impact":"Use-after-free in network route management (__dst_negative_advice) - actively exploited [KEV]","attack_vector":"Any tenant process in a container with CAP_NET_ADMIN","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2024-36971"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-06-10"},{"id":"CVE-2024-38080","cve":"CVE-2024-38080","aliases":[],"title":"Microsoft Hyper-V: Hyper-V elevation of privilege, exploited in the wild","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Microsoft Hyper-V","year":"2024","cvss_score":7.8,"severity":"high","kev":true,"impact":"Hyper-V elevation of privilege, exploited in the wild [KEV]","attack_vector":"Local user / guest on the host","remediation":"Windows update + host reboot with live-migration","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-38080"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-07-09"},{"id":"CVE-2024-38459","cve":"CVE-2024-38459","aliases":[],"title":"langchain-experimental (Python REPL): Python REPL exposed without an opt-in","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"langchain-experimental (Python REPL)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Python REPL exposed without an opt-in","attack_vector":"Untrusted agent input","remediation":"Upgrade past 0.0.61","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-38459"],"status":"curated","published":"2024-06-16"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-38545","cve":"CVE-2024-38545","aliases":[],"title":"Linux kernel (drivers/infiniband/hw/hns): The completion-queue refcount is not held under a lock, so a CQ asynchronous","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/hw/hns)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"The completion-queue refcount is not held under a lock, so a CQ asynchronous event that lands while the same CQ is being destroyed dereferences freed memory. A tenant that can provoke a CQ error event while tearing the CQ down gets a use-after-free in kernel slab memory shared with the rest of the node.","attack_vector":"Local: a tenant holding /dev/infiniband/uverbs* on a HiSilicon hns_roce adapter creates a CQ, provokes an asynchronous CQ event (for example a CQ overrun), and destroys the CQ concurrently. No fabric peer or root needed. Conditional on hns_roce hardware being the RDMA path on that node.","remediation":"No fixed version is recorded in this entry; boot a stable kernel carrying the xa_lock refcount fix (commits 330c825e66ef / 763780ef0336). Interim: drop /dev/infiniband device nodes from tenant containers on hns_roce nodes.","references":["https://git.kernel.org/stable/c/330c825e66ef65278e4ebe57fd49c1d6f3f4e34e","https://git.kernel.org/stable/c/763780ef0336a973e933e40e919339381732dcaf","https://nvd.nist.gov/vuln/detail/CVE-2024-38545"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-129","CWE-787"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-38556","cve":"CVE-2024-38556","aliases":[],"title":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core): A command that waits on the busy command-queue semaphore starts","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A command that waits on the busy command-queue semaphore starts its firmware completion timer before it owns a slot, so forced completion handling runs against index -22 and indexes the command array out of bounds. That is an attacker-influenced negative-index access in the NIC's single control path - kernel memory corruption or a panic that takes the whole shared node's fabric with it.","attack_vector":"Requires saturating the mlx5 firmware command queue so callers block on the semaphore. A tenant container holding /dev/infiniband/uverbs* can do this by opening verbs contexts and creating RDMA objects at high rate - the reported trace is literally a CREATE_UCTX, the command issued when userspace opens a verbs context - and a VF assigned into a tenant VM has the same reach. No host root and no fabric position needed.","remediation":"Boot a kernel carrying the command-semaphore timeout fix; kernel.org records it landing across the 5.5, 5.11, 5.16 and 5.17 stable branches, so take the latest point release on whichever branch you run and confirm your distro backported it. Interim controls: cap per-tenant RDMA object and context creation rates, and remove /dev/infiniband/* from containers that do not need verbs.","references":["https://git.kernel.org/stable/c/f9caccdd42e999b74303c9b0643300073ed5d319","https://git.kernel.org/stable/c/2d0962d05c93de391ce85f6e764df895f47c8918","https://nvd.nist.gov/vuln/detail/CVE-2024-38556"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-38581","cve":"CVE-2024-38581","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/mes): A use-after-free in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/mes)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the amdgpu firmware, ACPI and IP-block initialisation. Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amdgpu/mes: fix use-after-free issue","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-38581","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-06-19"},{"id":"CVE-2024-39291","cve":"CVE-2024-39291","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): An out-of-bounds access in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu firmware, ACPI and IP-block initialisation - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: Fix buffer size in gfx_v9_4_3_init_ cp_compute_microcode() and rlc_microcode()","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-39291","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-06-24"},{"id":"CVE-2024-39471","cve":"CVE-2024-39471","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): An out-of-bounds access in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu firmware, ACPI and IP-block initialisation - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: add error handle to avoid out-of-bounds","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-39471","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-06-25"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-39486","cve":"CVE-2024-39486","aliases":[],"title":"Linux kernel (drivers/gpu/drm): DRM core stores a pointer to the caller's struct pid before taking a reference on it","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"DRM core stores a pointer to the caller's struct pid before taking a reference on it, so two processes issuing ioctls on the same DRM fd can drive that pid to refcount zero while it is still referenced. This is a use-after-free of a core kernel object reachable from any GPU node, on any driver - a strong local privilege-escalation primitive out of a tenant container.","attack_vector":"Any tenant holding any /dev/dri node (renderD* is enough - the path is in drm_file, not a driver). Share the fd across two processes via fork() or SCM_RIGHTS and have both issue DRM ioctls in a loop. Requires CONFIG_PREEMPT_RCU, which is common on distro kernels. Driver-independent: amdgpu, xe, i915, nouveau, virtio-gpu are all exposed.","remediation":"Update to 6.6.37 or later (or the equivalent fix in your stable series). No workable interim control other than removing /dev/dri from tenant containers - this is core DRM, hit by every ioctl.","references":["https://git.kernel.org/stable/c/16682588ead4a593cf1aebb33b36df4d1e9e4ffa","https://git.kernel.org/stable/c/0acce2a5c619ef1abdee783d7fea5eac78ce4844","https://nvd.nist.gov/vuln/detail/CVE-2024-39486"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-40990","cve":"CVE-2024-40990","aliases":["RDMA/mlx5 add check for srq max_sge attribute"],"title":"Linux kernel mlx5_ib (shared receive queue): The max_sge attribute for a shared receive queue is taken from the user","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_ib (shared receive queue)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"The max_sge attribute for a shared receive queue is taken from the user and used unchecked. A tenant process creating an SRQ supplies a value the kernel trusts - memory corruption from an ordinary verbs call available to any RDMA workload on the node.","attack_vector":"Local, low-privileged process with RDMA verbs access on an mlx5 device.","remediation":"Upgrade the host kernel to 6.10 or a stable backport (5.10.221, 5.15.162, 6.1.96, 6.6.36, 6.9.7). Rolling reboot of the RDMA fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-40990","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2024/CVE-2024-40990.json"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-07-12"},{"id":"CVE-2024-41008","cve":"CVE-2024-41008","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A NULL pointer dereference in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A NULL pointer dereference in the amdgpu GEM/VM/command-submission ioctl surface. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: change vm->task_info handling","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-41008","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-07-16"},{"id":"CVE-2024-41011","cve":"CVE-2024-41011","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A correctness defect in the amdkfd (KFD compute","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdkfd (KFD compute driver, /dev/kfd) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdkfd: don't allow mapping the MMIO HDP page with large pages","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-41011","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-07-18"},{"id":"CVE-2024-41061","cve":"CVE-2024-41061","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Fix array-index-out-of-bounds in dml2/FCLKChangeSupport","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-41061","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-07-29"},{"id":"CVE-2024-41092","cve":"CVE-2024-41092","aliases":[],"title":"Linux i915 GPU kernel driver: A use-after-free in the i915 GPU kernel driver. The general shape is that a GPU object is","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux i915 GPU kernel driver","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the i915 GPU kernel driver. The general shape is that a GPU object is freed on one path while another path still holds a reference to it, so a local user with GPU access can get the kernel to read or write freed memory. Exploitability varies by heap layout, but on a GPU node every such bug is reachable from inside a container that was granted /dev/dri - the same boundary that is supposed to separate tenants. Specific trigger: revocation of fence registers racing with their use, leaving a dangling fence.","attack_vector":"Any local user or container with a DRM render node - i.e. any tenant that was scheduled a GPU. No privileged capability needed.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-41092","https://git.kernel.org/stable/c/06dec31a0a5112a91f49085e8a8fa1a82296d5c7"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-07-29"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-41096","cve":"CVE-2024-41096","aliases":[],"title":"Linux kernel (drivers/pci/msi): When MSI vector allocation for a PCI device fails, the core MSI code keeps using a","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/pci/msi)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"When MSI vector allocation for a PCI device fails, the core MSI code keeps using a descriptor that has already been freed, giving a use-after-free read (and a freed-slab pointer chase) inside the host kernel. On a passthrough node this is host-side memory corruption driven from the enable path of a device a tenant controls, which is the wrong side of the guest boundary for a dangling kernel pointer.","attack_vector":"Reachable on the host whenever pci_alloc_irq_vectors() takes the MSI failure path. The path that matters for a GPU cluster is device passthrough: a tenant VM holding an assigned GPU or NIC through /dev/vfio/* chooses when MSI is enabled and how many vectors are requested (VFIO_DEVICE_SET_IRQS -> vfio_msi_enable -> pci_alloc_irq_vectors), so a guest that asks for a vector count the host cannot satisfy drives the failing allocation itself. Requires vfio-pci passthrough (or any host driver bind that can fail MSI setup); a container with no device node cannot reach it.","remediation":"Boot a kernel carrying the msi_capability_init() fix (backported across the 5.10/5.15/6.1/6.6 stable lines - confirm your distro's build). Interim: keep host IRQ vector headroom so guest-requested MSI allocation does not fail, cap the vector count exposed to guests, and do not hand /dev/vfio/* to workloads that do not need passthrough.","references":["https://git.kernel.org/stable/c/0ae40b2d0a5de6b045504098e365d4fdff5bbeba","https://git.kernel.org/stable/c/ff1121d2214b794dc1772081f27bdd90721a84bc","https://nvd.nist.gov/vuln/detail/CVE-2024-41096"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-42064","cve":"CVE-2024-42064","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Skip pipe if the pipe idx not set properly","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-42064","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-07-29"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-190"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-42066","cve":"CVE-2024-42066","aliases":[],"title":"Linux kernel (drivers/gpu/drm/xe): The Xe VRAM manager computed a buffer object's minimum page size by shifting a","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/xe)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"The Xe VRAM manager computed a buffer object's minimum page size by shifting a 32-bit alignment value, which overflows before it is stored into a 64-bit field. A tenant that requests a large page alignment gets a wrong, truncated minimum size back into the VRAM allocator, which is how allocations end up smaller or worse-aligned than the driver believes - the setup for one tenant's buffer overlapping memory the allocator considers someone else's.","attack_vector":"Reachable from a container holding /dev/dri/renderD* on an Intel Xe node through ordinary buffer-object creation with attacker-chosen size and alignment; the overflow is in the TTM VRAM manager's allocation path, so no special ioctl or privilege is needed.","remediation":"Boot a kernel containing the fix commits below. No practical interim control other than withholding render-node access from untrusted tenants.","references":["https://git.kernel.org/stable/c/79d54ddf0e292b810887994bb04709c5ac0e1531","https://git.kernel.org/stable/c/4f4fcafde343a54465f85a2909fc684918507a4b","https://nvd.nist.gov/vuln/detail/CVE-2024-42066"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-42117","cve":"CVE-2024-42117","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: ASSERT when failing to find index by plane/stream id","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-42117","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-07-30"},{"id":"CVE-2024-42118","cve":"CVE-2024-42118","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Do not return negative stream id for array","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-42118","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-07-30"},{"id":"CVE-2024-42119","cve":"CVE-2024-42119","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Skip finding free audio for unknown engine_id","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-42119","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-07-30"},{"id":"CVE-2024-42120","cve":"CVE-2024-42120","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Check pipe offset before setting vblank","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-42120","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-07-30"},{"id":"CVE-2024-42121","cve":"CVE-2024-42121","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Check index msg_id before read or write","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-42121","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-07-30"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-125","CWE-787"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-42259","cve":"CVE-2024-42259","aliases":[],"title":"Linux kernel (drivers/gpu/drm/i915/gem): The size of a partial GEM mapping is computed without accounting for the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/i915/gem)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"The size of a partial GEM mapping is computed without accounting for the mapping offset, so the mapped window can extend past the end of the buffer object. A tenant that faults those pages reaches memory outside its own BO - read and write access to pages the driver never intended to expose to it.","attack_vector":"Tenant container holding /dev/dri/renderD* on an Intel i915 device: create a BO, request an mmap offset with a partial view / non-zero framebuffer offset, mmap it and touch pages past the object's end. Unprivileged, no display or master access required.","remediation":"Update to a stable kernel carrying the fix (commits below; no fixed_in published in the record). Interim: drop /dev/dri/renderD* from containers running untrusted code on i915 nodes.","references":["https://git.kernel.org/stable/c/3e06073d24807f04b4694108a8474decb7b99e60","https://git.kernel.org/stable/c/a256d019eaf044864c7e50312f0a65b323c24f39","https://nvd.nist.gov/vuln/detail/CVE-2024-42259"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-4467","cve":"CVE-2024-4467","aliases":[],"title":"QEMU (qemu-img): `qemu-img info` on an untrusted qcow2 image reaches arbitrary host file read/write","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"QEMU (qemu-img)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"`qemu-img info` on an untrusted qcow2 image reaches arbitrary host file read/write","attack_vector":"Tenant-supplied disk image processed by the control plane","remediation":"Update qemu-img and never run it untrusted-unsandboxed; no reboot. Directly relevant to any \"bring your own VM image\" feature","references":["https://access.redhat.com/security/cve/CVE-2024-4467"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2024-07-02"},{"id":"CVE-2024-44977","cve":"CVE-2024-44977","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): An out-of-bounds access in the amdgpu kernel driver core - a","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu kernel driver core - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: Validate TA binary size","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-44977","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-09-04"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-44978","cve":"CVE-2024-44978","aliases":[],"title":"Linux kernel (drivers/gpu/drm/xe): Freeing a scheduler job dereferences the VM it belongs to, but the final exec-queue","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/xe)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Freeing a scheduler job dereferences the VM it belongs to, but the final exec-queue put that happens first can already have destroyed that VM. The result is a use-after-free on the scheduler job path, reachable from an ordinary submit-then-close sequence inside a tenant container.","attack_vector":"Tenant holding /dev/dri/renderD* on Intel xe: submit jobs, then destroy the exec queue and VM so the last reference drops while jobs are still being freed. Unprivileged; the ordering is under the tenant's control.","remediation":"Update to a kernel carrying the fix (stable commits below; no fixed_in published). No interim control - this is the standard submission teardown path.","references":["https://git.kernel.org/stable/c/98aa0330f200b9b8fb9e1298e006eda57a13351c","https://git.kernel.org/stable/c/9e7f30563677fbeff62d368d5d2a5ac7aaa9746a","https://nvd.nist.gov/vuln/detail/CVE-2024-44978"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-45019","cve":"CVE-2024-45019","aliases":["net/mlx5e take state lock during tx timeout reporter"],"title":"Linux kernel mlx5_core TX timeout devlink health reporter: The TX timeout recovery path runs without the state lock, so","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_core TX timeout devlink health reporter","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"The TX timeout recovery path runs without the state lock, so it races channel teardown. The racing side is driven by unprivileged statistics reads - cat /proc/net/dev, ip -s link, or anything scraping /sys/class/net/*/statistics/*. That requeues the stats work while the channels are being freed, giving a reliable use-after-free read of groomable slab memory from an ordinary unprivileged account. Every Prometheus node-exporter on the fleet is polling exactly these files.","attack_vector":"Local unprivileged user reading standard network statistics, combined with a TX timeout or channel reconfiguration. No special capability needed.","remediation":"Upgrade the host kernel to a build carrying the fix (6.11 and the corresponding stable backports). Rolling reboot of the fleet. No practical config mitigation - you are not going to stop metrics collection.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-45019","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2024/CVE-2024-45019.json"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-11"},{"id":"CVE-2024-4610","cve":"CVE-2024-4610","aliases":[],"title":"Arm Mali GPU kernel driver: Use-after-free in the Bifrost/Valhall GPU kernel driver","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Arm Mali GPU kernel driver","year":"2024","cvss_score":7.8,"severity":"high","kev":true,"impact":"Use-after-free in the Bifrost/Valhall GPU kernel driver - a local non-privileged user reaches freed kernel memory; exploited in the wild [KEV]","attack_vector":"Local user with GPU device access","remediation":"Driver update + reboot. Not applicable to NVIDIA/AMD datacentre GPUs, but directly relevant to any Arm-based (Grace, Ampere Altra) node that ships a Mali display GPU, and it is the closest existing precedent for tenant-reachable GPU-driver privesc","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-4610"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-06-07"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-46683","cve":"CVE-2024-46683","aliases":[],"title":"Linux kernel (drivers/gpu/drm/xe): The preempt-fence lock lives inside the exec queue, but the queue reference is","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/xe)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"The preempt-fence lock lives inside the exec queue, but the queue reference is dropped as soon as the fence is signalled. A waiter woken afterwards takes a lock inside already-freed memory - a use-after-free of driver state driven by ordinary long-running-queue and VM-bind traffic, and a foothold for kernel memory corruption from a tenant container.","attack_vector":"Tenant holding /dev/dri/renderD* on Intel xe using long-running exec queues or VM binds (i.e. any compute workload): race queue teardown against a thread waiting on the preempt fence. Multiple independent reproducers are referenced in the upstream fix.","remediation":"Update to a kernel carrying the fix (stable commits below; no fixed_in published). No selective interim control - preempt fences are on the normal LR/compute submission path.","references":["https://git.kernel.org/stable/c/10081b0b0ed201f53e24bd92deb2e0f3c3e713d4","https://git.kernel.org/stable/c/730b72480e29f63fd644f5fa57c9d46109428953","https://nvd.nist.gov/vuln/detail/CVE-2024-46683"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-46725","cve":"CVE-2024-46725","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): An out-of-bounds access in the amdgpu kernel driver core - a","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu kernel driver core - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: Fix out-of-bounds write warning","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46725","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-09-18"},{"id":"CVE-2024-46729","cve":"CVE-2024-46729","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A division by zero in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A division by zero in the amdgpu display core (DC/DM), reachable with attacker-influenced parameters. The kernel takes a divide fault and the node goes down, taking every co-resident job with it. Upstream fix: drm/amd/display: Fix incorrect size calculation for loop","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46729","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-18"},{"id":"CVE-2024-46730","cve":"CVE-2024-46730","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Ensure array index tg_inst won't be -1","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46730","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-18"},{"id":"CVE-2024-46746","cve":"CVE-2024-46746","aliases":[],"title":"Linux HID/amd_sfh - driver_data freed after HID device destruction: A use-after-free in the AMD Sensor Fusion Hub HID","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux HID/amd_sfh - driver_data freed after HID device destruction","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the AMD Sensor Fusion Hub HID driver: driver_data is freed in the wrong order relative to hid_destroy_device(), so callbacks touch memory that is already gone. Kernel UAF is a privilege-escalation primitive. AMD SFH is a client-platform driver and unlikely to be loaded on an Instinct server - but it is compiled into stock distro kernels, and a driver that is present but unneeded is attack surface you are carrying for nothing.","attack_vector":"Local, on hosts where the amd_sfh driver is loaded.","remediation":"Fixed in the Linux kernel; take the distro update and reboot. Better: blacklist amd_sfh on server images. Auditing your GPU nodes for client-platform drivers that autoload and are never used is a cheap one-off that shrinks the kernel attack surface permanently.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46746"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-18"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-667","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-46750","cve":"CVE-2024-46750","aliases":[],"title":"Linux kernel (drivers/pci): Pci_bus_lock() locked every device on the bus except the bridge itself, so a secondary bus","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/pci)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Pci_bus_lock() locked every device on the bus except the bridge itself, so a secondary bus reset gets issued while config-space access to that bridge is still unblocked. Config reads and writes race against a live link reset, which means device state can be read or written mid-reset and left inconsistent - the CNA rated this high on confidentiality AND integrity, not just availability.","attack_vector":"A tenant container or VM holding /dev/vfio/* can drive the reset half of this itself: VFIO exposes device and bus reset to userspace (VFIO_DEVICE_RESET / VFIO_DEVICE_PCI_HOT_RESET), and those paths land in the pci_reset_bus() / pci_bus_lock() code that is missing the bridge lock. The upstream report came from vmd_probe() issuing an unlocked secondary bus reset. No host root is required for the tenant side once the device node is in the container; the exposure is real on any node where GPUs or NICs are passed through with VFIO.","remediation":"Update to a kernel carrying the fix (no fixed_in published; the stable commits below are backported into 6.1.y and 6.6.y). Interim: do not grant tenants VFIO groups that share a bridge with another tenant's device, and verify each passthrough device sits in its own IOMMU group behind its own downstream port before handing the node out.","references":["https://git.kernel.org/stable/c/0790b89c7e911003b8c50ae50e3ac7645de1fae9","https://git.kernel.org/stable/c/df77a678c33871a6e4ac5b54a71662f1d702335b","https://nvd.nist.gov/vuln/detail/CVE-2024-46750"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-46803","cve":"CVE-2024-46803","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): Missing or insufficient validation of user-supplied","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdkfd (KFD compute driver, /dev/kfd). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdkfd: Check debug trap enable before write dbg_ev_file","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46803","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-09-27"},{"id":"CVE-2024-46804","cve":"CVE-2024-46804","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Add array index check for hdcp ddc access","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46804","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-27"},{"id":"CVE-2024-46812","cve":"CVE-2024-46812","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Skip inactive planes within ModeSupportAndSystemConfiguration","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46812","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-27"},{"id":"CVE-2024-46813","cve":"CVE-2024-46813","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Check link_index before accessing dc->links[]","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46813","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-27"},{"id":"CVE-2024-46814","cve":"CVE-2024-46814","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Check msg_id before processing transcation","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46814","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-27"},{"id":"CVE-2024-46816","cve":"CVE-2024-46816","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Stop amdgpu_dm initialize when link nums greater than max_links","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46816","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-27"},{"id":"CVE-2024-46818","cve":"CVE-2024-46818","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Check gpio_id before used as array index","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46818","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-27"},{"id":"CVE-2024-46820","cve":"CVE-2024-46820","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn): A correctness defect in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/vcn: remove irq disabling in vcn 5 suspend","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46820","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-27"},{"id":"CVE-2024-46821","cve":"CVE-2024-46821","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): An out-of-bounds access in the amdgpu power","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu power management (SMU/powerplay) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/pm: Fix negative array index read","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46821","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-27"},{"id":"CVE-2024-46836","cve":"CVE-2024-46836","aliases":[],"title":"ASPEED USB device controller driver (drivers/usb/gadget/udc/aspeed_udc.c): The BMC presents itself to the host over USB","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ASPEED USB device controller driver (drivers/usb/gadget/udc/aspeed_udc.c)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"The BMC presents itself to the host over USB - that is how OpenBMC does virtual media, USB-network (the host-to-BMC management channel) and HID for iKVM. The endpoint index coming from the host is used without a bounds check, so the host walks past the endpoint array into adjacent BMC kernel memory. The direction of travel is what matters here: the host is the untrusted side on a rented bare-metal node, and this is a host-controlled index into BMC kernel structures. It is a genuine tenant-to-BMC crossing, not a BMC-local bug.","attack_vector":"Requires the ability to drive USB control traffic from the host to the BMC's USB gadget - root on the bare-metal server, which any tenant of that node has. The gadget is enabled by default on OpenBMC platforms that offer virtual media or a USB management NIC.","remediation":"Kernel patch backported across stable trees; on real fleets it arrives only in a new BMC firmware image, so per-node out-of-band flash with the usual ODM lag and brick risk. Partial config mitigation: unbind or do not compose the virtual-media and USB-gadget functions on nodes that do not need them. That is often viable for GPU nodes provisioned by network boot, less so where the ODM's own management path rides the USB NIC.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46836","https://git.kernel.org/stable/c/b2a50ffdd1a079869a62198a8d1441355c513c7c","https://lists.debian.org/debian-lts-announce/2025/01/msg00001.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-09-27"},{"id":"CVE-2024-46850","cve":"CVE-2024-46850","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Avoid race between dcn35_set_drr() and dc_state_destruct()","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46850","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-27"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-46866","cve":"CVE-2024-46866","aliases":[],"title":"Linux kernel (drivers/gpu/drm/xe): The per-client memory accounting walks buffer-object state (TTM resource, tt pages)","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/xe)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"The per-client memory accounting walks buffer-object state (TTM resource, tt pages) with no lock and no reference held. A tenant that keeps BOs churning while its fdinfo is read gets the kernel to follow freed pointers - use-after-free and NULL dereference in kernel context, the usual starting point for a container escape.","attack_vector":"Tenant holding /dev/dri/renderD* on Intel xe, plus anything that reads /proc/<pid>/fdinfo for that DRM fd - which the tenant can do to itself, and which fleet GPU-telemetry agents do continuously across tenants. One thread allocates and evicts BOs while the other reads fdinfo. No capabilities needed.","remediation":"Update to a kernel with the fix (stable commits below; no fixed_in published). Interim: stop any fleet agent that scrapes DRM fdinfo on xe nodes - it does not remove the self-triggered path but it removes the cross-tenant one.","references":["https://git.kernel.org/stable/c/abc8feacacf8fae10eecf6fea7865e8c1fee419c","https://git.kernel.org/stable/c/94c4aa266111262c96c98f822d1bccc494786fee","https://nvd.nist.gov/vuln/detail/CVE-2024-46866"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-46871","cve":"CVE-2024-46871","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Correct the defined value for AMDGPU_DMUB_NOTIFICATION_MAX","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46871","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-09"},{"id":"CVE-2024-48837","cve":"CVE-2024-48837","aliases":[],"title":"Dell SmartFabric OS10 (execution with unnecessary privileges): Low-privileged local attacker reaches command execution","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell SmartFabric OS10 (execution with unnecessary privileges)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Low-privileged local attacker reaches command execution through an over-privileged component of the switch OS.","attack_vector":"Local low-privilege switch access.","remediation":"Upgrade OS10 per DSA-2024-425.","references":["https://www.dell.com/support/kbdoc/en-us/000247217/dsa-2024-425-security-update-for-dell-networking-os10-vulnerabilities"],"status":"curated"},{"id":"CVE-2024-49557","cve":"CVE-2024-49557","aliases":[],"title":"Dell SmartFabric OS10 (command injection): Command injection from a low-privileged local account leading to code","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell SmartFabric OS10 (command injection)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Command injection from a low-privileged local account leading to code execution on the switch.","attack_vector":"Local low-privilege switch CLI access.","remediation":"Upgrade OS10 per DSA-2024-425.","references":["https://www.dell.com/support/kbdoc/en-us/000247217/dsa-2024-425-security-update-for-dell-networking-os10-vulnerabilities"],"status":"curated"},{"id":"CVE-2024-49558","cve":"CVE-2024-49558","aliases":[],"title":"Dell SmartFabric OS10 (improper privilege management): A low-privileged local attacker elevates privileges on the switch","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell SmartFabric OS10 (improper privilege management)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A low-privileged local attacker elevates privileges on the switch.","attack_vector":"Local low-privilege switch access.","remediation":"Upgrade OS10 per DSA-2024-425.","references":["https://www.dell.com/support/kbdoc/en-us/000247217/dsa-2024-425-security-update-for-dell-networking-os10-vulnerabilities"],"status":"curated"},{"id":"CVE-2024-49560","cve":"CVE-2024-49560","aliases":[],"title":"Dell SmartFabric OS10 (command injection): A low-privileged local attacker executes commands on the switch OS","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell SmartFabric OS10 (command injection)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A low-privileged local attacker executes commands on the switch OS.","attack_vector":"Local access to the switch CLI with a low-privilege account.","remediation":"Upgrade OS10 per DSA-2024-425. Switch reboot required.","references":["https://www.dell.com/support/kbdoc/en-us/000247217/dsa-2024-425-security-update-for-dell-networking-os10-vulnerabilities"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-49561","cve":"CVE-2024-49561","aliases":[],"title":"Dell SmartFabric OS10 (incorrect privilege assignment): Local low-privilege attacker escalates privileges on the switch","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell SmartFabric OS10 (incorrect privilege assignment)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local low-privilege attacker escalates privileges on the switch OS.","attack_vector":"Local low-privilege switch access.","remediation":"Upgrade OS10 per DSA-2025-070/069/079.","references":["https://www.dell.com/support/kbdoc/en-us/000289970/dsa-2025-070-security-update-for-dell-networking-os10-vulnerabilities"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-49865","cve":"CVE-2024-49865","aliases":[],"title":"Linux kernel (drivers/gpu/drm/xe): The GPU VM is published into the id table before the create ioctl finishes with it","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/xe)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"The GPU VM is published into the id table before the create ioctl finishes with it, so a tenant that guesses the id and calls VM destroy in parallel frees the VM while the create path is still using it. Use-after-free of a driver object under attacker timing control - the standard route from a tenant container to kernel code execution. The upstream fix names the attacker explicitly.","attack_vector":"Hostile tenant holding /dev/dri/renderD* on Intel xe: one thread spams VM_CREATE, another spams VM_DESTROY against the predictable next id. Unprivileged, no display access, no special platform features.","remediation":"Update to a kernel carrying the fix (stable commits below; no fixed_in published). Interim: remove /dev/dri/renderD* from containers running untrusted code on xe nodes - the ioctl pair cannot be filtered selectively.","references":["https://git.kernel.org/stable/c/09cf8901fc0225898311b375cfcc67bae37ed5da","https://git.kernel.org/stable/c/74231870cf4976f69e83aa24f48edb16619f652f","https://nvd.nist.gov/vuln/detail/CVE-2024-49865"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-49895","cve":"CVE-2024-49895","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Fix index out of bounds in DCN30 degamma hardware format translation","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49895","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"id":"CVE-2024-49969","cve":"CVE-2024-49969","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Fix index out of bounds in DCN30 color transformation","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49969","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"id":"CVE-2024-49989","cve":"CVE-2024-49989","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A double free in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A double free in the amdgpu display core (DC/DM). The same allocation is released twice, corrupting the slab allocator's freelist. This is a classic heap-corruption primitive: with slab grooming it becomes arbitrary kernel memory write and therefore host compromise from an unprivileged GPU workload. The cheap outcome is a node panic. Upstream fix: drm/amd/display: fix double free issue during amdgpu module unload","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49989","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"id":"CVE-2024-49991","cve":"CVE-2024-49991","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A use-after-free in the amdkfd (KFD compute driver","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the amdkfd (KFD compute driver, /dev/kfd). Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amdkfd: amdkfd_free_gtt_mem clear the correct pointer","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49991","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-10-21"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-763","CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-50001","cve":"CVE-2024-50001","aliases":[],"title":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core): When a DMA mapping fails on the multi-packet transmit path, the","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"When a DMA mapping fails on the multi-packet transmit path, the error handler tears down an unrelated, still-live mapping from the send queue's FIFO. The NIC then DMAs against an IOVA the kernel considers free - which the IOMMU can hand to a different mapping - so the device writes packet data into memory that now belongs to something else, or the PCI function is thrown into error state and the node loses its fabric link entirely.","attack_vector":"The failure trigger is a DMA mapping failure on the TX hot path, which any tenant can manufacture by putting the node under memory pressure while transmitting - no privilege, no device node, no fabric position required. It was originally observed in a plain stress-test environment, so it is not a theoretical corner. Impact is worst on systems with a strict IOMMU (s390 puts the function in error state; on x86 the freed IOVA can be recycled into another domain's mapping).","remediation":"Boot a kernel with the mlx5e MPWQE error-path fix (stable commits below; no fixed-version list published for this ID - match against your distro's mlx5_core backport). No useful interim control: the path is core packet transmit and cannot be disabled without taking the NIC out of service.","references":["https://git.kernel.org/stable/c/ce828b347cf1b3c1b12b091d02463c35ce5097f5","https://git.kernel.org/stable/c/fc357e78176945ca7bcacf92ab794b9ccd41b4f4","https://nvd.nist.gov/vuln/detail/CVE-2024-50001"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787","CWE-122"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-50090","cve":"CVE-2024-50090","aliases":[],"title":"Linux kernel (drivers/gpu/drm/xe): The observation/OA path reuses one batch buffer and appends a batch-end command on","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/xe)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"The observation/OA path reuses one batch buffer and appends a batch-end command on every reconfiguration, so repeated use runs the write past the end of the allocation. The kernel overwrites adjacent memory and the GPU then executes whatever follows the batch - kernel heap corruption driven entirely from userspace, plus GPU-side execution of unintended commands.","attack_vector":"Tenant holding /dev/dri/renderD* on Intel xe opens an observation (OA) stream against its own exec queue and repeatedly reconfigures the same metric set. Per-queue OA does not require CAP_PERFMON, so this is reachable from an ordinary tenant container.","remediation":"Update to a kernel carrying the fix (stable commits below; no fixed_in published). Interim: set perf_stream_paranoid to block observation streams, or deny the xe observation ioctl to tenant workloads.","references":["https://git.kernel.org/stable/c/bcb5be3421705e682b0b32073ad627056d6bc2a2","https://git.kernel.org/stable/c/6c10ba06bb1b48acce6d4d9c1e33beb9954f1788","https://nvd.nist.gov/vuln/detail/CVE-2024-50090"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-476","CWE-459"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-50101","cve":"CVE-2024-50101","aliases":[],"title":"Linux kernel (drivers/iommu/intel): VT-d walked the PCI DMA-alias list for devices that are not PCI at all while","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/intel)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"VT-d walked the PCI DMA-alias list for devices that are not PCI at all while clearing a domain's context entries, operating on the wrong structure. Upstream reports kernel hangs; the isolation consequence is that context-entry teardown on detach does not do what it was meant to, which is how stale translations survive a device being detached from its domain.","attack_vector":"Reached on domain detach and device release for non-PCI (ACPI namespace or platform) devices behind Intel VT-d. This is host-side device lifecycle, not a tenant ioctl - but it runs on the same detach path that is supposed to revoke a departing tenant's DMA access. The vendor scores it 7.8 with full confidentiality, integrity and availability impact.","remediation":"Update to 5.15.169, 6.1.114, or 6.6.58 or later. No practical interim control - the path runs whenever a non-PCI device behind VT-d is detached.","references":["https://git.kernel.org/stable/c/0bd9a30c22afb5da203386b811ec31429d2caa78","https://git.kernel.org/stable/c/cbfa3a83eba05240ce37839ed48280a05e8e8f6c","https://nvd.nist.gov/vuln/detail/CVE-2024-50101"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-50149","cve":"CVE-2024-50149","aliases":[],"title":"Linux kernel (drivers/gpu/drm/xe): Xe freed a job from inside timeout-detection-and-recovery while the submission","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/xe)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Xe freed a job from inside timeout-detection-and-recovery while the submission thread could still be running it, giving a use-after-free on the job object. The trigger is a GPU hang, and any tenant can cause a GPU hang on demand with a runaway shader - so this converts a self-inflicted hang into kernel memory corruption on a node shared with other tenants.","attack_vector":"An unprivileged container with /dev/dri/renderD* on an Intel Xe node submits a job that never completes; the TDR path then fires and races the run_job thread. Deliberately hanging the GPU is trivial from userspace, which makes this attacker-scheduled rather than incidental.","remediation":"Update to a kernel with the fix commits below, which defers the free to the scheduler. Interim: restrict render-node access to trusted workloads and monitor for repeated GPU resets, which are the signal that someone is exercising this path.","references":["https://git.kernel.org/stable/c/be8fe75e57f8fa3f87e3b1c283cc7cd9f9b80867","https://git.kernel.org/stable/c/82926f52d7a09c65d916c0ef8d4305fc95d68c0c","https://nvd.nist.gov/vuln/detail/CVE-2024-50149"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-50158","cve":"CVE-2024-50158","aliases":[],"title":"Linux kernel (drivers/infiniband/hw/bnxt_re): Collecting hardware counters writes doorbell-pacing statistics into a","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/hw/bnxt_re)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Collecting hardware counters writes doorbell-pacing statistics into a stats buffer that was never sized for them on adapters outside GenP5/P7, producing an 8-byte slab out-of-bounds write. KASAN caught it as a heap overflow; in practice an unprivileged reader of the counters corrupts neighbouring kernel slab objects on a node shared with other tenants.","attack_vector":"Local and unprivileged: reading the RDMA hardware counters (sysfs hw_counters, or an `rdma statistic` query) is enough to run the faulty parse. A tenant does not need a device node - counter exposure is wider than /dev/infiniband. Conditional on a Broadcom bnxt_re adapter that reports doorbell pacing while not being a GenP5/P7 part.","remediation":"No fixed version is recorded in this entry; boot a stable kernel carrying the out-of-bound check fix (commits 05c5fcc1869a / c11b9b03ea52). Interim: restrict access to RDMA counter surfaces (sysfs hw_counters, rdma netlink) from tenant namespaces on bnxt_re nodes.","references":["https://git.kernel.org/stable/c/05c5fcc1869a08e36a29691699b6534e5a00a82b","https://git.kernel.org/stable/c/c11b9b03ea5252898f91f3388c248f0dc47bda52","https://nvd.nist.gov/vuln/detail/CVE-2024-50158"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-50208","cve":"CVE-2024-50208","aliases":[],"title":"Linux kernel (drivers/infiniband/hw/bnxt_re): Building the two-level page list for a large RDMA resource assumes","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/hw/bnxt_re)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Building the two-level page list for a large RDMA resource assumes multiple page-directory pages where the hardware only has one, so past 256K page-list entries the driver writes into memory it does not own. A tenant that asks for a big enough queue or resource turns a normal allocation into kernel heap corruption on a shared node.","attack_vector":"Local: a tenant holding /dev/infiniband/uverbs* on a Broadcom bnxt_re adapter reaches it by creating a non-MR resource large enough to need more than 256K page-list entries (very deep queues / very large mappings). No fabric peer needed. Conditional on bnxt_re hardware.","remediation":"No fixed version is recorded in this entry; boot a stable kernel carrying the Level-2 PBL fix (commits df6fed0a2a1a / de5857fa7bcc). Interim: cap per-tenant RDMA resource limits, or remove /dev/infiniband from containers that do not need native verbs on bnxt_re nodes.","references":["https://git.kernel.org/stable/c/df6fed0a2a1a5e57f033bca40dc316b18e0d0ce6","https://git.kernel.org/stable/c/de5857fa7bcc9a496a914c7e21390be873109f26","https://nvd.nist.gov/vuln/detail/CVE-2024-50208"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-50221","cve":"CVE-2024-50221","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): An out-of-bounds access in the amdgpu power","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu power management (SMU/powerplay) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/pm: Vangogh: Fix kernel memory out of bounds write","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-50221","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-11-09"},{"id":"CVE-2024-50282","cve":"CVE-2024-50282","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): An out-of-bounds access in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu GEM/VM/command-submission ioctl surface - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: add missing size check in amdgpu_debugfs_gprwave_read()","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-50282","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-11-19"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-362","CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-53121","cve":"CVE-2024-53121","aliases":[],"title":"NVIDIA/Mellanox ConnectX flow steering core (mlx5 fs_core FTE active-flag race): Two-phase flow-table-entry deletion","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA/Mellanox ConnectX flow steering core (mlx5 fs_core FTE active-flag race)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Two-phase flow-table-entry deletion checks the entry's active flag without holding its lock, so adding a rule with the same match value concurrently with a delete makes fs_core clear the hardware deletion function early and panic on the next removal. The race is between adding and removing rules that match the same traffic - which is what happens when two tenants' network policies touch overlapping match criteria on a shared adapter, so this is reachable by ordinary concurrent activity, not just by a deliberate attacker.","attack_vector":"Local. Concurrent rule add and delete for the same match value on a shared mlx5 adapter, driven by CNI/network-policy churn.","remediation":"Kernel update checking the FTE active flag under its lock. No workaround.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=094d1a2121cee1e85ab07d74388f94809dcfb5b9","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2024/CVE-2024-53121.json"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-53133","cve":"CVE-2024-53133","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Handle dml allocation failure to avoid crash","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53133","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-12-04"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-125","CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-53214","cve":"CVE-2024-53214","aliases":[],"title":"Linux kernel (drivers/vfio/pci): Out-of-bounds read past the ecap_perms table when a tenant touches emulated PCIe","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/vfio/pci)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Out-of-bounds read past the ecap_perms table when a tenant touches emulated PCIe extended config space on a device whose first extended capability has to be hidden. That table is the policy that decides which config-space writes vfio lets through to real hardware, so indexing off the end means a tenant's config accesses are policed by whatever kernel bytes follow it - the emulation boundary that keeps a passthrough device inside its assignment.","attack_vector":"A tenant holding the vfio-pci device fd, doing ordinary reads/writes on the device's config-space region. No race, no special ioctl - the only precondition is a passthrough device whose first PCIe extended capability ID is above PCI_EXT_CAP_ID_MAX (an unknown or deliberately hidden capability), which vfio then has to mask by zeroing the ID in place. No host root required.","remediation":"The record lists no fixed release; boot a kernel carrying the stable fix commits below. Interim controls: audit the extended capability list of the device models you pass through, and drop /dev/vfio device nodes from containers that do not need them.","references":["https://git.kernel.org/stable/c/4464e5aa3aa4574063640f1082f7d7e323af8eb4","https://git.kernel.org/stable/c/7d121f66b67921fb3b95e0ea9856bfba53733e91","https://nvd.nist.gov/vuln/detail/CVE-2024-53214"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-56551","cve":"CVE-2024-56551","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): A use-after-free in the amdgpu kernel driver core. Freed kernel","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the amdgpu kernel driver core. Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amdgpu: fix usage slab after free","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-56551","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-12-27"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-362","CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-56552","cve":"CVE-2024-56552","aliases":[],"title":"Linux kernel (drivers/gpu/drm/xe): A tenant that suspends an exec queue and then closes it while the GuC","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/xe)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A tenant that suspends an exec queue and then closes it while the GuC scheduling-disable is still in flight desynchronises the driver's view of GuC state - the late response deregisters a queue that was never marked for destruction. That corrupts submission state for the whole GT (affecting co-tenants sharing the engine) and leaves the queue object referenced after teardown.","attack_vector":"Tenant container holding /dev/dri/renderD* on Intel xe: request suspend on an exec queue, then kill/close it before the GuC round-trip completes. Timing is easy to win - the trace in the fix shows the whole window is under a millisecond and it reproduces in ordinary test runs.","remediation":"Update to a kernel carrying the fix (stable commits below; no fixed_in published). No practical interim control - exec queue suspend and close are core submission ioctls.","references":["https://git.kernel.org/stable/c/5ddcb50b700221fa7d7be2adcb3d7d7afe8633dd","https://git.kernel.org/stable/c/87651f31ae4e6e6e7e6c7270b9b469405e747407","https://nvd.nist.gov/vuln/detail/CVE-2024-56552"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-56608","cve":"CVE-2024-56608","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Fix out-of-bounds access in 'dcn21_link_encoder_create'","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-56608","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-12-27"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-415","CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-56624","cve":"CVE-2024-56624","aliases":[],"title":"Linux kernel (drivers/iommu/iommufd): An error path releases the iommufd fault object and the iommufd context twice","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/iommufd)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"An error path releases the iommufd fault object and the iommufd context twice, driving their refcounts through zero. The object that owns a tenant's entire IOMMU address space gets freed while still referenced, giving a use-after-free on host kernel memory reachable from a tenant ioctl.","attack_vector":"A process holding /dev/iommu - the VMM (qemu) for a passthrough VM, or any container that was given the iommufd node. The tenant calls the fault-queue allocation ioctl and forces the tail of the handler to fail (e.g. by supplying a bad output pointer), which runs the double-release path. Conditional on iommufd being in use rather than the legacy vfio type1 container; iommufd is the default path for nested translation and I/O page-fault handling.","remediation":"The record lists no fixed release; boot a kernel carrying the stable fix commits below. Interim controls: do not expose /dev/iommu to tenant containers, and keep passthrough on the legacy vfio type1 container where the workload allows it.","references":["https://git.kernel.org/stable/c/2b3f30c8edbf9a122ce01f13f0f41fbca5f1d41d","https://git.kernel.org/stable/c/af7f4780514f850322b2959032ecaa96e4b26472","https://nvd.nist.gov/vuln/detail/CVE-2024-56624"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-56695","cve":"CVE-2024-56695","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): An out-of-bounds access in the amdkfd (KFD compute","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdkfd (KFD compute driver, /dev/kfd) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdkfd: Use dynamic allocation for CU occupancy array in 'kfd_get_cu_occupancy()'","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-56695","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-12-28"},{"id":"CVE-2024-56775","cve":"CVE-2024-56775","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A double free in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A double free in the amdgpu display core (DC/DM). The same allocation is released twice, corrupting the slab allocator's freelist. This is a classic heap-corruption primitive: with slab grooming it becomes arbitrary kernel memory write and therefore host compromise from an unprivileged GPU workload. The cheap outcome is a node panic. Upstream fix: drm/amd/display: Fix handling of plane refcount","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-56775","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-01-08"},{"id":"CVE-2024-56784","cve":"CVE-2024-56784","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Adding array index check to prevent memory corruption","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-56784","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-01-08"},{"id":"CVE-2024-57801","cve":"CVE-2024-57801","aliases":[],"title":"Linux kernel mlx5_core eswitch vport representors / IPsec FS: During driver unload the vport representor private struct","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel mlx5_core eswitch vport representors / IPsec FS","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"During driver unload the vport representor private struct is freed before unregister_netdev runs, so the kernel walks freed representor state across every VF on the box. Vport representors are the per-tenant hooks in switchdev mode; a use-after-free walking all of them on a shared host is a host-kernel corruption reachable through ordinary driver lifecycle events.","attack_vector":"Local - triggered on mlx5 driver unload/reload on a switchdev SR-IOV host. Reachable by anyone who can induce a driver reload (operator action, firmware reset flow, or a fault path an attacker provokes).","remediation":"Upgrade the host kernel to 6.13 or a stable backport (6.6.70, 6.12.9). Rolling reboot. Until then, avoid mlx5 driver unload/reload on live switchdev hosts - use full node reboots instead of in-place driver restarts.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-57801","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2024/CVE-2024-57801.json"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-01-15"},{"cwe":["CWE-190","CWE-787"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-57890","cve":"CVE-2024-57890","aliases":[],"title":"Linux kernel RDMA core (ib_uverbs post_send / post_recv command parsing): The uverbs write path multiplied two fully","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel RDMA core (ib_uverbs post_send / post_recv command parsing)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"The uverbs write path multiplied two fully user-controlled u32 fields (wqe_size x wr_count) without overflow checking, and fed the wrapped result to the request-buffer walker. A tenant picks the two values so the product wraps, the length check passes, and the kernel then walks past the end of the copied command buffer. Same shape applies to sge_count x sizeof(ib_uverbs_sge). This is the front door of the verbs API - it is what every RDMA application calls to post work - so no unusual capability or device is required beyond the uverbs node the tenant already has.","attack_vector":"Local write() on /dev/infiniband/uverbs*, i.e. any container doing RDMA. Unprivileged.","remediation":"Kernel update that reorders the bound check in uverbs_request_next_ptr() so the user-controlled length is compared alone, and converts the callers to size_mul(). No runtime toggle exists - the only alternative is denying uverbs to tenants, which disables RDMA.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=346db03e9926ab7117ed9bf19665699c037c773c","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2024/CVE-2024-57890.json"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-57918","cve":"CVE-2024-57918","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A race condition or locking defect in the amdgpu display","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A race condition or locking defect in the amdgpu display core (DC/DM). Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amd/display: fix page fault due to max surface definition mismatch","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-57918","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-01-19"},{"id":"CVE-2024-57921","cve":"CVE-2024-57921","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amdgpu): Missing or insufficient validation of user-supplied parameters","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amdgpu)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdgpu: Add a lock when accessing the buddy trim function","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-57921","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-01-19"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-125","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-57982","cve":"CVE-2024-57982","aliases":[],"title":"Linux kernel (net/xfrm): SA lookup can observe the new hash mask before the new bucket array is published, so it","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"SA lookup can observe the new hash mask before the new bucket array is published, so it indexes past the end of the old table - an out-of-bounds read in the code path that decides which security association handles a packet. Besides the read, a lookup that lands in garbage can miss or mismatch the SA for in-flight traffic.","attack_vector":"Driven by packet processing on any IPsec-enabled node while the state table is resized. The resize happens as SAs are added, i.e. as tenants' IPsec tunnels come and go, so the race window opens during normal churn on a per-tenant encrypted overlay. Anyone who can cause SA creation (the IKE daemon, or a container with CAP_NET_ADMIN in its netns) can force the rehash on demand.","remediation":"Boot a kernel carrying the linked stable commits. Interim: pre-size the xfrm state hash where possible and avoid rapid SA churn; drop CAP_NET_ADMIN from tenant containers so they cannot force rehashing.","references":["https://git.kernel.org/stable/c/b86dc510308d7a8955f3f47a4fea4bef887653e4","https://git.kernel.org/stable/c/a16871c7832ea6435abb6e0b58289ae7dcb7e4fc","https://nvd.nist.gov/vuln/detail/CVE-2024-57982"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-191","CWE-787"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-58019","cve":"CVE-2024-58019","aliases":[],"title":"Linux kernel (drivers/gpu/drm/nouveau/nvkm/subdev/gsp): The GSP message-queue read pointer is advanced by the wrong","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel (drivers/gpu/drm/nouveau/nvkm/subdev/gsp)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"The GSP message-queue read pointer is advanced by the wrong amount, so the driver parses message body bytes as a header, underflows the copy-length computation to a ~0xffffffff value and then dereferences a NULL message pointer. Oversized copy plus kernel panic - the node goes down for every tenant sharing it.","attack_vector":"Same GSP RPC path as CVE-2024-58018: two-page GSP event messages generated during normal tenant-driven GPU activity on a nouveau/GSP-firmware NVIDIA GPU. A tenant holding /dev/dri/renderD* supplies the workload that produces those messages; no capabilities required. Not applicable when the NVIDIA vendor module is in use.","remediation":"Update to a kernel carrying the fix (stable commits below; no fixed_in published). Interim: run the vendor kernel module rather than nouveau on shared GSP-class GPUs.","references":["https://git.kernel.org/stable/c/5185e63b45ea39339ed83f269e2ddfafb07e70d9","https://git.kernel.org/stable/c/67c9cf82f50236d9c000333b26b4f95eb2c3e1b2","https://nvd.nist.gov/vuln/detail/CVE-2024-58019"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-5998","cve":"CVE-2024-5998","aliases":[],"title":"LangChain (`FAISS.deserialize_from_bytes`): Pickle deserialization of an untrusted vector index","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LangChain (`FAISS.deserialize_from_bytes`)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Pickle deserialization of an untrusted vector index","attack_vector":"Customer-supplied FAISS index file from shared storage","remediation":"Upgrade; a vector index is a pickle and must be treated as a model artifact","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-5998"],"status":"curated","published":"2024-09-17"},{"id":"CVE-2025-0678","cve":"CVE-2025-0678","aliases":[],"title":"GRUB2 (squashfs): Integer overflow in the squash4 filesystem module leading to out-of-bounds write and possible Secure","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (squashfs)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Integer overflow in the squash4 filesystem module leading to out-of-bounds write and possible Secure Boot bypass","attack_vector":"Local, crafted filesystem image","remediation":"GRUB2 update plus dbx revocation; squashfs is used by live/netboot images, so this lands directly on the bare-metal provisioning path","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0678"],"status":"curated","published":"2025-03-03"},{"id":"CVE-2025-10155","cve":"CVE-2025-10155","aliases":[],"title":"picklescan: Improper input validation lets a crafted pickle evade scanning","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"picklescan","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Improper input validation lets a crafted pickle evade scanning","attack_vector":"Customer-supplied model file","remediation":"Upgrade past 0.0.30; layer scanning with format restriction","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-10155"],"status":"curated","published":"2025-09-17"},{"id":"CVE-2025-1753","cve":"CVE-2025-1753","aliases":[],"title":"LlamaIndex CLI: OS command injection via the `--files` argument","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LlamaIndex CLI","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"OS command injection via the `--files` argument","attack_vector":"Attacker-influenced filename in an automated pipeline","remediation":"Upgrade past 0.12.20","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-1753"],"status":"curated","published":"2025-05-28"},{"id":"CVE-2025-21333","cve":"CVE-2025-21333","aliases":[],"title":"Microsoft Hyper-V: Heap-based buffer overflow in the NT Kernel Integration VSP","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Microsoft Hyper-V","year":"2025","cvss_score":7.8,"severity":"high","kev":true,"impact":"Heap-based buffer overflow in the NT Kernel Integration VSP - elevation of privilege, exploited in the wild [KEV]","attack_vector":"Local user on the Hyper-V host (or a container using the VSP path)","remediation":"Windows update + host reboot with VM live-migration","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21333"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-01-14"},{"id":"CVE-2025-21334","cve":"CVE-2025-21334","aliases":[],"title":"Microsoft Hyper-V: Use-after-free in the NT Kernel Integration VSP - elevation of privilege, exploited in the wild","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Microsoft Hyper-V","year":"2025","cvss_score":7.8,"severity":"high","kev":true,"impact":"Use-after-free in the NT Kernel Integration VSP - elevation of privilege, exploited in the wild [KEV]","attack_vector":"Local user on the Hyper-V host","remediation":"Windows update + host reboot with VM live-migration","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21334"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-01-14"},{"id":"CVE-2025-21335","cve":"CVE-2025-21335","aliases":[],"title":"Microsoft Hyper-V: Use-after-free in the NT Kernel Integration VSP - elevation of privilege, exploited in the wild","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Microsoft Hyper-V","year":"2025","cvss_score":7.8,"severity":"high","kev":true,"impact":"Use-after-free in the NT Kernel Integration VSP - elevation of privilege, exploited in the wild [KEV]","attack_vector":"Local user on the Hyper-V host","remediation":"Windows update + host reboot with live-migration","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21335"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-01-14"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787","CWE-125"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-21687","cve":"CVE-2025-21687","aliases":[],"title":"Linux kernel (drivers/vfio/platform): Vfio-platform never bounds-checked the count and offset a caller passes to","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/vfio/platform)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Vfio-platform never bounds-checked the count and offset a caller passes to read()/write() on the device fd - only offset was masked to 40 bits. The holder of the device node gets an arbitrary out-of-bounds read and write past the end of the mapped device region, straight from the passthrough boundary that is supposed to contain them.","attack_vector":"Any tenant process or VMM holding /dev/vfio/* for a vfio-platform device issues a plain read() or write() on the device fd with an oversized count or a crafted offset. No ioctl gymnastics, no host root. Conditional on the vfio-platform (or vfio-amba) driver being in use - this is the non-PCI passthrough path, common on Arm SoC and embedded-style nodes rather than on vfio-pci GPU passthrough - so check whether your Arm hosts load it before deciding this is out of scope.","remediation":"No fixed release is listed in this record; apply the linked stable commits or move to a current stable/LTS kernel. Interim: blacklist vfio-platform / vfio-amba on nodes that only need PCI passthrough, and drop vfio-platform device nodes out of tenant containers.","references":["https://git.kernel.org/stable/c/f21636f24b6786c8b13f1af4319fa75ffcf17f38","https://git.kernel.org/stable/c/9377cdc118cf327248f1a9dde7b87de067681dc9","https://nvd.nist.gov/vuln/detail/CVE-2025-21687"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-21714","cve":"CVE-2025-21714","aliases":["RDMA/mlx5 fix implicit ODP use after free"],"title":"Linux kernel mlx5_ib on-demand paging (ODP): Implicit on-demand-paging memory region destroy work can be queued twice","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_ib on-demand paging (ODP)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Implicit on-demand-paging memory region destroy work can be queued twice, so the second run touches an already-freed MR - a refcount underflow and use-after-free. ODP is what lets RDMA register memory lazily, which is heavily used by large-model training frameworks, so this is reachable from ordinary tenant RDMA registration patterns on a GPU node.","attack_vector":"Local, low-privileged - a tenant process registering and tearing down ODP memory regions through libibverbs.","remediation":"Upgrade the host kernel to 6.14 or a stable backport (6.12.13, 6.13.2). Rolling reboot of nodes where ODP is enabled.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21714","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2025/CVE-2025-21714.json"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-02-27"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-21751","cve":"CVE-2025-21751","aliases":[],"title":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core/steering/hws): A matcher that fails to disconnect is reinserted","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core/steering/hws)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A matcher that fails to disconnect is reinserted into the shared hardware-steering matcher list and then freed by the caller, leaving a dangling entry in state that serves every tenant's offloaded flow rules on that NIC. The result is a use-after-free and node crash, and before the crash the steering graph can be left mis-wired so packets are evaluated against the wrong rule set.","attack_vector":"The HWS matcher list is shared state sitting behind tenant-visible flow offload - the tc/OVS rules programmed for VF representors on behalf of tenant VMs. Triggering the bad path requires a firmware command failure during matcher disconnect, which sustained rule churn or steering-resource exhaustion from a tenant can provoke. Requires mlx5 hardware steering with switchdev/eswitch offload enabled. Not reachable directly from the fabric.","remediation":"Boot a kernel carrying the HWS disconnect error-flow fix (stable commits below; no fixed-version list published for this ID). Interim controls: cap the number of offloaded flow rules per tenant representor, and prefer software steering (DMFS/SMFS) over HWS on nodes where tenants drive rule installation.","references":["https://git.kernel.org/stable/c/5682aad0276ff9b9b0eff3188eb6a1f504d6b436","https://git.kernel.org/stable/c/1ce840c7a659aa53a31ef49f0271b4fd0dc10296","https://nvd.nist.gov/vuln/detail/CVE-2025-21751"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-21756","cve":"CVE-2025-21756","aliases":[],"title":"Linux kernel (vsock): vsock binding not kept until socket destruction - use-after-free, local root with a public exploit","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (vsock)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"vsock binding not kept until socket destruction - use-after-free, local root with a public exploit","attack_vector":"Any tenant process in a container; tenant VM guest","remediation":"Livepatchable; otherwise drain + reboot. Blacklist vsock where unused","references":["https://access.redhat.com/security/cve/CVE-2025-21756"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-02-27"},{"id":"CVE-2025-21780","cve":"CVE-2025-21780","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu): An out-of-bounds access in the amdgpu power","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu power management (SMU/powerplay) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: avoid buffer overflow attach in smu_sys_set_pp_table()","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21780","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-02-27"},{"id":"CVE-2025-21842","cve":"CVE-2025-21842","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (amdkfd): A correctness defect in the amdkfd (KFD compute driver","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (amdkfd)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdkfd (KFD compute driver, /dev/kfd) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: amdkfd: properly free gang_ctx_bo when failed to init user queue","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21842","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-03-07"},{"id":"CVE-2025-21882","cve":"CVE-2025-21882","aliases":["net/mlx5 vport QoS cleanup on error"],"title":"Linux kernel mlx5_core eswitch vport QoS scheduling: When enabling per-vport QoS fails, the scheduling node is leaked","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel mlx5_core eswitch vport QoS scheduling","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"When enabling per-vport QoS fails, the scheduling node is leaked and the pointer left dangling. Per-VF QoS is the mechanism that stops one tenant's VF from starving the others' bandwidth, so failures here both leak host kernel memory and undermine the rate-limiting you sold as an isolation guarantee.","attack_vector":"Local, low-privileged - reached through the VF QoS configuration path on an SR-IOV host. An attacker who can make QoS enablement fail (resource exhaustion, invalid rates) reaches the bad path.","remediation":"Upgrade the host kernel to 6.14 or the 6.13.6 stable backport. Rolling reboot of SR-IOV hosts. No firmware flash.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21882","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2025/CVE-2025-21882.json"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-03-27"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-362","CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-21939","cve":"CVE-2025-21939","aliases":[],"title":"Linux kernel (drivers/gpu/drm/xe): Xe built its scatter-gather table from HMM page pointers without holding the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/xe)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Xe built its scatter-gather table from HMM page pointers without holding the notifier lock or validating the notifier sequence, so it dereferenced and dirtied struct pages the driver holds no reference to. A tenant that races a munmap or page migration against a userptr GPU mapping makes the driver touch pages that have already moved on to someone else - stale-page reference and corruption of memory the GPU was never entitled to.","attack_vector":"Unprivileged process in a container with /dev/dri/renderD* on an Intel Xe node: register a userptr GPU mapping, then unmap or migrate the backing memory concurrently with the driver's page-fault population path. No display node, no root, and userptr is exposed to any render-node client.","remediation":"Boot a kernel with the fix commits below, which builds the sg-table under a validated notifier seqno. Interim: restrict /dev/dri/renderD* to trusted workloads on xe hosts.","references":["https://git.kernel.org/stable/c/2a24c98f0e4cc994334598d4f3a851972064809d","https://git.kernel.org/stable/c/f9326f529da7298a95643c3267f1c0fdb0db55eb","https://nvd.nist.gov/vuln/detail/CVE-2025-21939"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-21968","cve":"CVE-2025-21968","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A use-after-free in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the amdgpu display core (DC/DM). Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amd/display: Fix slab-use-after-free on hdcp_work","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21968","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-04-01"},{"id":"CVE-2025-21985","cve":"CVE-2025-21985","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Fix out-of-bound accesses","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21985","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-04-01"},{"id":"CVE-2025-21991","cve":"CVE-2025-21991","aliases":[],"title":"Linux x86/microcode/AMD - out-of-bounds on CPU-less NUMA nodes: The AMD microcode loader iterated every NUMA node","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux x86/microcode/AMD - out-of-bounds on CPU-less NUMA nodes","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"The AMD microcode loader iterated every NUMA node and unconditionally touched per-CPU data for each node's first CPU - which does not exist on CPU-less NUMA nodes. The result is an out-of-bounds kernel access during microcode load. CPU-less NUMA nodes are not exotic on AI hardware: they are what you get with CXL memory expanders and certain GPU/HBM topologies, so this fires on exactly the machines an AI operator runs.","attack_vector":"Local; triggered during microcode loading on systems with CPU-less NUMA nodes. Not attacker-controlled so much as a crash you hit on affected topologies.","remediation":"Fixed in the Linux kernel microcode loader. Take the distro kernel update and reboot. No firmware or BIOS step. If you run CXL-attached memory or other configurations that produce memory-only NUMA nodes, treat this as a stability fix worth taking promptly rather than a security backlog item.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21991"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"published":"2025-04-02"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-668","CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-22089","cve":"CVE-2025-22089","aliases":[],"title":"Linux kernel RDMA core (hw_counters sysfs exposure across network namespaces): RDMA hardware counter sysfs attributes","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel RDMA core (hw_counters sysfs exposure across network namespaces)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"RDMA hardware counter sysfs attributes were exposed inside non-initial network namespaces, which is exactly where every containerized tenant lives. The immediate effect is a NULL-pointer oops on read - a tenant in its own netns reads a counter file and takes the node down, killing every co-tenant job on the machine. The structural problem is worse than the crash: device-wide hardware counters are a cross-tenant side channel, since they aggregate traffic from every queue pair on the adapter regardless of which namespace posted it, and this bug is evidence the namespace filter on that surface was not being enforced.","attack_vector":"Local read of the RDMA device's hw_counters sysfs files from inside a non-init network namespace. Any container on an RDMA node.","remediation":"Kernel update restoring the init-netns restriction on hw_counters. Until then, mask the RDMA sysfs counter paths out of tenant containers, and treat any fleet-wide RDMA counter dashboard as a surface tenants can read from, not just one you write to.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=0cf80f924aecb5b2bebd4f4ad11b2efc676a0b78","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2025/CVE-2025-22089.json"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-22472","cve":"CVE-2025-22472","aliases":[],"title":"Dell SmartFabric OS10 (command injection with elevated privileges): Local low-privilege attacker executes commands","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell SmartFabric OS10 (command injection with elevated privileges)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local low-privilege attacker executes commands at elevated privilege on the switch.","attack_vector":"Local low-privilege access to OS10.","remediation":"Upgrade OS10 per DSA-2025-070/069/079.","references":["https://www.dell.com/support/kbdoc/en-us/000289970/dsa-2025-070-security-update-for-dell-networking-os10-vulnerabilities"],"status":"curated"},{"id":"CVE-2025-22473","cve":"CVE-2025-22473","aliases":[],"title":"Dell SmartFabric OS10 (command injection, local): A low-privileged local attacker achieves code execution on the switch","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell SmartFabric OS10 (command injection, local)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A low-privileged local attacker achieves code execution on the switch operating system.","attack_vector":"Local low-privilege access to OS10 10.5.4.x-10.6.0.x.","remediation":"Upgrade OS10 per DSA-2025-070/069/079. Switch reboot.","references":["https://www.dell.com/support/kbdoc/en-us/000289970/dsa-2025-070-security-update-for-dell-networking-os10-vulnerabilities"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-22831","cve":"CVE-2025-22831","aliases":[],"title":"AMI AptioV BIOS (out-of-bounds write): Second local out-of-bounds write in the same AptioV advisory","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI AptioV BIOS (out-of-bounds write)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Second local out-of-bounds write in the same AptioV advisory.","attack_vector":"Local low-privilege access.","remediation":"AMI ships the fix to OEMs, not to you - obtain the updated BIOS from your board/server vendor (Supermicro, Gigabyte, ASRock Rack, Quanta, Tyan etc.) and flash it. Expect a lag of weeks to months between the AMI advisory and an OEM image for your exact SKU, and expect some SKUs never to get one. Cold reboot per node.","references":["https://go.ami.com/hubfs/Security%20Advisories/2025/AMI-SA-2025008.pdf"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-22832","cve":"CVE-2025-22832","aliases":[],"title":"AMI AptioV BIOS (out-of-bounds write): Local out-of-bounds write in firmware causing data corruption and loss","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI AptioV BIOS (out-of-bounds write)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local out-of-bounds write in firmware causing data corruption and loss of availability.","attack_vector":"Local low-privilege access.","remediation":"AMI ships the fix to OEMs, not to you - obtain the updated BIOS from your board/server vendor (Supermicro, Gigabyte, ASRock Rack, Quanta, Tyan etc.) and flash it. Expect a lag of weeks to months between the AMI advisory and an OEM image for your exact SKU, and expect some SKUs never to get one. Cold reboot per node.","references":["https://go.ami.com/hubfs/Security%20Advisories/2025/AMI-SA-2025008.pdf"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-22836","cve":"CVE-2025-22836","aliases":[],"title":"Intel ice driver (Ethernet 800 Series, Linux kernel mode): Kernel-mode flaw in the 800-series Ethernet Linux driver (an","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel ice driver (Ethernet 800 Series, Linux kernel mode)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Kernel-mode flaw in the 800-series Ethernet Linux driver (an integer overflow) reachable by an authenticated user for privilege escalation. Part of the same 2025 batch - patch them as one unit rather than individually.","attack_vector":"Authenticated local user; tenants on nodes exposing VFs or RDMA devices.","remediation":"Fixed in the Intel out-of-tree ice driver (or the equivalent in-kernel version). Updating the driver requires unloading and reloading the module, which drops every link on that NIC - on a node whose RDMA fabric carries collective traffic, that is a job-killing event, so drain first. If you take it via a distro kernel update instead, it is a reboot. No firmware flash for the driver-side fixes. Target ice 1.17.2 or later.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-22836","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01296.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-08-12"},{"id":"CVE-2025-22893","cve":"CVE-2025-22893","aliases":[],"title":"Intel ice driver (Ethernet 800 Series, Linux kernel mode): Kernel-mode flaw in the 800-series Ethernet Linux driver","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel ice driver (Ethernet 800 Series, Linux kernel mode)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Kernel-mode flaw in the 800-series Ethernet Linux driver (insufficient control-flow management) reachable by an authenticated user for privilege escalation. Part of the same 2025 batch - patch them as one unit rather than individually.","attack_vector":"Authenticated local user; tenants on nodes exposing VFs or RDMA devices.","remediation":"Fixed in the Intel out-of-tree ice driver (or the equivalent in-kernel version). Updating the driver requires unloading and reloading the module, which drops every link on that NIC - on a node whose RDMA fabric carries collective traffic, that is a job-killing event, so drain first. If you take it via a distro kernel update instead, it is a reboot. No firmware flash for the driver-side fixes. Target ice 1.17.2 or later.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-22893","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01296.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-08-12"},{"id":"CVE-2025-23244","cve":"CVE-2025-23244","aliases":[],"title":"GPU Display Driver: Local privesc (insufficient access control)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc (insufficient access control)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23244","https://github.com/NVIDIA/product-security/tree/main/2025/5630"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-863"],"fleet":{"pain_class":"node-reboot"},"published":"2025-05-01"},{"id":"CVE-2025-23264","cve":"CVE-2025-23264","aliases":[],"title":"Megatron-LM: Arbitrary code execution in the training job","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Megatron-LM","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Arbitrary code execution in the training job","attack_vector":"Malicious model/checkpoint or dataset","remediation":"Bump Megatron-LM in training images; rebuild; treat checkpoints as untrusted input","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23264","https://github.com/NVIDIA/product-security/tree/main/2025/5663"],"status":"curated","fleet":{"ubiquity":"Common - Megatron is the reference large-model training stack for customers doing pretraining on rented clusters","remediation_pain":"`hot-patch` (upgrade to v0.12.1, rebuild training images)","pain_class":"hot-patch","why_fleet_wide":"A malicious file supplied to the Python component triggers code injection; on a shared training cluster the attacker executes inside a job that already holds cluster-wide storage and fabric credentials"},"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"],"published":"2025-06-24"},{"id":"CVE-2025-23265","cve":"CVE-2025-23265","aliases":[],"title":"Megatron-LM: Arbitrary code execution","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Megatron-LM","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Arbitrary code execution","attack_vector":"Malicious model/checkpoint","remediation":"Bump Megatron-LM; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23265","https://github.com/NVIDIA/product-security/tree/main/2025/5663"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"],"published":"2025-06-24"},{"id":"CVE-2025-23276","cve":"CVE-2025-23276","aliases":[],"title":"GPU Display Driver: Local privesc via file permissions","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc via file permissions","attack_vector":"Local user on the node","remediation":"Driver upgrade; rolling reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23276","https://github.com/NVIDIA/product-security/tree/main/2025/5670"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-552"],"fleet":{"pain_class":"node-reboot"},"published":"2025-08-02"},{"id":"CVE-2025-23283","cve":"CVE-2025-23283","aliases":[],"title":"vGPU Manager: Guest-to-host escape (stack buffer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU Manager","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Guest-to-host escape (stack buffer overflow)","attack_vector":"Tenant VM guest","remediation":"vGPU Manager upgrade; evacuate all guest VMs, reboot host","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23283","https://github.com/NVIDIA/product-security/tree/main/2025/5670"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-121"],"fleet":{"pain_class":"node-reboot"},"published":"2025-08-02"},{"id":"CVE-2025-23284","cve":"CVE-2025-23284","aliases":[],"title":"vGPU Manager: Guest-to-host escape (stack buffer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU Manager","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Guest-to-host escape (stack buffer overflow)","attack_vector":"Tenant VM guest","remediation":"vGPU Manager upgrade; evacuate guest VMs","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23284","https://github.com/NVIDIA/product-security/tree/main/2025/5670"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-121"],"published":"2025-08-02"},{"id":"CVE-2025-23294","cve":"CVE-2025-23294","aliases":[],"title":"NVIDIA WebDataset: Arbitrary code execution with elevated permissions from the data-loading library","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA WebDataset","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Arbitrary code execution with elevated permissions from the data-loading library. WebDataset exists to stream shards from object storage, so the trust boundary is whoever can write to your bucket. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5658 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23294","https://github.com/NVIDIA/product-security/tree/main/2025/5658"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-78"],"published":"2025-08-13"},{"id":"CVE-2025-23295","cve":"CVE-2025-23295","aliases":[],"title":"NVIDIA Apex: A Python component injects code from a malicious file","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Apex","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A Python component injects code from a malicious file. Apex is a near-universal dependency in mixed-precision training images. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5680 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23295","https://github.com/NVIDIA/product-security/tree/main/2025/5680"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"],"published":"2025-08-13"},{"id":"CVE-2025-23296","cve":"CVE-2025-23296","aliases":[],"title":"NVIDIA Isaac-GR00T N1: A Python component injects code from crafted input into the robotics-model pipeline","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Isaac-GR00T N1","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A Python component injects code from crafted input into the robotics-model pipeline, which in practice runs on GPU cluster nodes. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5681 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23296","https://github.com/NVIDIA/product-security/tree/main/2025/5681"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"],"published":"2025-08-13"},{"id":"CVE-2025-23298","cve":"CVE-2025-23298","aliases":[],"title":"NVIDIA Merlin Transformers4Rec: A Python dependency permits code injection into the recommender training job","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Merlin Transformers4Rec","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A Python dependency permits code injection into the recommender training job. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5683 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23298","https://github.com/NVIDIA/product-security/tree/main/2025/5683"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"],"published":"2025-08-13"},{"id":"CVE-2025-23303","cve":"CVE-2025-23303","aliases":[],"title":"NVIDIA NeMo Framework: Deserialization of untrusted data reaches remote code execution when a crafted artifact","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NeMo Framework","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Deserialization of untrusted data reaches remote code execution when a crafted artifact is loaded. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5686 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23303","https://github.com/NVIDIA/product-security/tree/main/2025/5686"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2025-08-13"},{"id":"CVE-2025-23304","cve":"CVE-2025-23304","aliases":[],"title":"NVIDIA NeMo Framework: Loading a .nemo file with crafted metadata injects code at model-load time","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NeMo Framework","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Loading a .nemo file with crafted metadata injects code at model-load time - the model file itself is the payload. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5686 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23304","https://github.com/NVIDIA/product-security/tree/main/2025/5686"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-22"],"published":"2025-08-13"},{"id":"CVE-2025-23305","cve":"CVE-2025-23305","aliases":[],"title":"NVIDIA Megatron-LM: A code-injection flaw in the tools component executes attacker-controlled code inside the training","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Megatron-LM","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A code-injection flaw in the tools component executes attacker-controlled code inside the training job. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5685 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23305","https://github.com/NVIDIA/product-security/tree/main/2025/5685"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"],"published":"2025-08-13"},{"id":"CVE-2025-23306","cve":"CVE-2025-23306","aliases":[],"title":"NVIDIA Megatron-LM: megatron/training/arguments.py injects code from malicious input","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Megatron-LM","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"megatron/training/arguments.py injects code from malicious input - the argument-parsing path, which every training launch goes through. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5685 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23306","https://github.com/NVIDIA/product-security/tree/main/2025/5685"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"],"published":"2025-08-13"},{"id":"CVE-2025-23307","cve":"CVE-2025-23307","aliases":[],"title":"NVIDIA NeMo Curator: A malicious file processed by the data-curation pipeline injects code","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NeMo Curator","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A malicious file processed by the data-curation pipeline injects code. Curator exists to chew through large untrusted corpora, so the untrusted-input assumption is the product. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5690 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23307","https://github.com/NVIDIA/product-security/tree/main/2025/5690"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"],"published":"2025-08-26"},{"id":"CVE-2025-23312","cve":"CVE-2025-23312","aliases":[],"title":"NVIDIA NeMo Framework: Crafted data in the retrieval-services component injects code into the running job","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NeMo Framework","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Crafted data in the retrieval-services component injects code into the running job. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5689 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23312","https://github.com/NVIDIA/product-security/tree/main/2025/5689"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"],"published":"2025-08-26"},{"id":"CVE-2025-23313","cve":"CVE-2025-23313","aliases":[],"title":"NVIDIA NeMo Framework: Crafted data in the NLP component injects code into the running job","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NeMo Framework","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Crafted data in the NLP component injects code into the running job. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5689 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23313","https://github.com/NVIDIA/product-security/tree/main/2025/5689"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"],"published":"2025-08-26"},{"id":"CVE-2025-23314","cve":"CVE-2025-23314","aliases":[],"title":"NVIDIA NeMo Framework: A second NLP-component code-injection path with the same reach","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NeMo Framework","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A second NLP-component code-injection path with the same reach. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5689 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23314","https://github.com/NVIDIA/product-security/tree/main/2025/5689"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"],"published":"2025-08-26"},{"id":"CVE-2025-23315","cve":"CVE-2025-23315","aliases":[],"title":"NVIDIA NeMo Framework: Crafted data in the export-and-deploy component injects code","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NeMo Framework","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Crafted data in the export-and-deploy component injects code - notable because export/deploy typically runs with registry push credentials. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5689 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23315","https://github.com/NVIDIA/product-security/tree/main/2025/5689"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"],"published":"2025-08-26"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-276"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-23347","cve":"CVE-2025-23347","aliases":[],"title":"NVIDIA Project G-Assist (Windows display driver, R580/R570): A permissions flaw in the G-Assist component that ships","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Project G-Assist (Windows display driver, R580/R570)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A permissions flaw in the G-Assist component that ships inside the Windows display driver hands a low-privileged local account elevated code execution on the host. NVIDIA lists the full set of consequences: code execution, privilege escalation, data tampering, information disclosure and denial of service. This one matters beyond desktops because the affected branches cover Tesla-branded datacenter cards on Windows as well as GeForce, RTX, Quadro and NVS - so it is a driver-level escalation, not a consumer-app problem.","attack_vector":"A local, low-privileged account on a Windows host running an affected R580 or R570 display driver branch. No user interaction, no network access, and no need to touch the GPU directly - the vulnerable component is reachable from an ordinary user session.","remediation":"Move Windows display drivers to 581.42 or later on the R580 branch, or 573.76 or later on R570. Driver replacement on Windows takes the display stack down and requires a reboot to complete, so schedule it as a maintenance window rather than a live update.","references":["https://github.com/NVIDIA/product-security/tree/main/2025/5703","https://nvd.nist.gov/vuln/detail/CVE-2025-23347"],"status":"curated"},{"id":"CVE-2025-23348","cve":"CVE-2025-23348","aliases":[],"title":"NVIDIA Megatron-LM: The pretrain_gpt script injects code from crafted data","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Megatron-LM","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"The pretrain_gpt script injects code from crafted data. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5698 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23348","https://github.com/NVIDIA/product-security/tree/main/2025/5698"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"],"published":"2025-09-24"},{"id":"CVE-2025-23349","cve":"CVE-2025-23349","aliases":[],"title":"NVIDIA Megatron-LM: tasks/orqa/unsupervised/nq.py injects code from crafted data","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Megatron-LM","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"tasks/orqa/unsupervised/nq.py injects code from crafted data. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5698 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23349","https://github.com/NVIDIA/product-security/tree/main/2025/5698"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"],"published":"2025-09-24"},{"id":"CVE-2025-23352","cve":"CVE-2025-23352","aliases":[],"title":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko): A malicious guest drives the Virtual","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A malicious guest drives the Virtual GPU Manager into using an uninitialised pointer, reaching code execution, privilege escalation and information disclosure on the host. The attacker sits inside a tenant guest VM and the blast radius is the hypervisor host, which is running every other tenant's vGPU on the same physical GPU. This is precisely the boundary a vGPU-based multi-tenant offering is sold on. Uninitialised-pointer use under guest control is a strong escape primitive - prioritise this above the null-deref bugs in the same bulletin.","attack_vector":"A tenant inside their own guest VM, driving the paravirtualised vGPU control interface. Several of these need only an unprivileged process in the guest; the rest need guest root, which a tenant already has on a VM they rented. No host credentials are involved at any point.","remediation":"Patch the vGPU Manager on the hypervisor host per bulletin 5703. Cost: the highest of any class here. The host driver cannot be reloaded while vGPUs are attached, so every tenant VM on that hypervisor must be live-migrated or powered off - a full host drain. NVIDIA also enforces a supported host/guest driver skew, so budget a matching guest-driver campaign in the same window or tenants lose their vGPU on next boot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23352","https://github.com/NVIDIA/product-security/tree/main/2025/5703"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-824"],"fleet":{"pain_class":"node-drain"},"tags":["tenant-isolation"],"published":"2025-10-23"},{"id":"CVE-2025-23353","cve":"CVE-2025-23353","aliases":[],"title":"NVIDIA Megatron-LM: The msdp preprocessing script injects code from crafted data","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Megatron-LM","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"The msdp preprocessing script injects code from crafted data - preprocessing is where untrusted corpora arrive. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5698 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23353","https://github.com/NVIDIA/product-security/tree/main/2025/5698"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"],"published":"2025-09-24"},{"id":"CVE-2025-23354","cve":"CVE-2025-23354","aliases":[],"title":"NVIDIA Megatron-LM: The ensemble_classifier script injects code from crafted data","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Megatron-LM","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"The ensemble_classifier script injects code from crafted data. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5698 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23354","https://github.com/NVIDIA/product-security/tree/main/2025/5698"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"],"published":"2025-09-24"},{"id":"CVE-2025-23357","cve":"CVE-2025-23357","aliases":[],"title":"NVIDIA Megatron-LM: A script in the repository injects code from crafted data","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Megatron-LM","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A script in the repository injects code from crafted data. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5712 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23357","https://github.com/NVIDIA/product-security/tree/main/2025/5712"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"],"published":"2025-11-11"},{"id":"CVE-2025-23361","cve":"CVE-2025-23361","aliases":[],"title":"NVIDIA NeMo Framework: Malicious input causes improper control of code generation, reaching code execution","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NeMo Framework","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Malicious input causes improper control of code generation, reaching code execution. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5718 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23361","https://github.com/NVIDIA/product-security/tree/main/2025/5718"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"],"published":"2025-11-11"},{"id":"CVE-2025-24303","cve":"CVE-2025-24303","aliases":[],"title":"Intel ice driver (Ethernet 800 Series, Linux kernel mode): Kernel-mode flaw in the 800-series Ethernet Linux driver (a","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel ice driver (Ethernet 800 Series, Linux kernel mode)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Kernel-mode flaw in the 800-series Ethernet Linux driver (a missing check for an exceptional condition) reachable by an authenticated user for privilege escalation. Part of the same 2025 batch - patch them as one unit rather than individually.","attack_vector":"Authenticated local user; tenants on nodes exposing VFs or RDMA devices.","remediation":"Fixed in the Intel out-of-tree ice driver (or the equivalent in-kernel version). Updating the driver requires unloading and reloading the module, which drops every link on that NIC - on a node whose RDMA fabric carries collective traffic, that is a job-killing event, so drain first. If you take it via a distro kernel update instead, it is a reboot. No firmware flash for the driver-side fixes. Target ice 1.17.2 or later.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-24303","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01296.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-08-12"},{"id":"CVE-2025-24484","cve":"CVE-2025-24484","aliases":[],"title":"Intel ice driver (Ethernet 800 Series, Linux kernel mode): Kernel-mode flaw in the 800-series Ethernet Linux driver","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel ice driver (Ethernet 800 Series, Linux kernel mode)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Kernel-mode flaw in the 800-series Ethernet Linux driver (improper input validation) reachable by an authenticated user for privilege escalation. Part of the same 2025 batch - patch them as one unit rather than individually.","attack_vector":"Authenticated local user; tenants on nodes exposing VFs or RDMA devices.","remediation":"Fixed in the Intel out-of-tree ice driver (or the equivalent in-kernel version). Updating the driver requires unloading and reloading the module, which drops every link on that NIC - on a node whose RDMA fabric carries collective traffic, that is a job-killing event, so drain first. If you take it via a distro kernel update instead, it is a reboot. No firmware flash for the driver-side fixes. Target ice 1.17.2 or later.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-24484","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01296.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-08-12"},{"id":"CVE-2025-27689","cve":"CVE-2025-27689","aliases":[],"title":"Dell iDRAC Tools (improper access control): A low-privileged local attacker escalates privileges through the iDRAC","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC Tools (improper access control)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A low-privileged local attacker escalates privileges through the iDRAC Tools package on the management workstation or jump host where it is installed.","attack_vector":"Local low-privilege access to a host running iDRAC Tools below 11.3.0.0.","remediation":"Upgrade iDRAC Tools to 11.3.0.0. This lives on admin workstations and provisioning servers, not the GPU nodes - so the fix is a management-plane package update, but those hosts hold fleet-wide BMC credentials.","references":["https://www.dell.com/support/kbdoc/en-us/000323242/dsa-2025-169-security-update-for-dell-idrac-tools-vulnerabilities"],"status":"curated"},{"id":"CVE-2025-32463","cve":"CVE-2025-32463","aliases":[],"title":"sudo: Local privilege escalation via the `--chroot` option","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"sudo","year":"2025","cvss_score":7.8,"severity":"high","kev":true,"impact":"Local privilege escalation via the `--chroot` option - any local user reaches root on default sudoers, no sudo rights needed [KEV]","attack_vector":"Local user, incl. inside a container that ships sudo","remediation":"Package update (sudo >= 1.9.17p1); no reboot. Also rebuild every container base image that includes sudo - the host fix does not cover tenant images","references":["https://access.redhat.com/security/cve/CVE-2025-32463"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2025-06-30"},{"id":"CVE-2025-33044","cve":"CVE-2025-33044","aliases":[],"title":"AMI AptioV BIOS (out-of-bounds memory operation): Local attacker causes firmware memory corruption impacting integrity","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI AptioV BIOS (out-of-bounds memory operation)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local attacker causes firmware memory corruption impacting integrity and availability of the platform.","attack_vector":"Local low-privilege access.","remediation":"AMI ships the fix to OEMs, not to you - obtain the updated BIOS from your board/server vendor (Supermicro, Gigabyte, ASRock Rack, Quanta, Tyan etc.) and flash it. Expect a lag of weeks to months between the AMI advisory and an OEM image for your exact SKU, and expect some SKUs never to get one. Cold reboot per node.","references":["https://go.ami.com/hubfs/Security%20Advisories/2025/AMI-SA-2025008.pdf"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-33178","cve":"CVE-2025-33178","aliases":[],"title":"NVIDIA NeMo Framework: Crafted data in the BERT services component injects code","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NeMo Framework","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Crafted data in the BERT services component injects code. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5718 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33178","https://github.com/NVIDIA/product-security/tree/main/2025/5718"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"],"published":"2025-11-11"},{"id":"CVE-2025-33183","cve":"CVE-2025-33183","aliases":[],"title":"NVIDIA Isaac-GR00T N1.5: A Python component injects code from crafted input","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Isaac-GR00T N1.5","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A Python component injects code from crafted input. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5725 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33183","https://github.com/NVIDIA/product-security/tree/main/2025/5725"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"],"published":"2025-11-18"},{"id":"CVE-2025-33184","cve":"CVE-2025-33184","aliases":[],"title":"NVIDIA Isaac-GR00T N1.5: A second Python code-injection path in the same release","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Isaac-GR00T N1.5","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A second Python code-injection path in the same release. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5725 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33184","https://github.com/NVIDIA/product-security/tree/main/2025/5725"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"],"published":"2025-11-18"},{"id":"CVE-2025-33189","cve":"CVE-2025-33189","aliases":[],"title":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware: An out-of-bounds write in SROOT firmware reaches code","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds write in SROOT firmware reaches code execution and privilege escalation in the root-of-trust context. These sit in the GB10 root-of-trust chain, so a successful exploit undermines the platform's own attestation and secure-boot story rather than just the OS above it.","attack_vector":"Local access to the DGX Spark. Several of the set need no privileges at all; the rest need host root. This is a desk-side developer box, so physical and local access assumptions are much weaker than for a racked DGX.","remediation":"Apply the DGX Spark firmware update from bulletin 5720. Cost: flash plus reboot, low drain cost given the form factor, but not live-patchable and root-of-trust firmware cannot be rolled back once applied.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33189","https://github.com/NVIDIA/product-security/tree/main/2025/5720"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"firmware-flash"},"published":"2025-11-25"},{"id":"CVE-2025-33204","cve":"CVE-2025-33204","aliases":[],"title":"NVIDIA NeMo Framework: Crafted data in the NLP and LLM components injects code","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NeMo Framework","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Crafted data in the NLP and LLM components injects code. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5729 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33204","https://github.com/NVIDIA/product-security/tree/main/2025/5729"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"],"published":"2025-11-25"},{"id":"CVE-2025-33206","cve":"CVE-2025-33206","aliases":[],"title":"Nsight Graphics: Local code exec via command injection","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Nsight Graphics","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local code exec via command injection","attack_vector":"Unprivileged local user on a dev node","remediation":"Upgrade Nsight packages in dev images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33206","https://github.com/NVIDIA/product-security/tree/main/2026/5738"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-78"],"published":"2026-01-14"},{"id":"CVE-2025-33217","cve":"CVE-2025-33217","aliases":[],"title":"GPU Display Driver: Local privesc to host root (use-after-free)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc to host root (use-after-free)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33217","https://github.com/NVIDIA/product-security/tree/main/2026/5747"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"published":"2026-01-28"},{"id":"CVE-2025-33218","cve":"CVE-2025-33218","aliases":[],"title":"GPU Display Driver: Local privesc (integer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc (integer overflow)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33218","https://github.com/NVIDIA/product-security/tree/main/2026/5747"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-190"],"fleet":{"pain_class":"node-reboot"},"published":"2026-01-28"},{"id":"CVE-2025-33219","cve":"CVE-2025-33219","aliases":[],"title":"GPU Display Driver: Local privesc (integer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc (integer overflow)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33219","https://github.com/NVIDIA/product-security/tree/main/2026/5747"],"status":"curated","fleet":{"ubiquity":"Universal - the Linux kernel module is on every NVIDIA compute node","remediation_pain":"`node-drain` then `node-reboot` (kernel module replacement)","pain_class":"node-reboot","why_fleet_wide":"Integer overflow in the Linux kernel module reachable from a tenant process: code execution at elevated privilege, DoS, or access to another tenant's data on the same host"},"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-190"],"published":"2026-01-28"},{"id":"CVE-2025-33220","cve":"CVE-2025-33220","aliases":[],"title":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko): A malicious guest causes the Virtual","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A malicious guest causes the Virtual GPU Manager to access heap memory after it has been freed, reaching host code execution and privilege escalation. The attacker sits inside a tenant guest VM and the blast radius is the hypervisor host, which is running every other tenant's vGPU on the same physical GPU. This is precisely the boundary a vGPU-based multi-tenant offering is sold on. A guest-triggerable host-side use-after-free is the highest-value vGPU bug in the 2026 bulletin set; treat it as a presumed escape until proven otherwise.","attack_vector":"A tenant inside their own guest VM, driving the paravirtualised vGPU control interface. Several of these need only an unprivileged process in the guest; the rest need guest root, which a tenant already has on a VM they rented. No host credentials are involved at any point.","remediation":"Patch the vGPU Manager on the hypervisor host per bulletin 5747. Cost: the highest of any class here. The host driver cannot be reloaded while vGPUs are attached, so every tenant VM on that hypervisor must be live-migrated or powered off - a full host drain. NVIDIA also enforces a supported host/guest driver skew, so budget a matching guest-driver campaign in the same window or tenants lose their vGPU on next boot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33220","https://github.com/NVIDIA/product-security/tree/main/2026/5747"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-drain"},"tags":["tenant-isolation"],"published":"2026-01-28"},{"id":"CVE-2025-33226","cve":"CVE-2025-33226","aliases":[],"title":"NVIDIA NeMo Framework: Crafted data reaches code injection and privilege escalation in the job context","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NeMo Framework","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Crafted data reaches code injection and privilege escalation in the job context. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5736 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33226","https://github.com/NVIDIA/product-security/tree/main/2025/5736"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2025-12-16"},{"id":"CVE-2025-33233","cve":"CVE-2025-33233","aliases":[],"title":"Merlin Transformers4Rec: RCE via unsanitized input","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Merlin Transformers4Rec","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via unsanitized input","attack_vector":"Malicious model/dataset","remediation":"Bump the package; rebuild recsys images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33233","https://github.com/NVIDIA/product-security/tree/main/2026/5761"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"],"published":"2026-01-20"},{"id":"CVE-2025-33234","cve":"CVE-2025-33234","aliases":[],"title":"NVIDIA runx: Arbitrary command exec via shell metacharacter injection","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA runx","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Arbitrary command exec via shell metacharacter injection","attack_vector":"Local user invoking the tool","remediation":"Upgrade runx on nodes","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33234","https://github.com/NVIDIA/product-security/tree/main/2026/5764"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-78"],"published":"2026-01-27"},{"id":"CVE-2025-33235","cve":"CVE-2025-33235","aliases":[],"title":"NVIDIA Resiliency Extension: A race condition in the checkpointing core reaches information disclosure, data tampering","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Resiliency Extension","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A race condition in the checkpointing core reaches information disclosure, data tampering and privilege escalation. Checkpoints are the training run's crown jewels - a tampering primitive here means silently poisoned model state that survives every restart.","attack_vector":"Local, low privileges. An account on a node participating in the checkpoint write path.","remediation":"Update the Resiliency Extension per bulletin 5746 and rebuild training images. Cost: image rebuild and job restart. Verify checkpoint integrity out-of-band if you suspect exposure - a corrupted checkpoint does not announce itself.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33235","https://github.com/NVIDIA/product-security/tree/main/2025/5746"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-362"],"published":"2025-12-16"},{"id":"CVE-2025-33236","cve":"CVE-2025-33236","aliases":[],"title":"NeMo Framework: Arbitrary Python code exec via code injection","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NeMo Framework","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Arbitrary Python code exec via code injection","attack_vector":"Malicious model/config artifact","remediation":"Bump NeMo; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33236","https://github.com/NVIDIA/product-security/tree/main/2026/5762"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"],"published":"2026-02-18"},{"id":"CVE-2025-33239","cve":"CVE-2025-33239","aliases":[],"title":"Megatron-Bridge: RCE via unsafe pickle deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Megatron-Bridge","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via unsafe pickle deserialization","attack_vector":"Malicious checkpoint","remediation":"Bump Megatron-Bridge; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33239","https://github.com/NVIDIA/product-security/tree/main/2026/5781"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"],"published":"2026-02-18"},{"id":"CVE-2025-33240","cve":"CVE-2025-33240","aliases":[],"title":"Megatron-Bridge: RCE via unsafe pickle deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Megatron-Bridge","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via unsafe pickle deserialization","attack_vector":"Malicious checkpoint","remediation":"Bump Megatron-Bridge; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33240","https://github.com/NVIDIA/product-security/tree/main/2026/5781"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"],"published":"2026-02-18"},{"id":"CVE-2025-33241","cve":"CVE-2025-33241","aliases":[],"title":"NeMo Framework: RCE via unsafe deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NeMo Framework","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via unsafe deserialization","attack_vector":"Malicious checkpoint","remediation":"Bump NeMo; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33241","https://github.com/NVIDIA/product-security/tree/main/2026/5762"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2026-02-18"},{"id":"CVE-2025-33243","cve":"CVE-2025-33243","aliases":[],"title":"NeMo Framework: RCE via unsafe deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NeMo Framework","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via unsafe deserialization","attack_vector":"Malicious checkpoint","remediation":"Bump NeMo; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33243","https://github.com/NVIDIA/product-security/tree/main/2026/5762"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2026-02-18"},{"id":"CVE-2025-33246","cve":"CVE-2025-33246","aliases":[],"title":"NeMo Framework: Local command injection","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NeMo Framework","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local command injection","attack_vector":"Malicious config / local user","remediation":"Bump NeMo; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33246","https://github.com/NVIDIA/product-security/tree/main/2026/5762"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-77"],"published":"2026-02-18"},{"id":"CVE-2025-33247","cve":"CVE-2025-33247","aliases":[],"title":"Megatron-LM: Local privesc / RCE via unsafe pickle deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Megatron-LM","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc / RCE via unsafe pickle deserialization","attack_vector":"Malicious checkpoint","remediation":"Bump Megatron-LM; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33247","https://github.com/NVIDIA/product-security/tree/main/2026/5769"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2026-03-24"},{"id":"CVE-2025-33248","cve":"CVE-2025-33248","aliases":[],"title":"Megatron-LM: RCE via insecure deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Megatron-LM","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via insecure deserialization","attack_vector":"Malicious checkpoint","remediation":"Bump Megatron-LM; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33248","https://github.com/NVIDIA/product-security/tree/main/2026/5769"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2026-03-24"},{"id":"CVE-2025-33249","cve":"CVE-2025-33249","aliases":[],"title":"NeMo Framework: Local command injection","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NeMo Framework","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local command injection","attack_vector":"Malicious config","remediation":"Bump NeMo; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33249","https://github.com/NVIDIA/product-security/tree/main/2026/5762"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-77"],"published":"2026-02-18"},{"id":"CVE-2025-33250","cve":"CVE-2025-33250","aliases":[],"title":"NeMo Framework: RCE via unsafe object deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NeMo Framework","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via unsafe object deserialization","attack_vector":"Malicious checkpoint","remediation":"Bump NeMo; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33250","https://github.com/NVIDIA/product-security/tree/main/2026/5762"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"],"published":"2026-02-18"},{"id":"CVE-2025-33251","cve":"CVE-2025-33251","aliases":[],"title":"NeMo Framework: RCE via unsafe deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NeMo Framework","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via unsafe deserialization","attack_vector":"Malicious checkpoint","remediation":"Bump NeMo; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33251","https://github.com/NVIDIA/product-security/tree/main/2026/5762"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"],"published":"2026-02-18"},{"id":"CVE-2025-33252","cve":"CVE-2025-33252","aliases":[],"title":"NeMo Framework: RCE via pickle deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NeMo Framework","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via pickle deserialization","attack_vector":"Malicious checkpoint","remediation":"Bump NeMo; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33252","https://github.com/NVIDIA/product-security/tree/main/2026/5762"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2026-02-18"},{"id":"CVE-2025-33253","cve":"CVE-2025-33253","aliases":[],"title":"NeMo Framework: RCE via insecure deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NeMo Framework","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via insecure deserialization","attack_vector":"Malicious checkpoint","remediation":"Bump NeMo; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33253","https://github.com/NVIDIA/product-security/tree/main/2026/5762"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2026-02-18"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-404","CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-37756","cve":"CVE-2025-37756","aliases":[],"title":"Linux kernel (net/tls): KTLS never supported disconnect, but nothing stopped it. A connect(AF_UNSPEC) on a TLS socket","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/tls)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"KTLS never supported disconnect, but nothing stopped it. A connect(AF_UNSPEC) on a TLS socket leaves the ULP, the strparser anchor and any NIC offload context pointing at torn-down state, and the receive path then walks it - the reported symptom is a warning in tls_strp_msg_load, with the underlying state confusion reachable in several worse shapes.","attack_vector":"Any unprivileged local process on the node can do it against its own socket: enable kTLS with setsockopt, then disconnect and keep reading. No device node, no capability, no cooperating peer. Every tenant container with a normal socket API reaches this.","remediation":"Boot a kernel carrying the linked stable commits (which make disconnect on a TLS socket return an error). Interim: none at the tenant boundary - the syscall sequence is ordinary socket usage.","references":["https://git.kernel.org/stable/c/7bdcf5bc35ae59fc4a0fa23276e84b4d1534a3cf","https://git.kernel.org/stable/c/ac91c6125468be720eafde9c973994cb45b61d44","https://nvd.nist.gov/vuln/detail/CVE-2025-37756"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-190"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-37761","cve":"CVE-2025-37761","aliases":[],"title":"Linux kernel (drivers/gpu/drm/xe): A GPU TLB invalidation for a very large address range computes its length with a","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/xe)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A GPU TLB invalidation for a very large address range computes its length with a power-of-two roundup that overflows, producing an undefined shift and a bogus invalidation length. The operator-level consequence is that GPU TLB entries for pages the kernel believes it has unmapped may survive, letting a tenant's GPU keep touching host pages after they have been freed and potentially handed to another workload.","attack_vector":"Reachable from an unprivileged process in a container with /dev/dri/renderD* on an Intel Xe node with SVM/userptr in use: map a huge address range into the GPU, then tear it down (or simply exit) so the MMU notifier fires xe_svm_invalidate with a range larger than the roundup can represent. Observed in practice from a plain userspace exec test, no special privilege.","remediation":"Update to a kernel with the fix commits below, which falls back to a full TLB invalidation above a size threshold. Interim: deny /dev/dri render nodes to untrusted tenants on xe hosts; there is no knob to disable SVM range invalidation.","references":["https://git.kernel.org/stable/c/28477f701b63922ff88e9fb13f5519c11cd48b86","https://git.kernel.org/stable/c/e4715858f87b78ce58cfa03bbe140321edbbaf20","https://nvd.nist.gov/vuln/detail/CVE-2025-37761"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-37765","cve":"CVE-2025-37765","aliases":[],"title":"Linux kernel (drivers/gpu/drm/nouveau): A buffer object imported over PRIME leaves a dangling pointer behind, and the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/nouveau)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A buffer object imported over PRIME leaves a dangling pointer behind, and the deferred TTM delete worker later walks it - the oops trace shows classic freed-slab poison being dereferenced. A tenant driving dma-buf import and release controls when the freed object is touched, turning it into a use-after-free in kernel worker context.","attack_vector":"A tenant process holding /dev/dri/renderD* on a nouveau GPU imports and releases dma-bufs in a loop; the fault lands asynchronously in the TTM delayed-delete workqueue, so it is not confined to the tenant's own task. Conditional on the node using the upstream nouveau driver.","remediation":"Boot a kernel carrying the nouveau prime lifetime fix below. Interim: on nouveau nodes, withhold /dev/dri/renderD* from untrusted tenants, or forbid dma-buf import for workloads that do not need buffer sharing across devices.","references":["https://git.kernel.org/stable/c/706868a1a1072cffd8bd63f7e161d79141099849","https://git.kernel.org/stable/c/47761deabb69a5df0c2c4ec400d80bb3e072bd2e","https://nvd.nist.gov/vuln/detail/CVE-2025-37765"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-37854","cve":"CVE-2025-37854","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A use-after-free in the amdkfd (KFD compute driver","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the amdkfd (KFD compute driver, /dev/kfd). Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amdkfd: Fix mode1 reset crash issue","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-37854","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-05-09"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-37869","cve":"CVE-2025-37869","aliases":[],"title":"Linux kernel (drivers/gpu/drm/xe): The error path of the Xe VRAM clear helper waits on a fence pointer that is only","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/xe)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"The error path of the Xe VRAM clear helper waits on a fence pointer that is only stable while the job mutex is held, so a racing submission can free it underneath the waiter. A tenant that can drive migration/clear failures gets a use-after-free on a dma-fence in the buffer-migration path - kernel memory corruption from an unprivileged container.","attack_vector":"Requires a container with /dev/dri/renderD* on an Intel Xe node. VRAM clear runs when buffer objects are allocated and evicted, so a tenant reaches it just by churning BO allocations under memory pressure; the race window opens when the clear submission fails and another thread is concurrently replacing the migration fence.","remediation":"Boot a kernel containing the fix commits below. Interim mitigation is limited to removing render-node access from untrusted tenants; VRAM clear on allocation cannot be disabled.","references":["https://git.kernel.org/stable/c/2ac5f466f62892a7d1ac2d1a3eb6cd14efbe2f2d","https://git.kernel.org/stable/c/dc712938aa26b001f448d5e93f59d57fa80f2dbd","https://nvd.nist.gov/vuln/detail/CVE-2025-37869"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-37903","cve":"CVE-2025-37903","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A use-after-free in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the amdgpu display core (DC/DM). Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amd/display: Fix slab-use-after-free in hdcp","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-37903","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-05-20"},{"id":"CVE-2025-37911","cve":"CVE-2025-37911","aliases":[],"title":"Linux bnxt_en driver (ethtool coredump / bnxt_get_coredump): Out-of-bounds memcpy when retrieving a firmware coredump","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux bnxt_en driver (ethtool coredump / bnxt_get_coredump)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Out-of-bounds memcpy when retrieving a firmware coredump via ethtool — the returned DMA length can exceed the buffer the driver allocated, corrupting kernel memory. What makes this operationally awkward is that `ethtool -w` is exactly what your support workflow runs when a NIC misbehaves, so the diagnostic step is the trigger.","attack_vector":"Local privileged user running an ethtool coredump against the NIC, with firmware returning an over-long length. Relevant if tenants have root on bare metal, or if a compromised NIC firmware can influence the returned length.","remediation":"Kernel/driver upgrade plus host reboot. Until patched, avoid `ethtool -w` on Broadcom NICs in your automated diagnostics.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-37911"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-05-20"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-38024","cve":"CVE-2025-38024","aliases":[],"title":"Linux kernel Soft-RoCE completion queue (rdma_rxe, rxe_cq_cleanup on create failure): Syzkaller-found slab","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel Soft-RoCE completion queue (rdma_rxe, rxe_cq_cleanup on create failure)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Syzkaller-found slab use-after-free reached straight from ib_uverbs_create_cq - the error unwind in rxe_create_cq() cleans up a queue that has already been released. Reachable by a tenant simply asking for a completion queue it knows will fail to allocate, which makes it easy to trigger repeatedly and therefore easy to shape the heap around.","attack_vector":"Local, unprivileged, via the uverbs create-CQ command on a Soft-RoCE device.","remediation":"Kernel update fixing the cleanup ordering. Blacklist rdma_rxe where software RoCE is not in use.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=16c45ced0b3839d3eee72a86bb172bef6cf58980","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2025/CVE-2025-38024.json"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-38062","cve":"CVE-2025-38062","aliases":[],"title":"Linux kernel (drivers/iommu): With iommufd, a tenant can change a passthrough device's IOMMU domain while MSI","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"With iommufd, a tenant can change a passthrough device's IOMMU domain while MSI interrupts are being programmed, and the MSI descriptor still held a raw pointer into the old domain's cookie. Racing the two ioctls gives a use-after-free on the structure that decides where the device's interrupt writes land - kernel memory corruption reached from a VM's own device fd, and an interrupt address the attacker has influence over.","attack_vector":"A tenant VMM holding /dev/vfio/* plus /dev/iommu races VFIO_DEVICE_ATTACH_IOMMUFD_PT (which re-attaches the IOMMU domain) against VFIO_DEVICE_SET_IRQS (which composes the translated MSI message) on the same device. Both are ordinary unprivileged ioctls on fds a passthrough tenant already owns; the fix notes the unlocked iommu_get_domain_for_dev() on the MSI translation path is a second UAF in the same window. Requires the iommufd path (not legacy VFIO type1 containers, which could not swap domains at runtime).","remediation":"No fixed release is listed in this record; apply the linked stable commits or run a current stable/LTS kernel on every passthrough node. Interim: keep tenants on the legacy VFIO type1 container path rather than iommufd where the platform still allows it, since type1 cannot change the domain while the device is live, and do not expose /dev/iommu to untrusted containers.","references":["https://git.kernel.org/stable/c/e4d3763223c7b72ded53425207075e7453b4e3d5","https://git.kernel.org/stable/c/ba41e4e627db51d914444aee0b93eb67f31fa330","https://nvd.nist.gov/vuln/detail/CVE-2025-38062"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-38080","cve":"CVE-2025-38080","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Increase block_sequence array size","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38080","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-06-18"},{"id":"CVE-2025-38091","cve":"CVE-2025-38091","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: check stream id dml21 wrapper to get plane_id","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38091","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-07-02"},{"id":"CVE-2025-38098","cve":"CVE-2025-38098","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Don't treat wb connector as physical in create_validate_stream_for_sink","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38098","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-07-03"},{"id":"CVE-2025-38109","cve":"CVE-2025-38109","aliases":[],"title":"NVIDIA BlueField (mlx5 ECVF): Use-after-free during ECVF vport unload in the mlx5 driver","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA BlueField (mlx5 ECVF)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Use-after-free during ECVF vport unload in the mlx5 driver — kernel-level memory corruption on the host attached to the DPU","attack_vector":"Local, host kernel","remediation":"Host kernel update plus driver refresh; requires a node drain because the mlx5 driver carries the tenant's data path","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38109"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2025-07-03"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-193","CWE-617"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-38166","cve":"CVE-2025-38166","aliases":[],"title":"Linux kernel (net/tls): A BPF verdict that grows the scatterlist (bpf_msg_push_data) combined with a cork_bytes setting","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/tls)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A BPF verdict that grows the scatterlist (bpf_msg_push_data) combined with a cork_bytes setting makes kTLS roll back to the non-zerocopy path with a reset msg_iter but an already-grown sg.size. The revert walks past the end of the iterator and hits a hard BUG() in iov_iter_revert - an immediate kernel panic on a shared node.","attack_vector":"Needs a sockmap/sk_msg BPF program with cork_bytes set attached to a kTLS TX socket, then any sendmsg/sendto on that socket. Attaching needs CAP_BPF (CNI, service-mesh sidecar, or a tenant granted BPF); once attached, ordinary tenant sends trigger the panic. Not reachable from a plain unprivileged container with no BPF access.","remediation":"Boot a kernel carrying the linked stable commits. Interim: do not grant CAP_BPF to tenant containers, and remove sk_msg programs that push data on corked kTLS sockets.","references":["https://git.kernel.org/stable/c/328cac3f9f8ae394748485e769a527518a9137c8","https://git.kernel.org/stable/c/2e36a81d388ec9c3f78b6223f7eda2088cd40adb","https://nvd.nist.gov/vuln/detail/CVE-2025-38166"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-38187","cve":"CVE-2025-38187","aliases":[],"title":"Linux kernel (drivers/gpu/drm/nouveau/nvkm/subdev/gsp/rm/r535): Nouveau's GSP RPC layer frees the caller's message","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel (drivers/gpu/drm/nouveau/nvkm/subdev/gsp/rm/r535)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Nouveau's GSP RPC layer frees the caller's message container after sending the first fragment of a multi-fragment RPC, then keeps using it to send the rest. Every remaining fragment is a write through freed kernel memory on the control channel between the driver and GPU firmware - the path that carries object allocation and mapping requests on Turing-and-later NVIDIA GPUs.","attack_vector":"Reached indirectly by a tenant holding /dev/dri/renderD* on a nouveau-driven NVIDIA GPU: userspace ioctls that allocate objects or set up mappings generate GSP RPCs, and any request whose payload exceeds a single fragment takes the buggy path. Conditional on the open nouveau driver with GSP firmware (not the proprietary NVIDIA module).","remediation":"Update to a kernel carrying the fix commits below. Interim: on nodes using the proprietary NVIDIA driver, ensure nouveau is blacklisted so the vulnerable path is not loaded at all; otherwise restrict /dev/dri access.","references":["https://git.kernel.org/stable/c/cd4677407c0ee250fc21e36439c8a442ddd62cc1","https://git.kernel.org/stable/c/9802f0a63b641f4cddb2139c814c2e95cb825099","https://nvd.nist.gov/vuln/detail/CVE-2025-38187"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-38352","cve":"CVE-2025-38352","aliases":[],"title":"Linux kernel (posix-cpu-timers): TOCTOU race between handle_posix_cpu_timers() and posix_cpu_timer_del()","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (posix-cpu-timers)","year":"2025","cvss_score":7.8,"severity":"high","kev":true,"impact":"TOCTOU race between handle_posix_cpu_timers() and posix_cpu_timer_del() - local privilege escalation, exploited in the wild [KEV]","attack_vector":"Any tenant process in a container","remediation":"Livepatchable; otherwise drain + reboot. No capabilities or namespaces needed, so container hardening does not mitigate it","references":["https://access.redhat.com/security/cve/CVE-2025-38352"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-07-22"},{"id":"CVE-2025-38361","cve":"CVE-2025-38361","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Check dce_hwseq before dereferencing it","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38361","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-07-25"},{"id":"CVE-2025-38389","cve":"CVE-2025-38389","aliases":[],"title":"Linux i915 GPU kernel driver (GT timeline / VMA allocation): A timeline is left held when VMA allocation fails, so","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux i915 GPU kernel driver (GT timeline / VMA allocation)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A timeline is left held when VMA allocation fails, so the error path leaks a reference and the driver wedges. Shows up as hung GPU submissions and an unresponsive card after memory pressure - which on a busy training node is a common condition, not a rare one.","attack_vector":"Any local user or container with a DRM render node - i.e. any tenant that was scheduled a GPU. No privileged capability needed.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38389","https://git.kernel.org/stable/c/40e09506aea1fde1f3e0e04eca531bbb23404baf"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-07-25"},{"id":"CVE-2025-38455","cve":"CVE-2025-38455","aliases":[],"title":"Linux KVM/SVM - SEV/SEV-ES intra-host migration during vCPU creation: KVM permitted SEV/SEV-ES intra-host migration","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux KVM/SVM - SEV/SEV-ES intra-host migration during vCPU creation","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"KVM permitted SEV/SEV-ES intra-host migration while vCPU creation was still in flight, producing a race on confidential-VM state. Migration racing against vCPU setup means encrypted vCPU state can be moved or referenced while half-built - a route to host memory corruption driven from the VM lifecycle path.","attack_vector":"Through the KVM ioctl interface used for migration, reachable by the VMM process - so a compromised orchestrator or VMM.","remediation":"Fixed in the Linux kernel. Take the distro kernel update (RHEL/Rocky, Ubuntu, SLES) and reboot the host - no firmware, VBIOS or AGESA step. On a GPU fleet this is a cordon, drain and rolling reboot; plan it as normal kernel maintenance. Interim control: disable intra-host migration for SEV guests in your VMM configuration.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38455"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-07-25"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-843"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-38475","cve":"CVE-2025-38475","aliases":[],"title":"Linux kernel SMC (struct smc_sock type confusion with inet_sock): Struct smc_sock does not embed struct inet_sock as","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel SMC (struct smc_sock type confusion with inet_sock)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Struct smc_sock does not embed struct inet_sock as its first member but claims AF_INET/AF_INET6 in sk_family, so generic IPv4 socket code operates on it as if it were an inet_sock. syzbot's report shows cipso_v4_sock_setattr() freeing what it thinks is inet_opt but is actually smc_sock.clcsk_data_ready - a function pointer in the text segment. Type confusion that makes the kernel treat a function pointer as a heap pointer is a strong exploitation primitive, and it is reachable by any unprivileged process that can create an SMC socket.","attack_vector":"Local, unprivileged. Create an AF_SMC socket and drive it through generic IPv4 socket options (the reported path is CIPSO).","remediation":"Kernel update correcting the socket type handling. Immediate mitigation is real here: block SMC socket creation for tenants (seccomp on the socket family, or keep the module unloaded), which removes the confusion entirely without a reboot.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=5b02e397929e5b13b969ef1f8e43c7951e2864f5","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2025/CVE-2025-38475.json"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-415"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-38500","cve":"CVE-2025-38500","aliases":[],"title":"Linux kernel (net/xfrm): The guard that forbids changing a collect_md xfrm interface never fired, so a changelink puts","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"The guard that forbids changing a collect_md xfrm interface never fired, so a changelink puts the special interface into the per-netns hash while it is still referenced by the collect_md pointer. The netdevice is then freed twice when the namespace goes away - a kernel BUG and heap corruption in the host kernel, reached from inside a container's own network namespace.","attack_vector":"A container with CAP_NET_ADMIN in a user namespace and its own netns (the common configuration for CNI-managed pods and for anything running with NET_ADMIN) creates a collect_md xfrm interface via RTM_NEWLINK, then issues a changelink on it. The double free lands when that netns is torn down, which happens on every pod delete.","remediation":"Boot a kernel carrying the linked stable commits. Interim: drop CAP_NET_ADMIN from tenant containers, or block the xfrm interface link type (blacklist xfrm_interface) on nodes that do not need it.","references":["https://git.kernel.org/stable/c/a8d4748b954584ab7bd800f1a4e46d5b0eeb5ce4","https://git.kernel.org/stable/c/bfebdb85496e1da21d3cf05de099210915c3e706","https://nvd.nist.gov/vuln/detail/CVE-2025-38500"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-38594","cve":"CVE-2025-38594","aliases":[],"title":"Linux kernel (drivers/iommu/intel): VT-d tore the device off the I/O page-fault queue before the hardware had stopped","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/intel)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"VT-d tore the device off the I/O page-fault queue before the hardware had stopped generating faults and before in-flight ones were drained, so the fault worker kept running against freed fault groups. Result is a refcount underflow and use-after-free in kernel context - reachable by a tenant that simply closes an SVA context while its device still has page requests outstanding.","attack_vector":"Local, on Intel hosts with VT-d scalable mode and an SVA/PASID-capable device exposed to tenants (DSA/IAA accelerators, SVM-capable GPUs, PRI-capable NICs). The tenant binds SVA, drives the device to generate page requests against unmapped addresses, then unbinds or exits so the last IOPF-capable domain detaches while faults are still queued. No host root required; needs PRI/IOPF enabled on the device.","remediation":"No fixed release is listed in this record; apply the linked stable commits or run a current stable/LTS kernel. Interim: disable PRI/SVA on devices handed to tenants, or stop exposing SVA-capable accelerator device nodes into tenant containers.","references":["https://git.kernel.org/stable/c/c68332b7ee893292bba6e87d31ef2080c066c65d","https://git.kernel.org/stable/c/f0b9d31c6edd50a6207489cd1bd4ddac814b9cd2","https://nvd.nist.gov/vuln/detail/CVE-2025-38594"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-38598","cve":"CVE-2025-38598","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdgpu): A use-after-free in the amdkfd (KFD compute driver","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdgpu)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the amdkfd (KFD compute driver, /dev/kfd). Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amdgpu: fix use-after-free in amdgpu_userq_suspend+0x51a/0x5a0","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38598","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-08-19"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-125"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-38616","cve":"CVE-2025-38616","aliases":[],"title":"Linux kernel (net/tls): KTLS assumes it owns the TCP receive queue. When another reader drains bytes first, the old","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/tls)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"KTLS assumes it owns the TCP receive queue. When another reader drains bytes first, the old code hit a WARN and returned early leaving the strparser anchor pointing at a freed skb, and could read past the end of what is actually queued. Beyond the memory-safety hit, the parser can then decrypt something that is not a valid record - a missed alert or a missed attack on that stream.","attack_vector":"Local: a process that read from the TCP socket before the TLS ULP was installed, or that uses a non-standard/zerocopy read API on the same socket. That is a same-container or same-process condition rather than a cross-tenant one, but it is unprivileged and needs no device node; the peer supplies the record data that gets misparsed.","remediation":"Boot a kernel carrying the linked stable commits. Interim: install the TLS ULP before any read on the socket and do not mix zerocopy receive with kTLS.","references":["https://git.kernel.org/stable/c/f1fe99919f629f980d0b8a7ff16950bffe06a859","https://git.kernel.org/stable/c/eb0336f213fe88bbdb7d2b19c9c9ec19245a3155","https://nvd.nist.gov/vuln/detail/CVE-2025-38616"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-457","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-38675","cve":"CVE-2025-38675","aliases":[],"title":"Linux kernel (net/xfrm): If the task is preempted onto another CPU during SA lookup, a hit in the per-CPU state cache","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"If the task is preempted onto another CPU during SA lookup, a hit in the per-CPU state cache jumps to the found label while the acquire path is entered with an uninitialized state_ptrs structure. The SA selection path then works off stack garbage - a wild-pointer read/write in the code that decides which security association a flow gets, which is both a corruption primitive and a wrong-SA risk.","attack_vector":"Driven by any outbound flow that needs an SA on an IPsec-enabled node - tenant traffic on an encrypted overlay is enough; the race needs preemption between the cache lookup and the acquire branch, which ordinary scheduling on a busy node provides. Requires the per-CPU xfrm state cache (6.7 and later kernels) and configured xfrm policy.","remediation":"Update to 6.12.41 or later in the 6.12 series, or a kernel carrying the linked stable commits. Interim: no meaningful workaround short of disabling IPsec on the node.","references":["https://git.kernel.org/stable/c/6bf2daafc51bcb9272c0fdff2afd38217337d0d3","https://git.kernel.org/stable/c/463562f9591742be62ddde3b426a0533ed496955","https://nvd.nist.gov/vuln/detail/CVE-2025-38675"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-38703","cve":"CVE-2025-38703","aliases":[],"title":"Linux kernel (drivers/gpu/drm/xe): Xe frees data that its exported dma-fences still point at - notably the timeline","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/xe)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Xe frees data that its exported dma-fences still point at - notably the timeline name - as soon as the owning submit queue is closed. Because those fences can already have been handed to a different process as a sync_file fd, the holder reads freed kernel memory. One tenant closing a queue corrupts state observed by whoever it shared the fence with, which is a genuine cross-process use-after-free.","attack_vector":"A container with /dev/dri/renderD* on an Intel Xe node exports a fence (sync_file / syncobj / dma-buf) to another process, then destroys its exec queue; the receiving side's subsequent access to the fence hits freed memory. Both sides are unprivileged, and fence sharing across process boundaries is a normal, supported operation.","remediation":"Update to a kernel with the fix commits below, which adds RCU grace periods before freeing fence-referenced data. Interim: do not pass fence/sync fds across tenant boundaries, and restrict render-node access.","references":["https://git.kernel.org/stable/c/b17fcce70733c211cb5dabf54f4f9491920b1d92","https://git.kernel.org/stable/c/ba37807d08bae67de6139346a85650cab5f6145a","https://nvd.nist.gov/vuln/detail/CVE-2025-38703"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-38722","cve":"CVE-2025-38722","aliases":[],"title":"habanalabs kernel driver (dma-buf export path): A use-after-free in the habanalabs dma-buf export path: the driver","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"habanalabs kernel driver (dma-buf export path)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the habanalabs dma-buf export path: the driver installs a file descriptor into the process table and then keeps using the object, so a second thread that closes the fd first frees memory the driver is still writing. dma-buf is exactly the mechanism used to hand accelerator memory to another process or device, so a successful exploit is a kernel-memory write reachable from an unprivileged accelerator user.","attack_vector":"Any local user with access to the habanalabs device node - on Kubernetes that is any pod granted a Gaudi device. Requires a deliberate race, not a lucky one.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38722","https://git.kernel.org/stable/c/33927f3d0ecdcff06326d6e4edb6166aed42811c"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-09-04"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-415"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-38731","cve":"CVE-2025-38731","aliases":[],"title":"Linux kernel (drivers/gpu/drm/xe): A tenant that submits a deliberately malformed array bind to the Xe VM_BIND ioctl","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/xe)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A tenant that submits a deliberately malformed array bind to the Xe VM_BIND ioctl gets the driver to free the same kernel allocation twice, which is a slab-corruption primitive rather than a clean error return. On a shared Intel GPU node that is the classic path from an unprivileged container to kernel memory corruption and privilege escalation, or at minimum a panic that takes every co-tenant on the box down with it.","attack_vector":"Directly reachable from any container holding /dev/dri/renderD* on a host running the Intel Xe driver - no card* node, no DRM master, no display access required. The attacker just calls DRM_IOCTL_XE_VM_BIND with an array of bind ops whose argument validation fails; the failure path is the bug.","remediation":"Boot a kernel carrying the fix commits below. Until then, remove /dev/dri/renderD* from any container that does not genuinely need GPU access, and treat nodes running the xe driver with untrusted tenants as exposed - there is no runtime toggle that disables VM_BIND while keeping the GPU usable.","references":["https://git.kernel.org/stable/c/77a946bf1af0e8110ef6e243394217a17f9b7e33","https://git.kernel.org/stable/c/111fb43a557726079a67ce3ab51f602ddbf7097e","https://nvd.nist.gov/vuln/detail/CVE-2025-38731"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-38743","cve":"CVE-2025-38743","aliases":[],"title":"Dell iDRAC Service Module (iSM): Buffer access with incorrect length in the in-band agent gives a low-privileged local","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC Service Module (iSM)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Buffer access with incorrect length in the in-band agent gives a low-privileged local attacker code execution and elevation on the host OS. iSM is the bridge between host and BMC, so it is a stepping stone from tenant workload toward out-of-band control.","attack_vector":"Any low-privilege local account on the server OS - including a container that escaped to the host.","remediation":"Upgrade iSM to 6.0.3.0. Host package update plus service restart, no reboot and no firmware flash. Cheap - do it in the normal patch cycle. If you do not actually consume iSM telemetry, uninstall it instead.","references":["https://www.dell.com/support/kbdoc/en-us/000359617/dsa-2025-311-security-update-for-dell-idrac-service-module-vulnerabilities"],"status":"curated","fleet":{"pain_class":"daemon-restart"}},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-39740","cve":"CVE-2025-39740","aliases":[],"title":"Linux kernel (drivers/gpu/drm/xe): On the migration error path the previous fence is released before the code waits on","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/xe)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"On the migration error path the previous fence is released before the code waits on it, so the wait runs against freed memory. Migration is what moves buffer objects between VRAM and system memory under pressure, so a tenant that keeps the GPU memory-constrained can reach this use-after-free while other tenants' buffers are being evicted around it.","attack_vector":"A tenant container holding /dev/dri/renderD* on an Intel xe GPU triggers it through ordinary buffer-object migration/eviction, forced by allocating VRAM until eviction kicks in and then failing the copy. No special capability, no display path, no host root.","remediation":"Boot a kernel carrying the xe_migrate fence-ordering fix below. Interim: cap per-tenant VRAM so a single container cannot keep the device in constant eviction, which is the state that exercises this path.","references":["https://git.kernel.org/stable/c/7e46fa64a4b94208563c3a5bf1d7f4346f94abea","https://git.kernel.org/stable/c/145832fbdd17b1d77ffd6cdd1642259e101d1b7e","https://nvd.nist.gov/vuln/detail/CVE-2025-39740"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-39810","cve":"CVE-2025-39810","aliases":[],"title":"Linux bnxt_en driver (ring defaults vs traffic classes on ifdown): Memory corruption when firmware resources change","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux bnxt_en driver (ring defaults vs traffic classes on ifdown)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Memory corruption when firmware resources change while the interface is down, because the default-ring calculation assumes no traffic classes have been created. Relevant to RoCE clusters specifically: traffic classes are how you carve out lossless queues for RDMA with PFC, so any cluster running RoCEv2 with DCB has traffic classes configured and is in the affected configuration by construction.","attack_vector":"Local — an interface down/up cycle combined with a firmware resource change, on a host with traffic classes configured.","remediation":"Kernel/driver upgrade plus host reboot. Until then, avoid ifdown/ifup cycles on RoCE-configured Broadcom interfaces as a routine operational step — use link-level draining at the switch instead.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-39810"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-09-16"},{"id":"CVE-2025-39906","cve":"CVE-2025-39906","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: remove oem i2c adapter on finish","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-39906","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-10-01"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-39965","cve":"CVE-2025-39965","aliases":[],"title":"Linux kernel (net/xfrm): SPI 0 means 'no SPI assigned', but the duplicate-SPI rework started creating states with SPI 0","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"SPI 0 means 'no SPI assigned', but the duplicate-SPI rework started creating states with SPI 0 and putting them on the byspi list. Deleting such a state never removes it from that list, so the next walk of the list touches freed memory - a use-after-free in the security-association table that also leaves phantom entries in SA lookup.","attack_vector":"Reached through the normal IKE control path: XFRM_MSG_ALLOCSPI / SA creation over xfrm netlink, available to the node's IKE daemon (strongSwan, libreswan) and to any process with CAP_NET_ADMIN in its network namespace - which includes containers granted NET_ADMIN. On a node with churning child SAs the condition arises without an attacker.","remediation":"Update to 6.6.109 / 6.12.50 / 6.16.10 or later, or a kernel carrying the linked stable commits. Interim: drop CAP_NET_ADMIN from tenant containers so only the node's IKE daemon can create SAs.","references":["https://git.kernel.org/stable/c/0baf92d0b1590b903c1f4ead75e61715e50e8146","https://git.kernel.org/stable/c/9fcedabaae0096f712bbb4ccca6a8538af1cd1c8","https://nvd.nist.gov/vuln/detail/CVE-2025-39965"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-39966","cve":"CVE-2025-39966","aliases":[],"title":"Linux kernel (drivers/iommu/iommufd): Aborting an iommufd object allocation freed the object immediately while the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/iommufd)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Aborting an iommufd object allocation freed the object immediately while the paired file was still sitting on the fput workqueue, so the deferred release() ran against freed memory. Confirmed slab use-after-free found by syzkaller, reachable by anything holding /dev/iommu - the same fd a passthrough tenant needs.","attack_vector":"Any process with /dev/iommu open drives an object-allocation ioctl (the fault/event queue objects are the reported path) down a failure branch so the core aborts after the file has been created. No host root, no device needed beyond the iommufd character device. Conditional on iommufd being in use and /dev/iommu being visible to the tenant - which it is on any node using the modern VFIO/iommufd passthrough stack.","remediation":"No fixed release is listed in this record; apply the linked stable commits or run a current stable/LTS kernel on passthrough nodes. Interim: remove /dev/iommu from containers that do not do passthrough, and keep it restricted to the VMM process rather than the tenant workload.","references":["https://git.kernel.org/stable/c/17195a7d754a5c6a31888702ca93f6f08f3383ad","https://git.kernel.org/stable/c/e4825368285e33d6360c6c6a6a10d2d83da06e55","https://nvd.nist.gov/vuln/detail/CVE-2025-39966"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-39979","cve":"CVE-2025-39979","aliases":["net/mlx5 fs, fix UAF in flow counter release"],"title":"Linux kernel mlx5_core flow steering / flow counters (hardware steering): Use-after-free releasing","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_core flow steering / flow counters (hardware steering)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Use-after-free releasing the hardware-steering action of a local flow counter - the refcount and mutex were never initialized and the counter struct can already be freed when the rule is deleted. Notably it is reached through an ib_uverbs ioctl into mlx5_ib_destroy_flow, so a tenant holding an RDMA verbs handle can drive host kernel memory corruption in the NIC's flow-steering tables.","attack_vector":"Local, low-privileged - reachable from a userspace RDMA verbs handle destroying a flow, not only from privileged tc/devlink paths.","remediation":"Upgrade the host kernel to 6.17 or the 6.16.10 stable backport. Rolling reboot of the fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-39979","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2025/CVE-2025-39979.json"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-10-15"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-40012","cve":"CVE-2025-40012","aliases":[],"title":"Linux kernel (net/smc): SMC-D loopback registers DMBs (the direct memory buffers a peer reads and writes) out of","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/smc)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"SMC-D loopback registers DMBs (the direct memory buffers a peer reads and writes) out of kmalloc memory, which is not page-backed, so the receive path cannot take a page reference before handing the buffer to a pipe. splice() from an SMC-D socket therefore runs against a buffer that can be freed under it - a use-after-free on the remote-memory-buffer path reachable by an ordinary tenant.","attack_vector":"Local and unprivileged, conditional on the SMC-D loopback ISM device being available (the smc_loopback module - present on s390 and on x86 kernels that ship loopback-ism). A tenant opens an AF_SMC socket that negotiates SMC-D over loopback and splices from it; no capability, no device node, and no RDMA hardware is needed, and socket(AF_SMC, ...) autoloads the family itself.","remediation":"Boot a kernel carrying the fix commits (DMBs allocated with folio_alloc() so they are page-backed). Interim: blacklist the smc_loopback / loopback-ism module so SMC-D loopback is not offered, or blacklist smc entirely on nodes not using it.","references":["https://git.kernel.org/stable/c/14fc4fdae42e34d7ee871b292ac2ecc61c2c5de7","https://git.kernel.org/stable/c/a35c04de2565db191726b5741e6b66a35002c652","https://nvd.nist.gov/vuln/detail/CVE-2025-40012"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-362","CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-40061","cve":"CVE-2025-40061","aliases":[],"title":"Linux kernel (drivers/infiniband/sw/rxe): Use-after-free from a race between a busy soft-RoCE task and its own","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/sw/rxe)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Use-after-free from a race between a busy soft-RoCE task and its own teardown. A task that exhausts its work budget resets its state and reschedules itself, silently overwriting the draining state set by cleanup - so cleanup proceeds and frees objects the rescheduled task keeps using. Heap corruption reachable by an unprivileged tenant.","attack_vector":"A tenant container holding /dev/infiniband/uverbs* on a node with rdma_rxe loaded: keep a queue-pair saturated with work (which forces the task to hit its iteration limit) while destroying or disabling it. Inbound traffic from a peer on the fabric makes the busy state easy to sustain. rxe only.","remediation":"No fixed release is published in this record - apply the listed stable fix commits or run a current stable kernel. Interim: blacklist/unload rdma_rxe where soft-RoCE is not required, and drop /dev/infiniband/* from containers that do not use verbs.","references":["https://git.kernel.org/stable/c/85288bcf7ffe11e7b036edf91937bc62fd384076","https://git.kernel.org/stable/c/52edccfb555142678c836c285bf5b4ec760bd043","https://nvd.nist.gov/vuln/detail/CVE-2025-40061"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-40064","cve":"CVE-2025-40064","aliases":[],"title":"Linux kernel (net/smc): Connect() on an SMC socket takes the destination device pointer out of the dst cache without a","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/smc)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Connect() on an SMC socket takes the destination device pointer out of the dst cache without a reference and only then grabs RTNL, so a device that goes away in that window is dereferenced after free during ISM/RoCE device selection. KASAN confirms the use-after-free read in __pnet_find_base_ndev; a tenant that races connect() against interface teardown reads and can corrupt freed kernel memory.","attack_vector":"Local and unprivileged: loop connect() on AF_SMC sockets while a netdev is being removed - trivial for a tenant that controls its own veth or has a container being torn down alongside. socket(AF_SMC, ...) requires no capability and autoloads the smc module via the net-pf-43 alias, so no /dev/infiniband access is needed to reach the pnetid lookup.","remediation":"Boot a kernel carrying the fix commits (holds the device reference across smc_pnet_find_ism_resource / smc_pnet_find_roce_resource). Interim: blacklist the smc module or deny socket family 43 to tenants.","references":["https://git.kernel.org/stable/c/233927b645cb7a14bb98d23ac72e4c7243a9f0d9","https://git.kernel.org/stable/c/3d3466878afd8d43ec0ca2facfbc7f03e40d0f79","https://nvd.nist.gov/vuln/detail/CVE-2025-40064"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-415","CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-40096","cve":"CVE-2025-40096","aliases":[],"title":"Linux kernel (drivers/gpu/drm/scheduler): When adding reservation-object dependencies to a job, the helper already","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/scheduler)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"When adding reservation-object dependencies to a job, the helper already consumes the fence reference on failure, so the caller's cleanup drops it a second time - a double free of a dma-fence. Because this sits in the shared DRM scheduler used by amdgpu and xe, one tenant's job submission can corrupt fence objects the whole node's GPU scheduling depends on.","attack_vector":"A tenant process holding /dev/dri/renderD* submits jobs with many buffer-object dependencies while the internal xarray fails to expand - which a tenant induces simply by driving the node into memory pressure from its own container. Applies to any driver on the common DRM scheduler (amdgpu, xe, and others), so this is the mainline datacenter GPU path, not a niche driver.","remediation":"Boot a kernel carrying the drm/sched dependency-tracking fix below. Interim: enforce hard per-container memory limits so tenants cannot drive the node into the allocation-failure window, and cap job dependency counts where the userspace stack allows it.","references":["https://git.kernel.org/stable/c/4c38a63ae12ecc9370a7678077bde2d61aa80e9c","https://git.kernel.org/stable/c/57239762aa90ad768dac055021f27705dae73344","https://nvd.nist.gov/vuln/detail/CVE-2025-40096"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-40111","cve":"CVE-2025-40111","aliases":[],"title":"Linux kernel (drivers/gpu/drm/vmwgfx): A guest process can get a node left in the vmwgfx validation hash table after","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/vmwgfx)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A guest process can get a node left in the vmwgfx validation hash table after its backing resource has already been destroyed, so the next command submission dereferences freed memory. Use-after-free inside the execbuf validation path is directly weaponisable for guest kernel privilege escalation, and it lives in the code that every 3D command submission goes through.","attack_vector":"Reachable by any process inside a VMware guest that holds /dev/dri/renderD* or /dev/dri/card* - the bug is in vmw_execbuf_process's validation bookkeeping, so it is on the ordinary command-submission path. Relevant wherever tenant workloads run as VMware guests with the paravirtual vmwgfx GPU exposed; not applicable to bare-metal GPU nodes.","remediation":"Update guest kernels to a build carrying the fix commits below. Interim: drop /dev/dri device nodes from untrusted guest containers, or configure the VM without the vmwgfx 3D device so the render node is not present.","references":["https://git.kernel.org/stable/c/1822e5287b7dfa59d0af966756ebf1dc652b60ee","https://git.kernel.org/stable/c/fb7165e5f3b3b10721ff70553583ad12e90e447a","https://nvd.nist.gov/vuln/detail/CVE-2025-40111"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-40149","cve":"CVE-2025-40149","aliases":[],"title":"Linux kernel (net/tls): The kTLS device-offload setup resolved the socket's netdevice outside RCU, so the net_device it","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/tls)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"The kTLS device-offload setup resolved the socket's netdevice outside RCU, so the net_device it hands to the NIC offload path can be freed underneath it - a use-after-free on the device object reached from a plain setsockopt call.","attack_vector":"Any unprivileged socket owner enabling kTLS NIC offload: setsockopt(SOL_TLS, TLS_TX/TLS_RX, ...) on a socket whose route or lower device is changing (bonding failover, link churn, veth teardown). This is the exact configuration used for offloaded storage and control-plane TLS on ConnectX-class NICs, so it is live on nodes that lean on kTLS offload.","remediation":"Boot a kernel carrying the linked stable commits. Interim: disable kTLS device offload (ethtool -K <dev> tls-hw-tx-offload off / tls-hw-rx-offload off) so the software path is used, and avoid link churn on nodes with live kTLS offload sessions.","references":["https://git.kernel.org/stable/c/2b1bef126bbb8d0da51491357559126d567c1dee","https://git.kernel.org/stable/c/e37ca0092ddace60833790b4ad7a390408fb1be9","https://nvd.nist.gov/vuln/detail/CVE-2025-40149"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-40202","cve":"CVE-2025-40202","aliases":[],"title":"The Linux kernel's IPMI driver message-handling layer: A use-after-free in a kernel driver reachable from the host's","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"The Linux kernel's IPMI driver message-handling layer","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in a kernel driver reachable from the host's IPMI device nodes gives a local attacker a path to kernel memory corruption and, from there, to privilege escalation on the host. On a bare-metal GPU node the significance is the direction of travel: this is a route from an unprivileged tenant process, through the kernel, toward the interface that talks to the BMC. It is the host-side half of the out-of-band security story, and it is the half that operators tend not to inventory because it lives in the kernel rather than in firmware. The per-user message limit was miscounted in several paths, producing a use-after-free. This is the in-kernel code every host uses to talk to its own BMC over the KCS or SSIF interface.","attack_vector":"A local process on the host with access to the IPMI character devices (/dev/ipmi*). On many stock server images those permissions are looser than they should be, and any container or tenant workload given access to them is in position.","remediation":"Kernel update and reboot - which on a GPU node means draining long-running training jobs, so it lands in the same expensive maintenance window as everything else kernel-level. There is a genuinely effective config-only mitigation that costs nothing: unless a workload needs in-band IPMI, do not expose /dev/ipmi* to it, and consider blacklisting the ipmi_devintf module entirely on tenant-facing bare metal. Most operators poll their BMCs over the network anyway and do not need the in-band path at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-40202","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2025/40xxx/CVE-2025-40202.json"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-40274","cve":"CVE-2025-40274","aliases":[],"title":"Linux kernel (virt/kvm): Unbinding a memslot from a guest_memfd was skipped once the file was already dying, so if the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (virt/kvm)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Unbinding a memslot from a guest_memfd was skipped once the file was already dying, so if the memslot is freed first the release path writes eight bytes through a freed pointer. syzbot and KASAN confirm the slab use-after-free write in host kernel memory - a straightforward kernel corruption primitive on nodes that run confidential VMs.","attack_vector":"Reached through KVM ioctls on the VM fd (memslot deletion racing guest_memfd close), not from inside a guest. Under a trusted-VMM model this is not a tenant path; it becomes one wherever untrusted local users or tenant-controlled VMM processes hold /dev/kvm, where it is a local privilege-escalation primitive. Only nodes with guest_memfd in use (SEV-SNP / TDX confidential VMs) exercise this code.","remediation":"Update to a stable kernel carrying the linked fix (no fixed release enumerated; take the branch containing commit ae431059e75d). Interim control: keep /dev/kvm out of tenant containers and restricted to the operator's VMM account.","references":["https://git.kernel.org/stable/c/ae431059e75d36170a5ae6b44cc4d06d43613215","https://git.kernel.org/stable/c/393893693a523e053f84d69320d090b93503f79f","https://nvd.nist.gov/vuln/detail/CVE-2025-40274"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-190","CWE-787"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-40277","cve":"CVE-2025-40277","aliases":[],"title":"Linux kernel (drivers/gpu/drm/vmwgfx): The vmwgfx command-buffer parser trusted a size field taken straight from the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/vmwgfx)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"The vmwgfx command-buffer parser trusted a size field taken straight from the guest's command stream and used it in buffer offset arithmetic, so an oversized value overflows the calculation and walks the kernel off the end of the command buffer. That is an attacker-chosen out-of-bounds access inside the guest kernel, reachable from any process that can submit GPU commands.","attack_vector":"A container or process inside a VMware guest holding /dev/dri/renderD* submits an execbuf whose SVGA command header declares a size beyond SVGA_CMD_MAX_DATASIZE. Pure userspace-to-kernel input validation failure, no privileges beyond the render node. Applies only where tenants run as VMware guests with the vmwgfx device present.","remediation":"Update guest kernels to a build with the fix commits below. Interim: remove /dev/dri from untrusted guest containers, or remove the virtual 3D device from tenant VMs entirely.","references":["https://git.kernel.org/stable/c/e58559845021c3bad5e094219378b869157fad53","https://git.kernel.org/stable/c/54d458b244893e47bda52ec3943fdfbc8d7d068b","https://nvd.nist.gov/vuln/detail/CVE-2025-40277"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-40354","cve":"CVE-2025-40354","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: increase max link count and fix link->enc NULL pointer access","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-40354","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-12-16"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-427"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2025-40763","cve":"CVE-2025-40763","aliases":["SSA-514895"],"title":"Altair Grid Engine (shared library loading): Grid Engine does not sanitise the environment variables that control","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Altair Grid Engine (shared library loading)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Grid Engine does not sanitise the environment variables that control shared-library search, so a local user gets their own code loaded into a privileged Grid Engine process. That is a straight local privilege escalation on scheduler and execution hosts.","attack_vector":"A local user on a Grid Engine host who can set the environment of a Grid Engine binary - which includes anyone submitting jobs through the local commands. All versions before V2026.0.0.","remediation":"Upgrade Altair Grid Engine to V2026.0.0. This ships in the same Siemens advisory as CVE-2025-40760, so treat them as one upgrade.","references":["https://cert-portal.siemens.com/productcert/html/ssa-514895.html","https://nvd.nist.gov/vuln/detail/CVE-2025-40763"],"status":"curated"},{"id":"CVE-2025-41244","cve":"CVE-2025-41244","aliases":[],"title":"VMware Aria Operations / VMware Tools: Local privilege escalation to root inside a managed VM via SDMP service discovery","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"VMware Aria Operations / VMware Tools","year":"2025","cvss_score":7.8,"severity":"high","kev":true,"impact":"Local privilege escalation to root inside a managed VM via SDMP service discovery; exploited in the wild since Oct 2024 [KEV]","attack_vector":"Local non-admin user inside a managed guest VM","remediation":"Update VMware Tools / Aria Operations in guest images; no host reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-41244"],"status":"curated","published":"2025-09-29"},{"id":"CVE-2025-53000","cve":"CVE-2025-53000","aliases":[],"title":"nbconvert: Template-driven conversion executes attacker content","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"nbconvert","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Template-driven conversion executes attacker content","attack_vector":"Customer-supplied notebook converted by a pipeline","remediation":"Upgrade; report-generation pipelines ingest tenant notebooks","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-53000"],"status":"curated","published":"2025-12-17"},{"id":"CVE-2025-6018","cve":"CVE-2025-6018","aliases":[],"title":"PAM (pam-config): Local user is treated as `allow_active` in PAM","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"PAM (pam-config)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local user is treated as `allow_active` in PAM - first half of a public unprivileged-to-root chain on SUSE/openSUSE","attack_vector":"Local user (SSH session counts)","remediation":"Package update; no reboot. Chain with CVE-2025-6019 - patch both or neither is meaningful","references":["https://access.redhat.com/security/cve/CVE-2025-6018"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2025-07-23"},{"id":"CVE-2025-67450","cve":"CVE-2025-67450","aliases":["ETN-VA-2025-1027"],"title":"Eaton UPS Companion (EUC) executable - library loading: Insecure library loading in the shipped executable gives","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Eaton UPS Companion (EUC) executable - library loading","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Insecure library loading in the shipped executable gives arbitrary code execution to an attacker with access to the software package. Persistent code on hosts that talk to the UPS, running with the privileges power-management agents typically hold.","attack_vector":"Local, requires access to the software package or its directory on the host.","remediation":"Update to the fixed EUC version. Lock down the install directory permissions - power-management agents are installed to writable locations more often than they should be.","references":["https://www.eaton.com/content/dam/eaton/company/news-insights/cybersecurity/security-bulletins/etn-va-2025-1027.pdf"],"status":"curated","published":"2025-12-26"},{"id":"CVE-2025-68174","cve":"CVE-2025-68174","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (amd/amdkfd): A division by zero in the amdkfd (KFD compute driver","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (amd/amdkfd)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A division by zero in the amdkfd (KFD compute driver, /dev/kfd), reachable with attacker-influenced parameters. The kernel takes a divide fault and the node goes down, taking every co-resident job with it. Upstream fix: amd/amdkfd: enhance kfd process check in switch partition","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-68174","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-12-16"},{"id":"CVE-2025-68286","cve":"CVE-2025-68286","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Check NULL before accessing","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-68286","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-12-16"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-68379","cve":"CVE-2025-68379","aliases":[],"title":"Linux kernel (drivers/infiniband/sw/rxe): Two failed shared-receive-queue resizes in a row panic the node. The first","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/sw/rxe)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Two failed shared-receive-queue resizes in a row panic the node. The first failure leaves the queue pointer null, and the second call dereferences it while validating attributes - a deterministic, unprivileged kernel crash that takes the whole shared host down.","attack_vector":"A tenant container holding /dev/infiniband/uverbs* on a node with soft-RoCE (rdma_rxe) loaded: call the standard SRQ-modify verb twice with a size that makes the queue reallocation fail. Entirely tenant-controlled, no race to win, no fabric peer needed. Hardware HCAs do not run this code.","remediation":"No fixed release is published in this record - apply the listed stable fix commits or run a current stable kernel. Interim: blacklist/unload rdma_rxe unless soft-RoCE is deliberately in use, and keep /dev/infiniband/* out of containers that do not need verbs.","references":["https://git.kernel.org/stable/c/58aca869babd48cb9c3d6ee9e1452c4b9f5266a6","https://git.kernel.org/stable/c/b8f6eeb87a76b6fb1f6381b0b2894568e1b784f7","https://nvd.nist.gov/vuln/detail/CVE-2025-68379"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-68793","cve":"CVE-2025-68793","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): A use-after-free in the amdgpu RAS / GPU reset and","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the amdgpu RAS / GPU reset and recovery path. Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amdgpu: fix a job->pasid access race in gpu recovery","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-68793","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-01-13"},{"id":"CVE-2025-68801","cve":"CVE-2025-68801","aliases":["mlxsw spectrum_router fix neighbour use-after-free"],"title":"Linux kernel mlxsw (Spectrum switch router, neighbour table): The driver stored neighbour pointers without holding","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel mlxsw (Spectrum switch router, neighbour table)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"The driver stored neighbour pointers without holding a reference, taking one only when the neighbour was used by a nexthop - a slab use-after-free when updating neighbour entries. Reproduced on an NVIDIA SN5600. Neighbour churn on a busy leaf is normal operation, so this is a switch crash that arrives on its own schedule and takes a rack's uplinks with it.","attack_vector":"Local on the switch, driven by neighbour table churn - which an attacker on an attached network can amplify by cycling ARP/ND entries.","remediation":"Upgrade the switch OS to a build carrying kernel 6.19 or a stable backport (5.10.248, 5.15.198, 6.1.160, 6.6.120, 6.12.64, 6.18.3). Switch OS upgrade and reload - fabric rolling window, one switch at a time with ECMP draining.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-68801","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2025/CVE-2025-68801.json"],"status":"curated","published":"2026-01-13"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-68810","cve":"CVE-2025-68810","aliases":[],"title":"Linux kernel (virt/kvm): KVM blocked turning KVM_MEM_GUEST_MEMFD on for an existing memslot but not turning it off, and","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (virt/kvm)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"KVM blocked turning KVM_MEM_GUEST_MEMFD on for an existing memslot but not turning it off, and clearing the flag left the guest_memfd binding in place. Releasing the file then writes through a stale pointer - a KASAN-confirmed slab use-after-free in host kernel memory, i.e. a kernel memory-corruption primitive for whoever can drive VM lifecycle ioctls.","attack_vector":"Not reachable from inside a guest. It is reached from the process holding the VM file descriptor via KVM_SET_USER_MEMORY_REGION2 with the guest_memfd flag cleared. The VMM is trusted here, so this matters on nodes where untrusted local users or tenant-run VMM processes can open /dev/kvm - there it is a local privilege-escalation primitive. guest_memfd is the backing store for confidential VMs (SEV-SNP / TDX), so confidential-compute nodes are the ones carrying this code.","remediation":"Update to a stable kernel with the linked fix (no fixed release is enumerated; take the branch carrying commit 9935df5333aa). Interim control: keep /dev/kvm off tenant containers and restrict it to the operator's VMM service account; do not let tenants run their own VMM process on shared nodes.","references":["https://git.kernel.org/stable/c/9935df5333aa503a18de5071f53762b65c783c4c","https://git.kernel.org/stable/c/cb51bef465d8ec60a968507330e01020e35dc127","https://nvd.nist.gov/vuln/detail/CVE-2025-68810"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-71089","cve":"CVE-2025-71089","aliases":[],"title":"Linux kernel (drivers/iommu): In an SVA context the IOMMU walks and caches the CPU's page tables, and on x86 every","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"In an SVA context the IOMMU walks and caches the CPU's page tables, and on x86 every process page table maps the kernel half. The kernel had no way to tell the IOMMU when a kernel page-table page was freed and recycled, so the IOMMU kept walking stale entries into memory that now holds attacker-controlled data. That is a device DMA engine pointed at arbitrary physical memory - full host compromise potential from a device a tenant drives. This patch is the first of the series and hard-disables SVA on x86 until the invalidation mechanism (CVE-2025-71202) is in place.","attack_vector":"Local unprivileged, on x86 hosts with IOMMU SVA enabled and an SVA/PASID-capable device reachable by the tenant (Intel DSA/IAA, SVM-capable GPUs, PRI-capable NICs). The trigger is ordinary kernel page-table page recycling - the series notes vfree() is the common case and is reachable by unprivileged users - combined with an SVA-bound device that still has the stale entry cached. Requires CONFIG_IOMMU_SVA and a driver that offers SVA binding to userspace.","remediation":"No fixed release is listed in this record; apply the linked stable commits (this one plus CVE-2025-71202, which lands the actual kernel-VA invalidation) or run a current stable/LTS kernel. Interim: turn off SVA/PASID for tenant-facing devices, keep SVA-capable accelerator nodes out of tenant containers, and verify no tenant workload silently depends on shared virtual addressing before flipping it off.","references":["https://git.kernel.org/stable/c/b34289505180a83607fcfdce14b5a290d0528476","https://git.kernel.org/stable/c/7cad37e358970af1bb49030ff01f06a69fa7d985","https://nvd.nist.gov/vuln/detail/CVE-2025-71089"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-71092","cve":"CVE-2025-71092","aliases":[],"title":"Linux bnxt_re RoCE driver (bnxt_re_copy_err_stats out-of-bounds write): Out-of-bounds write in the Broadcom RoCE","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux bnxt_re RoCE driver (bnxt_re_copy_err_stats out-of-bounds write)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Out-of-bounds write in the Broadcom RoCE driver's error-statistics copy, introduced when three RoCE hardware counters were added past the end of the existing array. Reading RDMA counters is something monitoring agents do constantly on an AI cluster, so the vulnerable path runs on a schedule whether or not anyone attacks it.","attack_vector":"Triggered by reading RoCE hardware counters — reachable from any local process permitted to query RDMA statistics, including monitoring agents.","remediation":"Kernel/driver upgrade plus host reboot. Interim: stop polling RoCE hardware counters on Broadcom adapters, which costs you fabric observability — usually a worse trade than patching.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-71092"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-01-13"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-71099","cve":"CVE-2025-71099","aliases":[],"title":"Linux kernel (drivers/gpu/drm/xe): The observation-config ioctl dereferences the config object after releasing the lock","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/xe)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"The observation-config ioctl dereferences the config object after releasing the lock that guards its lifetime, so a second thread that guesses the freshly allocated config id and removes it wins the race and frees the object mid-use. The upstream description names the attacker explicitly - this is a deliberate ioctl-versus-ioctl race producing a kernel use-after-free.","attack_vector":"A tenant process holding /dev/dri/renderD* on an Intel xe GPU runs add-config and remove-config in two threads, guessing the id. Reachability is conditional on the container being able to use the xe observation/OA ioctls, which normally means CAP_PERFMON inside the container or a relaxed perf_stream_paranoid setting - if you hand tenants profiling access to the GPU, they have it.","remediation":"Boot a kernel carrying the xe_oa locking fix below. Interim: do not grant CAP_PERFMON to tenant containers, keep kernel.perf_event_paranoid/xe observation defaults strict, and drop GPU profiling access for untrusted workloads.","references":["https://git.kernel.org/stable/c/c6d30b65b7a44dac52ad49513268adbf19eab4a2","https://git.kernel.org/stable/c/7cdb9a9da935c687563cc682155461fef5f9b48d","https://nvd.nist.gov/vuln/detail/CVE-2025-71099"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-71130","cve":"CVE-2025-71130","aliases":[],"title":"Linux i915 GPU kernel driver (execbuffer VMA array): The execbuffer VMA array was not zero-initialised, so","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Linux i915 GPU kernel driver (execbuffer VMA array)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"The execbuffer VMA array was not zero-initialised, so uninitialised kernel stack/heap contents could be acted on or leaked back through the GPU submission path. Info-leak-grade on its own, and useful as the KASLR-defeating first stage for one of the neighbouring i915 use-after-frees.","attack_vector":"Any local user or container with a DRM render node - i.e. any tenant that was scheduled a GPU. No privileged capability needed.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-71130","https://git.kernel.org/stable/c/0336188cc85d0eab8463bd1bbd4ded4e9602de8b"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-01-14"},{"id":"CVE-2025-8747","cve":"CVE-2025-8747","aliases":[],"title":"Keras: Safe-mode bypass in Keras 3.0.0–3.10.0","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Keras","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Safe-mode bypass in Keras 3.0.0–3.10.0 → arbitrary code execution","attack_vector":"Customer-supplied `.keras` file","remediation":"Upgrade past 3.10.0; bypass of the CVE-2025-1550 fix","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-8747"],"status":"curated","fleet":{"ubiquity":"Very common - the bypass of the 3.9.0 fix, so operators who patched once are still exposed","remediation_pain":"**Image rebuild** again - the second round, which is exactly the fleet-wide-patch fatigue pattern","pain_class":"other","why_fleet_wide":"Reuse of internal functionality re-enables arbitrary code execution on `load_model`, proving the model-file-as-code class is not closable by a single patch"},"published":"2025-08-11"},{"id":"CVE-2025-8875","cve":"CVE-2025-8875","aliases":[],"title":"N-able N-central: Deserialization of untrusted data allowing local code execution on the RMM server","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"N-able N-central","year":"2025","cvss_score":7.8,"severity":"high","kev":true,"impact":"Deserialization of untrusted data allowing local code execution on the RMM server","attack_vector":"Local","remediation":"Control-plane: patch to 2025.3.1+; the RMM reaches every managed host","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-8875"],"status":"curated","published":"2025-08-14"},{"id":"CVE-2026-12484","cve":"CVE-2026-12484","aliases":[],"title":"Keras (`TorchModuleWrapper`): Unsafe deserialization of attacker-controlled PyTorch pickle inside a Keras model","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Keras (`TorchModuleWrapper`)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Unsafe deserialization of attacker-controlled PyTorch pickle inside a Keras model","attack_vector":"Customer-supplied model file","remediation":"Upgrade past 3.15.0; cross-framework pickle re-entry defeats Keras' own safe mode","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-12484"],"status":"curated","published":"2026-07-19"},{"id":"CVE-2026-1839","cve":"CVE-2026-1839","aliases":[],"title":"HuggingFace transformers (`Trainer._load_rng_state`): Arbitrary code execution when a training run resumes","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"HuggingFace transformers (`Trainer._load_rng_state`)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Arbitrary code execution when a training run resumes from a checkpoint","attack_vector":"Customer-supplied or poisoned checkpoint directory in shared storage","remediation":"Tenant-owned code; provider must prevent cross-tenant writes to checkpoint directories","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-1839"],"status":"curated","published":"2026-04-07"},{"id":"CVE-2026-23198","cve":"CVE-2026-23198","aliases":[],"title":"Linux KVM - irqfd routing type clobbered on deassign: Deassigning a KVM_IRQFD clobbers the irqfd's copy of the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux KVM - irqfd routing type clobbered on deassign","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Deassigning a KVM_IRQFD clobbers the irqfd's copy of the interrupt routing entry, leaving stale or wrong routing behind. Interrupt routing decides which guest receives which interrupt, so corrupting it on teardown is a cross-VM correctness failure on the interrupt path - and on a GPU host, irqfd is exactly how passed-through accelerator interrupts reach their guest.","attack_vector":"Through the KVM ioctl interface, from the VMM process managing guests - reachable when devices are hot-unplugged or VMs torn down.","remediation":"Fixed in the Linux kernel. Distro kernel update plus host reboot; no firmware step. Relevant to any fleet doing GPU passthrough with dynamic device attach/detach.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-23198"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-02-14"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-23239","cve":"CVE-2026-23239","aliases":[],"title":"Linux kernel (net/xfrm): Closing an ESP-in-TCP socket cancels its transmit work item, but the write-space callback can","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Closing an ESP-in-TCP socket cancels its transmit work item, but the write-space callback can re-schedule that work from softirq immediately afterwards. The worker then runs against a freed espintcp context or a freed socket - a use-after-free on teardown whose timing is supplied by the peer's acknowledgement behaviour. Same defect shape as the kTLS one in CVE-2026-23240, in the IPsec-over-TCP path.","attack_vector":"A local process that attaches the espintcp ULP to a TCP socket (no privilege beyond owning the socket) and closes it with data still in flight. The re-schedule comes from the delayed-ACK handler or ksoftirqd, so a remote peer on the fabric can widen the window by controlling when it acknowledges. Conditional on CONFIG_INET_ESPINTCP. Found by source audit.","remediation":"Boot a kernel carrying the fix commits below (no fixed stable version published). Interim control: disable CONFIG_INET_ESPINTCP on nodes that do not need IPsec-over-TCP so the ULP cannot be attached.","references":["https://git.kernel.org/stable/c/f7ad8b1d0e421c524604d5076b73232093490d5c","https://git.kernel.org/stable/c/664e9df53226b4505a0894817ecad2c610ab11d8","https://nvd.nist.gov/vuln/detail/CVE-2026-23239"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-191","CWE-787"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-23243","cve":"CVE-2026-23243","aliases":[],"title":"Linux kernel InfiniBand user MAD interface (ib_umad, /dev/infiniband/umad*): A process with access to the user MAD","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel InfiniBand user MAD interface (ib_umad, /dev/infiniband/umad*)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A process with access to the user MAD character device can craft a write() whose declared MAD header size and RMPP header length disagree, driving the computed data_len negative. The negative length propagates into ib_create_send_mad(), where the padding calculation overshoots the segment size and produces a slab out-of-bounds memset in alloc_send_rmpp_list(). That is an attacker-influenced heap write from a device node that fabric-management tooling routinely leaves accessible, and the umad path is also the channel that speaks to the subnet manager - so heap control here sits next to the code that configures the IB fabric itself.","attack_vector":"Local write() to /dev/infiniband/umad*. Any container or user that has been granted the umad device - common on nodes running opensm, ibdiagnet, perfquery or any vendor fabric agent - can reach it without extra privilege.","remediation":"Kernel update adding the explicit negative-data_len rejection in ib_umad_write(). Interim: audit which pods and which non-root users actually hold /dev/infiniband/umad* - in most clusters the answer should be 'only the fabric-management daemonset', and tightening that is a same-day change that does not need a reboot.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=1371ef6b1ecf3676b8942f5dfb3634fb0648128e","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-23243.json"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-23429","cve":"CVE-2026-23429","aliases":[],"title":"Linux kernel (drivers/iommu): Unbinding shared virtual addressing touches the mm's IOMMU state after the domain-free","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Unbinding shared virtual addressing touches the mm's IOMMU state after the domain-free path has already dropped the last reference to that mm, so the kernel reads a freed mm_struct. Any tenant process using SVA on an accelerator can turn an ordinary exit into a use-after-free on host memory, with host code execution as the realistic ceiling.","attack_vector":"A tenant process that bound SVA to a device - GPU SVM through /dev/dri/renderD*, or an accelerator via /dev/vfio/* or uacce - simply unbinds or exits. The upstream report was hit through the Intel Xe GPU driver's VM close path, so a container holding a DRM render node is enough. Conditional on PASID/SVA being enabled on the platform and device; no host root.","remediation":"Update to 6.18.20 or later (or your distro's backport of commits 58abeb7b / f5daaa2c). Interim: disable SVA/PASID on tenant nodes (intel_iommu=sm_off on VT-d, or the equivalent driver switch) where the workload does not require shared virtual addressing.","references":["https://git.kernel.org/stable/c/58abeb7b9562f25bdfa2f5ae5ce803eb02e74433","https://git.kernel.org/stable/c/f5daaa2c959d9f894fb5b1ab76da8612dd220a0d","https://nvd.nist.gov/vuln/detail/CVE-2026-23429"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-23856","cve":"CVE-2026-23856","aliases":["DSA-2026-077"],"title":"Dell iDRAC Service Module (iSM) for Windows and Linux: Improper access control in the host-side iDRAC Service Module","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC Service Module (iSM) for Windows and Linux","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Improper access control in the host-side iDRAC Service Module lets a low-privilege local user escalate on the host. iSM is the agent that bridges the operating system to the iDRAC over the internal USB-NIC / passthrough channel, so it is the component that deliberately crosses the host-to-BMC boundary. Escalating through it gets an attacker host-level privilege and puts them next to a channel that talks to the service processor - the pivot direction that turns a tenant workload compromise into an out-of-band one. Affects iSM for Windows before 6.0.3.1 and iSM for Linux before 5.4.1.1.","attack_vector":"Host-side, local, low privilege - a tenant workload or any account on the operating system where iSM is installed. Not reachable from the management VLAN; the exposure is entirely inside the host.","remediation":"Upgrade the iSM package on the host to 6.0.3.1 (Windows) or 5.4.1.1 (Linux) or later. This is a host-side package update and a service restart, not a firmware flash - no host reboot in the normal case, so no job drain. Config-only alternative worth considering on nodes that do not need it: uninstall iSM entirely, or disable the iDRAC host USB-NIC passthrough (iDRAC Settings > OS to iDRAC Pass-through), which removes the host-to-BMC channel at the cost of in-band iDRAC access and some OS-level telemetry.","references":["https://www.dell.com/support/kbdoc/en-us/000426282/dsa-2026-077-security-update-for-dell-idrac-service-module-vulnerability","https://nvd.nist.gov/vuln/detail/CVE-2026-23856"],"status":"curated","fleet":{"pain_class":"daemon-restart"},"published":"2026-02-12"},{"id":"CVE-2026-24141","cve":"CVE-2026-24141","aliases":[],"title":"NVIDIA Model Optimizer: RCE via unsafe deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Model Optimizer","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via unsafe deserialization","attack_vector":"Malicious model artifact","remediation":"Bump ModelOpt; rebuild optimization images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24141","https://github.com/NVIDIA/product-security/tree/main/2026/5798"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2026-03-24"},{"id":"CVE-2026-24149","cve":"CVE-2026-24149","aliases":[],"title":"NVIDIA Megatron-LM: A further script-level code-injection path, filed under the same class as the 2025 set","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Megatron-LM","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A further script-level code-injection path, filed under the same class as the 2025 set. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5712 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24149","https://github.com/NVIDIA/product-security/tree/main/2025/5712"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"],"published":"2026-02-03"},{"id":"CVE-2026-24150","cve":"CVE-2026-24150","aliases":[],"title":"NVIDIA Megatron-LM: Checkpoint loading reaches remote code execution when a user loads a crafted checkpoint","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Megatron-LM","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Checkpoint loading reaches remote code execution when a user loads a crafted checkpoint - the highest-value path in this family, since checkpoints move between organisations routinely. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5769 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24150","https://github.com/NVIDIA/product-security/tree/main/2026/5769"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2026-03-24"},{"id":"CVE-2026-24151","cve":"CVE-2026-24151","aliases":[],"title":"NVIDIA Megatron-LM: The inferencing path reaches remote code execution on crafted input","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Megatron-LM","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"The inferencing path reaches remote code execution on crafted input. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5769 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24151","https://github.com/NVIDIA/product-security/tree/main/2026/5769"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2026-03-24"},{"id":"CVE-2026-24152","cve":"CVE-2026-24152","aliases":[],"title":"NVIDIA Megatron-LM: A second checkpoint-loading remote code execution path","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Megatron-LM","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A second checkpoint-loading remote code execution path. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5769 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24152","https://github.com/NVIDIA/product-security/tree/main/2026/5769"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2026-03-24"},{"id":"CVE-2026-24155","cve":"CVE-2026-24155","aliases":[],"title":"NeMo Framework: RCE via malicious YAML deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NeMo Framework","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via malicious YAML deserialization","attack_vector":"Malicious config artifact","remediation":"Bump NeMo; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24155","https://github.com/NVIDIA/product-security/tree/main/2026/5839"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"],"published":"2026-06-16"},{"id":"CVE-2026-24157","cve":"CVE-2026-24157","aliases":[],"title":"NeMo Framework: RCE via unsafe deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NeMo Framework","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via unsafe deserialization","attack_vector":"Malicious checkpoint","remediation":"Bump NeMo; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24157","https://github.com/NVIDIA/product-security/tree/main/2026/5800"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2026-03-24"},{"id":"CVE-2026-24159","cve":"CVE-2026-24159","aliases":[],"title":"NeMo Framework: RCE via unsafe deserialization on model load","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NeMo Framework","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via unsafe deserialization on model load","attack_vector":"Malicious model","remediation":"Bump NeMo; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24159","https://github.com/NVIDIA/product-security/tree/main/2026/5800"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2026-03-24"},{"id":"CVE-2026-24162","cve":"CVE-2026-24162","aliases":[],"title":"NVIDIA Merlin Transformers4Rec: Improper deserialization of untrusted data reaches code execution and information","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Merlin Transformers4Rec","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Improper deserialization of untrusted data reaches code execution and information disclosure. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5838 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24162","https://github.com/NVIDIA/product-security/tree/main/2026/5838"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2026-05-26"},{"id":"CVE-2026-24165","cve":"CVE-2026-24165","aliases":[],"title":"BioNeMo Framework: RCE via malicious pickled data","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"BioNeMo Framework","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via malicious pickled data","attack_vector":"Malicious model artifact","remediation":"Bump BioNeMo; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24165","https://github.com/NVIDIA/product-security/tree/main/2026/5808"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2026-03-31"},{"id":"CVE-2026-24183","cve":"CVE-2026-24183","aliases":[],"title":"NVIDIA Cumulus Linux: Improper privilege management in the user-management component lets an unprivileged switch user","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Cumulus Linux","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Improper privilege management in the user-management component lets an unprivileged switch user escalate to full switch administration.","attack_vector":"Local on the switch, unprivileged. Anyone with any shell account on the Cumulus switch - including read-only monitoring accounts.","remediation":"Upgrade Cumulus Linux per bulletin 5817. Cost: switch reboot and link flap; sequence leaf-by-leaf so the fabric keeps ECMP paths. Audit who actually has shell on your switches while you are at it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24183","https://github.com/NVIDIA/product-security/tree/main/2026/5817"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-250"],"fleet":{"pain_class":"node-reboot"},"published":"2026-08-18"},{"id":"CVE-2026-24190","cve":"CVE-2026-24190","aliases":[],"title":"GPU Display Driver: Local privesc (insufficient permission checks)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc (insufficient permission checks)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24190","https://github.com/NVIDIA/product-security/tree/main/2026/5821"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-862"],"fleet":{"pain_class":"node-reboot"},"published":"2026-05-26"},{"id":"CVE-2026-24191","cve":"CVE-2026-24191","aliases":[],"title":"GPU Display Driver: Local privesc (synchronization issue)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc (synchronization issue)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24191","https://github.com/NVIDIA/product-security/tree/main/2026/5821"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-367"],"fleet":{"pain_class":"node-reboot"},"published":"2026-05-26"},{"id":"CVE-2026-24192","cve":"CVE-2026-24192","aliases":[],"title":"GPU Display Driver: Local privesc (integer overflow in address calc)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc (integer overflow in address calc)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24192","https://github.com/NVIDIA/product-security/tree/main/2026/5821"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-681"],"fleet":{"pain_class":"node-reboot"},"published":"2026-05-26"},{"id":"CVE-2026-24193","cve":"CVE-2026-24193","aliases":[],"title":"GPU Display Driver: Local privesc (buffer overflow in GPU command processing)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc (buffer overflow in GPU command processing)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24193","https://github.com/NVIDIA/product-security/tree/main/2026/5821"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-reboot"},"published":"2026-05-26"},{"id":"CVE-2026-24194","cve":"CVE-2026-24194","aliases":[],"title":"GPU Display Driver / GPU firmware: Local privesc via improper GPU firmware parameter validation","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver / GPU firmware","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc via improper GPU firmware parameter validation","attack_vector":"Any tenant with a container","remediation":"Driver + GPU firmware upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24194","https://github.com/NVIDIA/product-security/tree/main/2026/5821"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-281"],"fleet":{"pain_class":"firmware-flash"},"published":"2026-05-26"},{"id":"CVE-2026-24216","cve":"CVE-2026-24216","aliases":[],"title":"BioNeMo Framework: RCE via insecure deserialization on model load","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"BioNeMo Framework","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via insecure deserialization on model load","attack_vector":"Malicious model","remediation":"Bump BioNeMo; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24216","https://github.com/NVIDIA/product-security/tree/main/2026/5831"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2026-05-20"},{"id":"CVE-2026-24221","cve":"CVE-2026-24221","aliases":[],"title":"NVTabular: RCE via unsafe pickle deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVTabular","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via unsafe pickle deserialization","attack_vector":"Malicious dataset artifact","remediation":"Bump NVTabular; rebuild data images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24221","https://github.com/NVIDIA/product-security/tree/main/2026/5851"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2026-06-02"},{"id":"CVE-2026-24228","cve":"CVE-2026-24228","aliases":[],"title":"NeMo Framework: RCE via unsafe object deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NeMo Framework","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via unsafe object deserialization","attack_vector":"Malicious config artifact","remediation":"Bump NeMo; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24228","https://github.com/NVIDIA/product-security/tree/main/2026/5839"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2026-06-16"},{"id":"CVE-2026-24237","cve":"CVE-2026-24237","aliases":[],"title":"NVTabular: RCE via insecure deserialization in data processing","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVTabular","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via insecure deserialization in data processing","attack_vector":"Malicious dataset artifact","remediation":"Bump NVTabular; rebuild data images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24237","https://github.com/NVIDIA/product-security/tree/main/2026/5851"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2026-06-02"},{"id":"CVE-2026-24238","cve":"CVE-2026-24238","aliases":[],"title":"TensorRT: Info disclosure / code exec (OOB read in tensor parsing)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Info disclosure / code exec (OOB read in tensor parsing)","attack_vector":"Malicious engine/model file","remediation":"Bump TensorRT; rebuild serving images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24238","https://github.com/NVIDIA/product-security/tree/main/2026/5855"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-129"],"published":"2026-07-14"},{"id":"CVE-2026-24240","cve":"CVE-2026-24240","aliases":[],"title":"Megatron-Bridge: RCE via insecure deserialization of model configs","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Megatron-Bridge","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via insecure deserialization of model configs","attack_vector":"Malicious checkpoint","remediation":"Bump Megatron-Bridge; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24240","https://github.com/NVIDIA/product-security/tree/main/2026/5841"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2026-07-01"},{"id":"CVE-2026-24242","cve":"CVE-2026-24242","aliases":[],"title":"Megatron-Bridge: SSRF in file operations","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Megatron-Bridge","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"SSRF in file operations","attack_vector":"Tenant-supplied checkpoint URI","remediation":"Bump Megatron-Bridge; add egress policy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24242","https://github.com/NVIDIA/product-security/tree/main/2026/5841"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-918"],"published":"2026-07-01"},{"id":"CVE-2026-24243","cve":"CVE-2026-24243","aliases":[],"title":"Megatron-Bridge: RCE via unsafe deserialization in weight loading","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Megatron-Bridge","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via unsafe deserialization in weight loading","attack_vector":"Malicious checkpoint","remediation":"Bump Megatron-Bridge; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24243","https://github.com/NVIDIA/product-security/tree/main/2026/5841"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2026-07-01"},{"id":"CVE-2026-24244","cve":"CVE-2026-24244","aliases":[],"title":"Megatron-Bridge: RCE via deserialization of untrusted model data","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Megatron-Bridge","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via deserialization of untrusted model data","attack_vector":"Malicious checkpoint","remediation":"Bump Megatron-Bridge; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24244","https://github.com/NVIDIA/product-security/tree/main/2026/5841"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2026-07-01"},{"id":"CVE-2026-24245","cve":"CVE-2026-24245","aliases":[],"title":"Megatron-Bridge: RCE via insecure config deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Megatron-Bridge","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via insecure config deserialization","attack_vector":"Malicious checkpoint","remediation":"Bump Megatron-Bridge; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24245","https://github.com/NVIDIA/product-security/tree/main/2026/5841"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2026-07-01"},{"id":"CVE-2026-24246","cve":"CVE-2026-24246","aliases":[],"title":"Megatron-Bridge: Validation bypass via incorrect type comparison","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Megatron-Bridge","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Validation bypass via incorrect type comparison","attack_vector":"Malicious checkpoint","remediation":"Bump Megatron-Bridge; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24246","https://github.com/NVIDIA/product-security/tree/main/2026/5841"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-470"],"published":"2026-07-01"},{"id":"CVE-2026-24247","cve":"CVE-2026-24247","aliases":[],"title":"Megatron-Bridge: RCE via unsafe deserialization in module loading","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Megatron-Bridge","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via unsafe deserialization in module loading","attack_vector":"Malicious checkpoint","remediation":"Bump Megatron-Bridge; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24247","https://github.com/NVIDIA/product-security/tree/main/2026/5841"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2026-07-01"},{"id":"CVE-2026-24248","cve":"CVE-2026-24248","aliases":[],"title":"Megatron-Bridge: RCE via malicious deserialized objects","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Megatron-Bridge","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via malicious deserialized objects","attack_vector":"Malicious checkpoint","remediation":"Bump Megatron-Bridge; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24248","https://github.com/NVIDIA/product-security/tree/main/2026/5841"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-94"],"published":"2026-07-01"},{"id":"CVE-2026-24249","cve":"CVE-2026-24249","aliases":[],"title":"Megatron-Bridge: RCE via unsafe evaluation of loaded config","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Megatron-Bridge","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via unsafe evaluation of loaded config","attack_vector":"Malicious checkpoint","remediation":"Bump Megatron-Bridge; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24249","https://github.com/NVIDIA/product-security/tree/main/2026/5841"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"],"published":"2026-07-01"},{"id":"CVE-2026-24250","cve":"CVE-2026-24250","aliases":[],"title":"NeMo Framework: Command injection in script processing","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NeMo Framework","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Command injection in script processing","attack_vector":"Malicious config artifact","remediation":"Bump NeMo; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24250","https://github.com/NVIDIA/product-security/tree/main/2026/5841"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2026-07-01"},{"id":"CVE-2026-24251","cve":"CVE-2026-24251","aliases":[],"title":"Megatron-Bridge: RCE via deserialization of untrusted objects","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Megatron-Bridge","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via deserialization of untrusted objects","attack_vector":"Malicious checkpoint","remediation":"Bump Megatron-Bridge; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24251","https://github.com/NVIDIA/product-security/tree/main/2026/5841"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2026-07-01"},{"id":"CVE-2026-24252","cve":"CVE-2026-24252","aliases":[],"title":"NVIDIA NeMo Framework: OS command injection reaches code execution with the job's privileges","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NeMo Framework","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"OS command injection reaches code execution with the job's privileges. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5839 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24252","https://github.com/NVIDIA/product-security/tree/main/2026/5839"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-78"],"published":"2026-07-27"},{"id":"CVE-2026-24268","cve":"CVE-2026-24268","aliases":[],"title":"TensorRT: Code exec (OOB write in tensor manipulation)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Code exec (OOB write in tensor manipulation)","attack_vector":"Malicious engine/model file","remediation":"Bump TensorRT; rebuild serving images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24268","https://github.com/NVIDIA/product-security/tree/main/2026/5855"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-122"],"published":"2026-07-14"},{"id":"CVE-2026-24272","cve":"CVE-2026-24272","aliases":[],"title":"TensorRT: Code exec (buffer overflow in tensor dimension handling)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Code exec (buffer overflow in tensor dimension handling)","attack_vector":"Malicious engine/model file","remediation":"Bump TensorRT; rebuild serving images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24272","https://github.com/NVIDIA/product-security/tree/main/2026/5855"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-122"],"published":"2026-07-14"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-120","CWE-787"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2026-25506","cve":"CVE-2026-25506","aliases":["GHSA-r9cr-jf4v-75gh"],"title":"MUNGE (munged credential daemon): This is the root of trust under Slurm. A crafted message with an oversized","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"MUNGE (munged credential daemon)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"This is the root of trust under Slurm. A crafted message with an oversized address-length field overflows a buffer in munged and lets a local attacker read the MAC subkey out of the daemon's memory. With that key they forge MUNGE credentials for any identity, including root, to every service on the cluster that trusts MUNGE - which on a Slurm site is slurmctld, slurmd and slurmdbd. A researcher PoC demonstrated key extraction with ASLR, PIE, NX and RELRO all enabled.","attack_vector":"A local user on any node running munged - i.e. any tenant with a job on a compute node. The forged credentials are then usable cluster-wide, so a foothold on one node becomes authority over the whole scheduler.","remediation":"Upgrade MUNGE to 0.5.18 on every node. Then regenerate the MUNGE key: stop munged cluster-wide, run mungekey --create --force as the munge user on one node, push the new key to all nodes, restart munged. Key regeneration means an all-node munged stop, so running jobs that need to re-authenticate will break - schedule a window. Patching alone is not sufficient if you believe the key was already extracted.","references":["https://github.com/dun/munge/security/advisories/GHSA-r9cr-jf4v-75gh","https://github.com/dun/munge/releases/tag/munge-0.5.18","https://nvd.nist.gov/vuln/detail/CVE-2026-25506"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-27905","cve":"CVE-2026-27905","aliases":[],"title":"BentoML (`safe_extract_tarfile`): Tar extraction escape despite the \"safe\" helper","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"BentoML (`safe_extract_tarfile`)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Tar extraction escape despite the \"safe\" helper","attack_vector":"Customer-supplied bento archive","remediation":"Upgrade to 1.4.36+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-27905"],"status":"curated","published":"2026-03-03"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-31406","cve":"CVE-2026-31406","aliases":[],"title":"Linux kernel (net/xfrm): Flushing xfrm states during namespace cleanup re-arms the NAT-keepalive delayed work after it","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Flushing xfrm states during namespace cleanup re-arms the NAT-keepalive delayed work after it was cancelled, so the work can fire against a struct net that has already been freed and reallocated - a use-after-free in the host kernel driven by ordinary namespace teardown.","attack_vector":"Any workload that creates a network namespace, installs SAs with NAT keepalive, and exits - i.e. every pod delete on a node where containers hold CAP_NET_ADMIN in a user namespace and run IPsec. The race is between cleanup_net rounds, so high netns churn (normal Kubernetes behaviour) widens the window. Conditional on NAT-T keepalives being in use.","remediation":"Boot a kernel carrying the linked stable commits. Interim: drop CAP_NET_ADMIN from tenant containers so per-tenant SAs with keepalives cannot be created, or disable NAT keepalives where NAT is not in the path.","references":["https://git.kernel.org/stable/c/32d0f44c2f14d60fe8e920e69a28c11051543ec1","https://git.kernel.org/stable/c/2255ed6adbc3100d2c4a83abd9d0396d04b87792","https://nvd.nist.gov/vuln/detail/CVE-2026-31406"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-31431","cve":"CVE-2026-31431","aliases":[],"title":"Linux kernel (crypto algif_aead): Incorrect resource transfer between spheres in algif_aead (reverted to out-of-place","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (crypto algif_aead)","year":"2026","cvss_score":7.8,"severity":"high","kev":true,"impact":"Incorrect resource transfer between spheres in algif_aead (reverted to out-of-place operation) - local privilege escalation, actively exploited [KEV]","attack_vector":"Any tenant process in a container (AF_ALG socket access)","remediation":"Livepatchable; otherwise drain + reboot. Compensating control: block AF_ALG in the default container seccomp profile","references":["https://access.redhat.com/security/cve/CVE-2026-31431"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-04-22"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-415","CWE-911"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-31468","cve":"CVE-2026-31468","aliases":[],"title":"Linux kernel (drivers/vfio/pci): The error path of the vfio-pci dma-buf export falls through the whole unwind chain","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/vfio/pci)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"The error path of the vfio-pci dma-buf export falls through the whole unwind chain after the dma-buf was already exported, producing a double free of the allocated objects and an unbalanced reference count on the vfio device itself. A tenant that can steer the export into its error path gets a double free in the kernel - the classic starting point for host privilege escalation - and can drop the device's refcount below what is live.","attack_vector":"A tenant holding /dev/vfio/<group> calls the dma-buf export ioctl while its own file-descriptor table is exhausted, which it controls simply by opening files up to RLIMIT_NOFILE. That makes the 'unlikely' fd-allocation failure a deliberate, repeatable trigger rather than a rare accident. Conditional on the vfio-pci dma-buf export feature being available; no host privilege.","remediation":"Update to a stable kernel carrying commits 83ad334a / e98137f0. Interim: do not enable the vfio-pci dma-buf export feature for tenant devices, and cap per-container file-descriptor limits so the error path is harder to reach on demand.","references":["https://git.kernel.org/stable/c/83ad334afc9a645cef1062f5346526b1e36d6516","https://git.kernel.org/stable/c/e98137f0a874ab36d0946de4707aa48cb7137d1c","https://nvd.nist.gov/vuln/detail/CVE-2026-31468"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-31488","cve":"CVE-2026-31488","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A use-after-free in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the amdgpu display core (DC/DM). Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amd/display: Do not skip unrelated mode changes in DSC validation","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-31488","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-04-22"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-415","CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-31507","cve":"CVE-2026-31507","aliases":[],"title":"Linux kernel (net/smc): Tee(2) duplicates an SMC splice pipe buffer without duplicating the private state hanging off","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/smc)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Tee(2) duplicates an SMC splice pipe buffer without duplicating the private state hanging off it, so both pipes free the same object on release. The result is a double free plus a double sock_put, escalating from a KASAN slab-use-after-free in smc_rx_pipe_buf_release to a NULL dereference and panic - a clean, deterministic heap-corruption primitive for any tenant.","attack_vector":"Local and fully unprivileged, no race required: read from an AF_SMC socket with splice(), tee() the pipe, then close both pipes. socket(AF_SMC, ...) needs no capability and autoloads the smc module via the net-pf-43 alias, so a tenant container with nothing but a shell reaches it - this is the kind of bug that gets weaponised into a container escape.","remediation":"Boot a kernel carrying the fix commits (refcounts the per-buffer private state in the pipe_buf_operations .get handler). Interim: blacklist the smc module (`install smc /bin/false`) or deny socket family 43 in tenant seccomp profiles - this is worth doing pre-emptively on any node that does not deliberately use SMC.","references":["https://git.kernel.org/stable/c/7e8916f46c2f48607f907fd401590093753a6bc5","https://git.kernel.org/stable/c/98ba5cb274768146e25ffbfde47753652c1c20d3","https://nvd.nist.gov/vuln/detail/CVE-2026-31507"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-31516","cve":"CVE-2026-31516","aliases":[],"title":"Linux kernel (net/xfrm): An XFRM_MSG_NEWSPDINFO request queues a per-namespace work item on the global system","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"An XFRM_MSG_NEWSPDINFO request queues a per-namespace work item on the global system workqueue, and the callback recovers its enclosing namespace by pointer arithmetic with nothing holding that namespace alive. Teardown before the work runs leaves the policy-hash rebuild operating on freed namespace memory. Because the attacker chooses both halves - queue the work, then destroy the namespace - this is a controllable use-after-free on freed kernel objects, i.e. a container-to-host escalation primitive.","attack_vector":"Deterministic from a tenant container that holds CAP_NET_ADMIN in its own user+network namespace: send XFRM_MSG_NEWSPDINFO to set policy hash thresholds, then immediately exit the namespace, and repeat. Existing teardown only flushes the policy hash work, not this one. No fabric access, no host root, no device node.","remediation":"Boot a kernel carrying the fix commits below (no fixed stable version published). Interim control: deny CAP_NET_ADMIN inside tenant user namespaces, and where possible disable unprivileged user namespaces so tenants cannot create the namespaces to destroy.","references":["https://git.kernel.org/stable/c/56ea2257b83ee29a543f158159e3d1abc1e3e4fe","https://git.kernel.org/stable/c/8854e9367465d784046362698731c1111e3b39b8","https://nvd.nist.gov/vuln/detail/CVE-2026-31516"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-31566","cve":"CVE-2026-31566","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdgpu): A use-after-free in the amdkfd (KFD compute driver","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdgpu)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the amdkfd (KFD compute driver, /dev/kfd). Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amdgpu: Fix fence put before wait in amdgpu_amdkfd_submit_ib","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-31566","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-04-24"},{"id":"CVE-2026-31656","cve":"CVE-2026-31656","aliases":[],"title":"Linux i915 GPU kernel driver: A use-after-free in the i915 GPU kernel driver. The general shape is that a GPU object is","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux i915 GPU kernel driver","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the i915 GPU kernel driver. The general shape is that a GPU object is freed on one path while another path still holds a reference to it, so a local user with GPU access can get the kernel to read or write freed memory. Exploitability varies by heap layout, but on a GPU node every such bug is reachable from inside a container that was granted /dev/dri - the same boundary that is supposed to separate tenants. Specific trigger: a refcount underflow in the engine heartbeat park path, which frees an engine still referenced.","attack_vector":"Any local user or container with a DRM render node - i.e. any tenant that was scheduled a GPU. No privileged capability needed.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-31656","https://git.kernel.org/stable/c/2af8b200cae3fdd0e917ecc2753b28bb40c876c1"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-04-24"},{"id":"CVE-2026-32305","cve":"CVE-2026-32305","aliases":[],"title":"Traefik: mTLS bypass via SNI pre-sniffing on fragmented ClientHello packets","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Traefik","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"mTLS bypass via SNI pre-sniffing on fragmented ClientHello packets","attack_vector":"Unauthenticated network","remediation":"Rolling Traefik upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-32305"],"status":"curated","published":"2026-03-20"},{"id":"CVE-2026-33298","cve":"CVE-2026-33298","aliases":[],"title":"llama.cpp (`ggml_nbytes`): Integer overflow in the core ggml size calculation","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"llama.cpp (`ggml_nbytes`)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Integer overflow in the core ggml size calculation","attack_vector":"Customer-supplied model file","remediation":"Rebuild past b7824; affects every ggml-based downstream (whisper.cpp, stable-diffusion.cpp)","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-33298"],"status":"curated","published":"2026-03-24"},{"id":"CVE-2026-33744","cve":"CVE-2026-33744","aliases":[],"title":"BentoML (`docker.system_packages`): Command injection through the package list field","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"BentoML (`docker.system_packages`)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Command injection through the package list field","attack_vector":"Customer-supplied build config","remediation":"Upgrade to 1.4.37+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-33744"],"status":"curated","published":"2026-03-27"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-78"],"fleet":{"pain_class":"hot-patch"},"id":"CVE-2026-35043","cve":"CVE-2026-35043","aliases":["GHSA-fgv4-6jr3-jgfw"],"title":"BentoML (cloud deployment path, setup.sh generation in deployment.py): The March fix that added shlex.quote to the","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"BentoML (cloud deployment path, setup.sh generation in deployment.py)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"The March fix that added shlex.quote to the Dockerfile path never reached the cloud deployment code, so system_packages is still interpolated raw into a shell command. The generated setup.sh is uploaded to BentoCloud and executed on the build machine, giving the attacker command execution on the shared cloud builder rather than only on the developer's laptop.","attack_vector":"Anyone who can supply the bentofile a victim deploys - a contributed repo, a shared project, or a compromised branch. Requires the victim to run the cloud deployment flow.","remediation":"Upgrade BentoML to 1.4.38 or later. Treat the whole family of these injection bugs (CVE-2026-33744, CVE-2026-35044, CVE-2026-44345, CVE-2026-44346) as one patch decision and land the newest release rather than the individual minimum fix.","references":["https://github.com/bentoml/BentoML/security/advisories/GHSA-fgv4-6jr3-jgfw","https://nvd.nist.gov/vuln/detail/CVE-2026-35043"],"status":"curated"},{"id":"CVE-2026-35051","cve":"CVE-2026-35051","aliases":[],"title":"Traefik: Authentication bypass in ForwardAuth when trustForwardHeader=false","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Traefik","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Authentication bypass in ForwardAuth when trustForwardHeader=false","attack_vector":"Unauthenticated network","remediation":"Rolling Traefik upgrade to 2.11.43/3.6.14+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-35051"],"status":"curated","published":"2026-04-30"},{"id":"CVE-2026-39858","cve":"CVE-2026-39858","aliases":[],"title":"Traefik: Authentication bypass in ForwardAuth and snippet-based auth middleware","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Traefik","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Authentication bypass in ForwardAuth and snippet-based auth middleware","attack_vector":"Unauthenticated network","remediation":"Rolling Traefik upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-39858"],"status":"curated","published":"2026-04-30"},{"id":"CVE-2026-3989","cve":"CVE-2026-3989","aliases":[],"title":"SGLang (`replay_request_dump.py`): Insecure `pickle.load()` on a `.pkl` dump","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"SGLang (`replay_request_dump.py`)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Insecure `pickle.load()` on a `.pkl` dump","attack_vector":"Customer-supplied dump file replayed by an operator during debugging","remediation":"Upgrade; operator tooling that ingests tenant artifacts is a privilege-escalation path into the provider plane","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-3989"],"status":"curated","published":"2026-03-12"},{"id":"CVE-2026-40912","cve":"CVE-2026-40912","aliases":[],"title":"Traefik: Authentication bypass via StripPrefixRegex middleware","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Traefik","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Authentication bypass via StripPrefixRegex middleware","attack_vector":"Unauthenticated network","remediation":"Rolling Traefik upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-40912"],"status":"curated","published":"2026-04-30"},{"id":"CVE-2026-43128","cve":"CVE-2026-43128","aliases":["RDMA/umem fix double dma_buf_unpin in failure path"],"title":"Linux kernel InfiniBand core dmabuf umem (GPUDirect RDMA path): When mapping a dmabuf-backed RDMA memory region fails","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel InfiniBand core dmabuf umem (GPUDirect RDMA path)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"When mapping a dmabuf-backed RDMA memory region fails, the dmabuf is unpinned immediately but the pinned flag is left set, so release unpins it a second time. This is the dmabuf path that GPUDirect RDMA uses to register GPU memory for the NIC - a double-unpin here corrupts the pinning state of buffers that sit between GPU memory and the fabric, exactly the code an operator most wants to be boring.","attack_vector":"Local, low-privileged - a tenant process registering GPU memory for RDMA via dmabuf, able to induce a mapping failure (resource pressure, invalid parameters).","remediation":"Upgrade the host kernel to 7.0 or a stable backport (6.1.165, 6.6.128, 6.12.75, 6.18.16, 6.19.6). Rolling reboot of GPU nodes using GPUDirect RDMA over dmabuf.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43128","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-43128.json"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-05-06"},{"id":"CVE-2026-43206","cve":"CVE-2026-43206","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): An out-of-bounds access in the amdkfd (KFD compute","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdkfd (KFD compute driver, /dev/kfd) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdkfd: Fix out-of-bounds write in kfd_event_page_set()","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43206","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-05-06"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-832","CWE-667"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-43211","cve":"CVE-2026-43211","aliases":[],"title":"Linux kernel (drivers/pci): The PCI slot-lock failure path releases a lock the caller never took, which at best warns","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/pci)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"The PCI slot-lock failure path releases a lock the caller never took, which at best warns and at worst drops a lock another thread is holding. That thread then runs its slot reset with no serialisation against a second reset or a rescan on the same slot - concurrent unsynchronised reset of shared PCI hierarchy state on a node several tenants are sitting on.","attack_vector":"Local, and reachable from device passthrough: slot-level reset is what a tenant VM triggers through VFIO_DEVICE_PCI_HOT_RESET on /dev/vfio/*, and the racing trylock failure is exactly what you get when two callers contend for the same slot. So two tenants whose devices sit under the same bridge, or one tenant looping resets on its assigned GPU, can drive it. Also reachable by host root through the sysfs reset attributes. A container with no /dev/vfio/* and no PCI sysfs write access cannot reach it.","remediation":"Update to 5.10.252 / 5.15.202 / 6.1.165 / 6.6.128 or later (also fixed in 6.11+ and 4.20/5.5 lines). Interim: stop granting VFIO hot-reset-capable device access to untrusted tenants, and make sure passthrough devices are in single-device IOMMU groups so a tenant reset cannot contend on a slot shared with another tenant.","references":["https://git.kernel.org/stable/c/ebb27b7399ab8b9eb1f792b329aa5f6250c590d4","https://git.kernel.org/stable/c/fbe06a3058114bf95a17a4941b205f4b321c6f0a","https://nvd.nist.gov/vuln/detail/CVE-2026-43211"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-43237","cve":"CVE-2026-43237","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A use-after-free in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the amdgpu GEM/VM/command-submission ioctl surface. Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amdgpu: Refactor amdgpu_gem_va_ioctl for Handling Last Fence Update and Timeline Management v4","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43237","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-05-06"},{"id":"CVE-2026-43260","cve":"CVE-2026-43260","aliases":[],"title":"Linux bnxt_en driver (RSS context delete logic): RSS contexts are not always freed in firmware when the driver deletes","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux bnxt_en driver (RSS context delete logic)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"RSS contexts are not always freed in firmware when the driver deletes them, leaving stale VNIC state on the adapter. Stale receive-steering state on a NIC is worth flagging in a multi-tenant context: RSS contexts and their VNICs determine which queues — and therefore which owner — receives which packets, and leaked contexts are exactly the kind of residue that should not survive a tenant teardown.","attack_vector":"Local, via repeated RSS context create/delete cycles from the host.","remediation":"Kernel/driver upgrade plus host reboot. On bare-metal handoff, a cold power cycle of the NIC clears residual adapter state regardless of driver version — worth doing between tenants anyway.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43260"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-05-06"},{"id":"CVE-2026-43370","cve":"CVE-2026-43370","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): A use-after-free in the amdgpu firmware","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the amdgpu firmware, ACPI and IP-block initialisation. Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amdgpu: Fix use-after-free race in VM acquire","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43370","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-05-08"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-415"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-43494","cve":"CVE-2026-43494","aliases":[],"title":"Linux kernel (net/rds): When pinning user pages for a zerocopy RDS send fails, the pages are released but the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/rds)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"When pinning user pages for a zerocopy RDS send fails, the pages are released but the scatter-list count is left non-zero, so the message purge path walks the stale entries and frees the same pages a second time. A tenant that makes iov_iter_get_pages2() fail on purpose gets a double free of page structures - a strong heap-corruption primitive that can be aimed at memory belonging to other tenants.","attack_vector":"Local and unprivileged: socket(AF_RDS, SOCK_SEQPACKET, 0) - autoloaded via the net-pf-21 alias with no capability check - then sendmsg() with MSG_ZEROCOPY over an iovec crafted so page pinning fails partway (unmapped or unpinnable pages). No RDMA hardware, device node, or capability is needed; the failure is entirely under the tenant's control, which makes it far more attractive than a timing race.","remediation":"Boot a kernel carrying the fix commits (resets op_nents on the zerocopy pin failure path). Interim: blacklist the rds module family (`install rds /bin/false`) or deny socket family 21 in tenant seccomp profiles.","references":["https://git.kernel.org/stable/c/c6e51512a784c4a7b86e1a044988696e3b3721fa","https://git.kernel.org/stable/c/d84ce1786ce40fdd3dd98db47aec5527817e1ef6","https://nvd.nist.gov/vuln/detail/CVE-2026-43494"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-401"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-43502","cve":"CVE-2026-43502","aliases":[],"title":"Linux kernel (net/rds): A zerocopy RDS send that fails after pinning user pages but before the message reaches the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/rds)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A zerocopy RDS send that fails after pinning user pages but before the message reaches the socket queue is cleaned up as if it owned ordinary payload pages, because the purge path infers zerocopy ownership from the socket pointer rather than from the notifier. The pinned-page accounting and the page references are then handled wrongly - the tenant controls when the failure happens, so it controls which pages are mis-released.","attack_vector":"Local and unprivileged: an AF_RDS socket (family 21 autoloads on socket() with no capability check) doing a MSG_ZEROCOPY sendmsg that fails early - before the message is attached to the socket. No RDMA device or privileged access is required, and the failure timing is attacker-chosen rather than racy.","remediation":"Boot a kernel carrying the fix commits (uses op_mmp_znotifier as the cleanup discriminator in rds_message_purge). Interim: blacklist rds/rds_rdma/rds_tcp or deny socket family 21 to tenants.","references":["https://git.kernel.org/stable/c/e9aefdc5c53fe9aed108c14e3d155710a1bb14c9","https://git.kernel.org/stable/c/1e262db7675e27f42c3f3f47d6011855f4454f24","https://nvd.nist.gov/vuln/detail/CVE-2026-43502"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-43627","cve":"CVE-2026-43627","aliases":[],"title":"llama.cpp (`llama_batch_init`): Integer overflow from unchecked multiplication","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"llama.cpp (`llama_batch_init`)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Integer overflow from unchecked multiplication","attack_vector":"Tenant-controlled batch parameters","remediation":"Rebuild past b9058","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43627"],"status":"curated","published":"2026-08-06"},{"id":"CVE-2026-4372","cve":"CVE-2026-4372","aliases":[],"title":"HuggingFace transformers: Critical RCE in all versions before 5.3.0","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"HuggingFace transformers","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Critical RCE in all versions before 5.3.0","attack_vector":"Customer-supplied model repo loaded by `from_pretrained`","remediation":"Rebuild every image with transformers >= 5.3.0. Tenant-pinned versions are outside provider control","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-4372"],"status":"curated","published":"2026-05-24"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-415"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-45852","cve":"CVE-2026-45852","aliases":[],"title":"Linux kernel Soft-RoCE shared receive queue (rdma_rxe, rxe_srq_from_init): If the copy_to_user() that returns the SRQ","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel Soft-RoCE shared receive queue (rdma_rxe, rxe_srq_from_init)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"If the copy_to_user() that returns the SRQ number fails, the queue is freed but the stale pointer is left in srq->rq.queue, and the caller's error path frees it again. A tenant forces the copy to fail by pointing it at an unmapped address - which is entirely under its control - so this is a reliably reachable kernel double free, the classic starting point for heap grooming into privilege escalation on a shared node.","attack_vector":"Local, unprivileged. Create an SRQ on a Soft-RoCE device with a deliberately bad userspace response buffer.","remediation":"Kernel update clearing srq->rq.queue after the cleanup. Blacklist rdma_rxe where not needed.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=0beefd0e15d962f497aad750b2d5e9c3570b66d1","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-45852.json"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-45853","cve":"CVE-2026-45853","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A correctness defect in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu GEM/VM/command-submission ioctl surface reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu: Use kvfree instead of kfree in amdgpu_gmc_get_nps_memranges()","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-45853","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-05-27"},{"cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-362","CWE-908"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-45862","cve":"CVE-2026-45862","aliases":[],"title":"Linux kernel (drivers/iommu/intel): VT-d publishes the address of a freshly allocated PASID table into the PASID","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/intel)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"VT-d publishes the address of a freshly allocated PASID table into the PASID directory before that table's zeroed contents have been flushed to memory the IOMMU can see. In that window non-coherent IOMMU hardware walks whatever stale data occupied the page, treating old bytes as PASID entries - meaning a device translates through page-table roots that are not its own. Uncontrolled DMA from a passthrough device into host or other-tenant memory, which the vendor scores scope-changed.","attack_vector":"Any path that allocates a PASID table: enabling SVA/PASID for a tenant's device, or a VMM attaching a PASID-capable device through vfio/iommufd. Conditional on Intel VT-d scalable mode on a platform where the IOMMU is not cache-coherent for page-table walks. It is a timing window (the vendor vector marks AC:H), not a deterministic primitive.","remediation":"Update to 5.10.252, 5.15.202, or 6.1.165 or later (or a newer stable series carrying the fix). Interim controls: do not enable SVA/PASID for tenant workloads on affected non-coherent VT-d platforms; keep PASID-capable device attach mediated by the host VMM.","references":["https://git.kernel.org/stable/c/cd75e77125c8a51754ca4cd60b4ca083ed735d1d","https://git.kernel.org/stable/c/0616137b70e6d9a547d4b60df8e1b64e36d83661","https://nvd.nist.gov/vuln/detail/CVE-2026-45862"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-45878","cve":"CVE-2026-45878","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): An out-of-bounds access in the amdkfd (KFD compute","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdkfd (KFD compute driver, /dev/kfd) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdkfd: Fix watch_id bounds checking in debug address watch v2","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-45878","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-05-27"},{"cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-362","CWE-665"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-45894","cve":"CVE-2026-45894","aliases":[],"title":"Linux kernel (drivers/iommu/intel): The 512-bit VT-d PASID entry is zeroed all at once while still marked present, and","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/intel)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"The 512-bit VT-d PASID entry is zeroed all at once while still marked present, and the hardware fetches it in several bursts, so the IOMMU can read a mixture of live and cleared fields. A device's PASID context - the thing that decides which address space its DMA lands in - is momentarily undefined while the device is still issuing traffic.","attack_vector":"Runs on PASID teardown: a tenant detaching an SVA context or releasing a PASID-attached device through /dev/vfio/* or /dev/iommu. Conditional on VT-d scalable mode with PASID; the race is between the CPU's zeroing writes and a hardware fetch driven by the tenant's own in-flight DMA, so the attacker controls one side of it. No host root.","remediation":"Update to a stable kernel carrying commits a84d30e8 / 821807c1. Interim: stop tenant DMA before tearing down a PASID context, and disable scalable mode/PASID on nodes that do not need SVA.","references":["https://git.kernel.org/stable/c/a84d30e8d2bacd21782a6481158b7c9c552f4868","https://git.kernel.org/stable/c/821807c167b7b48a41b95b6607c6b9f97600f7d9","https://nvd.nist.gov/vuln/detail/CVE-2026-45894"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-45910","cve":"CVE-2026-45910","aliases":[],"title":"Linux kernel (drivers/infiniband/sw/rxe): The soft-RoCE retransmit and ack timers race against queue-pair destruction","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/sw/rxe)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"The soft-RoCE retransmit and ack timers race against queue-pair destruction: the QP refcount hits zero while a timer handler is still running and then schedules work on the freed object. That is a refcount underflow and use-after-free reachable without any special hardware.","attack_vector":"Any unprivileged process on a node with the rdma_rxe module loaded - soft-RoCE needs no HCA and is commonly available inside containers that were given RDMA access. The attacker creates a QP, keeps retransmit timers armed by stalling or dropping peer acknowledgements from the network side, and destroys the QP concurrently. A fabric peer can help by withholding acks.","remediation":"No fixed version is listed in the record - take the stable kernel carrying 756c93d6df7c (or 3c2ae79fb19d / 5ae9da022ee3) and reboot. Interim: unload and blacklist rdma_rxe on nodes that do not actually need soft-RoCE; this removes the whole attack surface cheaply.","references":["https://git.kernel.org/stable/c/756c93d6df7c3bc599f6590b8e5afead6a41de1c","https://git.kernel.org/stable/c/3c2ae79fb19dfd67341c14f1e78a5f1744eacfe2","https://nvd.nist.gov/vuln/detail/CVE-2026-45910"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-46036","cve":"CVE-2026-46036","aliases":[],"title":"Linux kernel (drivers/vfio/cdx): VFIO_DEVICE_SET_IRQS was not serialized, so two concurrent interrupt-configuration","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/vfio/cdx)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"VFIO_DEVICE_SET_IRQS was not serialized, so two concurrent interrupt-configuration ioctls can have one caller operating on the MSI interrupt array while the other frees it. That is a use-after-free reached directly through the device fd's interrupt-setup ioctl - the textbook shape of a vfio interrupt-index bug, and it lands with tenant-controlled timing.","attack_vector":"Two threads inside a tenant container holding the vfio device fd issue VFIO_DEVICE_SET_IRQS concurrently - one enabling MSI, one disabling. No host privilege and no guest needed. Hardware-conditional: vfio-cdx binds to the AMD/Xilinx CDX bus on Versal-class SoCs, so this is not reachable on x86 or standard Arm GPU nodes. Carry it as a pattern to check for in the PCI SET_IRQS path rather than as a live exposure on a GPU fleet.","remediation":"Update to a stable kernel carrying commits ddf96e23 / 7b436ade on any CDX-bus platform. No action needed on x86/Arm GPU nodes that do not build or load vfio-cdx; confirm with lsmod that vfio-cdx is absent.","references":["https://git.kernel.org/stable/c/ddf96e23c366c566283fce8377928851fa7f5e81","https://git.kernel.org/stable/c/7b436ade16cc81095d79b79f8efa3af0a4f5c5a2","https://nvd.nist.gov/vuln/detail/CVE-2026-46036"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-415"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-46053","cve":"CVE-2026-46053","aliases":[],"title":"Linux kernel RDS RDMA (memory-region cleanup on cookie copy failure): Once __rds_rdma_map() has handed the","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel RDS RDMA (memory-region cleanup on cookie copy failure)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Once __rds_rdma_map() has handed the scatter-gather list and pinned pages to the transport, ownership has moved - but the error path taken when copying the resulting cookie back to userspace fails would unpin and free them again, while the normal teardown path also frees them. The tenant forces the copy failure by supplying an unmapped destination, so this is an on-demand double free of pinned DMA memory: pages that a device may still be able to address get returned to the allocator and handed to someone else.","attack_vector":"Local, unprivileged. Issue an RDS RDMA map request with a deliberately invalid userspace output pointer.","remediation":"Kernel update removing the duplicate unpin/free from the put_user() failure branch. Blacklist rds/rds_rdma if unused.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=033370ffb3c9c0264d19f8ba9ef769523266589a","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-46053.json"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-46116","cve":"CVE-2026-46116","aliases":[],"title":"Linux kernel (net/xfrm): SA deletion decided whether to unhash from the by-SPI and by-sequence chains using field","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"SA deletion decided whether to unhash from the by-SPI and by-sequence chains using field values rather than actual list membership, so a state can be unlinked twice or left linked after free. The reporter clusters nine distinct KASAN signatures on the same slab object, including out-of-bounds and use-after-free writes reached from SA lookup and SPI allocation. That is a corruptible kernel heap object sitting on the IPsec state hash chains, which are walked by the packet decrypt path.","attack_vector":"Needs the ability to churn xfrm states - create, allocate SPIs for, and delete SAs concurrently. That is host root or a tenant container with CAP_NET_ADMIN in its own user+network namespace; the syzkaller reproduction runs entirely from netlink. The corrupted chains are the same ones __xfrm_state_lookup walks for every inbound ESP packet, so a peer sending fabric traffic can help land the freed object on a hot path.","remediation":"Boot a kernel carrying the fix commits below (no fixed stable version published; the report is against 6.12.y stable and reproduces on mainline). Interim control: deny CAP_NET_ADMIN in tenant user namespaces so tenants cannot drive the SA create/delete churn.","references":["https://git.kernel.org/stable/c/6b4dc3181b4bfc5f5fc33ab33b1dc6e15759f4b6","https://git.kernel.org/stable/c/3943fcad7694a7d0b15aeabe7d3cc2a2eb8e92e8","https://nvd.nist.gov/vuln/detail/CVE-2026-46116"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-20","CWE-787"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-46117","cve":"CVE-2026-46117","aliases":[],"title":"Linux kernel (drivers/infiniband/hw/mana): The userspace ABI lets a tenant point several work queues at the same","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/hw/mana)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"The userspace ABI lets a tenant point several work queues at the same completion queue, which the driver only flagged with a warning before continuing on to corrupt kernel state. A tenant gets kernel memory corruption from a legal-looking QP creation request rather than a rejection.","attack_vector":"A tenant container holding /dev/infiniband/uverbs* on an Azure MANA node creates an RSS QP whose work queues share one completion queue. No fabric peer or host root required.","remediation":"No fixed version is listed in the record - take the stable kernel carrying 9cc0c6b1ba8c (or 9ef65af26b2a / db991ba50087) and reboot. Interim: drop /dev/infiniband/* from untrusted containers on MANA nodes.","references":["https://git.kernel.org/stable/c/9cc0c6b1ba8cd5c55aef043e1384de0a8b4efa71","https://git.kernel.org/stable/c/9ef65af26b2a6738bf15812042e84b3112402d3a","https://nvd.nist.gov/vuln/detail/CVE-2026-46117"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787","CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-46145","cve":"CVE-2026-46145","aliases":[],"title":"Linux kernel (drivers/infiniband/hw/mana): The RSS hash-key length arrived from the userspace ABI structure and went","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/hw/mana)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"The RSS hash-key length arrived from the userspace ABI structure and went straight into a memcpy with no bounds check, so a tenant chooses how many bytes the kernel copies into a fixed destination. That is an attacker-controlled kernel heap overwrite - the cleanest container-to-host escalation shape in this batch.","attack_vector":"A tenant container holding /dev/infiniband/uverbs* on an Azure MANA node creates an RSS QP with an oversized rx_hash_key_len. No fabric peer, no host root, single syscall.","remediation":"No fixed version is listed in the record - take the stable kernel carrying 7d7c9f0fcd19 (or 11c1431d641e / 012796f9541f) and reboot. Interim: remove /dev/infiniband/* from untrusted containers on MANA nodes.","references":["https://git.kernel.org/stable/c/7d7c9f0fcd19c4d2f0164347c58d49cafa961b72","https://git.kernel.org/stable/c/11c1431d641e0e4e0529e96957995820600c7287","https://nvd.nist.gov/vuln/detail/CVE-2026-46145"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-415"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-46176","cve":"CVE-2026-46176","aliases":[],"title":"Linux kernel (drivers/infiniband/hw/mlx5): If the second of the two device-wide shared SRQs fails to allocate, the","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/hw/mlx5)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"If the second of the two device-wide shared SRQs fails to allocate, the error path frees the first one but still publishes both the freed pointer and an error pointer as the device's shared resources. The lock-free fast path then treats the device as initialised forever, so every later QP creation - by any tenant on that adapter - dereferences freed memory or an error pointer, and teardown double-frees. One transient allocation failure permanently converts the shared mlx5 device into a use-after-free generator for the whole node.","attack_vector":"Needs the shared-SRQ allocation to fail once, which a tenant can push for with memory or device-resource pressure from inside its container. After that the corruption is reached by ordinary ibv_create_qp calls from any container holding /dev/infiniband/uverbs*. mlx5 is the NIC in most GPU clusters, so this is not a niche driver.","remediation":"Update to 6.6.140 or later, or a stable kernel carrying a13c2ac4d480 / bc2cf5935b46, and reboot. Interim: none that is real - once the device state is poisoned only a reboot clears it, so patch and drain rather than mitigate.","references":["https://git.kernel.org/stable/c/a13c2ac4d480b734342c6fbf8249fc48afd675f3","https://git.kernel.org/stable/c/bc2cf5935b4665172235341163315905197ae91d","https://nvd.nist.gov/vuln/detail/CVE-2026-46176"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-46181","cve":"CVE-2026-46181","aliases":[],"title":"Linux kernel (drivers/net/ethernet/mellanox/mlx4): Shared receive queue objects are looked up from an asynchronous","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/net/ethernet/mellanox/mlx4)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Shared receive queue objects are looked up from an asynchronous event handler under RCU, but nothing frees them with RCU and the handler can run against an SRQ that is only half-constructed or already gone. A tenant that creates and destroys SRQs while events land gets a use-after-free on a kernel object - the memory-corruption primitive is inside RDMA connection state, reachable without host privileges.","attack_vector":"A tenant container holding /dev/infiniband/uverbs* on a ConnectX-3 (mlx4) adapter can create and tear down SRQs in a loop while a remote peer on the fabric drives SRQ async events (limit-reached, catastrophic error) against those queues. Requires the mlx4_ib/mlx4_core stack in use; no host root and no VFIO passthrough needed.","remediation":"Update to a kernel carrying the fix on your stream. Interim: withhold /dev/infiniband/* from untrusted tenants on mlx4-based nodes, or retire ConnectX-3 hardware from multi-tenant duty.","references":["https://git.kernel.org/stable/c/1e2a44875b6afb4add1115f7f3351dcbeb6f273d","https://git.kernel.org/stable/c/8b7833f3bce35cb0d01c1503781523c099c675f0","https://nvd.nist.gov/vuln/detail/CVE-2026-46181"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-46197","cve":"CVE-2026-46197","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): An out-of-bounds access in the amdkfd (KFD compute","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdkfd (KFD compute driver, /dev/kfd) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdkfd: validate SVM ioctl nattr against buffer size","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-46197","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-05-28"},{"id":"CVE-2026-46263","cve":"CVE-2026-46263","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Fix out-of-bounds stream encoder index v3","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-46263","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-06-03"},{"id":"CVE-2026-46311","cve":"CVE-2026-46311","aliases":[],"title":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu/userq): A correctness defect in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu/userq)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu user-mode queues (doorbell submission path) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/userq: fix access to stale wptr mapping","attack_vector":"Local. Reachable by any process with a render node open that can create user-mode queues - the normal ROCm submission path, reachable from an unprivileged container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-46311","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-06-08"},{"id":"CVE-2026-47472","cve":"CVE-2026-47472","aliases":[],"title":"TensorRT-LLM: RCE via insecure deserialization on model load","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT-LLM","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via insecure deserialization on model load","attack_vector":"Malicious model artifact","remediation":"Bump TensorRT-LLM; rebuild serving images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47472","https://github.com/NVIDIA/product-security/tree/main/2026/5840"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2026-07-14"},{"id":"CVE-2026-47749","cve":"CVE-2026-47749","aliases":[],"title":"stable-diffusion.cpp: Memory-safety flaw in model loading","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"stable-diffusion.cpp","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Memory-safety flaw in model loading","attack_vector":"Customer-supplied diffusion model file","remediation":"Rebuild; inherits the ggml parser class","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47749"],"status":"curated","published":"2026-06-16"},{"id":"CVE-2026-48020","cve":"CVE-2026-48020","aliases":[],"title":"Traefik: StripPrefix middleware allows route-level authentication bypass","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Traefik","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"StripPrefix middleware allows route-level authentication bypass","attack_vector":"Unauthenticated network","remediation":"Rolling Traefik upgrade to 2.11.48/3.6.19/3.7.3+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-48020"],"status":"curated","published":"2026-06-23"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-281","CWE-732"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-52908","cve":"CVE-2026-52908","aliases":[],"title":"Linux kernel RDMA core (ib_umem / IB_MR_REREG_ACCESS re-registration): An RDMA memory region registered read-only can","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel RDMA core (ib_umem / IB_MR_REREG_ACCESS re-registration)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"An RDMA memory region registered read-only can be re-registered writable without the kernel ever re-pinning the underlying pages for write. The pin taken at first registration lacked FOLL_WRITE, so the NIC ends up with a writable rkey/lkey pointing at pages the kernel still believes are read-only - copy-on-write mappings, and page-cache pages backing shared files. This is a Dirty-COW-shaped primitive delivered over the RDMA verbs path: a tenant process inside a container with /dev/infiniband/uverbs* mapped in can write to memory it only ever had read access to, including page-cache pages backing binaries shared with other workloads on the node.","attack_vector":"Local to the node, but 'local' in a GPU cluster means any tenant container that was given the RDMA device nodes - which is every container that does RDMA, since GPUDirect and NCCL over IB require /dev/infiniband to be mapped in. No CAP_SYS_ADMIN, no CAP_NET_ADMIN. Sequence is deterministic: register an MR read-only, reregister with IB_MR_REREG_ACCESS adding write, then RDMA-write through it.","remediation":"Kernel update carrying the ib_umem_check_rereg() helper and the per-driver call sites (mlx5, mlx4, hns, irdma). There is no configuration mitigation short of removing RDMA device access from tenant containers, which turns off GPUDirect RDMA and collapses collective bandwidth - so on a training fleet this is a drain-and-reboot, not a tunable.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=09dc18894148381d3bfc550083b1236043870dce","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-52908.json"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-125","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-52935","cve":"CVE-2026-52935","aliases":[],"title":"Linux kernel (net/xfrm): ESP-in-TCP keeps a single in-flight transmit. For a blocking caller the flush of that state","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"ESP-in-TCP keeps a single in-flight transmit. For a blocking caller the flush of that state can report success while the previous partial send is still live, so the socket rebuilds its scatter-gather message over state the old transfer still owns. Upstream calls this a memory-safety fix explicitly: a stale offset carried into a fresh message produces an out-of-bounds read on the send path, leaking adjacent kernel memory into the encrypted stream or faulting.","attack_vector":"A local process that attaches the espintcp ULP to a TCP socket - setsockopt(TCP_ULP, \"espintcp\"), which needs nothing beyond owning the socket, so any unprivileged tenant process qualifies - and then issues blocking sendmsg calls that only partially complete. Partial completion is easy to force by shrinking the send buffer or by having a slow peer. Conditional on CONFIG_INET_ESPINTCP being enabled in the kernel, which it is in most distro builds.","remediation":"Boot a kernel carrying the fix commits below (no fixed stable version published). Interim control: build or ship kernels without CONFIG_INET_ESPINTCP where IPsec-over-TCP is not needed, which removes the ULP entirely.","references":["https://git.kernel.org/stable/c/6564e9c7af7e1dc7bfe7f3093b728abe484d7630","https://git.kernel.org/stable/c/1777ceac4bea5e568a5ad44b7f9bb219c1db21b6","https://nvd.nist.gov/vuln/detail/CVE-2026-52935"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-52987","cve":"CVE-2026-52987","aliases":[],"title":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu): A double free in the amdgpu user-mode","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A double free in the amdgpu user-mode queues (doorbell submission path). The same allocation is released twice, corrupting the slab allocator's freelist. This is a classic heap-corruption primitive: with slab grooming it becomes arbitrary kernel memory write and therefore host compromise from an unprivileged GPU workload. The cheap outcome is a node panic. Upstream fix: drm/amdgpu: avoid double drm_exec_fini() in userq validate","attack_vector":"Local. Reachable by any process with a render node open that can create user-mode queues - the normal ROCm submission path, reachable from an unprivileged container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-52987","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-06-24"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-668"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-53077","cve":"CVE-2026-53077","aliases":[],"title":"Linux kernel RDS over InfiniBand (use outside the initial network namespace): The RDS/IB transport was never written to","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel RDS over InfiniBand (use outside the initial network namespace)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"The RDS/IB transport was never written to be namespace-aware, but nothing stopped it being used from a non-initial namespace - so containerised workloads could drive an RDMA transport whose internal state assumes a single global namespace. The fix is the honest one: forbid it outright. For an operator the finding is that RDS/IB reachable from a container was never a supported configuration, and any cluster where tenants can open RDS sockets has been running the transport outside its design envelope.","attack_vector":"Local, unprivileged. Any container able to create an RDS socket over an IB device.","remediation":"Kernel update restricting RDS/IB to the init namespace. Blacklist rds_rdma and rds now - it is almost certainly unused on a GPU cluster, and unloading it is immediate.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=07035306bf722f4676a1aee35cbeb3732c76194e","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-53077.json"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-190","CWE-787"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-53133","cve":"CVE-2026-53133","aliases":[],"title":"Linux kernel (drivers/infiniband/core): When the IOMMU coalesces a large memory registration into one block spanning","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/core)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"When the IOMMU coalesces a large memory registration into one block spanning several scatter entries, the umem block iterator reassembles it with 32-bit arithmetic and computes wrong DMA addresses for everything past the 4GB boundary. The HCA is then programmed to DMA into physical memory the tenant never registered, so a remote peer's RDMA writes against a legitimate rkey land in someone else's pages.","attack_vector":"Reachable by any tenant container holding /dev/infiniband/uverbs* that registers a memory region larger than 4GB (ibv_reg_mr) on a node where the IOMMU linearises the mapping - which is ordinary behaviour for GPU training jobs with large pinned buffers. No fabric peer is needed to create the bad mapping; once it exists, any RDMA peer writing to that rkey writes to the wrong physical memory.","remediation":"No fixed version is listed in the record - take the stable kernel carrying commit 2ff4b7817e5b (or the backports dee2a49adeeb / cc644d5608e3) and reboot the node. There is no realistic interim control: capping registrations below 4GB is not workable for GPU workloads, so patch rather than mitigate.","references":["https://git.kernel.org/stable/c/2ff4b7817e5b78070c30f5fb5e678e452a2628b3","https://git.kernel.org/stable/c/dee2a49adeeb2a5e16a3fc858fa21b841c519802","https://nvd.nist.gov/vuln/detail/CVE-2026-53133"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-53136","cve":"CVE-2026-53136","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Clamp VBIOS HDMI retimer register count to array size","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53136","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-06-25"},{"id":"CVE-2026-53137","cve":"CVE-2026-53137","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Clamp HDMI HDCP2 rx_id_list read to buffer size","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53137","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-06-25"},{"id":"CVE-2026-53143","cve":"CVE-2026-53143","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): An out-of-bounds access in the amdkfd (KFD compute","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdkfd (KFD compute driver, /dev/kfd) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdkfd: Fix buffer overflow in SDMA queue checkpoint/restore on GFX11","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53143","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-06-25"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-53239","cve":"CVE-2026-53239","aliases":[],"title":"Linux kernel (net/xfrm): Policy deletion dropped the policy lock before pruning the inexact-policy bin, and a","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Policy deletion dropped the policy lock before pruning the inexact-policy bin, and a concurrent policy-hash rebuild can free that bin in the gap. The delete then walks and prunes freed memory - a use-after-free in the structure that decides which IPsec policy a flow matches. Corrupting the inexact policy tables is the worst place to get memory damage in this subsystem, because the same tables decide whether a tenant's traffic is encrypted at all.","attack_vector":"Two concurrent netlink operations, both of which need CAP_NET_ADMIN: XFRM_MSG_DELPOLICY on one thread and XFRM_MSG_NEWSPDINFO on another. A tenant container holding CAP_NET_ADMIN in its own user+network namespace can run both in a tight loop against its own policy set; no fabric access and no host root are required. The upstream commit publishes the exact interleaving.","remediation":"Boot a kernel carrying the fix commits below (no fixed stable version published). Interim control: remove CAP_NET_ADMIN from tenant user namespaces so tenant workloads cannot issue XFRM policy netlink at all.","references":["https://git.kernel.org/stable/c/8fc536e9f6856230f19c7d13e71af064b6a77b22","https://git.kernel.org/stable/c/c4c1ea36d83bf3c4569468ca5b8b614fda1bf821","https://nvd.nist.gov/vuln/detail/CVE-2026-53239"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-53622","cve":"CVE-2026-53622","aliases":[],"title":"Traefik: HTTP/3 QUIC TLS configuration selection lets clients bypass router-specific mTLS enforcement","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Traefik","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"HTTP/3 QUIC TLS configuration selection lets clients bypass router-specific mTLS enforcement","attack_vector":"Unauthenticated network","remediation":"Rolling Traefik upgrade to 3.7.3+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53622"],"status":"curated","published":"2026-06-23"},{"id":"CVE-2026-58659","cve":"CVE-2026-58659","aliases":[],"title":"PyTorch Lightning (`_load_state`): RCE by importing and executing classes named in the checkpoint","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"PyTorch Lightning (`_load_state`)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE by importing and executing classes named in the checkpoint","attack_vector":"Customer-supplied checkpoint","remediation":"No format-level fix — the checkpoint format is the vulnerability. Enforce safetensors","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-58659"],"status":"curated","published":"2026-07-15"},{"cwe":["CWE-787"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-63794","cve":"CVE-2026-63794","aliases":[],"title":"Linux kernel (arch/x86/kvm/svm): The SEV debug-encrypt path bounds each iteration by the source page offset but not the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (arch/x86/kvm/svm)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"The SEV debug-encrypt path bounds each iteration by the source page offset but not the destination offset, so the PSP and a following memcpy write past the end of a single-page kernel buffer - KASAN reports a 4095-byte out-of-bounds write. That is host kernel heap corruption, i.e. root on the node, driven by whoever drives the SEV VM.","attack_vector":"Reachable by any process that can open /dev/kvm and create an SEV VM, via KVM_MEMORY_ENCRYPT_OP with a debug-encrypt request whose destination offset exceeds the source offset. It is not the guest that triggers it - it is the process owning the VM, so this matters wherever tenants hold /dev/kvm (nested virtualization exposed, or bare-metal tenants). AMD SEV hosts running kvm_amd only.","remediation":"Update to a kernel with the referenced stable commits. Interim: do not expose /dev/kvm to tenants, and disable SEV on nodes where you cannot patch promptly (kvm_amd sev=0).","references":["https://git.kernel.org/stable/c/f701ae476cb92a3a3d8844bb39bb63b4512684c8","https://git.kernel.org/stable/c/64f2449841ffc7d203183aa4c748c9c77951ecc5","https://nvd.nist.gov/vuln/detail/CVE-2026-63794"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-63840","cve":"CVE-2026-63840","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu/jpeg): A correctness defect in the amdgpu kernel driver core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu/jpeg)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu kernel driver core reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/jpeg: set no_user_fence for JPEG v5.3.0 ring","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63840","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-07-19"},{"id":"CVE-2026-63841","cve":"CVE-2026-63841","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/jpeg): A correctness defect in the amdgpu RAS / GPU","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/jpeg)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/jpeg: set no_user_fence for JPEG v5.0.1 ring","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63841","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-07-19"},{"id":"CVE-2026-63842","cve":"CVE-2026-63842","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu/jpeg): A correctness defect in the amdgpu kernel driver core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu/jpeg)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu kernel driver core reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/jpeg: set no_user_fence for JPEG v5.0.0 ring","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63842","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-07-19"},{"id":"CVE-2026-63843","cve":"CVE-2026-63843","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu/jpeg): A correctness defect in the amdgpu kernel driver core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu/jpeg)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu kernel driver core reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/jpeg: set no_user_fence for JPEG v4.0.5 ring","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63843","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-07-19"},{"id":"CVE-2026-63844","cve":"CVE-2026-63844","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu/jpeg): A correctness defect in the amdgpu kernel driver core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu/jpeg)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu kernel driver core reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/jpeg: set no_user_fence for JPEG v4.0.3 ring","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63844","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-07-19"},{"id":"CVE-2026-63845","cve":"CVE-2026-63845","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu/jpeg): A correctness defect in the amdgpu kernel driver core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu/jpeg)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu kernel driver core reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/jpeg: set no_user_fence for JPEG v4.0 ring","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63845","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-07-19"},{"id":"CVE-2026-63846","cve":"CVE-2026-63846","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu/jpeg): A correctness defect in the amdgpu kernel driver core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu/jpeg)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu kernel driver core reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/jpeg: set no_user_fence for JPEG v3.0 ring","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63846","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-07-19"},{"id":"CVE-2026-63847","cve":"CVE-2026-63847","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu/jpeg): A correctness defect in the amdgpu kernel driver core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu/jpeg)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu kernel driver core reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/jpeg: set no_user_fence for JPEG v2.5 ring","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63847","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-07-19"},{"id":"CVE-2026-63848","cve":"CVE-2026-63848","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu/jpeg): A correctness defect in the amdgpu kernel driver core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu/jpeg)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu kernel driver core reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/jpeg: set no_user_fence for JPEG v2.0 ring","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63848","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-07-19"},{"id":"CVE-2026-63849","cve":"CVE-2026-63849","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn): A correctness defect in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/vcn: set no_user_fence for VCN v5.0.1 enc ring","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63849","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-07-19"},{"id":"CVE-2026-63850","cve":"CVE-2026-63850","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn): A correctness defect in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/vcn: set no_user_fence for VCN v5.0.0 enc ring","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63850","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-07-19"},{"id":"CVE-2026-63851","cve":"CVE-2026-63851","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn): A correctness defect in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/vcn: set no_user_fence for VCN v4.0.5 enc ring","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63851","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-07-19"},{"id":"CVE-2026-63852","cve":"CVE-2026-63852","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn): A correctness defect in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/vcn: set no_user_fence for VCN v4.0.3 enc ring","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63852","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-07-19"},{"id":"CVE-2026-63853","cve":"CVE-2026-63853","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn): A correctness defect in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/vcn: set no_user_fence for VCN v4.0 enc ring","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63853","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-07-19"},{"id":"CVE-2026-63854","cve":"CVE-2026-63854","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn): A correctness defect in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/vcn: set no_user_fence for VCN v3.0 enc/dec rings","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63854","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-07-19"},{"id":"CVE-2026-63855","cve":"CVE-2026-63855","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn): A correctness defect in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/vcn: set no_user_fence for VCN v2.5 enc/dec rings","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63855","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-07-19"},{"id":"CVE-2026-63856","cve":"CVE-2026-63856","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn): A correctness defect in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/vcn: set no_user_fence for VCN v2.0 enc/dec rings","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63856","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-07-19"},{"id":"CVE-2026-63879","cve":"CVE-2026-63879","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A correctness defect in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu GEM/VM/command-submission ioctl surface reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu: fix amdgpu_hmm_range_get_pages","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63879","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-07-19"},{"id":"CVE-2026-63881","cve":"CVE-2026-63881","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): An out-of-bounds access in the amdkfd (KFD compute","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdkfd (KFD compute driver, /dev/kfd) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdkfd: fix a vulnerability of integer overflow in kfd debugger","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63881","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-07-19"},{"id":"CVE-2026-63884","cve":"CVE-2026-63884","aliases":[],"title":"Linux i915 GPU kernel driver: A use-after-free in the i915 GPU kernel driver. The general shape is that a GPU object is","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux i915 GPU kernel driver","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the i915 GPU kernel driver. The general shape is that a GPU object is freed on one path while another path still holds a reference to it, so a local user with GPU access can get the kernel to read or write freed memory. Exploitability varies by heap layout, but on a GPU node every such bug is reachable from inside a container that was granted /dev/dri - the same boundary that is supposed to separate tenants. Specific trigger: the TTM buffer object being swapped out from under a purge, so the purge frees the wrong object.","attack_vector":"Any local user or container with a DRM render node - i.e. any tenant that was scheduled a GPU. No privileged capability needed.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63884","https://git.kernel.org/stable/c/073bcbc95e9648c976da1654c7590a8d6ee12c2d"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-07-19"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-415"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-63911","cve":"CVE-2026-63911","aliases":[],"title":"Linux kernel (net/xfrm): Cloning an IPTFS security association kmemdups the mode data, so the clone shares the original","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Cloning an IPTFS security association kmemdups the mode data, so the clone shares the original SA's skb queue, hrtimers, spinlock and in-flight reassembly state. If migration fails before re-init, destroying the clone splices and frees skbs still owned by the original SA - use-after-free and double-free of packet buffers belonging to a live encrypted association.","attack_vector":"Driven by XFRM_MSG_MIGRATE (or SA clone during migration) over xfrm netlink, requiring CAP_NET_ADMIN in the network namespace - available to the node's IKE daemon and to any container granted NET_ADMIN with its own netns. Conditional on IPTFS-mode SAs (RFC 9347 aggregation) being in use, and the corruption is most reachable when the SA has packets queued, which an attacker arranges by migrating under load.","remediation":"Boot a kernel carrying the linked stable commits. Interim: avoid IPTFS mode SAs, do not run SA migration on nodes using IPTFS, and drop CAP_NET_ADMIN from tenant containers.","references":["https://git.kernel.org/stable/c/9327252e04626d4bb02ca8c0c108fbe8eabf0c5a","https://git.kernel.org/stable/c/dfb9f6cbfa9826655a49698cf90eb800fce2178e","https://nvd.nist.gov/vuln/detail/CVE-2026-63911"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-362","CWE-787"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-64005","cve":"CVE-2026-64005","aliases":[],"title":"Linux kernel (net/smc): The SMC socket hashtables are re-initialised at the end of module init, after the protocol and","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/smc)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"The SMC socket hashtables are re-initialised at the end of module init, after the protocol and socket family have already been registered. Sockets created in that window get their hash-list heads zeroed out from under them, leaving a corrupted list the kernel keeps walking and writing - memory corruption seeded at module load time.","attack_vector":"Local and unprivileged, and the trigger is the module autoload itself: socket(AF_SMC, ...) from an unprivileged process makes the kernel request the smc module through the net-pf-43 alias with no capability check, and a second thread racing socket() calls against that load lands in the window between sock_register() and the hashtable re-init. A tenant container can arrange this deliberately.","remediation":"Boot a kernel carrying the fix commits (drops the redundant INIT_HLIST_HEAD calls). Interim: pre-load the smc module at boot on nodes that need it so no tenant can race a cold autoload, or blacklist it outright (`install smc /bin/false`) where SMC is unused.","references":["https://git.kernel.org/stable/c/cdc79c05cc375f68ae87b0c74fdaac1a5c93155a","https://git.kernel.org/stable/c/64c96e497d5ada0b90e99bf58f893aa2b73dcfbc","https://nvd.nist.gov/vuln/detail/CVE-2026-64005"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-191","CWE-787"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-64009","cve":"CVE-2026-64009","aliases":[],"title":"Linux kernel (net/xfrm): An unprivileged user who can create IPsec SAs turns one outbound datagram into a multi-exabyte","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"An unprivileged user who can create IPsec SAs turns one outbound datagram into a multi-exabyte memset past the end of a socket buffer. In practice that is instant kernel heap destruction and a node panic; with more care over the SA parameters it is a controllable kernel write, i.e. a path from a tenant container to full host compromise on a shared GPU node.","attack_vector":"Reachable by any principal that can install an xfrm state: host root, or - as the upstream report describes - a 'nobody' user inside a container that holds CAP_NET_ADMIN in its own user+network namespace, which is normal for CNI plugins, VPN sidecars and any pod running with NET_ADMIN. The recipe is entirely local: add an IPv4 ESP tunnel SA with a long truncated auth key, set a tiny interface MTU and a large XFRMA_TFCPAD, then send one UDP datagram. No fabric access and no cooperating peer are needed.","remediation":"Boot a kernel carrying the fix commits below (the CVE record publishes no fixed stable version, so match by commit against your vendor kernel). Interim control: stop granting CAP_NET_ADMIN inside tenant user namespaces, and block the XFRM netlink family from tenant workloads that do not genuinely need to program IPsec.","references":["https://git.kernel.org/stable/c/8014f70c4e6e5ab101ae3860a614e65e988372e3","https://git.kernel.org/stable/c/1021d2877b689a648b27815c854557a917122e93","https://nvd.nist.gov/vuln/detail/CVE-2026-64009"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-64097","cve":"CVE-2026-64097","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Validate GPIO pin LUT table size before iterating","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-64097","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-07-19"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-415","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-64581","cve":"CVE-2026-64581","aliases":[],"title":"Linux kernel (net/xfrm): Setting a per-socket IPsec policy reset the socket's destination cache non-atomically while","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Setting a per-socket IPsec policy reset the socket's destination cache non-atomically while the UDP transmit fast path resets the same field with an atomic exchange. Racing the two drops the single reference twice and frees the xfrm destination bundle while it is still in use. The upstream report shows a KASAN use-after-free write with a working exploit binary in the trace - a freed-object write primitive in the kernel, which on a shared GPU node means a tenant escaping its container onto the host.","attack_vector":"The commit states it plainly: reachable by an unprivileged user via a user+network namespace. Concretely, a tenant process creates a connected UDP socket, loops IP_XFRM_POLICY setsockopt against a concurrent sendmsg, and wins the race. No CAP_NET_ADMIN on the host, no fabric access, no device node - just an unprivileged process inside a container with unprivileged user namespaces enabled, which is the default on most container hosts.","remediation":"Boot a kernel carrying the fix commits below (the version list in the record marks introduction points, not fixes - match by commit). Interim control: disable unprivileged user namespaces on shared nodes (kernel.unprivileged_userns_clone=0 or user.max_user_namespaces=0 for tenant cgroups), which removes the stated reachability path.","references":["https://git.kernel.org/stable/c/96b678d08268b5f5c6fc99d4289d9b7e334fc683","https://git.kernel.org/stable/c/c283e9ada7fcb7dd4b10592623086b2e6d2f9925","https://nvd.nist.gov/vuln/detail/CVE-2026-64581"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-64582","cve":"CVE-2026-64582","aliases":[],"title":"Linux kernel Soft-RoCE mmap path (rdma_rxe, rxe_mmap vs concurrent DESTROY_CQ): Rxe_mmap() removes the mmap-info object","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel Soft-RoCE mmap path (rdma_rxe, rxe_mmap vs concurrent DESTROY_CQ)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Rxe_mmap() removes the mmap-info object from the pending list and drops the lock while its kref is still 1, then calls remap_vmalloc_range(), which walks page tables with no lock held. A concurrent DESTROY_CQ ioctl on another CPU drops the last reference, vfree()s the object mid-walk and frees the tracking struct. The tenant controls both threads, so this is a deterministic-enough race giving a use-after-free plus page-table manipulation on memory being torn down - the strongest primitive class in this driver.","attack_vector":"Local, unprivileged. One tenant thread mmaps a Soft-RoCE completion queue while another destroys it.","remediation":"Kernel update holding the reference across the remap. Blacklist rdma_rxe where software RoCE is not needed.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=3371f2036e0f970166bbb624e25bba46d32fa18e","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-64582.json"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-65600","cve":"CVE-2026-65600","aliases":[],"title":"Traefik: Authentication bypass via path traversal in ReplacePathRegex","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Traefik","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Authentication bypass via path traversal in ReplacePathRegex","attack_vector":"Unauthenticated network","remediation":"Rolling Traefik upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-65600"],"status":"curated","published":"2026-07-22"},{"id":"CVE-2026-68104","cve":"CVE-2026-68104","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): A NULL pointer dereference in the amdgpu kernel driver core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A NULL pointer dereference in the amdgpu kernel driver core. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: invoke pm_genpd_remove() before freeing genpd","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68104","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-10"},{"id":"CVE-2026-68106","cve":"CVE-2026-68106","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): A division by zero in the amdgpu firmware","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A division by zero in the amdgpu firmware, ACPI and IP-block initialisation, reachable with attacker-influenced parameters. The kernel takes a divide fault and the node goes down, taking every co-resident job with it. Upstream fix: drm/amdgpu: fix division by zero with invalid uvd dimensions","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68106","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-10"},{"id":"CVE-2026-68236","cve":"CVE-2026-68236","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A use-after-free in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the amdgpu display core (DC/DM). Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amd/display: set new_stream to NULL after release","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68236","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-10"},{"id":"CVE-2026-68245","cve":"CVE-2026-68245","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A use-after-free in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the amdgpu GEM/VM/command-submission ioctl surface. Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amdgpu: fix lifetime issue of amdgpu_vm_get_task_info_pasid()","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68245","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-08-10"},{"id":"CVE-2026-68257","cve":"CVE-2026-68257","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): An out-of-bounds access in the amdkfd (KFD compute","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdkfd (KFD compute driver, /dev/kfd) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdkfd: fix 32-bit overflow in CWSR total size calculation","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68257","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-08-10"},{"id":"CVE-2026-68273","cve":"CVE-2026-68273","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): A use-after-free in the amdgpu RAS / GPU reset and","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the amdgpu RAS / GPU reset and recovery path. Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amdgpu: Fix context pstate override handling","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68273","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-08-10"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-68290","cve":"CVE-2026-68290","aliases":[],"title":"Linux kernel (net/rds): Network-namespace teardown frees the per-netns RDS/TCP listen socket before unregistering the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/rds)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Network-namespace teardown frees the per-netns RDS/TCP listen socket before unregistering the matching sysctl table, and the sysctl handler derives its namespace from that socket. A write to the sysctl racing namespace exit dereferences the freed socket - KASAN confirms a slab-use-after-free, giving a tenant that owns a namespace a host-kernel corruption primitive.","attack_vector":"Local, and reachable by a tenant rather than only by host root wherever unprivileged user namespaces are enabled: a container that creates its own user+network namespace holds CAP_NET_ADMIN inside it, so it can write /proc/sys/net/rds/tcp/* and simultaneously tear the namespace down. On hosts with unprivileged userns disabled this needs real root in a netns. The rds_tcp module autoloads from an unprivileged socket(AF_RDS, ...).","remediation":"Update to 6.12.101 or later on that branch, or any kernel carrying the fix commits (unregisters the sysctl table before killing the listen socket). Interim: blacklist rds/rds_tcp, or set kernel.unprivileged_userns_clone=0 / user.max_user_namespaces=0 for tenant workloads that do not need namespaces.","references":["https://git.kernel.org/stable/c/80fffed08dc1c10e971066941d2daa56253f1552","https://git.kernel.org/stable/c/3aa13fe0c1bb7bc5312f878e61523e5d8cf3f85d","https://nvd.nist.gov/vuln/detail/CVE-2026-68290"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-668","CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-68335","cve":"CVE-2026-68335","aliases":[],"title":"Linux kernel RDS (rds_find_bound socket lookup ignores network namespace): This is a literal cross-tenant delivery bug.","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel RDS (rds_find_bound socket lookup ignores network namespace)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"This is a literal cross-tenant delivery bug. RDS looks sockets up in one global hash table keyed only on address, port and scope - the network namespace is not part of the key - so a sender in namespace A delivers a message to a socket living in namespace B. Container isolation on Linux is network namespaces; a protocol whose demultiplexing ignores them is not isolating anything. The memory-safety consequence follows: the received message points at a connection owned by namespace A, and when that namespace is torn down the connection is freed while the surviving socket in namespace B still references it.","attack_vector":"Local, unprivileged. A tenant creates a network namespace (or is given one, as every container is), binds an RDS socket on an address that collides with another namespace's, and sends.","remediation":"Kernel update making namespace part of the RDS bind lookup. Immediate and effective: blacklist the rds and rds_rdma modules. RDS is rarely used outside specific Oracle database deployments and is almost never required on a GPU cluster, so removing it is cheap and needs no reboot.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=03c574112e5d066df0ddce36d7438e850bcf3050","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-68335.json"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-362","CWE-908"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-68417","cve":"CVE-2026-68417","aliases":[],"title":"Linux kernel (drivers/infiniband/sw/siw): Soft-iWARP published a new queue pair into the lookup table before its","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/sw/siw)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Soft-iWARP published a new queue pair into the lookup table before its queues, completion-queue pointers and state were set up, so a concurrent QP-number lookup reaches a half-built object and follows uninitialised pointers. An attacker who wins the window gets kernel memory corruption from unprivileged userspace.","attack_vector":"An unprivileged process on a node with the siw module loaded creates a QP in one thread while a second thread - or an inbound iWARP segment from a fabric peer that resolves the freshly allocated QP number - looks it up. No hardware RDMA adapter needed, which is exactly why siw is dangerous to leave loaded in tenant containers.","remediation":"No fixed version is listed in the record - take the stable kernel carrying 3c9d12821996 (or 3ff82e3841ec / 36e91a58397c) and reboot. Interim: unload and blacklist siw on nodes that do not need soft-iWARP.","references":["https://git.kernel.org/stable/c/3c9d128219964dcea897bf6139b88242e987be8f","https://git.kernel.org/stable/c/3ff82e3841ecab1ff38d5817c969a019d266c83c","https://nvd.nist.gov/vuln/detail/CVE-2026-68417"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-20","CWE-863"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-68419","cve":"CVE-2026-68419","aliases":[],"title":"Linux kernel (drivers/infiniband/hw/irdma): The pseudo memory regions that back a QP/CQ/SRQ have no real hardware key","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/hw/irdma)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"The pseudo memory regions that back a QP/CQ/SRQ have no real hardware key, but the core still exposes them as normal MRs. A tenant that calls re-register on one drives a control-plane command against key 0 - an unvalidated operation on adapter-global state that other tenants on the same NIC depend on.","attack_vector":"A tenant container holding /dev/infiniband/uverbs* on an Intel irdma node calls ibv_rereg_mr on the MR it registered for its own QP/CQ/SRQ buffers. Purely local to the tenant; no fabric peer or host root.","remediation":"Update to 6.6.148 or later, or a stable kernel carrying fb46d134e1b8 / b5029e91c634, and reboot. Interim: drop /dev/infiniband/* from untrusted containers on irdma nodes.","references":["https://git.kernel.org/stable/c/fb46d134e1b8690bed2da9005b36d32d2efd34ac","https://git.kernel.org/stable/c/b5029e91c63406e4f4c8d58161048b41b6f0bd8c","https://nvd.nist.gov/vuln/detail/CVE-2026-68419"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-72072","cve":"CVE-2026-72072","aliases":[],"title":"Linux kernel mlx5_core MACsec offload: Deleting an offloaded MACsec RX secure channel frees the per-SC metadata_dst","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_core MACsec offload","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Deleting an offloaded MACsec RX secure channel frees the per-SC metadata_dst with a call that ignores the reference count, while the RX datapath is concurrently taking a reference on it under RCU - a use-after-free reachable from the packet path. The kernel CNA notes it is reachable by cloud tenants with SR-IOV VFs, containers, or user/network namespaces without init-namespace root, so this is a container-to-host kernel corruption on nodes doing link-layer encryption.","attack_vector":"A tenant with an SR-IOV VF, a container with network-namespace capability, or a local user in a user namespace - combined with MACsec RX secure channel churn on an mlx5 interface.","remediation":"Upgrade the host kernel to 7.2 or a stable backport (6.1.178, 6.6.145, 6.12.97, 6.18.40, 7.1.5). Rolling reboot. Interim: restrict unprivileged user namespaces and do not delegate MACsec configuration into tenant namespaces.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-72072","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-72072.json"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-08-15"},{"cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-459"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-72287","cve":"CVE-2026-72287","aliases":[],"title":"Linux kernel (arch/x86/kvm/vmx): The nested vTPR versus TPR-threshold consistency check ran only after KVM had already","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (arch/x86/kvm/vmx)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"The nested vTPR versus TPR-threshold consistency check ran only after KVM had already loaded state, and the failure path did not unwind vmcs01.GUEST_CR3 back to KVM's own value. With EPT disabled, L1 then runs with a CR3 the guest chose rather than the shadow page-table root KVM installed - the guest picks the page tables the CPU walks, which is the shortest path there is to reading and writing host memory from inside a VM.","attack_vector":"Guest-driven: the tenant executes VMLAUNCH/VMRESUME with a vmcs12 whose tpr_threshold fails the check. Needs three things to line up - nested VMX exposed to the guest, shadow paging in use (EPT off or unavailable), and the off-by-default early consistency check enabled (kvm_intel.nested_early_check=1) - which is why the CNA rates complexity high.","remediation":"Update to a stable kernel with the linked fix (no fixed release enumerated; take the branch carrying commit ebdac7554abb). Interim controls: keep EPT enabled (kvm_intel.ept=1), leave kvm_intel.nested_early_check off, and do not expose nested virtualization to tenants.","references":["https://git.kernel.org/stable/c/ebdac7554abb347ca4197be241116842161acd9b","https://git.kernel.org/stable/c/7d066368f72e6192af7e21c5817626f6da666991","https://nvd.nist.gov/vuln/detail/CVE-2026-72287"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-72449","cve":"CVE-2026-72449","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A use-after-free in the amdkfd (KFD compute driver","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the amdkfd (KFD compute driver, /dev/kfd). Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amdkfd: fix list_del corruption in kfd_criu_resume_svm","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-72449","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-08-15"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-20","CWE-125"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-72450","cve":"CVE-2026-72450","aliases":[],"title":"Linux kernel (net/xfrm): Xfrm_selector_match() compared selectors without checking that the selector family matches the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Xfrm_selector_match() compared selectors without checking that the selector family matches the flow family or that the prefix length fits the address size. An AF_UNSPEC selector carrying a 128-bit prefix length matched against an IPv4 flow shifts out of bounds and reads off the stack. Beyond the memory-safety read, this is a policy-matching bug: a selector can be made to match flows it has no business matching, which is how one tenant's IPsec policy ends up deciding whether another tenant's traffic is encrypted, dropped, or sent in clear.","attack_vector":"Reachable by any principal that can install an xfrm policy and then send a packet - host root, or a tenant container holding CAP_NET_ADMIN in its own user+network namespace, which is common for pods running a CNI or VPN sidecar. The syzbot reproducer is a policy add with an AF_UNSPEC selector plus ordinary traffic; no fabric peer and no device node are involved.","remediation":"Update to 5.10.261 / 5.15.212 / 6.1.178 / 6.6.145 or later on those stable series, or a vendor kernel carrying the fix commits below. Interim control: do not grant CAP_NET_ADMIN inside tenant user namespaces, and audit installed policies for AF_UNSPEC selectors with oversized prefix lengths.","references":["https://git.kernel.org/stable/c/87a5bbccc7ff4edb3f42fea387124237d2ba91ee","https://git.kernel.org/stable/c/bd7f202cf77556cff59f68dc30e4cdf40cb6e33b","https://nvd.nist.gov/vuln/detail/CVE-2026-72450"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-763"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-74296","cve":"CVE-2026-74296","aliases":[],"title":"NVIDIA/Mellanox ConnectX driver (mlx5_ib user access region index release): The driver released the software-side UAR","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA/Mellanox ConnectX driver (mlx5_ib user access region index release)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"The driver released the software-side UAR index rather than the index the hardware actually handed back. User access regions are the doorbell pages the driver maps into each tenant's address space; returning the wrong index to the allocator desynchronises the allocator from the hardware, so an index still owned by one context can be handed to the next one that asks. Two tenants sharing a doorbell page is a direct isolation failure on the ConnectX adapter, not merely a leak.","attack_vector":"Local. Occurs on the ordinary allocate/free cycle of RDMA user contexts, so a tenant that repeatedly creates and destroys contexts drives the desynchronisation.","remediation":"Kernel update freeing the hardware-provided UAR index. Nothing to tune - patch and reboot. The kernel CNA record is terse on exploitation detail; treat the doorbell-aliasing consequence described here as the operator-facing reading of the fix, and confirm against your own driver version before ranking it.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=449ae7927152e46acbe5f19f97eafdae6d3a96b1","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-74296.json"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-74297","cve":"CVE-2026-74297","aliases":["RDMA/mlx5 fix undefined shift of user RQ WQE size"],"title":"Linux kernel mlx5_ib (queue pair sizing): set_rq_size() computes the receive-queue work-entry size as 1 << rq_wqe_shift","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_ib (queue pair sizing)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"set_rq_size() computes the receive-queue work-entry size as 1 << rq_wqe_shift from a user-supplied shift that is only checked against being greater than 32, so shifts of 31 and 32 pass and overflow. A tenant process controls a shift that the kernel then uses to size an allocation - the classic setup for heap corruption from an RDMA verbs call.","attack_vector":"Local, low-privileged - any process that can create an RDMA queue pair on an mlx5 device.","remediation":"Upgrade the host kernel to 7.2 or a stable backport (5.10.261, 5.15.212, 6.1.178, 6.6.145, 6.12.97, 6.18.40, 7.1.5). Rolling reboot across the RDMA fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-74297","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-74297.json"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-15"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-74306","cve":"CVE-2026-74306","aliases":[],"title":"Linux kernel (drivers/vfio/pci/qat): Two concurrent writes to the QAT VF migration-resume file both pass the bounds","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/vfio/pci/qat)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Two concurrent writes to the QAT VF migration-resume file both pass the bounds check against a stale file offset, then copy past the end of the kernel migration-state buffer. Whoever can write that fd gets a controlled kernel heap overflow with attacker-supplied bytes - a straightforward path from a device-holder to host privilege.","attack_vector":"Whoever holds the vfio migration file descriptor for a QAT VF: the VMM restoring a migrated guest, or a tenant handed the device directly. Two threads writing the fd concurrently is the whole exploit. Conditional on the qat_vfio_pci variant driver being bound to an Intel QuickAssist VF - not reachable on nodes that do not pass through QAT VFs.","remediation":"Update to a stable kernel carrying commits 6465af00 / d416dcef. Interim: if you do not live-migrate QAT VFs, bind them to plain vfio-pci instead of qat_vfio_pci, or stop passing QAT VFs to tenants until the kernel is patched.","references":["https://git.kernel.org/stable/c/6465af0004dc1b067129a26ef44f19cdf13bbce6","https://git.kernel.org/stable/c/d416dcefdbac90d96b22485fd93f28229ad9984b","https://nvd.nist.gov/vuln/detail/CVE-2026-74306"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-362","CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-74334","cve":"CVE-2026-74334","aliases":[],"title":"Linux kernel (drivers/infiniband/core): Because re-registration can swap the protection domain behind a memory region","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/core)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Because re-registration can swap the protection domain behind a memory region, every place the resource-tracking netlink interface reached through mr->pd was racing. A tenant that re-registers an MR while the operator's tooling walks RDMA resources makes the host dereference a stale PD pointer - kernel memory corruption triggered from inside a container against host-context code.","attack_vector":"A tenant container holding /dev/infiniband/uverbs* calls ibv_rereg_mr in a loop; the racing reader is the host side running `rdma resource show mr` or any monitoring agent that polls nldev. Both halves are routine, so an operator with RDMA telemetry collection is exposed continuously rather than only under attack.","remediation":"Update to a stable kernel carrying 1a132ee4e655 (or 50d5c02ab8e6) and reboot. Interim: stop polling nldev MR resources from monitoring agents on nodes running untrusted tenants until patched.","references":["https://git.kernel.org/stable/c/1a132ee4e655288d9a0937ea5109a0d038431ae9","https://git.kernel.org/stable/c/50d5c02ab8e62325548bd3a6e6b758a9dcd6e7c3","https://nvd.nist.gov/vuln/detail/CVE-2026-74334"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-74357","cve":"CVE-2026-74357","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): An out-of-bounds access in the amdgpu RAS / GPU","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu RAS / GPU reset and recovery path - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: fix KASAN slab-out-of-bounds in amdgpu_coredump ring dump","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-74357","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-08-15"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-367","CWE-125"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-74377","cve":"CVE-2026-74377","aliases":[],"title":"Linux kernel Soft-RoCE responder (rdma_rxe, non-SRQ receive WQE handling): A textbook time-of-check-to-time-of-use","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel Soft-RoCE responder (rdma_rxe, non-SRQ receive WQE handling)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A textbook time-of-check-to-time-of-use across the userspace/kernel RDMA boundary. For non-SRQ queue pairs the responder read work-queue-entry fields - including num_sge and the SGE array itself - directly out of the queue buffer that is mapped shared into the tenant's address space. The tenant flips num_sge or an SGE length after the kernel validates it and before it uses it, producing out-of-bounds reads in rxe_resp_check_length() and copy_data(). The attacker controls both the trigger and the timing, and the read lands wherever the forged SGE points.","attack_vector":"Local, unprivileged: a tenant with a Soft-RoCE device races its own shared receive-queue memory against the kernel responder while inbound traffic is being processed.","remediation":"Kernel update introducing get_recv_wqe(), which validates num_sge and copies the WQE into a kernel-private buffer before use - the same discipline the SRQ path already had. If Soft-RoCE is not actually needed (it usually is not on nodes with real ConnectX hardware), blacklisting rdma_rxe removes the entire rxe surface without a reboot and is the fastest real mitigation.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=2e60378fb3c8b51c94103bb40014c4fe38fa5033","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-74377.json"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-74378","cve":"CVE-2026-74378","aliases":["RDMA/rxe get_srq_wqe TOCTOU","Soft-RoCE shared receive queue heap overflow"],"title":"Linux kernel - RDMA/rxe (Soft-RoCE) responder, drivers/infiniband/sw/rxe/rxe_resp.c: The shared receive queue buffer is","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel - RDMA/rxe (Soft-RoCE) responder, drivers/infiniband/sw/rxe/rxe_resp.c","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"The shared receive queue buffer is mapped into userspace, and get_srq_wqe() validates the num_sge field from that buffer against max_sge but then re-reads the same field to size the memcpy. A local process racing a second thread flips num_sge between the check and the copy and overflows the kernel heap. On a multi-tenant node this is a container-to-kernel escalation reachable by any tenant given an RDMA device - which, for RDMA to be useful at all, is every tenant. Full loss of confidentiality, integrity and availability on the host, meaning escape from the container onto the GPU node and everything else scheduled on it.","attack_vector":"Local, low privilege. The attacker maps the SRQ, posts a work queue entry, and has a concurrent thread rewrite num_sge in the shared mapping in the window between validation and use. Classic time-of-check-to-time-of-use on a userspace-writable structure the kernel reads twice. Requires the rdma_rxe module to be loaded and an rxe device present.","remediation":"Host reboot / kernel upgrade. Strong interim mitigation: rxe (Soft-RoCE) is a software RoCE emulation used for development and for nodes without a real RNIC - on a GPU cluster with ConnectX or BlueField hardware it is almost never needed. Unload and blacklist rdma_rxe (config change, no downtime) and this and the other rxe findings in this database all disappear at once. Audit which nodes have it loaded; it is frequently pulled in by test tooling and left behind.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2026/CVE-2026-74378.json","https://nvd.nist.gov/vuln/detail/CVE-2026-74378"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-08-15"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787","CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-74390","cve":"CVE-2026-74390","aliases":[],"title":"Linux kernel (drivers/infiniband/hw/irdma): The page-address copy loop only honoured its bound when the bound was","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/hw/irdma)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"The page-address copy loop only honoured its bound when the bound was non-zero, so in the flat (level 0) case it kept writing until it ran out of user DMA blocks - past the end of a fixed four-entry array. A tenant that registers more pages than it declared gets a controlled kernel out-of-bounds write, which is a container-to-host escalation primitive.","attack_vector":"A tenant container holding /dev/infiniband/uverbs* on an Intel E810/irdma node creates a CQ, QP or SRQ declaring a small page count (e.g. req.cq_pages) while supplying a user memory region made of more DMA blocks. Entirely local to the tenant - no fabric peer, no host root.","remediation":"Update to a stable kernel carrying 4780f58672ee (or 79a20a8e201a / 9f8f0d2099e3) and reboot. Interim: remove /dev/infiniband/* from untrusted containers on irdma nodes, or blacklist the irdma module where RDMA is not required.","references":["https://git.kernel.org/stable/c/4780f58672ee6328accd54a95f9c00683477e499","https://git.kernel.org/stable/c/79a20a8e201a779224b4bf115250a7713bde72c0","https://nvd.nist.gov/vuln/detail/CVE-2026-74390"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-74446","cve":"CVE-2026-74446","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A race condition or locking defect in the amdkfd (KFD","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A race condition or locking defect in the amdkfd (KFD compute driver, /dev/kfd). Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdkfd: hold event_mutex while checkpointing CRIU events","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-74446","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-15"},{"id":"CVE-2026-74447","cve":"CVE-2026-74447","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): An out-of-bounds access in the amdkfd (KFD compute","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdkfd (KFD compute driver, /dev/kfd) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdkfd: fix uint32_t overflow in EOP ring buffer size alignment","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-74447","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-08-15"},{"id":"CVE-2026-74449","cve":"CVE-2026-74449","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Fix divide-by-zero in calculate_mcache_setting on zero viewport","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-74449","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-15"},{"id":"CVE-2026-74450","cve":"CVE-2026-74450","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): A use-after-free in the amdgpu power management","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the amdgpu power management (SMU/powerplay). Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amd/pm: fix pptable use-after-free","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-74450","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-15"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-74563","cve":"CVE-2026-74563","aliases":[],"title":"Linux kernel (net/rds): Bind() on an RDS socket with a scoped IPv6 address looks up the interface under RCU, drops the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/rds)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Bind() on an RDS socket with a scoped IPv6 address looks up the interface under RCU, drops the RCU lock, and then passes the bare net_device pointer into the address check. A concurrent interface delete frees the device in that window, so the bind path reads freed memory - KASAN confirms the slab-use-after-free, and the published reproducer is literally named 'exploit'.","attack_vector":"Local and unprivileged: bind an AF_RDS socket to a scoped IPv6 address while an interface is being removed. A tenant with its own network namespace can delete its own veth to supply the race partner, and socket(AF_RDS, ...) autoloads rds/rds_tcp through the net-pf-21 alias with no capability check. No RDMA hardware or device node needed.","remediation":"Boot a kernel carrying the fix commits (keeps the RCU read-side lock held across ipv6_chk_addr). Interim: blacklist the rds module family (`install rds /bin/false`) or deny socket family 21 in tenant seccomp profiles.","references":["https://git.kernel.org/stable/c/f8a8977af2134a1d91e5f9773cb7d9d53278c830","https://git.kernel.org/stable/c/c4933624a6f416ecfcc31ab58d585da1207a0597","https://nvd.nist.gov/vuln/detail/CVE-2026-74563"],"status":"curated","tags":["tenant-isolation"]},{"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2025-019-pytorch-flatbuffer-model-parsing","cve":null,"aliases":["GHSA-g6v3-crfc-cggj"],"title":"PyTorch (flatbuffer model parsing, torch::load / parse_and_initialize_mobile_module): MALICIOUS MODEL FILE TO MEMORY","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"PyTorch (flatbuffer model parsing, torch::load / parse_and_initialize_mobile_module)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"MALICIOUS MODEL FILE TO MEMORY CORRUPTION: loading an untrusted model through torch::load reaches an arbitrary-address write. When the input matches the flatbuffer file format, GetMutableRoot in module.h adds a pointer to a value dereferenced from that same attacker-controlled buffer, so the resulting flatbuffer_module address can be corrupted; the same pattern recurs in flatbuffer_loader.cpp where IndirectHelper::Read applies an attacker-supplied offset. The corrupted pointer chain ends up in func, which parseFunction then writes through. Found by fuzzing with sydr-fuzz. The operator-facing point is that this sits below the pickle discussion everyone knows about: teams that switched to non-pickle model formats specifically to avoid arbitrary code execution still have a memory-safety parser in the load path, and in a GPU cluster the process doing the loading is a training or serving pod with the accelerator, the weights and a cluster identity.","attack_vector":"Local / artifact delivery: the victim process calls torch::load on an attacker-supplied file in flatbuffer format. Reaching the victim usually means write access to a model registry, artifact store or dataset mount that the cluster loads from.","remediation":"Upgrade PyTorch to 2.1.0 or later and restart the workloads that load models. Treat model files as untrusted input regardless of format: restrict who can write to model registries and artifact stores per tenant, verify artifact signatures before load, and run model-loading pods with a scoped service account so a corruption foothold does not inherit cluster credentials.","references":["https://github.com/pytorch/pytorch/security/advisories/GHSA-g6v3-crfc-cggj"],"status":"curated"},{"id":"CVE-2020-15114","cve":"CVE-2020-15114","aliases":[],"title":"etcd: Gateway can be pointed at itself, causing an infinite loop and control-plane DoS","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"etcd","year":"2020","cvss_score":7.7,"severity":"high","kev":false,"impact":"Gateway can be pointed at itself, causing an infinite loop and control-plane DoS","attack_vector":"Anyone who can influence etcd gateway config or DNS","remediation":"Rolling etcd upgrade; no GPU drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-15114"],"status":"curated","published":"2020-08-06"},{"id":"CVE-2022-24348","cve":"CVE-2022-24348","aliases":[],"title":"Argo CD: Directory traversal via Helm charts discloses credentials from other Applications' value files","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2022","cvss_score":7.7,"severity":"high","kev":false,"impact":"Directory traversal via Helm charts discloses credentials from other Applications' value files; cross-tenant secret leak","attack_vector":"Any user with repo write access","remediation":"Rolling Argo CD upgrade; rotate any credentials stored in chart values","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-24348"],"status":"curated","published":"2022-02-04"},{"id":"CVE-2022-24730","cve":"CVE-2022-24730","aliases":[],"title":"Argo CD: Path traversal plus improper access control in the repo-server","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2022","cvss_score":7.7,"severity":"high","kev":false,"impact":"Path traversal plus improper access control in the repo-server","attack_vector":"Any user with repo access","remediation":"Rolling Argo CD upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-24730"],"status":"curated","published":"2022-03-23"},{"id":"CVE-2022-28183","cve":"CVE-2022-28183","aliases":[],"title":"NVIDIA GPU Display Driver - kernel mode layer (Windows nvlddmkm.sys and Linux nvidia.ko): An out-of-bounds read","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - kernel mode layer (Windows nvlddmkm.sys and Linux nvidia.ko)","year":"2022","cvss_score":7.7,"severity":"high","kev":false,"impact":"An out-of-bounds read in the kernel mode layer leaks kernel memory to an unprivileged local user. Both the Windows and Linux datacenter drivers are affected, so a mixed fleet needs two separate rollouts.","attack_vector":"Local and unprivileged on either OS. On Linux it is reachable from any GPU container via /dev/nvidia*; on Windows from any session holding a GPU handle.","remediation":"Upgrade both the Linux and the Windows datacenter driver branches listed in bulletin 5353. Cost: Linux needs a drain and nvidia.ko reload per node; Windows needs a reboot per node. Two change windows unless your fleet is homogeneous.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28183","https://github.com/NVIDIA/product-security/tree/main/2022/5353"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:H","cwe":["CWE-125"],"fleet":{"pain_class":"node-reboot"},"published":"2022-05-17"},{"id":"CVE-2022-31666","cve":"CVE-2022-31666","aliases":[],"title":"Harbor: Missing permission validation on Webhook policies","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Harbor","year":"2022","cvss_score":7.7,"severity":"high","kev":false,"impact":"Missing permission validation on Webhook policies; view, update and delete another tenant's webhooks","attack_vector":"Authenticated registry user","remediation":"Upgrade Harbor","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31666"],"status":"curated","published":"2024-11-14"},{"id":"CVE-2022-31670","cve":"CVE-2022-31670","aliases":[],"title":"Harbor: Missing permission validation on tag retention policies across projects","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Harbor","year":"2022","cvss_score":7.7,"severity":"high","kev":false,"impact":"Missing permission validation on tag retention policies across projects","attack_vector":"Authenticated registry user","remediation":"Upgrade Harbor","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31670"],"status":"curated","published":"2024-11-14"},{"id":"CVE-2022-42275","cve":"CVE-2022-42275","aliases":[],"title":"DGX servers BMC: Improper access control on BMC","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX servers BMC","year":"2022","cvss_score":7.7,"severity":"high","kev":false,"impact":"Improper access control on BMC","attack_vector":"Network-adjacent","remediation":"Flash BMC 2.09.00+ out-of-band","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42275","https://github.com/NVIDIA/product-security/tree/main/2024/5513"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:U/C:N/I:H/A:H","cwe":["CWE-288","CWE-120"],"published":"2023-01-13"},{"id":"CVE-2023-34335","cve":"CVE-2023-34335","aliases":["AMI-SA-2023005","NVIDIA OSR review"],"title":"AMI MegaRAC SPx 13 (IPMI handler / host SPI flash path): The multi-tenant bare-metal nightmare","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx 13 (IPMI handler / host SPI flash path)","year":"2023","cvss_score":7.7,"severity":"high","kev":false,"impact":"The multi-tenant bare-metal nightmare. An unauthenticated host - meaning code running on the server's own OS with no BMC credentials at all - can write the host's SPI flash through the BMC's IPMI handler, bypassing Secure Boot. A tenant who rents a bare-metal GPU node for an hour can leave a BIOS-resident implant behind that the next tenant inherits, and that no OS reimage, disk wipe or node rebuild will find. For anyone selling bare-metal GPU capacity this is a tenant-isolation break, not just a firmware bug.","attack_vector":"Requires code execution on the tenant/host OS (root or equivalent), then pivots inboard over the host-BMC interface - KCS/LPC or the in-band IPMI channel - to reach the BMC's flash write path. No BMC password, no management-network access, and no physical presence. The BMC VLAN being airtight does not help you here, because the attack comes from the host side.","remediation":"BMC firmware flash to MegaRAC SPx_13.5 or later; out-of-band, per node, ODM-gated. Because the attack path is in-band, network segmentation buys you nothing - the only other lever is disabling the in-band host-to-BMC IPMI interface (KCS) in BIOS setup, which is a BIOS config change plus a reboot and will break in-band ipmitool, node health agents and most vendor management agents running on the host. If you sell bare-metal, treat this as a wipe-and-reflash-BIOS-between-tenants question rather than a patch question.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2023005.pdf","https://nvd.nist.gov/vuln/detail/CVE-2023-34335"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-06-12"},{"id":"CVE-2024-3056","cve":"CVE-2024-3056","aliases":[],"title":"Podman: Crafted container sharing IPC creates unbounded IPC resources in /dev/shm","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Podman","year":"2024","cvss_score":7.7,"severity":"high","kev":false,"impact":"Crafted container sharing IPC creates unbounded IPC resources in /dev/shm; node resource exhaustion","attack_vector":"Any tenant workload sharing an IPC namespace","remediation":"Upgrade Podman; disable shared IPC across tenants","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-3056"],"status":"curated","published":"2024-08-02"},{"id":"CVE-2024-3095","cve":"CVE-2024-3095","aliases":[],"title":"LangChain (Web Research Retriever): SSRF","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LangChain (Web Research Retriever)","year":"2024","cvss_score":7.7,"severity":"high","kev":false,"impact":"SSRF","attack_vector":"Attacker-supplied URL or retrieved content","remediation":"Upgrade; block internal egress","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-3095"],"status":"curated","published":"2024-06-06"},{"id":"CVE-2024-31410","cve":"CVE-2024-31410","aliases":[],"title":"CyberPower PowerPanel managed devices - shared device certificates: Every managed device uses an identical certificate","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"CyberPower PowerPanel managed devices - shared device certificates","year":"2024","cvss_score":7.7,"severity":"high","kev":false,"impact":"Every managed device uses an identical certificate derived from a hardcoded key, so any device can impersonate any other. An attacker who compromises one PDU in one rack can pose as every other device in the estate and feed the DCIM whatever telemetry they like - including telling it everything is fine while a hall overheats, or triggering automated responses that are themselves PHYSICAL actions.","attack_vector":"Anyone who obtains the key - which means anyone who obtains any single managed device, including a unit bought secondhand or pulled from an RMA pile.","remediation":"Vendor firmware and platform upgrade that issues per-device certificates. Until then, the device identity layer provides no assurance and you should not build automated power actions on top of it. Note this also breaks the trust assumption in your decommissioning process: a device leaving your estate carries the fleet key with it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-31410"],"status":"curated","published":"2024-05-15"},{"id":"CVE-2024-8698","cve":"CVE-2024-8698","aliases":[],"title":"Keycloak: SAML signature scope determined by position, not Reference","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Keycloak","year":"2024","cvss_score":7.7,"severity":"high","kev":false,"impact":"SAML signature scope determined by position, not Reference -> signature validation bypass and impersonation","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade the IdP; revalidate every SAML client config","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-8698"],"status":"curated","published":"2024-09-19"},{"id":"CVE-2025-3937","cve":"CVE-2025-3937","aliases":["CVE-2025-3936","CVE-2025-3944","CVE-2025-3945","CVE-2025-3938","CVE-2025-3943"],"title":"Tridium Niagara Framework and Niagara Enterprise Security (before 4.10.11 / 4.14.2 / 4.15.1): A chain, not a single","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Tridium Niagara Framework and Niagara Enterprise Security (before 4.10.11 / 4.14.2 / 4.15.1)","year":"2025","cvss_score":7.7,"severity":"high","kev":false,"impact":"A chain, not a single bug, and the chain is what matters. Password hashes are stored with insufficient computational effort so they crack offline; a missing cryptographic step and an observable response discrepancy help an attacker recover or confirm credentials; incorrect permission assignment on the QNX-based JACE controllers allows file manipulation; and an argument-injection path on QNX turns that into command execution. Put together, an attacker with a foothold anywhere near the station escalates to control of the Niagara supervisor and the JACE field controllers underneath it - which is direct write access to cooling commands and setpoints for the whole site. Niagara Enterprise Security is affected too, so on sites that use it the same chain reaches door control. The operator-facing consequence is loss of thermal control over the hall plus, potentially, loss of the physical access boundary around the cages, from one credential-recovery weakness.","attack_vector":"Requires reaching the Niagara station or capturing its authentication material - so a compromised facilities workstation, a foothold on the building network, an integrator's remote path, or an internet-exposed station. The QNX permission and argument-injection pieces then apply on the JACE hardware controllers themselves, which sit on the facility VLAN and are rarely monitored by anyone.","remediation":"Upgrade the framework to 4.10.11, 4.14.2 or 4.15.1 (or later) per the Tridium/Honeywell tech bulletins. This is a supervisor software upgrade plus a firmware push to every JACE, done by the integrator - a real project, typically a scheduled outage of supervisory control while field controllers keep running standalone. After upgrading, rotate every Niagara credential, because the old hashes are assumed compromised. Compensating controls while you wait: restrict the station to a jump host, enforce MFA on the path to it, and put the JACEs behind an allow-list. If your Niagara estate is integrator-managed under a service contract, the upgrade is a billable engagement - price it now rather than discovering it during an incident.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-3937","https://nvd.nist.gov/vuln/detail/CVE-2025-3945","https://docs.niagara-community.com/category/tech_bull","https://www.honeywell.com/us/en/product-security"],"status":"curated"},{"id":"CVE-2026-24177","cve":"CVE-2026-24177","aliases":[],"title":"KAI Scheduler: Missing authentication on API endpoints","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"KAI Scheduler","year":"2026","cvss_score":7.7,"severity":"high","kev":false,"impact":"Missing authentication on API endpoints -> scheduler manipulation","attack_vector":"Network-adjacent attacker inside the cluster","remediation":"Upgrade KAI Scheduler chart; add network policy in front of the API","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24177","https://github.com/NVIDIA/product-security/tree/main/2026/5818"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:C/C:H/I:N/A:N","cwe":["CWE-306"],"published":"2026-04-21"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:C/C:N/I:N/A:H","cwe":["CWE-129"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2026-40886","cve":"CVE-2026-40886","aliases":["GHSA-5jv8-h7qh-rf5p"],"title":"Argo Workflows (controller pod informer, pod-gc-strategy annotation parsing): A malformed","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo Workflows (controller pod informer, pod-gc-strategy annotation parsing)","year":"2026","cvss_score":7.7,"severity":"high","kev":false,"impact":"A malformed workflows.argoproj.io/pod-gc-strategy annotation makes the pod informer index past the end of a split, and the panic happens in an informer goroutine outside the controller's recover, killing the whole process. The poisoned pod survives restarts, so the controller crash-loops and every tenant's workflow scheduling stops until an operator manually finds and deletes that one pod.","attack_vector":"Anyone who can create or annotate a pod in a namespace the controller watches - which includes any tenant with normal pod-create rights, not just workflow submitters.","remediation":"Upgrade the controller to 3.7.14 or 4.0.5 and restart it. To recover a cluster already in the crash loop, find the pod carrying the malformed annotation and delete it before or during the upgrade, otherwise the new controller crashes on the same object.","references":["https://github.com/argoproj/argo-workflows/security/advisories/GHSA-5jv8-h7qh-rf5p","https://nvd.nist.gov/vuln/detail/CVE-2026-40886"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-56793","cve":"CVE-2026-56793","aliases":[],"title":"Dell OpenManage Server Administrator (improper authentication): An unauthenticated remote attacker gets unauthorized","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Dell OpenManage Server Administrator (improper authentication)","year":"2026","cvss_score":7.7,"severity":"high","kev":false,"impact":"An unauthenticated remote attacker gets unauthorized access to OMSA - the in-band management agent that runs on the server itself, with hardware-level visibility and control.","attack_vector":"Unauthenticated network access to OMSA below 11.1.0.2.","remediation":"Upgrade OMSA to 11.1.0.2. Host agent update plus service restart. If you manage servers exclusively out-of-band via iDRAC/Redfish, removing OMSA is the better answer than patching it.","references":["https://www.dell.com/support/kbdoc/en-us/000494958/dsa-2026-326-security-update-for-dell-openmanage-server-administrator-omsa-network-access-vulnerabilities"],"status":"curated","fleet":{"pain_class":"daemon-restart"}},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:H","cwe":["CWE-125","CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-72487","cve":"CVE-2026-72487","aliases":[],"title":"Linux kernel (drivers/pci): The option-ROM parser trusts the header and data-structure offsets it reads out of the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/pci)","year":"2026","cvss_score":7.7,"severity":"high","kev":false,"impact":"The option-ROM parser trusts the header and data-structure offsets it reads out of the device's own ROM, so a device presenting a malformed ROM makes the kernel read past the end of the mapped ROM window. Confirmed to page-fault the kernel on x86_64 and to take an alignment fault on arm64 - a panic on a shared node, and the same out-of-bounds read can pull kernel memory adjacent to the ROM mapping.","attack_vector":"The parse runs whenever something reads the device ROM - the PCI sysfs rom attribute, or the ROM region of a passthrough device. That second path is the one that matters here: a tenant VM with a GPU or NIC assigned through /dev/vfio/* reads the VFIO ROM region and the host runs pci_map_rom()/pci_get_rom_size() on ROM content the device supplies. Any device whose ROM contents are attacker-influenced (flashed firmware, a rehosted card, a malicious add-in device) turns this into a host panic. Vendor scores it as needing no privileges.","remediation":"Boot a kernel with the ROM header and data-structure address/alignment checks in pci_get_rom_size(). Interim: do not expose the ROM region to guests (mask the VFIO ROM region / disable ROM BAR on assigned devices), keep the PCI sysfs rom attribute out of containers, and refuse nodes with cards whose firmware provenance you cannot vouch for.","references":["https://git.kernel.org/stable/c/cd5b242d5848369b9e957340a7e30adee7b6763d","https://git.kernel.org/stable/c/4997873e3abbad47a9e1e4abd05f045f77a74f22","https://nvd.nist.gov/vuln/detail/CVE-2026-72487"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-9804","cve":"CVE-2026-9804","aliases":[],"title":"KubeVirt: Symlink path traversal in the virt-exportserver VMExport directory endpoint","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"KubeVirt","year":"2026","cvss_score":7.7,"severity":"high","kev":false,"impact":"Symlink path traversal in the virt-exportserver VMExport directory endpoint","attack_vector":"Cluster user with namespace access","remediation":"Upgrade KubeVirt","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-9804"],"status":"curated","published":"2026-05-28"},{"cvss_vector":"CVSS:3.0/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:L/A:L","cwe":["CWE-20"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2018-10923","cve":"CVE-2018-10923","aliases":[],"title":"GlusterFS (brick, mknod): Mknod can create device nodes that point at real devices on the storage server, so a client","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"GlusterFS (brick, mknod)","year":"2018","cvss_score":7.6,"severity":"high","kev":false,"impact":"Mknod can create device nodes that point at real devices on the storage server, so a client creates a block-device node inside the volume and reads raw disk. That bypasses the file layer entirely and exposes every tenant's data sitting on the same physical device.","attack_vector":"Any authenticated gluster client that can mount a volume and call mknod.","remediation":"Upgrade glusterfs server and restart the bricks. Mount client-side with nodev where the workload allows, and confirm the brick process is not running with the capabilities needed to open raw devices.","references":["https://access.redhat.com/security/cve/CVE-2018-10923","https://nvd.nist.gov/vuln/detail/CVE-2018-10923"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2018-3652","cve":"CVE-2018-3652","aliases":["INTEL-SA-00127"],"title":"Intel DCI (Direct Connect Interface) UEFI setting restrictions - Xeon E3 v5/v6, Xeon Scalable, Xeon D: The UEFI setting","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel DCI (Direct Connect Interface) UEFI setting restrictions - Xeon E3 v5/v6, Xeon Scalable, Xeon D","year":"2018","cvss_score":7.6,"severity":"high","kev":false,"impact":"The UEFI setting that is supposed to lock out DCI can be bypassed, re-enabling Intel's closed-chassis debug path. DCI exposes JTAG-class control of the CPU over a USB port: halt cores, read and write all of physical memory including anything a tenant left resident, single-step SMM, and modify the boot chain. This is the strongest possible below-the-OS position on a node and everything it plants survives a reimage. It is squarely a tenant-handoff and colocation problem - a departing tenant with physical or smart-hands access to the chassis can leave the debug door open for the next occupant's data.","attack_vector":"Physical or near-physical access to a USB3 port on the node (or to a KVM/USB-over-IP appliance wired to one). Relevant wherever your halls have shared cages, third-party remote-hands, or hardware that transits an untrusted logistics chain.","remediation":"BIOS update from the OEM (Dell, HPE, Supermicro, Lenovo, Quanta, Wiwynn) that correctly enforces the DCI lock - reboot and drain required, and OEM releases for Xeon Scalable server boards lagged Intel's July 2018 advisory considerably. Alongside the flash: set and lock the BIOS admin password, disable DCI and USB debug explicitly in the BIOS profile, physically disable or block front-panel USB on production nodes, and treat any node that has left your physical custody as needing a full firmware re-flash plus measured-boot re-attestation before it goes back into a tenant pool.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-3652","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00127.html","https://security.netapp.com/advisory/ntap-20180802-0001/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-25742","cve":"CVE-2021-25742","aliases":[],"title":"ingress-nginx: Custom nginx snippets in an Ingress annotation retrieve the ingress-nginx service-account token","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2021","cvss_score":7.6,"severity":"high","kev":false,"impact":"Custom nginx snippets in an Ingress annotation retrieve the ingress-nginx service-account token and therefore every Secret in the cluster","attack_vector":"Cluster user with namespace access who can create Ingress objects","remediation":"Rolling controller upgrade and set `allow-snippet-annotations: false`; rotate all cluster Secrets if exploited","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-25742"],"status":"curated","published":"2021-10-29"},{"id":"CVE-2021-25745","cve":"CVE-2021-25745","aliases":[],"title":"ingress-nginx: Ingress `path` can be pointed at the service-account token file","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2021","cvss_score":7.6,"severity":"high","kev":false,"impact":"Ingress `path` can be pointed at the service-account token file","attack_vector":"Cluster user who can create Ingress objects","remediation":"Rolling controller upgrade; enable path validation","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-25745"],"status":"curated","published":"2022-05-06"},{"id":"CVE-2021-25746","cve":"CVE-2021-25746","aliases":[],"title":"ingress-nginx: Directive injection through Ingress annotations obtains controller credentials","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2021","cvss_score":7.6,"severity":"high","kev":false,"impact":"Directive injection through Ingress annotations obtains controller credentials","attack_vector":"Cluster user who can create Ingress objects","remediation":"Rolling controller upgrade; enable annotation validation","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-25746"],"status":"curated","published":"2022-05-06"},{"id":"CVE-2021-25748","cve":"CVE-2021-25748","aliases":[],"title":"ingress-nginx: Newline character bypasses `path` sanitization","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2021","cvss_score":7.6,"severity":"high","kev":false,"impact":"Newline character bypasses `path` sanitization","attack_vector":"Cluster user who can create Ingress objects","remediation":"Rolling controller upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-25748"],"status":"curated","published":"2023-05-24"},{"id":"CVE-2022-39388","cve":"CVE-2022-39388","aliases":[],"title":"Istio: Localhost access to the istiod pod lets a user impersonate any workload identity in the mesh","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Istio","year":"2022","cvss_score":7.6,"severity":"high","kev":false,"impact":"Localhost access to the istiod pod lets a user impersonate any workload identity in the mesh","attack_vector":"An attacker with a foothold in the istiod pod","remediation":"Rolling istiod upgrade; restrict exec into istio-system","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-39388"],"status":"curated","published":"2022-11-10"},{"id":"CVE-2022-43758","cve":"CVE-2022-43758","aliases":[],"title":"Rancher: OS command injection through an untrusted Helm catalog URL","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Rancher","year":"2022","cvss_score":7.6,"severity":"high","kev":false,"impact":"OS command injection through an untrusted Helm catalog URL","attack_vector":"A user who can add a Helm catalog","remediation":"Upgrade Rancher; restrict catalog sources","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-43758"],"status":"curated","published":"2023-02-07"},{"id":"CVE-2023-25531","cve":"CVE-2023-25531","aliases":[],"title":"DGX H100 BMC (IPMI): Credential exposure","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX H100 BMC (IPMI)","year":"2023","cvss_score":7.6,"severity":"high","kev":false,"impact":"Credential exposure","attack_vector":"Network-adjacent IPMI client","remediation":"Flash BMC 23.08.18; rotate IPMI credentials; disable IPMI where possible","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5473/5473.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:H/PR:L/UI:R/S:C/C:H/I:H/A:H","cwe":["CWE-522"],"published":"2023-09-20"},{"id":"CVE-2023-34337","cve":"CVE-2023-34337","aliases":["AMI-SA-2023006","Nozomi Labs BMC audit"],"title":"AMI MegaRAC SPx (BMC cryptography / HMAC): The BMC uses inadequate HMAC strength, so an attacker positioned","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx (BMC cryptography / HMAC)","year":"2023","cvss_score":7.6,"severity":"high","kev":false,"impact":"The BMC uses inadequate HMAC strength, so an attacker positioned on the management network can forge or replay authenticated material rather than having to break in. The practical outcome is impersonating a legitimate management session and issuing privileged BMC operations - power, virtual media, firmware update - without ever holding a valid password. AMI scores the scope as changed, i.e. the consequences land outside the BMC.","attack_vector":"Adjacent network with a low-privilege foothold and some user interaction, at high attack complexity. Realistically this is an attacker already sitting on the management VLAN who can observe or interpose on BMC traffic, e.g. after compromising a management jump host or a switch on that segment.","remediation":"Firmware flash to SPx_12.2 / SPx_13.0 or later - this one has been fixed for a long time, so the operator question is whether your ODM image is actually from a fixed branch, not whether AMI shipped a fix. Audit the running BMC build across the fleet before assuming it. Until then, force TLS everywhere on the BMC, disable the plain-HTTP and legacy management listeners, and keep the management plane on its own switched segment so there is nowhere to interpose.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2023006.pdf","https://nvd.nist.gov/vuln/detail/CVE-2023-34337"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-07-05"},{"id":"CVE-2023-39347","cve":"CVE-2023-39347","aliases":[],"title":"Cilium: An attacker able to update pod labels causes Cilium to apply the wrong network policy","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2023","cvss_score":7.6,"severity":"high","kev":false,"impact":"An attacker able to update pod labels causes Cilium to apply the wrong network policy","attack_vector":"Cluster user with namespace access","remediation":"Rolling Cilium upgrade; restrict pod label patch RBAC","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-39347"],"status":"curated","published":"2023-09-27"},{"id":"CVE-2023-5043","cve":"CVE-2023-5043","aliases":[],"title":"ingress-nginx: Annotation injection causes arbitrary command execution in the controller pod","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2023","cvss_score":7.6,"severity":"high","kev":false,"impact":"Annotation injection causes arbitrary command execution in the controller pod","attack_vector":"Cluster user who can create Ingress objects","remediation":"Rolling controller upgrade; disable snippet annotations","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-5043"],"status":"curated","published":"2023-10-25"},{"id":"CVE-2023-5044","cve":"CVE-2023-5044","aliases":[],"title":"ingress-nginx: Code injection via the permanent-redirect annotation","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2023","cvss_score":7.6,"severity":"high","kev":false,"impact":"Code injection via the permanent-redirect annotation","attack_vector":"Cluster user who can create Ingress objects","remediation":"Rolling controller upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-5044"],"status":"curated","published":"2023-10-25"},{"id":"CVE-2023-5077","cve":"CVE-2023-5077","aliases":[],"title":"HashiCorp Vault: GCP secrets engine drops existing IAM Conditions when creating/updating rolesets","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HashiCorp Vault","year":"2023","cvss_score":7.6,"severity":"high","kev":false,"impact":"GCP secrets engine drops existing IAM Conditions when creating/updating rolesets -> over-broad cloud grants","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade + re-apply IAM conditions on all GCP rolesets","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-5077"],"status":"curated","published":"2023-09-29"},{"id":"CVE-2024-0122","cve":"CVE-2024-0122","aliases":[],"title":"NVIDIA License System - Delegated Licensing Service (DLS): An unauthorised action against the DLS reaches partial","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA License System - Delegated Licensing Service (DLS)","year":"2024","cvss_score":7.6,"severity":"high","kev":false,"impact":"An unauthorised action against the DLS reaches partial denial of service and disclosure of confidential information. A DLS outage eventually strands vGPU guests when cached licences expire, so this is an availability problem for the whole vGPU estate, not just a management-plane nuisance.","attack_vector":"Adjacent network, no privileges required. Anyone with a route to the licensing appliance - which is usually the same management VLAN as everything else.","remediation":"Update the DLS appliance per bulletin 5570. Cost: appliance restart only. Confirm your licence lease duration so you know how long guests survive a DLS outage before this becomes a tenant-visible incident.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0122","https://github.com/NVIDIA/product-security/tree/main/2024/5570"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:N/UI:N/S:U/C:H/I:L/A:L","cwe":["CWE-862"],"published":"2024-11-23"},{"id":"CVE-2024-0135","cve":"CVE-2024-0135","aliases":[],"title":"Container Toolkit / GPU Operator: Container escape to host root (insufficient input validation)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Container Toolkit / GPU Operator","year":"2024","cvss_score":7.6,"severity":"high","kev":false,"impact":"Container escape to host root (insufficient input validation)","attack_vector":"Any tenant that can run a container image on a GPU node","remediation":"Bump nvidia-container-toolkit package + restart containerd/docker; upgrade GPU Operator Helm chart; evict running tenant containers","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0135","https://github.com/NVIDIA/product-security/tree/main/2025/5599"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:H/UI:R/S:C/C:H/I:H/A:H","cwe":["CWE-653"],"fleet":{"pain_class":"node-drain"},"published":"2025-01-28"},{"id":"CVE-2024-0136","cve":"CVE-2024-0136","aliases":[],"title":"Container Toolkit / GPU Operator: Container escape to host (insufficient input validation)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Container Toolkit / GPU Operator","year":"2024","cvss_score":7.6,"severity":"high","kev":false,"impact":"Container escape to host (insufficient input validation)","attack_vector":"Any tenant with a container","remediation":"Bump toolkit + restart runtime; upgrade GPU Operator chart; evict tenants","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0136","https://github.com/NVIDIA/product-security/tree/main/2025/5599"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:H/UI:R/S:C/C:H/I:H/A:H","cwe":["CWE-653"],"fleet":{"pain_class":"node-drain"},"published":"2025-01-28"},{"id":"CVE-2024-0148","cve":"CVE-2024-0148","aliases":[],"title":"IGX Orin bootloader: Improper access control in bootloader","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"IGX Orin bootloader","year":"2024","cvss_score":7.6,"severity":"high","kev":false,"impact":"Improper access control in bootloader -> persistent compromise","attack_vector":"Local attacker with device access","remediation":"Flash bootloader out-of-band","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0148","https://github.com/NVIDIA/product-security/tree/main/2025/5617"],"status":"curated","cvss_vector":"CVSS:3.1/AV:P/AC:L/PR:N/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-447"],"published":"2025-02-25"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:L/I:L/A:H","cwe":["CWE-269"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-21985","cve":"CVE-2024-21985","aliases":[],"title":"NetApp ONTAP 9 role-based access control: A user holding several remote accounts with different roles performs actions","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"NetApp ONTAP 9 role-based access control","year":"2024","cvss_score":7.6,"severity":"high","kev":false,"impact":"A user holding several remote accounts with different roles performs actions none of those roles should permit, which defeats the separation between an operator who can read and one who can destroy.","attack_vector":"An authenticated ONTAP user with more than one remote account on a system below 9.9.1P18, 9.10.1P16, 9.11.1P13, 9.12.1P10 or 9.13.1P4.","remediation":"Upgrade to the fixed patch level. In the meantime, avoid granting the same person multiple ONTAP accounts with different roles, since that is the precondition.","references":["https://security.netapp.com/advisory/ntap-20240126-0001/","https://nvd.nist.gov/vuln/detail/CVE-2024-21985"],"status":"curated"},{"id":"CVE-2024-25943","cve":"CVE-2024-25943","aliases":["DSA-2024-099"],"title":"Dell iDRAC9 (IPMI 2.0 over LAN): iDRAC9 generates predictable IPMI 2.0 session IDs, so an attacker can hijack somebody","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC9 (IPMI 2.0 over LAN)","year":"2024","cvss_score":7.6,"severity":"high","kev":false,"impact":"iDRAC9 generates predictable IPMI 2.0 session IDs, so an attacker can hijack somebody else's live IPMI session rather than authenticate as themselves. The stolen session carries whatever privilege the legitimate user had - typically chassis power control, boot device selection, and sensor/SEL access. For a GPU fleet the concrete outcomes are unauthorised power cycling of running training nodes and boot-order manipulation that sets up an attacker-controlled boot on the next restart. Spans 14G, 15G and 16G PowerEdge, which is most of the current GPU-server installed base.","attack_vector":"Anything that can reach UDP 623 on the iDRAC address, plus the ability to observe or race a legitimate IPMI session. That means the OOB management VLAN, and in practice any automation host, monitoring collector, or DCIM system that already speaks IPMI to the fleet.","remediation":"Flash iDRAC9 to 7.00.00.172 (14G) or 7.10.50.00 (15G/16G) or later - out-of-band, per-node, no host reboot and no drain of running jobs. The strong config-only mitigation here is to disable IPMI over LAN entirely: Redfish and racadm cover everything modern tooling needs, and turning IPMI off removes an entire legacy attack surface rather than patching one bug in it. Budget for the tooling migration if any of your automation still speaks ipmitool.","references":["https://www.dell.com/support/kbdoc/en-us/000226503/dsa-2024-099-security-update-for-dell-idrac9-ipmi-session-vulnerability","https://nvd.nist.gov/vuln/detail/CVE-2024-25943"],"status":"curated","published":"2024-06-29"},{"id":"CVE-2024-43805","cve":"CVE-2024-43805","aliases":[],"title":"JupyterLab: XSS via untrusted notebook content","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"JupyterLab","year":"2024","cvss_score":7.6,"severity":"high","kev":false,"impact":"XSS via untrusted notebook content → session compromise","attack_vector":"Customer-supplied notebook opened by another user","remediation":"Upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-43805"],"status":"curated","published":"2024-08-28"},{"id":"CVE-2025-23249","cve":"CVE-2025-23249","aliases":[],"title":"NeMo Framework: RCE via insecure deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NeMo Framework","year":"2025","cvss_score":7.6,"severity":"high","kev":false,"impact":"RCE via insecure deserialization","attack_vector":"Malicious model/checkpoint","remediation":"Bump NeMo in training images; rebuild; restrict checkpoint provenance","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23249","https://github.com/NVIDIA/product-security/tree/main/2025/5641"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:R/S:U/C:L/I:H/A:L","cwe":["CWE-502"],"published":"2025-04-22"},{"id":"CVE-2025-23250","cve":"CVE-2025-23250","aliases":[],"title":"NeMo Framework: Arbitrary file write/read via path traversal","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NeMo Framework","year":"2025","cvss_score":7.6,"severity":"high","kev":false,"impact":"Arbitrary file write/read via path traversal","attack_vector":"Malicious model artifact","remediation":"Bump NeMo; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23250","https://github.com/NVIDIA/product-security/tree/main/2025/5641"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:R/S:U/C:L/I:H/A:L","cwe":["CWE-22"],"published":"2025-04-22"},{"id":"CVE-2025-23251","cve":"CVE-2025-23251","aliases":[],"title":"NeMo Framework: RCE","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NeMo Framework","year":"2025","cvss_score":7.6,"severity":"high","kev":false,"impact":"RCE","attack_vector":"Malicious model artifact","remediation":"Bump NeMo; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23251","https://github.com/NVIDIA/product-security/tree/main/2025/5641"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:R/S:U/C:L/I:H/A:L","cwe":["CWE-94"],"published":"2025-04-22"},{"id":"CVE-2025-23263","cve":"CVE-2025-23263","aliases":[],"title":"Mellanox OFED: Authentication bypass in the host networking stack","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Mellanox OFED","year":"2025","cvss_score":7.6,"severity":"high","kev":false,"impact":"Authentication bypass in the host networking stack","attack_vector":"Network-adjacent attacker","remediation":"Upgrade MLNX_OFED / DOCA-Host on all nodes; driver reload requires node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23263","https://github.com/NVIDIA/product-security/tree/main/2025/5654"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:L","cwe":["CWE-279"],"fleet":{"pain_class":"node-drain"},"published":"2025-07-17"},{"id":"CVE-2025-23343","cve":"CVE-2025-23343","aliases":[],"title":"NVIDIA NVDebug tool: NVDebug can be induced to write files into restricted components, reaching data tampering","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NVDebug tool","year":"2025","cvss_score":7.6,"severity":"high","kev":false,"impact":"NVDebug can be induced to write files into restricted components, reaching data tampering and information disclosure with a changed scope on a DGX/HGX platform host.","attack_vector":"Adjacent network, low privileges, user interaction, high complexity. Narrow, but the target is a platform-management host.","remediation":"Update NVDebug per bulletin 5696. Cost: trivial - replace the tool bundle. No node drain.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23343","https://github.com/NVIDIA/product-security/tree/main/2025/5696"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:H/PR:L/UI:R/S:C/C:H/I:H/A:H","cwe":["CWE-22"],"published":"2025-09-09"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:H/UI:N/S:C/C:N/I:H/A:L","cwe":["CWE-862"],"fleet":{"pain_class":"firmware-flash"},"id":"CVE-2025-33182","cve":"CVE-2025-33182","aliases":[],"title":"NVIDIA Jetson Linux (UEFI): UEFI accepts a Linux Device Tree without checking authorization, so anyone who reaches","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Jetson Linux (UEFI)","year":"2025","cvss_score":7.6,"severity":"high","kev":false,"impact":"UEFI accepts a Linux Device Tree without checking authorization, so anyone who reaches a privileged account on the device can rewrite the hardware description the kernel boots against. That is persistence below the operating system: the tampered DTB survives an OS reinstall, and a reimaged device comes back still owned. For a fleet of edge or embedded GPU devices this is the case where your standard recovery play - wipe and redeploy - does not actually clear the attacker.","attack_vector":"Requires an already-privileged account, but NVIDIA scores it as network-reachable, meaning the vulnerable update path is exposed to the device's privileged management surface rather than needing hands on the hardware. Realistically it is a second-stage move: chain any root-level bug into it and you convert a temporary foothold into firmware-level persistence.","remediation":"Update Jetson Linux to 35.6.3 or later on Xavier and Orin. Separately, any device you suspect was reached at root needs a full firmware reflash and DTB verification, not an OS update - an in-place upgrade will not remove an already-planted Device Tree. Restrict who can reach the privileged management interface on these devices in the meantime.","references":["https://github.com/NVIDIA/product-security/tree/main/2025/5716","https://nvd.nist.gov/vuln/detail/CVE-2025-33182"],"status":"curated"},{"id":"CVE-2025-33203","cve":"CVE-2025-33203","aliases":[],"title":"NVIDIA NeMo Agent Toolkit (Web UI): The chat API endpoint is vulnerable to server-side request forgery, so an attacker","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NeMo Agent Toolkit (Web UI)","year":"2025","cvss_score":7.6,"severity":"high","kev":false,"impact":"The chat API endpoint is vulnerable to server-side request forgery, so an attacker makes the toolkit issue requests to internal endpoints - cloud metadata services and in-cluster APIs being the obvious targets. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5726 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33203","https://github.com/NVIDIA/product-security/tree/main/2025/5726"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:L/A:L","cwe":["CWE-918"],"published":"2025-11-25"},{"id":"CVE-2025-3744","cve":"CVE-2025-3744","aliases":[],"title":"HashiCorp Nomad Enterprise: Jobs using the policy-override option bypass mandatory Sentinel policies","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HashiCorp Nomad Enterprise","year":"2025","cvss_score":7.6,"severity":"high","kev":false,"impact":"Jobs using the policy-override option bypass mandatory Sentinel policies","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade Nomad 1.10.1/1.9.9/1.8.13; audit recently submitted jobs","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-3744"],"status":"curated","published":"2025-05-13"},{"id":"CVE-2025-4123","cve":"CVE-2025-4123","aliases":[],"title":"Grafana: Client path traversal + open redirect","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Grafana","year":"2025","cvss_score":7.6,"severity":"high","kev":false,"impact":"Client path traversal + open redirect -> load an attacker-hosted frontend plugin and run arbitrary JS","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; disable anonymous access","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-4123"],"status":"curated","published":"2025-05-22"},{"id":"CVE-2025-47290","cve":"CVE-2025-47290","aliases":[],"title":"containerd: TOCTOU during image unpack: a crafted image can arbitrarily modify the host filesystem","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2025","cvss_score":7.6,"severity":"high","kev":false,"impact":"TOCTOU during image unpack: a crafted image can arbitrarily modify the host filesystem","attack_vector":"Malicious image","remediation":"Rolling containerd upgrade with node drain; restrict tenant registries","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-47290"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2025-05-20"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:L/A:L","cwe":["CWE-266"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2026-18621","cve":"CVE-2026-18621","aliases":[],"title":"Kubeflow Pipelines (Data Science Pipelines V1 API Argo Workflow spec path): The V1 API path accepts an arbitrary Argo","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubeflow Pipelines (Data Science Pipelines V1 API Argo Workflow spec path)","year":"2026","cvss_score":7.6,"severity":"high","kev":false,"impact":"The V1 API path accepts an arbitrary Argo Workflow spec and skips the V2 security hardening entirely, so the pipelines API server acts as a confused deputy and creates privileged pods on the submitter's behalf. A tenant holding only namespace edit rights escalates to root on the underlying node, which on a GPU box means access to every other tenant's containers and to the device nodes they are using.","attack_vector":"A user with namespace editor privileges in a Data Science Pipelines / Kubeflow Pipelines namespace, submitting through the legacy V1 API endpoint.","remediation":"Apply the Red Hat OpenShift AI errata (RHSA-2026:53261/53262/53263) and restart the pipelines API server. If you run upstream Kubeflow Pipelines, disable or gate the V1 API path and enforce pod restrictions with an admission policy rather than relying on the API server's own hardening.","references":["https://access.redhat.com/security/cve/CVE-2026-18621","https://nvd.nist.gov/vuln/detail/CVE-2026-18621"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:P/AC:L/PR:N/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-78"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-24154","cve":"CVE-2026-24154","aliases":[],"title":"NVIDIA Jetson Linux (initrd command-line handling): An attacker with physical access and no credentials at all can","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Jetson Linux (initrd command-line handling)","year":"2026","cvss_score":7.6,"severity":"high","kev":false,"impact":"An attacker with physical access and no credentials at all can inject command-line arguments that initrd consumes, giving them code execution in early boot before the OS trust boundary exists. Everything downstream inherits the compromise - disk contents, any secrets the device unseals, and persistence that a normal reinstall will not clear. NVIDIA rates the scope as changed and all three impacts high, which is the correct read: this is total ownership of the device, not an escalation within it.","attack_vector":"Physical access to the device's boot path or console. No account, no prior foothold, no user interaction. Pair it with the nvluks issue in the same bulletin and a physically obtained Jetson gives up both its code execution path and its disk encryption.","remediation":"Update Jetson Linux to 35.6.4, 36.5 or 38.4 for your branch and reboot. For devices already deployed in locations you do not physically control, assume tampering is possible until the update lands, and prioritise those over lab and datacentre units - the datacentre units were never really exposed to this one.","references":["https://github.com/NVIDIA/product-security/tree/main/2026/5797","https://nvd.nist.gov/vuln/detail/CVE-2026-24154"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:L/UI:N/S:C/C:H/I:L/A:N","fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2026-045-cilium-mutual-authentication-tls","cve":null,"aliases":["GHSA-33qq-jq9c-6gcc"],"title":"Cilium (mutual authentication, TLS certificate chain handling): Mutual authentication, the control an operator turns on","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium (mutual authentication, TLS certificate chain handling)","year":"2026","cvss_score":7.6,"severity":"high","kev":false,"impact":"Mutual authentication, the control an operator turns on precisely to get cryptographic workload identity between tenants, can be defeated by anyone holding any valid client certificate. The attacker obtains the target identity's public certificate — public by construction — appends it to their own chain during the TLS handshake, and Cilium accepts them as that identity. Every policy decision keyed on mutual auth then evaluates against the impersonated workload. In a shared GPU cluster this means a low-value pod in one tenant's namespace can present as another tenant's service and reach whatever that service is allowed to reach: model registries, feature stores, training data endpoints, control-plane APIs. The failure is total rather than partial, and the feature is beta and already slated for deprecation, so operators who adopted it for tenant separation should assume it never provided the boundary they scoped it for.","attack_vector":"Adjacent network, from inside the cluster. The attacker needs any valid client certificate issued in the mesh plus the target identity's public certificate. Only clusters with Cilium mutual authentication enabled are affected.","remediation":"Upgrade to Cilium 1.19.6, 1.18.12 or 1.17.18 and roll the agent DaemonSet. There is no workaround short of upgrading. Given the feature is deprecated, plan the migration off mutual authentication onto a boundary you intend to keep, and in the meantime do not count mutual auth as the control separating tenants in your threat model or your customer-facing claims.","references":["https://github.com/cilium/cilium/security/advisories/GHSA-33qq-jq9c-6gcc","https://github.com/cilium/cilium/commit/168d16627bfb7b96717cf8dff54e33ceac09192c"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2011-3997","cve":"CVE-2011-3997","aliases":[],"title":"Opengear console server: Authentication bypass in the console server allowing remote attackers to modify settings","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Opengear console server","year":"2011","cvss_score":7.5,"severity":"high","kev":false,"impact":"Authentication bypass in the console server allowing remote attackers to modify settings and reach every attached device's serial console — switches, BMCs, storage controllers","attack_vector":"Network / OOB LAN, unauthenticated","remediation":"Console-server firmware upgrade; older units are EOL and the practical fix is replacing the appliance, which means a scheduled loss of out-of-band access to the rack","references":["https://nvd.nist.gov/vuln/detail/CVE-2011-3997"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2011-11-09"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N","fleet":{"pain_class":"daemon-restart"},"id":"CVE-2013-20001","cve":"CVE-2013-20001","aliases":[],"title":"OpenZFS (sharenfs export generation): When an NFS share is exported to IPv6 addresses via sharenfs, OpenZFS silently","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"OpenZFS (sharenfs export generation)","year":"2013","cvss_score":7.5,"severity":"high","kev":false,"impact":"When an NFS share is exported to IPv6 addresses via sharenfs, OpenZFS silently fails to parse the address and exports the dataset to everyone instead. The operator sees a restriction in the config that does not exist on the wire, so any host on the network mounts the dataset.","attack_vector":"Any host that can reach the NFS server, on any ZFS dataset whose sharenfs restriction was written with IPv6 addresses.","remediation":"Upgrade OpenZFS past 2.0.3 and re-export the datasets, then verify the actual export list with exportfs -v rather than trusting the sharenfs property. Prefer expressing restrictions in /etc/exports directly, and audit any dataset that was shared with an IPv6 restriction.","references":["https://github.com/openzfs/zfs/issues/1894","https://nvd.nist.gov/vuln/detail/CVE-2013-20001"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-798","CWE-522"],"fleet":{"pain_class":"firmware-flash"},"id":"CVE-2013-3620","cve":"CVE-2013-3620","aliases":[],"title":"Supermicro IPMI BMC firmware - hardcoded WSMAN credentials (X9 before SMT_X9_315, X8 before SMT X8 312): The BMC","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro IPMI BMC firmware - hardcoded WSMAN credentials (X9 before SMT_X9_315, X8 before SMT X8 312)","year":"2013","cvss_score":7.5,"severity":"high","kev":false,"impact":"The BMC firmware ships with a static WSMAN credential pair that is identical on every board of that generation and cannot be changed by the operator. Recover it once - from a firmware image anyone can download - and you hold a working management login on every affected BMC in the world, including every one in your competitors' racks and every one in yours. Rotating your BMC passwords does nothing: one of the credential sets is a digest-auth account with an immutable password, and the other is a basic-auth account that simply fails to follow the admin password when you change it. For an operator this breaks the assumption underneath all BMC access control, which is that credentials are something you own.","attack_vector":"Network, pre-auth in effect - the credential is public knowledge, so possession of it is not a privilege the attacker had to earn. Any reachability to the BMC management interface is sufficient.","remediation":"Firmware flash to SMT_X9_315 / SMT X8 312 or later; there is no configuration change that removes a hardcoded credential. Until then treat every affected BMC as having a permanent open account and rely entirely on network isolation - management VLAN, no tenant routability, jump-host-only access - because per-device credential hygiene provides zero protection here. This is also the item to check first when acquiring second-hand or colocated hardware, since the previous operator's firmware level is now your exposure.","references":["https://www.rapid7.com/blog/post/2013/11/06/supermicro-ipmi-firmware-vulnerabilities/","https://www.kb.cert.org/vuls/id/648646","https://nvd.nist.gov/vuln/detail/CVE-2013-3620"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-522","CWE-916"],"fleet":{"pain_class":"unpatchable / mitigate-only"},"id":"CVE-2013-4037","cve":"CVE-2013-4037","aliases":[],"title":"IBM Integrated Management Module (IMM/IMM2) IPMI 2.0 RAKP implementation: The vendor-acknowledged instance of the IPMI","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"IBM Integrated Management Module (IMM/IMM2) IPMI 2.0 RAKP implementation","year":"2013","cvss_score":7.5,"severity":"high","kev":false,"impact":"The vendor-acknowledged instance of the IPMI 2.0 RAKP design flaw. The BMC hands the salted HMAC of any named account's password to an unauthenticated client as part of the session-setup exchange - this is what the specification tells it to do, not a coding mistake. An attacker asks for the hash of the admin account, cracks it offline at leisure with no lockout, no rate limit and no log entry on the BMC, and comes back with a real credential. In a datacenter where BMC passwords are provisioned from a common template or a per-rack scheme, cracking one recovers many, which turns a single reachable BMC into fleet-wide out-of-band control.","attack_vector":"Network, pre-auth, UDP/623 (IPMI-over-LAN). Requires only the ability to send RAKP messages to the BMC and knowledge or guessing of an account name.","remediation":"Firmware update mitigates the vendor-specific handling but does not repair the protocol - RAKP hash disclosure is inherent to IPMI 2.0 authentication, so any BMC still speaking IPMI-over-LAN retains the exposure. The durable controls are structural: disable IPMI-over-LAN entirely and manage out-of-band through the vendor's authenticated HTTPS/Redfish interface; where IPMI must stay on, use long random per-device passwords generated by a secrets manager so offline cracking fails and one recovered credential unlocks exactly one node; and keep UDP/623 unreachable from anything but the management bastion. Treat this as mitigate-only rather than patchable.","references":["https://exchange.xforce.ibmcloud.com/vulnerabilities/86173","https://www.rapid7.com/blog/post/2013/07/02/a-penetration-testers-guide-to-ipmi/","https://www.cve.org/CVERecord?id=CVE-2013-4037"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2013-4786","cve":"CVE-2013-4786","aliases":[],"title":"IPMI 2.0 RAKP (all vendors): Protocol design flaw","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"IPMI 2.0 RAKP (all vendors)","year":"2013","cvss_score":7.5,"severity":"high","kev":false,"impact":"Protocol design flaw — RAKP message 2 returns an HMAC over the password hash to any unauthenticated requester, enabling offline cracking of every BMC account on the fleet","attack_vector":"Network / IPMI over LAN, unauthenticated","remediation":"Cannot be patched — it is the IPMI 2.0 spec. Only real remediation is disabling IPMI-over-LAN entirely and moving to Redfish with strong per-node unique credentials, which breaks legacy provisioning tooling","references":["https://nvd.nist.gov/vuln/detail/CVE-2013-4786"],"status":"curated","fleet":{"ubiquity":"universal - IPMI-over-LAN is enabled on essentially every server BMC unless deliberately disabled","remediation_pain":"unpatchable-mitigate-only - **this is a flaw in the IPMI 2.0 specification itself**, so no firmware fixes it; the only remediation is disabling IPMI-over-LAN fleet-wide or hard-isolating UDP/623, which breaks tooling that depends on it","pain_class":"unpatchable / mitigate-only","why_fleet_wide":"A vulnerable BMC hands out a password-derived HMAC-SHA1 before authentication, so any host that can reach the management network harvests offline-crackable BMC credentials for every node at once - and BMC passwords are typically identical across a fleet built from one golden config."},"published":"2013-07-08"},{"cvss_vector":"CVSS:3.0/AV:L/AC:H/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-119"],"fleet":{"pain_class":"node-drain"},"id":"CVE-2015-8554","cve":"CVE-2015-8554","aliases":["XSA-164"],"title":"Xen qemu-xen-traditional device model hw/pt-msi.c (MSI-X passthrough): Buffer overflow on the MSI-X table write path","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen qemu-xen-traditional device model hw/pt-msi.c (MSI-X passthrough)","year":"2015","cvss_score":7.5,"severity":"high","kev":false,"impact":"Buffer overflow on the MSI-X table write path for a passed-through device. A tenant with an MSI-X-capable device - which is every modern GPU and every modern NIC - writes past the end of the device model's MSI-X entry array and takes over the QEMU process. Without a device-model stub domain that process runs in dom0, so this is a direct guest-to-host escape reached through the ordinary act of a driver configuring its own interrupts. It is the most GPU-passthrough-specific escape of the pre-2018 era: the vulnerable code exists only because the device is passed through.","attack_vector":"Guest administrator - the tenant - with an assigned MSI-X-capable physical PCI device. Triggered from the guest's own device driver writing its MSI-X table.","remediation":"Patch per XSA-164 and restart the affected device models, which means stopping and restarting every guest holding a passed-through device on that host - a full node drain, not a live migration, since migrating a VM with an assigned device is not generally possible. The structural mitigation worth adopting is running the device model in a stub domain so a device-model compromise lands in a deprivileged domain rather than dom0; on qemu-xen-traditional that was frequently not the default.","references":["https://xenbits.xen.org/xsa/advisory-164.html","https://security.gentoo.org/glsa/201604-03","https://nvd.nist.gov/vuln/detail/CVE-2015-8554"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.0/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-908","CWE-200"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2016-5244","cve":"CVE-2016-5244","aliases":[],"title":"Linux kernel RDS net/rds/recv.c - rds_inc_info_copy: A structure member is left uninitialised before the RDS message","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel RDS net/rds/recv.c - rds_inc_info_copy","year":"2016","cvss_score":7.5,"severity":"high","kev":false,"impact":"A structure member is left uninitialised before the RDS message info block is copied out, so reading an RDS message returns kernel stack bytes to the reader. MITRE scopes this to remote attackers, which on an RDS-over-InfiniBand fabric means a peer node - another tenant's machine - harvesting the target's kernel stack over the wire. It is small per message and unlimited in repetition, so an attacker collects continuously until they have the pointers they need. In practice this is the reconnaissance half of a chain: leak enough kernel addresses to defeat KASLR, then fire one of the RDS or uverbs corruption bugs with a known target layout.","attack_vector":"Network / adjacent fabric, pre-auth from the perspective of the leaking host - the attacker reads RDS messages it is entitled to receive and gets kernel memory as a side effect.","remediation":"Kernel upgrade including commit 4116def2337991b39919f3b448326e21c40e0dbb (rds: fix an infoleak in rds_inc_info_copy); rolling reboot. As with the other RDS entries here, the zero-cost control most operators should apply first is to blacklist the rds module on every node that does not deliberately use it, which removes this and the RDS memory-corruption bugs at once and needs no reboot.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=4116def2337991b39919f3b448326e21c40e0dbb","https://access.redhat.com/security/cve/CVE-2016-5244","https://www.openwall.com/lists/oss-security/2016/06/03/5"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2017-5926","cve":"CVE-2017-5926","aliases":["AnC-style MMU side channel"],"title":"AMD processors - page table walk traces in the last-level cache: The MMU's page table walks during address translation","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD processors - page table walk traces in the last-level cache","year":"2017","cvss_score":7.5,"severity":"high","kev":false,"impact":"The MMU's page table walks during address translation leave traces in the last-level cache, which is shared across cores on an AMD socket. A side-channel attack on the MMU recovers information about a victim's virtual address layout - and since the LLC is shared, this reaches across cores, not just across SMT threads, so core pinning does not help.","attack_vector":"Local, co-resident on the same socket as the victim. Works across cores because the last-level cache is the shared resource.","remediation":"**Effectively unpatchable in hardware** - shared last-level cache is a design property, not a bug. Mitigation is architectural: do not co-schedule mutually untrusted tenants on the same socket, and where the threat model demands it, allocate whole nodes rather than slices. On a GPU fleet that maps naturally onto whole-node allocation for sensitive customers, which most operators already offer as a premium tier.","references":["https://nvd.nist.gov/vuln/detail/CVE-2017-5926"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"tags":["tenant-isolation"],"published":"2017-02-27"},{"id":"CVE-2018-0090","cve":"CVE-2018-0090","aliases":[],"title":"Cisco NX-OS (management interface ACL): The ACL you put on the management interface is not enforced, so traffic you","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco NX-OS (management interface ACL)","year":"2018","cvss_score":7.5,"severity":"high","kev":false,"impact":"The ACL you put on the management interface is not enforced, so traffic you believe is being dropped reaches the switch's control plane anyway. Every 'we restricted the mgmt interface to the jump host' assumption becomes false. It does not by itself grant access, but it silently removes the compensating control that most operators rely on for every other switch CVE in this list.","attack_vector":"Unauthenticated, remote — any host with IP reachability to mgmt0, even one the ACL was supposed to block.","remediation":"NX-OS upgrade plus reload. Do not treat management-interface ACLs as a substitute for real network segmentation — put the management interface on a physically or VRF-separated OOB network, which is a config/topology change and the durable fix.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-0090"],"status":"curated","published":"2018-01-18"},{"cvss_vector":"CVSS:3.0/AV:A/AC:H/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-294"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2018-1128","cve":"CVE-2018-1128","aliases":[],"title":"Ceph CephX authentication protocol: An attacker who sniffs the storage network can replay a CephX authentication","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ceph CephX authentication protocol","year":"2018","cvss_score":7.5,"severity":"high","kev":false,"impact":"An attacker who sniffs the storage network can replay a CephX authentication exchange and obtain a session as the client it copied, gaining that client's read and write rights against RADOS pools. The victim's identity is fully assumed - there is no distinct attacker identity to audit.","attack_vector":"Passive-then-active attacker on the Ceph public/cluster network. Any tenant node sharing the storage L2 domain qualifies.","remediation":"Upgrade to a Ceph release carrying the cephx replay fix (12.2.6+/13.2.x) and restart all daemons. Move to msgr2 secure mode and isolate the storage fabric from tenant-controlled interfaces.","references":["https://access.redhat.com/security/cve/CVE-2018-1128","https://nvd.nist.gov/vuln/detail/CVE-2018-1128"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2018-1211","cve":"CVE-2018-1211","aliases":[],"title":"Dell iDRAC7 / iDRAC8 (web server URI parser): Directory traversal in the BMC's own HTTP front end lets an attacker","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC7 / iDRAC8 (web server URI parser)","year":"2018","cvss_score":7.5,"severity":"high","kev":false,"impact":"Directory traversal in the BMC's own HTTP front end lets an attacker with no credentials read files off the iDRAC filesystem. In practice that is reconnaissance that turns into access: configuration, session material, and stored secrets that let the attacker come back as an authenticated user. Relevant to fleets that still run 12G/13G PowerEdge as head nodes, storage, or staging boxes alongside the GPU racks - the older tier tends to be the one nobody re-flashed.","attack_vector":"Anything routable to the iDRAC web port on the out-of-band management VLAN, unauthenticated. No host access and no valid account required.","remediation":"Flash iDRAC7/iDRAC8 to 2.52.52.52 or later - out-of-band, per-node, no host reboot. On a fleet this old the real cost is inventory: finding which nodes are still on pre-2.52 firmware. Config-only interim: ACL the iDRAC web port to the jump hosts only. Note the original Dell TechCenter advisory URL is dead; NVD carries the authoritative version data.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-1211","https://www.dell.com/support/kbdoc/en-us/000131409/dsa-2018-004-idrac-vulnerabilities"],"status":"curated","published":"2018-03-23"},{"id":"CVE-2018-12922","cve":"CVE-2018-12922","aliases":[],"title":"Emerson/Vertiv Liebert IntelliSlot Web Card (config/configUser.htm, config/configTelnet.htm): The IntelliSlot card","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Emerson/Vertiv Liebert IntelliSlot Web Card (config/configUser.htm, config/configTelnet.htm)","year":"2018","cvss_score":7.5,"severity":"high","kev":false,"impact":"The IntelliSlot card is the network brain bolted into Liebert CRAC/CRAH units, condensers and thermal management gear. A remote attacker can reconfigure its access control and telnet settings by hitting the config pages directly, which means they can create their own administrative access and keep it. From there the card exposes the unit's operating parameters - fan command, compressor enable, setpoints, alarm relays. Turning off or de-rating the CRAHs serving a GPU hall is a fleet-wide availability kill: the racks do not have hours of ride-through, they have minutes. Equally damaging and much quieter is suppressing the alarm path, so the operations team's first indication of a thermal event is GPUs dropping off the fabric rather than a temperature alarm.","attack_vector":"Remote HTTP to the card, no authentication. These cards live on the facility monitoring VLAN alongside the PDU network cards and environmental sensors. Many operators inherited that VLAN from the building and treat it as trusted. Note that IntelliSlot cards are frequently also reachable from the DCIM/monitoring server, so compromise of a monitoring host is a direct path in.","remediation":"Replace the card. The IntelliSlot Web Card generation covered here is end-of-life; Vertiv's answer is migration to a current Unity/IS-UNITY card, which is a hardware swap per cooling unit and needs the unit taken off network (not necessarily off cooling). Until then: isolate the card VLAN, block telnet outright at the switch, and put an ACL so only the DCIM collector can reach the card's HTTP port. If you lease, this is landlord equipment - ask for the card model and firmware inventory in writing and treat 'IntelliSlot Web Card' in the answer as a finding.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-12922","https://www.seebug.org/vuldb/ssvid-97372"],"status":"curated"},{"id":"CVE-2018-15664","cve":"CVE-2018-15664","aliases":[],"title":"Docker / moby: `docker cp` symlink-exchange TOCTOU gives arbitrary host read/write as root","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2018","cvss_score":7.5,"severity":"high","kev":false,"impact":"`docker cp` symlink-exchange TOCTOU gives arbitrary host read/write as root","attack_vector":"Any tenant workload on a node whose operator uses docker cp","remediation":"Upgrade Docker Engine; restart daemon (drains containers unless live-restore is on)","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-15664"],"status":"curated","fleet":{"pain_class":"daemon-restart"},"published":"2019-05-23"},{"id":"CVE-2018-5254","cve":"CVE-2018-5254","aliases":[],"title":"Arista EOS (BGP UPDATE): Malformed path attribute in a BGP UPDATE from a peer causes denial of service","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (BGP UPDATE)","year":"2018","cvss_score":7.5,"severity":"high","kev":false,"impact":"Malformed path attribute in a BGP UPDATE from a peer causes denial of service","attack_vector":"Network, from a BGP peer","remediation":"EOS upgrade; relevant where the fabric peers with a transit provider or a customer-controlled router","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-5254"],"status":"curated","published":"2018-04-12"},{"id":"CVE-2018-7185","cve":"CVE-2018-7185","aliases":["Zero Origin timestamp / 'Zero-o'"],"title":"ntpd (protocol engine, zero-origin timestamp): Continually sending packets with a zero-origin timestamp lets a remote","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"ntpd (protocol engine, zero-origin timestamp)","year":"2018","cvss_score":7.5,"severity":"high","kev":false,"impact":"Continually sending packets with a zero-origin timestamp lets a remote attacker disrupt an ntpd peer association. Cheap, stateless, and it needs nothing but the ability to send UDP to port 123 — which is open on far more cluster nodes than operators realise, because NTP is usually configured once at image-build time and never reviewed.","attack_vector":"Remote, unauthenticated — UDP packets to the NTP port.","remediation":"Upgrade ntp to 4.2.8p11 or later and restart. Also firewall UDP/123 so only your internal time servers can reach cluster nodes — a host or fabric ACL change, applied live, that removes most of the NTP attack surface at once.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-7185"],"status":"curated","published":"2018-03-06"},{"cvss_vector":"CVSS:3.0/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-755"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2019-10222","cve":"CVE-2019-10222","aliases":[],"title":"Ceph RADOS Gateway (RGW, Beast frontend): An unauthenticated client can crash radosgw by sending valid headers followed","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ceph RADOS Gateway (RGW, Beast frontend)","year":"2019","cvss_score":7.5,"severity":"high","kev":false,"impact":"An unauthenticated client can crash radosgw by sending valid headers followed by an abrupt termination. Repeating it keeps the S3 endpoint down, which stalls every dataset loader and checkpoint writer that goes through object storage.","attack_vector":"Any host that can open a TCP connection to the RGW port. No credentials required, so a single tenant container with egress to the gateway is enough.","remediation":"Upgrade RGW to a release with the Beast frontend fix and restart radosgw. Run multiple gateways behind a load balancer with aggressive health checking and per-source connection rate limits.","references":["https://access.redhat.com/security/cve/CVE-2019-10222","https://nvd.nist.gov/vuln/detail/CVE-2019-10222"],"status":"curated"},{"id":"CVE-2019-11253","cve":"CVE-2019-11253","aliases":[],"title":"Kubernetes (kube-apiserver): \"Billion laughs\": malicious YAML/JSON payload consumes all apiserver memory","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-apiserver)","year":"2019","cvss_score":7.5,"severity":"high","kev":false,"impact":"\"Billion laughs\": malicious YAML/JSON payload consumes all apiserver memory; control-plane DoS","attack_vector":"Any authorized cluster user, and unauthenticated if anonymous auth is on","remediation":"Rolling control-plane upgrade; no GPU drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-11253"],"status":"curated","published":"2019-10-17"},{"id":"CVE-2019-13509","cve":"CVE-2019-13509","aliases":[],"title":"Docker / moby: Docker Engine in debug mode writes secrets into the debug log","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2019","cvss_score":7.5,"severity":"high","kev":false,"impact":"Docker Engine in debug mode writes secrets into the debug log","attack_vector":"Anyone with node or log-pipeline read access","remediation":"Turn off daemon debug mode; rotate leaked secrets; scrub log store","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-13509"],"status":"curated","published":"2019-07-18"},{"id":"CVE-2019-16884","cve":"CVE-2019-16884","aliases":[],"title":"runc: AppArmor restriction bypass via mount-target check flaw","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"runc","year":"2019","cvss_score":7.5,"severity":"high","kev":false,"impact":"AppArmor restriction bypass via mount-target check flaw; container can mount over /proc","attack_vector":"Malicious image","remediation":"Replace runc binary; drain node to restart existing containers","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-16884"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2019-09-25"},{"id":"CVE-2019-16919","cve":"CVE-2019-16919","aliases":[],"title":"Harbor: Broken access control allows creating robot accounts with push/pull rights to projects the user does not own","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Harbor","year":"2019","cvss_score":7.5,"severity":"high","kev":false,"impact":"Broken access control allows creating robot accounts with push/pull rights to projects the user does not own; cross-tenant image poisoning","attack_vector":"Project administrator of any project","remediation":"Upgrade Harbor; audit robot accounts","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-16919"],"status":"curated","published":"2019-10-18"},{"id":"CVE-2019-18948","cve":"CVE-2019-18948","aliases":[],"title":"Arista EOS (VxLAN agent): Malformed ARP packets crash the VxLAN software forwarding agent","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (VxLAN agent)","year":"2019","cvss_score":7.5,"severity":"high","kev":false,"impact":"Malformed ARP packets crash the VxLAN software forwarding agent — a tenant VM can take down overlay forwarding","attack_vector":"Network, from inside a tenant VLAN","remediation":"EOS upgrade with failover; notable because the trigger comes from tenant traffic, not the management plane","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-18948"],"status":"curated","published":"2020-04-16"},{"cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-269"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2019-19728","cve":"CVE-2019-19728","aliases":[],"title":"Slurm (srun --uid): Srun --uid drops privileges in the wrong order, so a step launched through it can end up running","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Slurm (srun --uid)","year":"2019","cvss_score":7.5,"severity":"high","kev":false,"impact":"Srun --uid drops privileges in the wrong order, so a step launched through it can end up running with more privilege than the target user should have. On a shared cluster this is a path from an admin-adjacent account to code execution as, or above, another tenant.","attack_vector":"A local user able to invoke srun with --uid on a login or submit node.","remediation":"Upgrade to Slurm 18.08.9 or 19.05.5 and restart slurmctld and slurmd. If you cannot upgrade immediately, remove --uid from any operator tooling and wrapper scripts that run as root.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-19728","https://www.debian.org/security/2021/dsa-4841","https://bugzilla.suse.com/show_bug.cgi?id=1159692"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2019-20423","cve":"CVE-2019-20423","aliases":[],"title":"Lustre ptlrpc / mdt modules (client-driven server panic family): The head of a family of ten Lustre defects","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Lustre ptlrpc / mdt modules (client-driven server panic family)","year":"2019","cvss_score":7.5,"severity":"high","kev":false,"impact":"The head of a family of ten Lustre defects (CVE-2019-20423 through CVE-2019-20432) that all share one shape — the server does not validate fields in packets sent by a client, so any client can panic the metadata or object storage server with a malformed RPC. Ten separate ways for one tenant's node to take down the filesystem that every other tenant's training job is reading from. In a shared-storage AI cluster this is the cheapest available cross-tenant denial of service: one machine, one packet, everyone's jobs stall.","attack_vector":"Any mounted Lustre client. No credentials, no escalation — the client is inherently trusted by the protocol.","remediation":"Upgrade Lustre servers to 2.12.3 or later and restart the MDS/OSS nodes; failover pairs limit but do not eliminate the I/O pause. There is no config workaround for the parsing itself. What you can do immediately is network-level: restrict which hosts can reach LNet, and stop treating 'the tenant's compute node' as a trusted peer of your storage servers.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-20423","https://nvd.nist.gov/vuln/detail/CVE-2019-20425","https://nvd.nist.gov/vuln/detail/CVE-2019-20432"],"status":"curated","tags":["fabric-dos"],"published":"2020-01-27"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2019-20424","cve":"CVE-2019-20424","aliases":["LU-12615","DDN EXAScaler"],"title":"Lustre (mdt module, mdt_object_remote): A client sends a packet with unvalidated fields and the metadata server","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Lustre (mdt module, mdt_object_remote)","year":"2019","cvss_score":7.5,"severity":"high","kev":false,"impact":"A client sends a packet with unvalidated fields and the metadata server dereferences NULL and panics. Losing the MDS takes the entire filesystem namespace offline, which stalls every training job that touches shared storage - not just the tenant that sent the packet.","attack_vector":"Anything with LNet reachability to the MDS. On most clusters that is every compute node, so any tenant with a job can reach it.","remediation":"Upgrade Lustre servers to 2.12.3 or later. The fix is in a kernel module on the MDS, so it means unloading and reloading Lustre modules - in practice an MDS failover or a reboot of the metadata server pair. DDN EXAScaler ships this Lustre code, so EXAScaler fleets inherit the issue and need DDN's corresponding release rather than an upstream build.","references":["https://jira.whamcloud.com/browse/LU-12615","http://wiki.lustre.org/Lustre_2.12.3_Changelog","https://nvd.nist.gov/vuln/detail/CVE-2019-20424"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2019-20425","cve":"CVE-2019-20425","aliases":["LU-12613","DDN EXAScaler"],"title":"Lustre (ptlrpc module): Out-of-bounds write in the RPC layer, reachable by a client that lies about packet field sizes.","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Lustre (ptlrpc module)","year":"2019","cvss_score":7.5,"severity":"high","kev":false,"impact":"Out-of-bounds write in the RPC layer, reachable by a client that lies about packet field sizes. The server panics; on a shared parallel filesystem that is a cluster-wide storage outage, and an out-of-bounds write is always a candidate for something worse than a crash.","attack_vector":"Any host with LNet reachability to a Lustre server, i.e. any compute node in the fabric.","remediation":"Upgrade Lustre servers to 2.12.3 or later and reload the Lustre modules, which means failing over or rebooting the affected OSS/MDS. Restricting LNet to trusted networks limits exposure but does not fix anything when tenants have their own nodes. DDN EXAScaler ships this Lustre code, so EXAScaler fleets inherit the issue and need DDN's corresponding release rather than an upstream build.","references":["https://jira.whamcloud.com/browse/LU-12613","http://wiki.lustre.org/Lustre_2.12.3_Changelog","https://nvd.nist.gov/vuln/detail/CVE-2019-20425"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2019-20426","cve":"CVE-2019-20426","aliases":["LU-12614","DDN EXAScaler"],"title":"Lustre (ptlrpc module): A second out-of-bounds access in ptlrpc triggered by unvalidated client packet fields, ending","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Lustre (ptlrpc module)","year":"2019","cvss_score":7.5,"severity":"high","kev":false,"impact":"A second out-of-bounds access in ptlrpc triggered by unvalidated client packet fields, ending in a server panic and loss of the shared filesystem for every tenant on it.","attack_vector":"Any host that can speak LNet to a Lustre server. No authentication step stands between a compute node and this code path in a default deployment.","remediation":"Upgrade Lustre servers to 2.12.3 or later. Plan an MDS/OSS failover or reboot - Lustre server fixes are kernel-module changes and cannot be hot-applied. DDN EXAScaler ships this Lustre code, so EXAScaler fleets inherit the issue and need DDN's corresponding release rather than an upstream build.","references":["https://jira.whamcloud.com/browse/LU-12614","http://wiki.lustre.org/Lustre_2.12.3_Changelog","https://nvd.nist.gov/vuln/detail/CVE-2019-20426"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-125"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2019-20428","cve":"CVE-2019-20428","aliases":["LU-12603","DDN EXAScaler"],"title":"Lustre (ptlrpc module): Out-of-bounds read in ptlrpc leading to a server panic. The read primitive also means server","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Lustre (ptlrpc module)","year":"2019","cvss_score":7.5,"severity":"high","kev":false,"impact":"Out-of-bounds read in ptlrpc leading to a server panic. The read primitive also means server memory adjacent to the RPC buffer can influence behaviour before the crash, so treat it as more than a pure availability bug.","attack_vector":"Any LNet peer - in practice every compute node and every tenant running on one.","remediation":"Upgrade Lustre servers to 2.12.3 or later, with an MDS/OSS failover or reboot to load the fixed modules. DDN EXAScaler ships this Lustre code, so EXAScaler fleets inherit the issue and need DDN's corresponding release rather than an upstream build.","references":["https://jira.whamcloud.com/browse/LU-12603","http://wiki.lustre.org/Lustre_2.12.3_Changelog","https://nvd.nist.gov/vuln/detail/CVE-2019-20428"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-125"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2019-20429","cve":"CVE-2019-20429","aliases":["LU-12590","DDN EXAScaler"],"title":"Lustre (ptlrpc module, lm_bufcount handling): A client that modifies the lm_bufcount field walks the server off the end","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Lustre (ptlrpc module, lm_bufcount handling)","year":"2019","cvss_score":7.5,"severity":"high","kev":false,"impact":"A client that modifies the lm_bufcount field walks the server off the end of a buffer and panics it. One line of client-side code takes down storage for the whole cluster.","attack_vector":"Any Lustre client on the fabric. Trivial to trigger once you know the field - no privileged position needed beyond mounting the filesystem.","remediation":"Upgrade Lustre servers to 2.12.3 or later and fail over or reboot the affected servers to load the new modules. DDN EXAScaler ships this Lustre code, so EXAScaler fleets inherit the issue and need DDN's corresponding release rather than an upstream build.","references":["https://jira.whamcloud.com/browse/LU-12590","http://wiki.lustre.org/Lustre_2.12.3_Changelog","https://nvd.nist.gov/vuln/detail/CVE-2019-20429"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-20","CWE-670"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2019-20430","cve":"CVE-2019-20430","aliases":["LU-12602","DDN EXAScaler"],"title":"Lustre (mdt module, MDT Body eadatasize): An oversized eadatasize field in an MDT request drives the metadata server","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Lustre (mdt module, MDT Body eadatasize)","year":"2019","cvss_score":7.5,"severity":"high","kev":false,"impact":"An oversized eadatasize field in an MDT request drives the metadata server into an LBUG panic. The MDS is the single point every file open goes through, so this is a full namespace outage for all tenants.","attack_vector":"Any Lustre client able to send MDT requests, i.e. any node with the filesystem mounted.","remediation":"Upgrade Lustre to 2.12.3 or later on the metadata servers. Requires an MDS failover or reboot. If you run an active/passive MDS pair, patch the passive side first and fail over rather than taking a cold outage. DDN EXAScaler ships this Lustre code, so EXAScaler fleets inherit the issue and need DDN's corresponding release rather than an upstream build.","references":["https://jira.whamcloud.com/browse/LU-12602","http://wiki.lustre.org/Lustre_2.12.3_Changelog","https://nvd.nist.gov/vuln/detail/CVE-2019-20430"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2019-20431","cve":"CVE-2019-20431","aliases":["LU-12612","DDN EXAScaler"],"title":"Lustre (ptlrpc, osd_map_remote_to_local): Out-of-bounds access in the object-storage mapping path, reachable from a","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Lustre (ptlrpc, osd_map_remote_to_local)","year":"2019","cvss_score":7.5,"severity":"high","kev":false,"impact":"Out-of-bounds access in the object-storage mapping path, reachable from a crafted client packet, ending in a panic on the object storage server. Loses the OSTs behind it, which is where the training data and checkpoints live.","attack_vector":"Any Lustre client on the LNet fabric.","remediation":"Upgrade Lustre servers to 2.12.3 or later; fail over or reboot the OSS nodes to load the fixed modules. DDN EXAScaler ships this Lustre code, so EXAScaler fleets inherit the issue and need DDN's corresponding release rather than an upstream build.","references":["https://jira.whamcloud.com/browse/LU-12612","http://wiki.lustre.org/Lustre_2.12.3_Changelog","https://nvd.nist.gov/vuln/detail/CVE-2019-20431"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2019-20432","cve":"CVE-2019-20432","aliases":["LU-12604","DDN EXAScaler"],"title":"Lustre (mdt module): Another unvalidated-field out-of-bounds access in the metadata server, ending in a panic. Same","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Lustre (mdt module)","year":"2019","cvss_score":7.5,"severity":"high","kev":false,"impact":"Another unvalidated-field out-of-bounds access in the metadata server, ending in a panic. Same blast radius as the rest of the 2.12.3 batch - the shared namespace goes away for every tenant at once.","attack_vector":"Any Lustre client with LNet reachability to the MDS.","remediation":"Upgrade Lustre servers to 2.12.3 or later. Treat the whole CVE-2019-20423 through CVE-2019-20432 family as one upgrade - they were all fixed in the same release and patching only the ones you have heard of leaves the rest live. DDN EXAScaler ships this Lustre code, so EXAScaler fleets inherit the issue and need DDN's corresponding release rather than an upstream build.","references":["https://jira.whamcloud.com/browse/LU-12604","http://wiki.lustre.org/Lustre_2.12.3_Changelog","https://nvd.nist.gov/vuln/detail/CVE-2019-20432"],"status":"curated"},{"cvss_vector":"CVSS:3.0/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N","fleet":{"pain_class":"node-reboot"},"id":"CVE-2019-5491","cve":"CVE-2019-5491","aliases":[],"title":"NetApp Clustered Data ONTAP (unauthenticated information disclosure): An attacker with no account extracts sensitive","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"NetApp Clustered Data ONTAP (unauthenticated information disclosure)","year":"2019","cvss_score":7.5,"severity":"high","kev":false,"impact":"An attacker with no account extracts sensitive information from the storage controller, which is useful both directly and as reconnaissance for a follow-on attack against the cluster.","attack_vector":"Network reach to a Clustered Data ONTAP system earlier than 9.1P15 or 9.3P7. No credentials required.","remediation":"Upgrade to 9.1P15 / 9.3P7 or later. Restrict which networks can reach the controller's management and data LIFs while the upgrade is scheduled.","references":["https://security.netapp.com/advisory/ntap-20190227-0001/","https://nvd.nist.gov/vuln/detail/CVE-2019-5491"],"status":"curated"},{"id":"CVE-2019-6193","cve":"CVE-2019-6193","aliases":["LEN-29477"],"title":"Lenovo XClarity Administrator (LXCA) - unauthenticated config file access: Unauthenticated access to LXCA configuration","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lenovo XClarity Administrator (LXCA) - unauthenticated config file access","year":"2019","cvss_score":7.5,"severity":"high","kev":false,"impact":"Unauthenticated access to LXCA configuration files, which contain usernames, license keys and IP addresses. LXCA is the fleet management console that talks to every XCC it manages, so an unauthenticated read of its configuration hands an attacker a map of the whole management estate - which addresses to hit, which account names to try - without touching a single node. Combined with any of the XCC authorisation bypasses of the same era, that map is the difference between a blind scan and a targeted walk through the fleet. Affects LXCA before 2.6.6.","attack_vector":"Anything routable to the LXCA appliance, unauthenticated. No credentials, no host access, no user interaction.","remediation":"Upgrade LXCA to 2.6.6 or later. Note the upgrade path: you must be on 2.6.0 before you can install the 2.6.6 fix bundle, so this is a two-step appliance upgrade, not a single jump. It is still a single appliance rather than a per-node campaign - no node reboots, no job drain, only LXCA's own downtime. Rotate anything the exposed configuration named, and put LXCA on a restricted segment rather than the general management VLAN.","references":["https://support.lenovo.com/us/en/product_security/LEN-29477","https://nvd.nist.gov/vuln/detail/CVE-2019-6193"],"status":"curated","published":"2020-02-14"},{"id":"CVE-2019-9946","cve":"CVE-2019-9946","aliases":[],"title":"CNI portmap plugin: portmap inserts rules ahead of the KUBE-SERVICES chain, so hostPort traffic bypasses NetworkPolicy","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"CNI portmap plugin","year":"2019","cvss_score":7.5,"severity":"high","kev":false,"impact":"portmap inserts rules ahead of the KUBE-SERVICES chain, so hostPort traffic bypasses NetworkPolicy","attack_vector":"Any pod on the cluster network","remediation":"Upgrade the CNI plugins package on every node; block hostPort for tenants","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-9946"],"status":"curated","published":"2019-04-02"},{"id":"CVE-2020-11487","cve":"CVE-2020-11487","aliases":[],"title":"NVIDIA DGX BMC (AMI firmware): A hard-coded RSA-1024 key with weak ciphers in the BMC firmware means the encryption","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVIDIA DGX BMC (AMI firmware)","year":"2020","cvss_score":7.5,"severity":"high","kev":false,"impact":"A hard-coded RSA-1024 key with weak ciphers in the BMC firmware means the encryption protecting BMC sessions and data is decryptable by anyone holding the firmware image - so passively captured management traffic can be read. Notably this one lists all DGX A100 BMC firmware versions as affected, not just older DGX-1/DGX-2 builds.","attack_vector":"Anyone who can capture traffic on the management network, plus anyone who can download the firmware - which is everyone.","remediation":"Flash the DGX BMC firmware from NVIDIA's DGX firmware update container (DGX-1 to 3.38.30 or later, DGX-2 to 1.06.06 or later; DGX A100 per the bulletin's table). A BMC flash does not require the host OS to reboot but drops out-of-band management for several minutes and NVIDIA recommends a host power cycle afterwards, so treat it as a per-node maintenance window. Rotate every BMC and IPMI credential after the flash - flashing does not invalidate secrets an attacker already pulled. Keep BMCs on an isolated management VLAN with no route from tenant or job networks.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-11487"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2020-10-29"},{"id":"CVE-2020-11489","cve":"CVE-2020-11489","aliases":[],"title":"NVIDIA DGX BMC (AMI firmware): Default SNMP community strings on the DGX BMC","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVIDIA DGX BMC (AMI firmware)","year":"2020","cvss_score":7.5,"severity":"high","kev":false,"impact":"Default SNMP community strings on the DGX BMC. Anyone who can reach the BMC over SNMP enumerates hardware inventory, sensor data and configuration without authenticating - cheap reconnaissance that maps your fleet before a real attack. DGX-1 before 3.38.30, DGX-2 before 1.06.06.","attack_vector":"Anyone with UDP reach to the BMC's SNMP port on the management network.","remediation":"Flash the DGX BMC firmware from NVIDIA's DGX firmware update container (DGX-1 to 3.38.30 or later, DGX-2 to 1.06.06 or later; DGX A100 per the bulletin's table). A BMC flash does not require the host OS to reboot but drops out-of-band management for several minutes and NVIDIA recommends a host power cycle afterwards, so treat it as a per-node maintenance window. Rotate every BMC and IPMI credential after the flash - flashing does not invalidate secrets an attacker already pulled. Keep BMCs on an isolated management VLAN with no route from tenant or job networks. Also change or disable SNMP community strings explicitly after flashing - a firmware update does not necessarily reset a configured default.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-11489"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2020-10-29"},{"id":"CVE-2020-11615","cve":"CVE-2020-11615","aliases":[],"title":"NVIDIA DGX BMC (AMI firmware): Hard-coded RC4 key in the DGX BMC firmware","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVIDIA DGX BMC (AMI firmware)","year":"2020","cvss_score":7.5,"severity":"high","kev":false,"impact":"Hard-coded RC4 key in the DGX BMC firmware. RC4 is broken to begin with and the key is public, so anything the BMC protects with it should be treated as cleartext. DGX-1 before BMC 3.38.30.","attack_vector":"Anyone who can capture BMC traffic or read BMC-stored data, plus anyone who downloads the firmware image.","remediation":"Flash the DGX BMC firmware from NVIDIA's DGX firmware update container (DGX-1 to 3.38.30 or later, DGX-2 to 1.06.06 or later; DGX A100 per the bulletin's table). A BMC flash does not require the host OS to reboot but drops out-of-band management for several minutes and NVIDIA recommends a host power cycle afterwards, so treat it as a per-node maintenance window. Rotate every BMC and IPMI credential after the flash - flashing does not invalidate secrets an attacker already pulled. Keep BMCs on an isolated management VLAN with no route from tenant or job networks.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-11615"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2020-10-29"},{"id":"CVE-2020-11616","cve":"CVE-2020-11616","aliases":[],"title":"NVIDIA DGX BMC (AMI firmware): The PRNG used by the IPMI implementation in the BMC's JSOL package","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVIDIA DGX BMC (AMI firmware)","year":"2020","cvss_score":7.5,"severity":"high","kev":false,"impact":"The PRNG used by the IPMI implementation in the BMC's JSOL package is not cryptographically strong, so session identifiers and other 'random' values are predictable. Combined with the rest of this bulletin it removes the guesswork from hijacking BMC/IPMI sessions. DGX-1 before BMC 3.38.30.","attack_vector":"Anyone with network reach to the BMC's IPMI service.","remediation":"Flash the DGX BMC firmware from NVIDIA's DGX firmware update container (DGX-1 to 3.38.30 or later, DGX-2 to 1.06.06 or later; DGX A100 per the bulletin's table). A BMC flash does not require the host OS to reboot but drops out-of-band management for several minutes and NVIDIA recommends a host power cycle afterwards, so treat it as a per-node maintenance window. Rotate every BMC and IPMI credential after the flash - flashing does not invalidate secrets an attacker already pulled. Keep BMCs on an isolated management VLAN with no route from tenant or job networks.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-11616"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2020-10-29"},{"id":"CVE-2020-11868","cve":"CVE-2020-11868","aliases":[],"title":"ntpd (NTP.org reference implementation): An off-path attacker can block a node's unauthenticated time synchronization","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"ntpd (NTP.org reference implementation)","year":"2020","cvss_score":7.5,"severity":"high","kev":false,"impact":"An off-path attacker can block a node's unauthenticated time synchronization by spoofing the source IP in a server-mode packet. Clock drift is an underrated cluster failure mode: it breaks TLS validity windows, Kerberos, distributed tracing correlation, checkpoint ordering, and any scheduler that reasons about deadlines. An attacker who can freeze your clocks without touching your data plane has a quiet, hard-to-attribute lever.","attack_vector":"Off-path attacker able to spoof source addresses toward the NTP client. Does not require being on the path between client and server.","remediation":"Upgrade ntp to 4.2.8p14 or later (or migrate to chrony, which most modern distributions default to) and restart the service. Package upgrade, no reboot. The structural fix is authenticated time — NTS or symmetric-key NTP — plus internal stratum-1 sources rather than public pools; that is a config and topology change and it is what actually removes this class.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-11868","https://nvd.nist.gov/vuln/detail/CVE-2020-13817"],"status":"curated","fleet":{"pain_class":"daemon-restart"},"published":"2020-04-17"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2020-12059","cve":"CVE-2020-12059","aliases":[],"title":"Ceph RADOS Gateway (RGW): A POST carrying malformed object-tagging XML dereferences a NULL pointer and kills the","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ceph RADOS Gateway (RGW)","year":"2020","cvss_score":7.5,"severity":"high","kev":false,"impact":"A POST carrying malformed object-tagging XML dereferences a NULL pointer and kills the gateway. One tenant can knock the shared S3 service over on demand with a single request.","attack_vector":"Any client able to send an HTTP POST to the RGW S3 endpoint.","remediation":"Upgrade RGW past 13.2.9 and restart the radosgw daemons. Keep more than one gateway in the pool so a targeted crash does not become a full outage.","references":["https://access.redhat.com/security/cve/CVE-2020-12059","https://nvd.nist.gov/vuln/detail/CVE-2020-12059"],"status":"curated"},{"id":"CVE-2020-12965","cve":"CVE-2020-12965","aliases":["Transient Execution of Non-canonical Accesses"],"title":"AMD processors - transient non-canonical loads and stores using lower 48 address bits: Combined with specific software","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD processors - transient non-canonical loads and stores using lower 48 address bits","year":"2020","cvss_score":7.5,"severity":"high","kev":false,"impact":"Combined with specific software sequences, AMD CPUs transiently execute non-canonical loads and stores using only the lower 48 address bits - so an access that should fault instead speculatively reads a truncated address. That truncation can land inside another security domain's memory, and the result is observable through the usual cache channels. Data leakage across the boundaries the address canonicality check was supposed to enforce.","attack_vector":"Local, needs the victim to contain a specific software sequence, so exploitability depends on what is running - but on a shared node you do not control what your tenants run.","remediation":"Mitigated by AMD microcode plus, on most of these, a kernel-side change - and the durable delivery vehicle is the OEM SBIOS/AGESA package, which carries **one to six months of OEM lag** and needs a drained node and a full power cycle. The linux-firmware amd-ucode blobs get you the microcode sooner via initramfs early-load and a reboot, but AMD does not support late-loading microcode on a running EPYC host, so either way this is reboot-required, not a live patch. Kernel-side mitigations exist for the known sequences; take the distro kernel update as well as the firmware.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-12965","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2022-02-04"},{"id":"CVE-2020-15383","cve":"CVE-2020-15383","aliases":[],"title":"Brocade Fabric OS (config and secnotify processes): Running a routine security scan against the SAN switch crashes","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Brocade Fabric OS (config and secnotify processes)","year":"2020","cvss_score":7.5,"severity":"high","kev":false,"impact":"Running a routine security scan against the SAN switch crashes the config and secnotify processes. Worth carrying because it inverts the usual advice: the compliance activity you are required to perform is itself the outage. Operators who scan their storage fabric on a schedule have been taking unexplained SAN switch faults from their own tooling.","attack_vector":"Any security scanner reaching the switch's management services — no attacker required.","remediation":"Fabric OS upgrade to v9.0.0 / v8.2.2d / v8.2.1e or later, firmware install plus reboot. Until then, exclude FOS management addresses from automated vulnerability scans or scan them only in a maintenance window.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-15383"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-06-09"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-200"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2020-1699","cve":"CVE-2020-1699","aliases":[],"title":"Ceph dashboard (ceph-mgr dashboard module): An unauthenticated HTTP request with traversal sequences reads arbitrary","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ceph dashboard (ceph-mgr dashboard module)","year":"2020","cvss_score":7.5,"severity":"high","kev":false,"impact":"An unauthenticated HTTP request with traversal sequences reads arbitrary files off the manager node. On a real cluster that means pulling /etc/ceph/ceph.client.admin.keyring and taking full administrative control of all storage, so it is effectively a pre-auth path to cluster admin.","attack_vector":"Anyone who can reach the dashboard's HTTP port. Dashboards are frequently left reachable from the management or tenant VLAN, which is what makes this severe.","remediation":"Upgrade ceph-mgr to 14.2.7 / 15.1.0 or later and restart the mgr. Assume the admin keyring leaked and rotate it. Bind the dashboard to a management-only interface behind authentication rather than exposing it on a shared network.","references":["https://access.redhat.com/security/cve/CVE-2020-1699","https://nvd.nist.gov/vuln/detail/CVE-2020-1699"],"status":"curated"},{"id":"CVE-2020-25632","cve":"CVE-2020-25632","aliases":[],"title":"GRUB2 (rmmod command): Use-after-free in the rmmod command","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (rmmod command)","year":"2020","cvss_score":7.5,"severity":"high","kev":false,"impact":"Use-after-free in the rmmod command. Unloading a module whose dependencies are still live leaves dangling pointers GRUB will later call through, which is a clean primitive for arbitrary pre-boot execution and another Secure Boot bypass.","attack_vector":"Local, via GRUB command line or a controlled grub.cfg. On a bare-metal fleet, any tenant who had console or root on the node.","remediation":"grub2 package update + reboot per node. If you leave the GRUB command line unlocked on your image, set a GRUB password as a stopgap - it does not fix the bug but it removes the easiest path to it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-25632","https://access.redhat.com/security/cve/CVE-2020-25632"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-03-03"},{"id":"CVE-2020-27174","cve":"CVE-2020-27174","aliases":[],"title":"Firecracker: Unbounded serial console buffer growth leaks host memory","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Firecracker","year":"2020","cvss_score":7.5,"severity":"high","kev":false,"impact":"Unbounded serial console buffer growth leaks host memory; host exhaustion from inside a guest","attack_vector":"Any tenant guest VM","remediation":"Upgrade Firecracker","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-27174"],"status":"curated","published":"2020-10-16"},{"id":"CVE-2020-27749","cve":"CVE-2020-27749","aliases":[],"title":"GRUB2 (grub_parser_split_cmdline): Stack buffer overflow from variable expansion in the GRUB command line","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (grub_parser_split_cmdline)","year":"2020","cvss_score":7.5,"severity":"high","kev":false,"impact":"Stack buffer overflow from variable expansion in the GRUB command line. Attacker gets code execution before the kernel and can therefore load an unsigned kernel or plant a bootkit that no in-OS EDR will see.","attack_vector":"Anyone who can type at the GRUB prompt or supply grub.cfg - local console, serial console server, or BMC KVM.","remediation":"grub2 package update + reboot per node. Set a GRUB password and lock the serial/BMC console as a partial mitigation in the meantime.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-27749","https://access.redhat.com/security/cve/CVE-2020-27749"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-03-03"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","fleet":{"pain_class":"node-reboot"},"id":"CVE-2020-5015","cve":"CVE-2020-5015","aliases":[],"title":"IBM Elastic Storage System / Elastic Storage Server (UDP request handling): An unauthenticated attacker who can send","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"IBM Elastic Storage System / Elastic Storage Server (UDP request handling)","year":"2020","cvss_score":7.5,"severity":"high","kev":false,"impact":"An unauthenticated attacker who can send UDP to an ESS node knocks storage service over with malformed packets. This is the whole appliance that the GPU fleet reads training data from, so a single spoofable UDP flow stalls the cluster.","attack_vector":"Network reach to the ESS management or data interfaces. UDP, so it is spoofable and does not need a completed handshake or any account.","remediation":"Upgrade ESS to 6.0.1.3 / 5.3.6.3 or later. In the meantime, filter the affected UDP ports at the fabric so only known cluster members can send to them.","references":["https://www.ibm.com/support/pages/node/6434155","https://nvd.nist.gov/vuln/detail/CVE-2020-5015"],"status":"curated"},{"id":"CVE-2021-23201","cve":"CVE-2021-23201","aliases":[],"title":"NVIDIA GPU firmware microcontroller (Falcon): A privileged user can craft microcode that the GPU's internal","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU firmware microcontroller (Falcon)","year":"2021","cvss_score":7.5,"severity":"high","kev":false,"impact":"A privileged user can craft microcode that the GPU's internal microcontroller accepts as valid. That means loading attacker-controlled firmware onto the GPU, below the driver and below the OS - persistent, invisible to host tooling, and NVIDIA explicitly notes the scope may extend to other components. On a bare-metal GPU rental this is the tenant-persistence scenario: tenant A leaves code on the card that outlives the reprovision and is there when tenant B arrives. Verify with the vendor whether a full VBIOS and firmware reflash between tenants is sufficient.","attack_vector":"A user with elevated privileges on the host. On rented bare metal that is the tenant; on a managed cluster it is anyone who got root on a node.","remediation":"NVIDIA shipped the fix in GPU firmware/microcode delivered with the R470 and R450 driver branches and, on some SKUs, in an updated VBIOS. On most datacenter parts the microcontroller image is loaded by the driver at GPU init, so a driver upgrade plus a node reboot applies it; check the bulletin's product table, because a subset of boards also needs an out-of-band VBIOS/InfoROM update, which is an offline per-node flash with the GPU idle. Either way the node has to be drained. Because this is a firmware-persistence risk, also add a firmware/VBIOS attestation or reflash step to your bare-metal tenant handoff - patching alone does not tell you whether a previous tenant already loaded something.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-23201"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2021-11-20"},{"id":"CVE-2021-23217","cve":"CVE-2021-23217","aliases":[],"title":"NVIDIA GPU firmware microcontroller (Falcon): A privileged user can time a DMA write from the GPU's internal","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU firmware microcontroller (Falcon)","year":"2021","cvss_score":7.5,"severity":"high","kev":false,"impact":"A privileged user can time a DMA write from the GPU's internal microcontroller to land inside a specific window and corrupt code execution, with impact NVIDIA says may extend to other components. A DMA engine writing outside its lane is the mechanism by which a GPU compromises its host.","attack_vector":"A user with elevated privileges on the GPU host, able to time operations precisely.","remediation":"NVIDIA shipped the fix in GPU firmware/microcode delivered with the R470 and R450 driver branches and, on some SKUs, in an updated VBIOS. On most datacenter parts the microcontroller image is loaded by the driver at GPU init, so a driver upgrade plus a node reboot applies it; check the bulletin's product table, because a subset of boards also needs an out-of-band VBIOS/InfoROM update, which is an offline per-node flash with the GPU idle. Either way the node has to be drained.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-23217"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2021-11-20"},{"id":"CVE-2021-26406","cve":"CVE-2021-26406","aliases":[],"title":"AMD SEV / SEV-ES - Owner's Certificate Authority (OCA) certificate parsing: Insufficient validation when parsing OCA","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV / SEV-ES - Owner's Certificate Authority (OCA) certificate parsing","year":"2021","cvss_score":7.5,"severity":"high","kev":false,"impact":"Insufficient validation when parsing OCA certificates in the SEV and SEV-ES user application crashes the host. OCA certificates come from whoever owns the platform's SEV identity, so this is a malformed-input crash on the certificate path that underpins SEV ownership and attestation - a guest or tooling that supplies a bad certificate takes the machine down.","attack_vector":"Reachable by whatever supplies OCA certificates to the SEV stack, which in a managed confidential-computing service is the control plane or the tenant-facing provisioning path.","remediation":"Fixed in AMD reference firmware (AGESA / PSP / SEV firmware) and delivered only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo and the ODMs each rebuild and requalify AMD's AGESA drop before shipping. **Expect one to six months of OEM lag**, and on end-of-support platforms expect nothing. Applying it is a drain plus full power cycle, not a driver reload. Verify by reading back the PSP/SMU firmware version afterwards rather than trusting the BIOS version string. This sits inside the SEV-SNP trust boundary, so the update moves the platform's reported TCB version: refresh VCEK certificates from AMD's KDS and update any attestation policy your tenants pin, or confidential guest launches will start failing right after the BIOS lands.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26406","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-05-09"},{"id":"CVE-2021-26923","cve":"CVE-2021-26923","aliases":[],"title":"Argo CD: /api/version leaks internal system information without authentication","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2021","cvss_score":7.5,"severity":"high","kev":false,"impact":"/api/version leaks internal system information without authentication","attack_vector":"Unauthenticated network","remediation":"Rolling Argo CD upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26923"],"status":"curated","published":"2021-03-15"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","fleet":{"pain_class":"node-reboot"},"id":"CVE-2021-27005","cve":"CVE-2021-27005","aliases":[],"title":"NetApp Clustered Data ONTAP httpd: A remote attacker with no credentials crashes the ONTAP web server, removing","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"NetApp Clustered Data ONTAP httpd","year":"2021","cvss_score":7.5,"severity":"high","kev":false,"impact":"A remote attacker with no credentials crashes the ONTAP web server, removing management and API access to the array until it recovers.","attack_vector":"Network path to httpd on Clustered Data ONTAP 9.6 and later below 9.6P16, 9.7P16, 9.8P7 or 9.9.1P3.","remediation":"Upgrade to the fixed patch level and limit which subnets can reach the management LIF.","references":["https://security.netapp.com/advisory/NTAP-20211029-0002/","https://nvd.nist.gov/vuln/detail/CVE-2021-27005"],"status":"curated"},{"id":"CVE-2021-28361","cve":"CVE-2021-28361","aliases":["CVE-2019-9547"],"title":"SPDK iSCSI target (before 20.01.01) and SPDK vhost target (before 19.01): A zero-length PDU sent where data is expected","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"SPDK iSCSI target (before 20.01.01) and SPDK vhost target (before 19.01)","year":"2021","cvss_score":7.5,"severity":"high","kev":false,"impact":"A zero-length PDU sent where data is expected crashes the SPDK iSCSI target on a NULL pointer dereference. The companion vhost defect lets a guest VM build a circular descriptor chain that partially wedges the SPDK vhost target. Both are single-packet or single-guest denial of service against a userspace process that is serving block storage to many tenants at once - the iSCSI one needs no authentication and no valid target, and the vhost one is reachable from inside any guest whose virtio-blk/virtio-scsi device SPDK is backing. On a node running SPDK as the storage datapath for a rack of VMs, one guest kills storage for all of them.","attack_vector":"iSCSI variant: any host that can connect to the SPDK iSCSI target port (3260) on the storage network, unauthenticated. vhost variant: a malicious or compromised guest VM whose virtio block device is served by SPDK vhost - i.e. a paying tenant.","remediation":"Upgrade SPDK past 20.01.01 (iSCSI) and 19.01 (vhost) and restart the target process, which disconnects all sessions on that node. Anyone still on an SPDK that old is likely running a vendored fork inside a storage appliance image, so the practical action is to identify the SPDK version compiled into your storage service, not to check a package manifest. For the vhost issue there is no configuration mitigation - the attacker is inside the guest by definition, which is the whole point of the exposure.","references":["https://github.com/spdk/spdk/commit/eca42c66092b9031711afe215fbc1891ee55f143","https://nvd.nist.gov/vuln/detail/CVE-2021-28361","https://nvd.nist.gov/vuln/detail/CVE-2019-9547"],"status":"curated","published":"2021-03-13"},{"id":"CVE-2021-28505","cve":"CVE-2021-28505","aliases":[],"title":"Arista EOS (VXLAN match rule in IPv4 ACL): If an IPv4 access list contains a VXLAN match rule, that rule and every rule","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (VXLAN match rule in IPv4 ACL)","year":"2021","cvss_score":7.5,"severity":"high","kev":false,"impact":"If an IPv4 access list contains a VXLAN match rule, that rule and every rule after it in the list ignore the IP protocol you specified. Your ACL silently permits or denies far more than you wrote. Anyone using ACLs to keep tenant overlays apart, or to fence off a storage VLAN, is enforcing something other than what is in the config — and `show access-list` will not tell you.","attack_vector":"Any traffic subject to the affected ACL. No attacker capability needed beyond being on a path the ACL was supposed to control.","remediation":"EOS upgrade plus reload. Immediate mitigation: reorder access lists so VXLAN match rules come last, or split them into a separate list — a live config change that restores correct enforcement of the remaining rules. Audit every ACL in the fabric for VXLAN match rules before assuming you are unaffected.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-28505"],"status":"curated","tags":["tenant-isolation"],"published":"2022-04-14"},{"id":"CVE-2021-28508","cve":"CVE-2021-28508","aliases":[],"title":"Arista EOS (TerminAttr / IPsec): TerminAttr leaks IPsec sensitive material in plaintext to authorized users","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (TerminAttr / IPsec)","year":"2021","cvss_score":7.5,"severity":"high","kev":false,"impact":"TerminAttr leaks IPsec sensitive material in plaintext to authorized users — fabric encryption keys exposed through the telemetry agent","attack_vector":"Local/network","remediation":"EOS + TerminAttr upgrade plus rotation of any IPsec keys that were live on the affected switches","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-28508"],"status":"curated","published":"2022-05-26"},{"id":"CVE-2021-3695","cve":"CVE-2021-3695","aliases":[],"title":"GRUB2 (PNG reader): A crafted PNG in the boot splash path causes an out-of-bounds write in GRUB","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (PNG reader)","year":"2021","cvss_score":7.5,"severity":"high","kev":false,"impact":"A crafted PNG in the boot splash path causes an out-of-bounds write in GRUB. Boot logos and themes are attacker-writable data on almost every image, so this is a low-effort way to turn cosmetic files into pre-boot code execution.","attack_vector":"Anyone who can replace a theme/splash image on the boot partition - local root, previous tenant, or a tampered golden image.","remediation":"grub2 package update + reboot. Strip custom boot themes from your golden image if you do not need them; it removes the attack surface entirely at zero operational cost.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-3695","https://access.redhat.com/security/cve/CVE-2021-3695"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2022-07-06"},{"id":"CVE-2021-3697","cve":"CVE-2021-3697","aliases":[],"title":"GRUB2 (JPEG reader): Crafted JPEG in the boot path drives a heap out-of-bounds write in GRUB","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (JPEG reader)","year":"2021","cvss_score":7.5,"severity":"high","kev":false,"impact":"Crafted JPEG in the boot path drives a heap out-of-bounds write in GRUB. This is the GRUB-side sibling of the firmware image-parser problem that LogoFAIL exploited a year later - same idea, different layer.","attack_vector":"Attacker-writable splash/theme file on the boot partition.","remediation":"grub2 package update + reboot. Removing boot theme images from the image is a real mitigation here.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-3697","https://access.redhat.com/security/cve/CVE-2021-3697"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2022-07-06"},{"id":"CVE-2021-3748","cve":"CVE-2021-3748","aliases":[],"title":"QEMU (virtio-net): Heap use-after-free in virtio_net_receive_rcu - guest-to-host code execution in the QEMU process","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"QEMU (virtio-net)","year":"2021","cvss_score":7.5,"severity":"high","kev":false,"impact":"Heap use-after-free in virtio_net_receive_rcu - guest-to-host code execution in the QEMU process","attack_vector":"Tenant VM guest","remediation":"QEMU package update + restart each VM's qemu process. Live-migrate to patched hosts to avoid tenant downtime; GPU-passthrough VMs cannot live-migrate, so this becomes a scheduled drain","references":["https://access.redhat.com/security/cve/CVE-2021-3748"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2022-03-23"},{"id":"CVE-2021-38960","cve":"CVE-2021-38960","aliases":["IBM X-Force 212047"],"title":"IBM OpenBMC OP920 / OP930 / OP940: An unauthenticated caller retrieves sensitive information from the BMC","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"IBM OpenBMC OP920 / OP930 / OP940","year":"2021","cvss_score":7.5,"severity":"high","kev":false,"impact":"An unauthenticated caller retrieves sensitive information from the BMC. Same operational shape as the 2024 bmcweb URI disclosure three years earlier and on the same product line, which is the useful signal: pre-auth information exposure on this BMC stack is recurring, not a one-off. For a fleet operator the leaked material is inventory and configuration detail that lets an attacker pick which nodes to attack and with what.","attack_vector":"Unauthenticated network access to the BMC's management interface.","remediation":"Fixed in later OP920/OP930/OP940 firmware - a per-node system firmware update requiring a maintenance window. Because this class keeps recurring on the same stack, the durable control is the network boundary: BMCs on an isolated management VLAN reachable only from bastion hosts, so pre-auth disclosure bugs have no audience. That is config-only and it covers the next one too.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-38960","https://www.ibm.com/support/pages/node/6529322","https://exchange.xforce.ibmcloud.com/vulnerabilities/212047"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-02-04"},{"id":"CVE-2021-39295","cve":"CVE-2021-39295","aliases":["GHSA-gg9x-v835-m48q"],"title":"OpenBMC phosphor-net-ipmid (IPMI LAN+): Sibling finding to the authentication bypass, from the same Google report","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"OpenBMC phosphor-net-ipmid (IPMI LAN+)","year":"2021","cvss_score":7.5,"severity":"high","kev":false,"impact":"Sibling finding to the authentication bypass, from the same Google report. Crafted IPMI messages take the BMC's IPMI daemon down without any credentials. Losing IPMI on its own is survivable if you run Redfish, but the operational shape is bad: an unauthenticated packet source on the management VLAN can knock out out-of-band management across every ASPEED node simultaneously, which is precisely when you would want it - during an incident, or to blind an operator while something else happens on the hosts.","attack_vector":"Unauthenticated, network, UDP 623 on the BMC. Same reachability precondition as the authentication bypass.","remediation":"Same fix and same delivery cost as the authentication bypass: post-2.9 OpenBMC via a per-node out-of-band BMC firmware flash. Config-only mitigation is the same and is the right first move: turn off IPMI over LAN and run Redfish, or ACL UDP 623 to your management jump hosts. If you are already flashing for CVE-2021-39296 you get this one in the same image.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-39295","https://github.com/google/security-research/security/advisories/GHSA-gg9x-v835-m48q","https://github.com/openbmc/openbmc/issues/3811"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-04-15"},{"id":"CVE-2021-43798","cve":"CVE-2021-43798","aliases":[],"title":"Grafana: Unauthenticated directory traversal via /public/plugins/<id>/","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Grafana","year":"2021","cvss_score":7.5,"severity":"high","kev":true,"impact":"Unauthenticated directory traversal via /public/plugins/<id>/ -> read local files incl. grafana.db","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; rotate every datasource credential stored in grafana.db","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-43798"],"status":"curated","published":"2021-12-07"},{"id":"CVE-2021-45454","cve":"CVE-2021-45454","aliases":["AMP-SB-0003","PLATYPUS on Ampere","power telemetry side channel"],"title":"Ampere Altra before SRP 1.08b and Altra Max before SRP 2.05","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ampere Altra before SRP 1.08b and Altra Max before SRP 2.05 - power telemetry exposed through the Linux HWmon interface","year":"2021","cvss_score":7.5,"severity":"high","kev":false,"impact":"Unprivileged readers get fine-grained CPU power telemetry, which is a data-dependent side channel: power correlates with the operands being processed, so it leaks key material and other secrets from workloads sharing the socket. On a multi-tenant Arm node this is a cross-tenant leak that needs no memory access at all - just a file read in sysfs. It is also a nuisance for anyone selling confidential inference, because the power trace of a model serving run is itself informative about the workload.","attack_vector":"Any unprivileged local user or container on an Altra / Altra Max host with the HWmon power sensors exposed. Containers that inherit the host sysfs make this trivially available to tenants.","remediation":"Update to Altra SRP 1.08b / Altra Max SRP 2.05 or later, which restricts the telemetry. Flash + reboot + drain. Cheaper interim control that works today: restrict access to the HWmon power sensors (root-only permissions, do not bind-mount host /sys into tenant containers, drop the sensor nodes from the container's device allowlist). Losing per-core power telemetry costs you some capacity-planning visibility - decide whether your scheduler actually consumes it before turning it off fleet-wide.","references":["https://amperecomputing.com/products/security-bulletins/platypus.html","https://nvd.nist.gov/vuln/detail/CVE-2021-45454"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2022-08-17"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2021-46983","cve":"CVE-2021-46983","aliases":[],"title":"Linux kernel NVMe-oF RDMA target (nvmet-rdma error completion handling with shared CQ): After the switch to shared","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel NVMe-oF RDMA target (nvmet-rdma error completion handling with shared CQ)","year":"2021","cvss_score":7.5,"severity":"high","kev":false,"impact":"After the switch to shared completion queues the cq_context no longer identifies the queue, but the SEND error handler still used it - so a transport-level error such as retry-counter-exceeded dereferences a stale pointer and crashes the target. The trigger is link disruption on the RDMA fabric, which a co-tenant can induce (link flap, congestion, or simply disconnecting mid-transfer) without any access to the target itself.","attack_vector":"Remote/fabric. Any condition that produces a SEND completion error on the target's RDMA queue pairs - reachable by a peer able to disturb the fabric path.","remediation":"Kernel update obtaining the queue from wc->qp instead of cq_context. Fabric-level: keep storage RDMA traffic on its own partition/VLAN so tenant-induced congestion does not reach the target's queue pairs.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=17fb6dfa5162b89ecfa07df891a53afec321abe8","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2021/CVE-2021-46983.json"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-763"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2021-47130","cve":"CVE-2021-47130","aliases":[],"title":"Linux kernel (drivers/nvme/target): When the target's peer-to-peer memory pool runs dry, it still tries to return the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/nvme/target)","year":"2021","cvss_score":7.5,"severity":"high","kev":false,"impact":"When the target's peer-to-peer memory pool runs dry, it still tries to return the request's scatterlist to the P2P pool instead of the regular one, hitting a kernel BUG() in the allocator and taking the storage node down. Client I/O volume alone decides when the pool runs dry, so a peer can drive the panic.","attack_vector":"Driven by a connected NVMe-oF client's I/O against a target configured with a P2P memory device - enough concurrent requests to exhaust the p2pmem pool is sufficient, no crafted command needed. The observed crash path is nvmet-rdma completion, i.e. a peer on the RDMA fabric. Conditional on the target having p2pmem enabled (CMB / GPUDirect-Storage style setups), which is exactly the configuration an AI-storage node is likely to run.","remediation":"No fixed version is listed on this record - boot a kernel carrying the linked stable commits. Interim: disable p2pmem on the nvmet port/subsystem so allocations fall back to the regular SGL pool, or size the P2P pool so it cannot be exhausted by expected client concurrency.","references":["https://git.kernel.org/stable/c/c440cd080761b18a52cac20f2a42e5da1e3995af","https://git.kernel.org/stable/c/8a452d62e7cea3c8a2676a3b89a9118755a1a271","https://nvd.nist.gov/vuln/detail/CVE-2021-47130"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2022-1708","cve":"CVE-2022-1708","aliases":[],"title":"CRI-O: Unbounded ExecSync output exhausts node memory or disk","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"CRI-O","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"Unbounded ExecSync output exhausts node memory or disk; node DoS","attack_vector":"Anyone with kube API exec access","remediation":"Upgrade CRI-O; drain node","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-1708"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2022-06-07"},{"id":"CVE-2022-23635","cve":"CVE-2022-23635","aliases":[],"title":"Istio: Crafted message crashes istiod","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Istio","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"Crafted message crashes istiod; mesh-wide config distribution stops","attack_vector":"Any pod that can reach istiod, i.e. any tenant pod","remediation":"Rolling istiod upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-23635"],"status":"curated","published":"2022-02-22"},{"id":"CVE-2022-23648","cve":"CVE-2022-23648","aliases":[],"title":"containerd: Crafted image config allows arbitrary host file read by containers launched via the CRI plugin","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"Crafted image config allows arbitrary host file read by containers launched via the CRI plugin","attack_vector":"Malicious image","remediation":"Rolling containerd upgrade with node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-23648"],"status":"curated","fleet":{"ubiquity":"Universal - containerd is the runtime under nearly every Kubernetes-based GPU cloud; affects <1.6.1 / 1.5.10 / 1.4.12","remediation_pain":"`daemon-restart` - containerd upgrade; with `--restart` semantics containers may survive, but the fleet-wide rollout still means touching every node","pain_class":"daemon-restart","why_fleet_wide":"A specially crafted *image config* - i.e. something a customer supplies - mounts read-only copies of arbitrary host files into the container, bypassing Pod Security Policy; any tenant who can push an image reads host secrets on every node they land on"},"published":"2022-03-03"},{"id":"CVE-2022-23818","cve":"CVE-2022-23818","aliases":[],"title":"AMD SEV-SNP - VM_HSAVE_PA MSR validation: Insufficient validation of the VM_HSAVE_PA model-specific register lets a","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV-SNP - VM_HSAVE_PA MSR validation","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"Insufficient validation of the VM_HSAVE_PA model-specific register lets a malicious hypervisor point the host-state save area somewhere it should not be, breaking SEV-SNP guest memory integrity. The RMP is supposed to make it impossible for the host to write guest pages; this is a way around that using an MSR the host legitimately controls.","attack_vector":"Requires host/hypervisor privilege - the SEV-SNP adversary model exactly. No guest cooperation needed.","remediation":"Fixed in AMD SEV firmware / AGESA and reaches you as an OEM SBIOS package - AMD hands AGESA to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before shipping BIOS. **Budget one to six months of OEM lag**, longer on older platforms and sometimes never on end-of-support SKUs. Applying it means draining the host and doing a full power cycle. Because the fix moves the platform's reported SEV-SNP TCB version, you must also pull fresh VCEK certificates from AMD's Key Distribution Service and update any attestation policy your tenants pin - otherwise guests will start failing launch validation the moment the BIOS lands. Some SEV firmware can alternatively be staged from linux-firmware (amd/amd_sev_*.sbin) and committed via the ccp driver at boot, which is faster than waiting on BIOS - check whether your platform supports firmware hot-load before assuming the OEM is the only route.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-23818","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2023-05-09"},{"id":"CVE-2022-25882","cve":"CVE-2022-25882","aliases":[],"title":"ONNX: Directory traversal via `external_data` field in the tensor proto","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"ONNX","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"Directory traversal via `external_data` field in the tensor proto","attack_vector":"Customer-supplied ONNX model file","remediation":"Upgrade onnx >= 1.13.0. Every ONNX ingest path must resolve external-data paths against a jail","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-25882"],"status":"curated","published":"2023-01-26"},{"id":"CVE-2022-26353","cve":"CVE-2022-26353","aliases":[],"title":"QEMU (virtio-net): Map leaking on error during receive - guest-triggered host memory exhaustion / DoS","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"QEMU (virtio-net)","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"Map leaking on error during receive - guest-triggered host memory exhaustion / DoS","attack_vector":"Tenant VM guest","remediation":"QEMU update + VM restart or live-migration","references":["https://access.redhat.com/security/cve/CVE-2022-26353"],"status":"curated","published":"2022-03-16"},{"id":"CVE-2022-27649","cve":"CVE-2022-27649","aliases":[],"title":"Podman: Containers started with non-empty default inheritable capabilities","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Podman","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"Containers started with non-empty default inheritable capabilities","attack_vector":"Any tenant workload","remediation":"Upgrade Podman; restart containers","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-27649"],"status":"curated","published":"2022-04-04"},{"id":"CVE-2022-2809","cve":"CVE-2022-2809","aliases":["GHSA-g3qc-375m-h66j"],"title":"OpenBMC bmcweb multipart_parser (Redfish / web UI HTTP front end): bmcweb is the single process behind Redfish, the web","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"OpenBMC bmcweb multipart_parser (Redfish / web UI HTTP front end)","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"bmcweb is the single process behind Redfish, the web UI, serial-over-LAN and the KVM websocket - kill it and you lose every out-of-band control path at once. A multipart form body containing a long header line with no colon writes one byte off the end of a heap buffer, and the write can be repeated in a loop. The reported outcome is denial of service, but the advisory classifies it as both heap and stack out-of-bounds writes reachable without authentication, which is a materially worse posture than the DoS framing suggests. Fleet impact: an unauthenticated source on the management VLAN can flatten Redfish on every ASPEED node.","attack_vector":"Unauthenticated HTTP(S) to bmcweb on the BMC's management interface. Multipart upload endpoints are reachable pre-auth in affected versions, so no credentials and no host access are needed.","remediation":"Fixed in bmcweb 2.13 / OpenBMC 2.13 (Gerrit 56796 and 56868). Delivery is a BMC firmware flash - per node, out-of-band, and gated on when your ODM last rebased bmcweb, which for many server vendors is a long time. Check the bmcweb version your image reports before assuming you are covered. Config-only mitigation: restrict which hosts can reach the BMC's HTTPS port to your management jump boxes, which is the same control that limits half of this cluster.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-2809","https://github.com/openbmc/bmcweb/security/advisories/GHSA-g3qc-375m-h66j"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-10-27"},{"id":"CVE-2022-2827","cve":"CVE-2022-2827","aliases":[],"title":"AMI MegaRAC: User enumeration — lets an attacker map valid BMC accounts before credential attack","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"User enumeration — lets an attacker map valid BMC accounts before credential attack","attack_vector":"Network","remediation":"BMC firmware update; low individual severity but it is the reconnaissance step for the rest of the MegaRAC chain","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-2827"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-12-05"},{"id":"CVE-2022-29179","cve":"CVE-2022-29179","aliases":[],"title":"Cilium: After a container escape, an attacker can install eBPF programs and take over the node dataplane","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"After a container escape, an attacker can install eBPF programs and take over the node dataplane","attack_vector":"An attacker who already escaped a container","remediation":"Rolling Cilium upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-29179"],"status":"curated","published":"2022-05-20"},{"id":"CVE-2022-29225","cve":"CVE-2022-29225","aliases":[],"title":"Envoy: Decompressor accumulates unbounded data","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"Decompressor accumulates unbounded data; memory exhaustion of the proxy","attack_vector":"Unauthenticated network","remediation":"Upgrade Envoy","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-29225"],"status":"curated","published":"2022-06-09"},{"id":"CVE-2022-30283","cve":"CVE-2022-30283","aliases":["INSYDE-SA-2022063"],"title":"Insyde InsydeH2O (UsbCoreDxe USB working buffer, DMA TOCTOU): UsbCoreDxe builds its USB transaction working buffer","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (UsbCoreDxe USB working buffer, DMA TOCTOU)","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"UsbCoreDxe builds its USB transaction working buffer outside SMRAM while the code consuming it runs inside SMM. The driver tries to sanitise pointers against a list of known-good buffer locations, but a pointer that misses the list is used anyway - so DMA tampering mid-transaction corrupts SMRAM and escalates to ring -2. A more subtle failure than the rest of the batch: the validation exists, it just does not fail closed.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in the kernel releases named in the advisory (Insyde does not enumerate per-kernel versions for this one). Disabling USB legacy support on headless nodes reduces how often the vulnerable transaction path runs. The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-30283","https://www.insyde.com/security-pledge/SA-2022063"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-11-15"},{"id":"CVE-2022-31600","cve":"CVE-2022-31600","aliases":[],"title":"NVIDIA DGX A100 - SBIOS / SMM firmware: An integer overflow in SmmCore, chainable from another bug, reaches SMM code","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX A100 - SBIOS / SMM firmware","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"An integer overflow in SmmCore, chainable from another bug, reaches SMM code execution and secure-boot-level compromise. This is firmware-level persistence: it survives OS reinstall, image re-flash and tenant handoff, and it is invisible to anything running above it. On a bare-metal GPU rental business it is the difference between wiping a node between tenants and not actually being able to.","attack_vector":"Local and already privileged - host root, or code that has reached the platform firmware/SMM path. It is not a first foothold; it is what turns a one-time root compromise into something you cannot remediate by reimaging.","remediation":"Flash the fixed SBIOS from bulletin 5367. Cost: not live-patchable. Full node drain, host power cycle, and on DGX the SBIOS ships inside a firmware bundle alongside BMC and CPLD components, so budget 30-60 minutes of node downtime plus a post-flash health check. Firmware rollback protection means you cannot cleanly revert - stage on one node before the fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31600","https://github.com/NVIDIA/product-security/tree/main/2022/5367"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-190"],"fleet":{"pain_class":"firmware-flash"},"published":"2022-07-04"},{"id":"CVE-2022-3409","cve":"CVE-2022-3409","aliases":["GHSA-g3qc-375m-h66j"],"title":"OpenBMC bmcweb multipart_parser (second variant found during the CVE-2022-2809 fix): The second bug the fuzzer found","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"OpenBMC bmcweb multipart_parser (second variant found during the CVE-2022-2809 fix)","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"The second bug the fuzzer found while the first one was being patched - same parser, same unauthenticated reachability, same result of taking down Redfish, KVM and SoL together. Its real value to an operator is as evidence about the code: bmcweb's multipart handling was not hardened, it was patched twice under fuzzing pressure, and the 2026 disclosures below show the same pattern repeating in the HTTP/2 and Expect-header paths. Treat bmcweb version currency as a standing fleet metric rather than a per-CVE chase.","attack_vector":"Unauthenticated HTTP(S) request to bmcweb on the management interface. No credentials, no host access.","remediation":"Same patch train as CVE-2022-2809 - bmcweb 2.13 and later, arriving as a BMC firmware flash per node, out-of-band, ODM-lagged. There is no separate action for this one. The operator-level control that actually pays: know the bmcweb version on every node in your fleet and put a floor on it in your acceptance criteria for ODM firmware drops.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-3409","https://github.com/openbmc/bmcweb/security/advisories/GHSA-g3qc-375m-h66j"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-10-27"},{"id":"CVE-2022-34425","cve":"CVE-2022-34425","aliases":[],"title":"Dell Enterprise SONiC OS (SSH cryptographic key): A cryptographic key weakness in SONiC's SSH implementation lets","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell Enterprise SONiC OS (SSH cryptographic key)","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"A cryptographic key weakness in SONiC's SSH implementation lets an unauthenticated remote attacker exploit the switch. Shared or predictable SSH host keys across a product line mean an attacker can impersonate any switch to your automation, harvesting the credentials your config-management pushes. CVE-2025-38741 is the same class recurring in SONiC 4.5.0, which tells you it is a build-pipeline problem, not a one-off.","attack_vector":"Unauthenticated, remote — anyone able to interpose on or reach the switch's SSH service.","remediation":"NOS image upgrade plus reboot, and then **regenerate the switch's SSH host keys** — the upgrade alone does not replace a key that was already weak or shared. Update your automation's known_hosts afterwards. Verify host-key uniqueness across the fleet as a standing check.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34425","https://nvd.nist.gov/vuln/detail/CVE-2025-38741"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2022-10-10"},{"id":"CVE-2022-35729","cve":"CVE-2022-35729","aliases":["INTEL-SA-00737"],"title":"Intel OpenBMC firmware (before version 0.72) - network-facing service: An unauthenticated caller reads out of bounds","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel OpenBMC firmware (before version 0.72) - network-facing service","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"An unauthenticated caller reads out of bounds and crashes the BMC. Intel shipped this in the same SA-00737 bundle as the OpenBMC IPMI authentication bypass, which is the useful context: Intel-badged OpenBMC below 0.72 carried both a pre-auth takeover and a pre-auth crash at the same time. Losing the BMC on Intel-platform GPU nodes means losing remote power and console; an attacker who wants an operator blind while working on hosts starts here.","attack_vector":"Unauthenticated, over the network, against the BMC management interface.","remediation":"Fixed in Intel OpenBMC 0.72 and later - a BMC firmware update per node, out-of-band, through Intel's platform update packages. If you are on an affected Intel platform you are almost certainly also exposed to CVE-2021-39296 from the same advisory, so treat this as one campaign rather than two. Config-only first move remains the same: ACL the management network and disable IPMI over LAN.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-35729","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00737.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-02-16"},{"id":"CVE-2022-3602","cve":"CVE-2022-3602","aliases":[],"title":"OpenSSL 3.0: X.509 email-address punycode buffer overflow (4-byte stack overflow)","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"OpenSSL 3.0","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"X.509 email-address punycode buffer overflow (4-byte stack overflow)","attack_vector":"Unauthenticated network (TLS peers/clients)","remediation":"Package update + restart consuming services; no reboot. Low real risk for a neocloud control plane but audit any mTLS front door","references":["https://access.redhat.com/security/cve/CVE-2022-3602"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2022-11-01"},{"id":"CVE-2022-37122","cve":"CVE-2022-37122","aliases":["ZSL-2022-5709"],"title":"Carel pCOWeb HVAC BACnet gateway 2.1.0 (logdownload.cgi): Unauthenticated arbitrary file read off the gateway","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Carel pCOWeb HVAC BACnet gateway 2.1.0 (logdownload.cgi)","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"Unauthenticated arbitrary file read off the gateway that bridges your BACnet field bus to IP. On its own that is information disclosure, but on this class of device the files worth reading are the ones that turn into control: stored web credentials, the BACnet device map, SNMP community strings, and the config that tells an attacker exactly which object instance is the CRAH fan-speed command and which is the chilled-water setpoint. This is the reconnaissance step that makes a subsequent thermal attack precise instead of a guess. Carel pCOWeb cards sit inside a very long list of OEM cooling products (chillers, CRAC/CRAH units, close-coupled cooling), so operators frequently have these in the hall without knowing the Carel name appears anywhere in their asset list.","attack_vector":"Unauthenticated HTTP GET on the facility network. No credentials, no user interaction. Reachability is the whole question: if your mechanical VLAN is flat with anything an attacker can phish into, this is a one-request win. These gateways also turn up on remote-access boxes that the mechanical contractor installed for support, which is how they end up internet-reachable.","remediation":"Carel firmware updates for pCOWeb are distributed through the OEM that embedded the card, not directly, so the practical path is: identify which of your cooling units carry a pCOWeb card (check the cooling vendor's BOM, not your CMDB), then ask that OEM for a firmware level that closes it. Expect a field technician and a maintenance window on live cooling, which most operators will not schedule for a file-read bug - so plan on segmentation as the actual control. Deny the gateway's web port from everything except the BMS supervisor, and log all access to it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-37122","https://www.zeroscience.mk/en/vulnerabilities/ZSL-2022-5709.php","https://packetstormsecurity.com/files/167684/"],"status":"curated"},{"id":"CVE-2022-3786","cve":"CVE-2022-3786","aliases":[],"title":"OpenSSL 3.0: X.509 email-address variable-length buffer overflow (DoS)","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"OpenSSL 3.0","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"X.509 email-address variable-length buffer overflow (DoS)","attack_vector":"Unauthenticated network","remediation":"Package update + service restart; no reboot","references":["https://access.redhat.com/security/cve/CVE-2022-3786"],"status":"curated","fleet":{"pain_class":"daemon-restart"},"published":"2022-11-01"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-798"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2022-39273","cve":"CVE-2022-39273","aliases":["GHSA-67x4-qr35-qvrm"],"title":"FlyteAdmin (built-in OAuth authorization server, default client secret hashes): Turning on Flyte's built-in","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"FlyteAdmin (built-in OAuth authorization server, default client secret hashes)","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"Turning on Flyte's built-in authorization server without replacing the shipped default client ID hashes leaves publicly known credentials in place. Anyone who reads the docs authenticates as FlytePropeller and reaches the FlyteAdmin control plane, which means enumerating and manipulating every tenant's workflow executions on a deployment the operator believes is authenticated.","attack_vector":"Any internet or network host that can reach FlyteAdmin, using the documented default client secrets. No prior access needed.","remediation":"Rotate the client ID hashes to values you generated, then upgrade FlyteAdmin to 1.1.44 or later and restart. Verify the running config does not still carry the shipped defaults after the upgrade - this is a configuration flaw that a version bump alone will not correct.","references":["https://github.com/flyteorg/flyteadmin/security/advisories/GHSA-67x4-qr35-qvrm","https://nvd.nist.gov/vuln/detail/CVE-2022-39273"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2022-39278","cve":"CVE-2022-39278","aliases":[],"title":"Istio: Crafted message DoSes istiod","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Istio","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"Crafted message DoSes istiod","attack_vector":"Any pod on the cluster network","remediation":"Rolling istiod upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-39278"],"status":"curated","published":"2022-10-13"},{"id":"CVE-2022-40242","cve":"CVE-2022-40242","aliases":[],"title":"AMI MegaRAC: Default credentials for the `sysadmin` account, shell access to the BMC","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"Default credentials for the `sysadmin` account, shell access to the BMC","attack_vector":"Network / SSH to BMC","remediation":"Same intake-time credential rotation; add a fleet scan asserting no default BMC accounts remain","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-40242"],"status":"curated","published":"2022-12-05"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:H/A:N","cwe":["CWE-287"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2022-41738","cve":"CVE-2022-41738","aliases":[],"title":"IBM Storage Scale Container Native Storage Access (network namespace exposure): Hosts outside the cluster can open","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"IBM Storage Scale Container Native Storage Access (network namespace exposure)","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"Hosts outside the cluster can open connections directly to CNSA containers, bypassing whatever ingress policy the operator thought was in force. Container-internal services that were only ever meant to be cluster-local become externally addressable.","attack_vector":"Network position outside the Kubernetes cluster with a route to the node network. No credentials required.","remediation":"Upgrade Storage Scale CNSA to the fixed level from IBM's bulletin, then confirm with an external port scan that the driver containers are no longer answering. Add explicit NetworkPolicy denying ingress to the storage namespace as defence in depth.","references":["https://www.ibm.com/support/pages/node/7095312","https://nvd.nist.gov/vuln/detail/CVE-2022-41738"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2022-42276","cve":"CVE-2022-42276","aliases":[],"title":"NVIDIA DGX A100 - SBIOS / SMM firmware: The SmiFlash SMM handler lets a privileged local user read, write and erase","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX A100 - SBIOS / SMM firmware","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"The SmiFlash SMM handler lets a privileged local user read, write and erase the SPI flash directly - a direct write primitive to the platform firmware, with scope extending to other components. This is firmware-level persistence: it survives OS reinstall, image re-flash and tenant handoff, and it is invisible to anything running above it. On a bare-metal GPU rental business it is the difference between wiping a node between tenants and not actually being able to.","attack_vector":"Local and already privileged - host root, or code that has reached the platform firmware/SMM path. It is not a first foothold; it is what turns a one-time root compromise into something you cannot remediate by reimaging.","remediation":"Flash the fixed SBIOS from bulletin 5435. Cost: not live-patchable. Full node drain, host power cycle, and on DGX the SBIOS ships inside a firmware bundle alongside BMC and CPLD components, so budget 30-60 minutes of node downtime plus a post-flash health check. Firmware rollback protection means you cannot cleanly revert - stage on one node before the fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42276","https://github.com/NVIDIA/product-security/tree/main/2022/5435"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-288"],"fleet":{"pain_class":"firmware-flash"},"published":"2023-01-13"},{"id":"CVE-2022-42277","cve":"CVE-2022-42277","aliases":[],"title":"NVIDIA DGX Station - SBIOS / SMM firmware: The same SmiFlash read/write/erase primitive on DGX Station, giving","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX Station - SBIOS / SMM firmware","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"The same SmiFlash read/write/erase primitive on DGX Station, giving a privileged local user arbitrary control of the platform flash. This is firmware-level persistence: it survives OS reinstall, image re-flash and tenant handoff, and it is invisible to anything running above it. On a bare-metal GPU rental business it is the difference between wiping a node between tenants and not actually being able to.","attack_vector":"Local and already privileged - host root, or code that has reached the platform firmware/SMM path. It is not a first foothold; it is what turns a one-time root compromise into something you cannot remediate by reimaging.","remediation":"Flash the fixed SBIOS from bulletin 5435. Cost: not live-patchable. Full node drain, host power cycle, and on DGX the SBIOS ships inside a firmware bundle alongside BMC and CPLD components, so budget 30-60 minutes of node downtime plus a post-flash health check. Firmware rollback protection means you cannot cleanly revert - stage on one node before the fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42277","https://github.com/NVIDIA/product-security/tree/main/2022/5435"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-288"],"fleet":{"pain_class":"firmware-flash"},"published":"2023-01-13"},{"id":"CVE-2022-43377","cve":"CVE-2022-43377","aliases":["CVE-2022-43376","CVE-2022-43378","SEVD-2022-312-01"],"title":"Schneider Electric APC NetBotz 4 environmental appliances (355/450/455/550/570, V4.7.0 and prior): No rate limiting","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Schneider Electric APC NetBotz 4 environmental appliances (355/450/455/550/570, V4.7.0 and prior)","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"No rate limiting on authentication, so an attacker brute-forces the account and takes over the appliance, with stored XSS and clickjacking issues alongside it that let them hit an administrator's browser session. NetBotz appliances are the environmental eyes of the room - temperature, humidity, airflow, leak detection under liquid-cooled racks, door sensors, and camera pods. Control of one means the attacker decides what the operations team sees. During a cooling attack that is decisive: the temperature curve on the wall display stays flat while inlet temperatures climb toward the GPU thermal-shutdown threshold, and the first real signal is accelerators dropping off the fabric. The camera pods are a second concern - a NetBotz with camera modules is a video feed into the hall and the cage aisles, which is both a surveillance-evasion tool for someone about to walk in and a privacy exposure. Leak detection matters specifically for direct-liquid-cooled GPU racks, where a suppressed leak alarm is a route to real hardware destruction.","attack_vector":"Network access to the appliance's web interface, unauthenticated for the brute-force. NetBotz units sit on the facility monitoring VLAN, are polled by DCIM, and are frequently reachable from the corporate network because facilities staff want the camera feed. Weak or default passwords on these appliances are extremely common because they are installed once and never revisited.","remediation":"Firmware update to a fixed NetBotz 4 release per Schneider's SEVD-2022-312-01 - straightforward appliance firmware, no cooling impact, so schedule it. Then do the part that actually closes the brute-force risk regardless of version: set a strong unique password per appliance (they are usually all identical across a site), disable unused accounts, and put the appliance behind an ACL permitting only the DCIM collector and a jump host. Treat environmental telemetry integrity as a control objective in its own right - if the only temperature data you have comes from appliances on the same VLAN as everything else, you have no independent way to detect a cooling attack, so consider a second, separately-networked temperature source for critical GPU rows.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-43377","https://nvd.nist.gov/vuln/detail/CVE-2022-43376","https://download.schneider-electric.com/files?p_Doc_Ref=SEVD-2022-312-01"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-43945","cve":"CVE-2022-43945","aliases":[],"title":"Linux nfsd (NFS server): NFSD buffer overflow - a client can force the send buffer to overflow the page array","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux nfsd (NFS server)","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"NFSD buffer overflow - a client can force the send buffer to overflow the page array","attack_vector":"Network (remote)","remediation":"Data-plane: kernel patch + rolling reboot of every NFS server - tenant-visible","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-43945"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2022-11-04"},{"id":"CVE-2022-46463","cve":"CVE-2022-46463","aliases":[],"title":"Harbor: Public and private image repositories accessible without authentication","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Harbor","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"Public and private image repositories accessible without authentication","attack_vector":"Unauthenticated network","remediation":"Upgrade Harbor and explicitly disable anonymous pull; audit whether tenant images were exposed","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-46463"],"status":"curated","published":"2023-01-13"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2022-48340","cve":"CVE-2022-48340","aliases":[],"title":"GlusterFS (dht translator, dht_setxattr_mds_cbk): A use-after-free in the distributed-hash translator crashes the brick","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"GlusterFS (dht translator, dht_setxattr_mds_cbk)","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"A use-after-free in the distributed-hash translator crashes the brick process. With brick multiplexing enabled one crash takes down several volumes at once, so a single tenant can stall storage for unrelated jobs.","attack_vector":"Reachable over the network against GlusterFS 11.0 without authentication per the CVSS assessment; in practice any client that can drive setxattr traffic to a brick.","remediation":"Upgrade past GlusterFS 11.0 to a release carrying the dht fix and restart the bricks. Consider disabling brick multiplexing on multi-tenant volumes so a crash is contained to one volume.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-48340"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-833"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-48694","cve":"CVE-2022-48694","aliases":[],"title":"Linux kernel (drivers/infiniband/hw/irdma): A permanent kernel hang once any queue-pair goes to error.","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/hw/irdma)","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"A permanent kernel hang once any queue-pair goes to error. Software-generated completions for outstanding work are posted to the wrong completion queue, so the send-queue drain never finishes: kernel workers wedge in uninterruptible sleep, and the RDMA storage transports layered on top (NFS/RDMA, nvme-rdma) stall for the entire node, not just the tenant that triggered it.","attack_vector":"Reachable from the fabric: any event that puts a QP into error - a peer disconnecting, a link transition, a cancelled connection - is enough, and the subsequent drain is issued by in-kernel consumers rather than the tenant. No credentials on the host are needed. Requires the irdma module (Intel E810-class RDMA NICs).","remediation":"No fixed release is published in this record - apply the listed stable fix commits or run a current stable kernel. There is no clean interim control since the trigger is normal connection teardown; drain workloads off affected irdma nodes and reboot into a patched kernel.","references":["https://git.kernel.org/stable/c/14d148401c5202fec3a071e24785481d540b22c3","https://git.kernel.org/stable/c/5becc531a3fa8da75158a8993f56cc3e0717716e","https://nvd.nist.gov/vuln/detail/CVE-2022-48694"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-617","CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-50136","cve":"CVE-2022-50136","aliases":[],"title":"Linux kernel (drivers/infiniband/sw/siw): A remote peer crashes the node during connection setup. When the MPA","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/sw/siw)","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"A remote peer crashes the node during connection setup. When the MPA handshake reply arrives incomplete, siw reports the connect-reply event twice and the iWARP CM hits a hard BUG() - an immediate kernel panic driven entirely from the wire, before any application-level authentication.","attack_vector":"Network-reachable and pre-authentication: a peer the node connects to (or that controls TCP segmentation toward it) sends the MPA reply split so siw sees a partial read. The upstream reproducer is stock ib_send_lat between two hosts, so this fires on ordinary traffic patterns, not just hostile crafting. Requires the siw (soft-iWARP) module to be loaded.","remediation":"No fixed release is published in this record - apply the listed stable fix commits or run a current stable kernel. Interim: unload/blacklist siw unless soft-iWARP is genuinely in use, and restrict which peers may complete iWARP connections to the node.","references":["https://git.kernel.org/stable/c/11edf0bba15ea9df49478affec7974f351bb2f6e","https://git.kernel.org/stable/c/9ade92ddaf2347fb34298c02080caaa3cdd7c27b","https://nvd.nist.gov/vuln/detail/CVE-2022-50136"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-50885","cve":"CVE-2022-50885","aliases":[],"title":"Linux kernel (drivers/infiniband/sw/rxe): A null-pointer dereference panics the node whenever queue-pair creation fails","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/sw/rxe)","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"A null-pointer dereference panics the node whenever queue-pair creation fails because the underlying UDP socket could not be set up. Cleanup runs before the pointer check, so the failure path itself is the crash - one tenant's failed QP takes the host away from every other tenant.","attack_vector":"Reachable two ways on a node with soft-RoCE (rdma_rxe) loaded: a tenant container holding /dev/infiniband/uverbs* calling create_qp, or an in-kernel consumer - the upstream report is a plain 'mount.cifs' over RDMA. Socket creation is made to fail through namespace/resource conditions the caller influences. Unprivileged; rxe only.","remediation":"No fixed release is published in this record - apply the listed stable fix commits or run a current stable kernel. Interim: blacklist/unload rdma_rxe if soft-RoCE is not required, and avoid SMB-Direct/RDMA mounts from untrusted contexts until patched.","references":["https://git.kernel.org/stable/c/ee24de095569935eba600f7735e8e8ddea5b418e","https://git.kernel.org/stable/c/7340ca9f782be6fbe3f64a134dc112772764f766","https://nvd.nist.gov/vuln/detail/CVE-2022-50885"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2023-0200","cve":"CVE-2023-0200","aliases":[],"title":"NVIDIA DGX-2 - SBIOS / SMM firmware: An out-of-bounds access in the OFBD SMM handler against a preconditioned heap","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX-2 - SBIOS / SMM firmware","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"An out-of-bounds access in the OFBD SMM handler against a preconditioned heap reaches SMM code execution with a changed scope. This is firmware-level persistence: it survives OS reinstall, image re-flash and tenant handoff, and it is invisible to anything running above it. On a bare-metal GPU rental business it is the difference between wiping a node between tenants and not actually being able to.","attack_vector":"Local and already privileged - host root, or code that has reached the platform firmware/SMM path. It is not a first foothold; it is what turns a one-time root compromise into something you cannot remediate by reimaging.","remediation":"Flash the fixed SBIOS from bulletin 5449. Cost: not live-patchable. Full node drain, host power cycle, and on DGX the SBIOS ships inside a firmware bundle alongside BMC and CPLD components, so budget 30-60 minutes of node downtime plus a post-flash health check. Firmware rollback protection means you cannot cleanly revert - stage on one node before the fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0200","https://github.com/NVIDIA/product-security/tree/main/2023/5449"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-788"],"fleet":{"pain_class":"firmware-flash"},"published":"2023-04-22"},{"id":"CVE-2023-0202","cve":"CVE-2023-0202","aliases":[],"title":"NVIDIA DGX A100 - SBIOS / SMM firmware: The GenericSio and LegacySmmSredir SMM APIs allow arbitrary modification","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX A100 - SBIOS / SMM firmware","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"The GenericSio and LegacySmmSredir SMM APIs allow arbitrary modification of SMRAM, which is full System Management Mode compromise. This is firmware-level persistence: it survives OS reinstall, image re-flash and tenant handoff, and it is invisible to anything running above it. On a bare-metal GPU rental business it is the difference between wiping a node between tenants and not actually being able to.","attack_vector":"Local and already privileged - host root, or code that has reached the platform firmware/SMM path. It is not a first foothold; it is what turns a one-time root compromise into something you cannot remediate by reimaging.","remediation":"Flash the fixed SBIOS from bulletin 5449. Cost: not live-patchable. Full node drain, host power cycle, and on DGX the SBIOS ships inside a firmware bundle alongside BMC and CPLD components, so budget 30-60 minutes of node downtime plus a post-flash health check. Firmware rollback protection means you cannot cleanly revert - stage on one node before the fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0202","https://github.com/NVIDIA/product-security/tree/main/2023/5449"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-123"],"fleet":{"pain_class":"firmware-flash"},"published":"2023-04-22"},{"id":"CVE-2023-0206","cve":"CVE-2023-0206","aliases":[],"title":"NVIDIA DGX A100 - SBIOS / SMM firmware: The NVME SMM API allows arbitrary SMRAM modification, again reaching full SMM","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX A100 - SBIOS / SMM firmware","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"The NVME SMM API allows arbitrary SMRAM modification, again reaching full SMM compromise. This is firmware-level persistence: it survives OS reinstall, image re-flash and tenant handoff, and it is invisible to anything running above it. On a bare-metal GPU rental business it is the difference between wiping a node between tenants and not actually being able to.","attack_vector":"Local and already privileged - host root, or code that has reached the platform firmware/SMM path. It is not a first foothold; it is what turns a one-time root compromise into something you cannot remediate by reimaging.","remediation":"Flash the fixed SBIOS from bulletin 5449. Cost: not live-patchable. Full node drain, host power cycle, and on DGX the SBIOS ships inside a firmware bundle alongside BMC and CPLD components, so budget 30-60 minutes of node downtime plus a post-flash health check. Firmware rollback protection means you cannot cleanly revert - stage on one node before the fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0206","https://github.com/NVIDIA/product-security/tree/main/2023/5449"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-119"],"fleet":{"pain_class":"firmware-flash"},"published":"2023-04-22"},{"id":"CVE-2023-0207","cve":"CVE-2023-0207","aliases":[],"title":"NVIDIA DGX-2 - SBIOS / SMM firmware: Privileged code can modify the ServerSetup NVRAM variable at runtime, altering","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX-2 - SBIOS / SMM firmware","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"Privileged code can modify the ServerSetup NVRAM variable at runtime, altering platform configuration below the OS. This is firmware-level persistence: it survives OS reinstall, image re-flash and tenant handoff, and it is invisible to anything running above it. On a bare-metal GPU rental business it is the difference between wiping a node between tenants and not actually being able to.","attack_vector":"Local and already privileged - host root, or code that has reached the platform firmware/SMM path. It is not a first foothold; it is what turns a one-time root compromise into something you cannot remediate by reimaging.","remediation":"Flash the fixed SBIOS from bulletin 5449. Cost: not live-patchable. Full node drain, host power cycle, and on DGX the SBIOS ships inside a firmware bundle alongside BMC and CPLD components, so budget 30-60 minutes of node downtime plus a post-flash health check. Firmware rollback protection means you cannot cleanly revert - stage on one node before the fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0207","https://github.com/NVIDIA/product-security/tree/main/2023/5449"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-732"],"fleet":{"pain_class":"firmware-flash"},"published":"2023-04-22"},{"id":"CVE-2023-20578","cve":"CVE-2023-20578","aliases":[],"title":"AMD SMM communications buffer - TOCTOU (AMD-SB-3003): A time-of-check-to-time-of-use race on the SMM communications","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD SMM communications buffer - TOCTOU (AMD-SB-3003)","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"A time-of-check-to-time-of-use race on the SMM communications buffer lets ring-0 code swap the buffer contents between validation and use, reaching arbitrary code execution in System Management Mode. Disclosed in the same August 2024 AMD server bulletin as SinkClose, and the same practical outcome: host root becomes ring -2, which is a compromise you cannot clean by reimaging.","attack_vector":"Local, requires ring-0 plus access to the BIOS menu or a UEFI shell - a meaningfully higher bar than plain root, but well within reach of anyone with physical or BMC-mediated console access to the node.","remediation":"Fixed in AMD PI/AGESA firmware and delivered only as an OEM SBIOS package - AMD ships the PI drop to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before releasing BIOS. **Budget one to six months of OEM lag**, and note that several CVEs in this batch are marked 'no fix planned' on Naples (EPYC 7001) - for those the only remediation is retiring the hardware. Applying it means cordon, drain and a full power cycle per node; there is no driver reload, no live patch and no VBIOS step. Because the extra prerequisite is console/UEFI-shell access, BMC hardening is a genuine compensating control: restrict who can reach the virtual console and who can reboot a node into firmware setup.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-20578","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3003.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"tags":["tenant-isolation"],"published":"2024-08-13"},{"id":"CVE-2023-24574","cve":"CVE-2023-24574","aliases":[],"title":"Dell Enterprise SONiC OS (authentication component): Uncontrolled resource consumption in SONiC's authentication","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell Enterprise SONiC OS (authentication component)","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"Uncontrolled resource consumption in SONiC's authentication component, reachable by an unauthenticated remote attacker. Exhausting the authentication path is a good denial of service because it locks *you* out of the switch at the same time it stays up forwarding — you lose the ability to respond.","attack_vector":"Unauthenticated, remote to the switch management services on Enterprise SONiC 3.5.3, 4.0.0, 4.0.1, 4.0.2.","remediation":"NOS image upgrade plus reboot. Interim: rate-limit and ACL the management interface so only your jump hosts can reach the authentication endpoints — a live config change that also preserves your own access.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-24574"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["fabric-dos"],"published":"2023-02-02"},{"id":"CVE-2023-25191","cve":"CVE-2023-25191","aliases":[],"title":"AMI MegaRAC SPx (Redfish): Password disclosure through Redfish","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx (Redfish)","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"Password disclosure through Redfish; credentials frequently reused across a whole fleet of identically-provisioned nodes","attack_vector":"Network / Redfish","remediation":"Firmware update to SPx_12-update-7.00 / SPx_13-update-5.00 plus a fleet-wide BMC credential rotation","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25191"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-02-15"},{"id":"CVE-2023-25506","cve":"CVE-2023-25506","aliases":[],"title":"NVIDIA DGX-1 - SBIOS / SMM firmware: An out-of-bounds access in the Ofbd handler in the AMI SBIOS reaches SMM code","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX-1 - SBIOS / SMM firmware","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"An out-of-bounds access in the Ofbd handler in the AMI SBIOS reaches SMM code execution against a preconditioned heap, with scope extending to other components. This is firmware-level persistence: it survives OS reinstall, image re-flash and tenant handoff, and it is invisible to anything running above it. On a bare-metal GPU rental business it is the difference between wiping a node between tenants and not actually being able to.","attack_vector":"Local and already privileged - host root, or code that has reached the platform firmware/SMM path. It is not a first foothold; it is what turns a one-time root compromise into something you cannot remediate by reimaging.","remediation":"Flash the fixed SBIOS from bulletin 5458. Cost: not live-patchable. Full node drain, host power cycle, and on DGX the SBIOS ships inside a firmware bundle alongside BMC and CPLD components, so budget 30-60 minutes of node downtime plus a post-flash health check. Firmware rollback protection means you cannot cleanly revert - stage on one node before the fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25506","https://github.com/NVIDIA/product-security/tree/main/2023/5458"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-788"],"fleet":{"pain_class":"firmware-flash"},"published":"2023-04-22"},{"id":"CVE-2023-25521","cve":"CVE-2023-25521","aliases":[],"title":"DGX A100 / A800 SBIOS: Code execution + privesc in SBIOS","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX A100 / A800 SBIOS","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"Code execution + privesc in SBIOS","attack_vector":"Local operator / compromised host OS","remediation":"Flash SBIOS to 1.21 via DGX firmware update container; full node drain + power cycle","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5461/5461.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-250"],"fleet":{"pain_class":"firmware-flash"},"published":"2023-07-04"},{"id":"CVE-2023-25522","cve":"CVE-2023-25522","aliases":[],"title":"DGX A100 / A800 SBIOS: DoS / data tampering / info disclosure","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX A100 / A800 SBIOS","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"DoS / data tampering / info disclosure","attack_vector":"Local operator / compromised host OS","remediation":"Flash SBIOS 1.21; node drain + power cycle","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5461/5461.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-20"],"fleet":{"pain_class":"firmware-flash"},"published":"2023-07-04"},{"id":"CVE-2023-25525","cve":"CVE-2023-25525","aliases":[],"title":"Cumulus Linux (switch OS): Cross-tenant info disclosure (VxLAN IPv6 mis-forwarding on SVI)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Cumulus Linux (switch OS)","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"Cross-tenant info disclosure (VxLAN IPv6 mis-forwarding on SVI)","attack_vector":"Network-adjacent unauthenticated on the fabric","remediation":"Upgrade Cumulus Linux to 5.6.0+; rolling switch upgrade, fabric redundancy required","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5480/5480.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-284"],"published":"2023-09-20"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-400"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-27314","cve":"CVE-2023-27314","aliases":[],"title":"NetApp ONTAP 9 HTTP service: An unauthenticated attacker crashes the ONTAP HTTP service, taking down the management and","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"NetApp ONTAP 9 HTTP service","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"An unauthenticated attacker crashes the ONTAP HTTP service, taking down the management and REST interfaces. Automation that provisions volumes or rotates snapshots for the GPU fleet stops working, and so does the operator's ability to respond.","attack_vector":"Network reach to the HTTP service on an ONTAP 9 system below 9.8P19, 9.9.1P16, 9.10.1P12, 9.11.1P8, 9.12.1P2 or 9.13.1. No account needed.","remediation":"Upgrade to the fixed ONTAP patch level. Meanwhile restrict the management LIF to an admin network - the data path stays up when the HTTP service dies, but you lose control of it.","references":["https://security.netapp.com/advisory/ntap-20231009-0001/","https://nvd.nist.gov/vuln/detail/CVE-2023-27314"],"status":"curated"},{"id":"CVE-2023-27532","cve":"CVE-2023-27532","aliases":[],"title":"Veeam Backup & Replication: Encrypted credentials in the configuration database can be obtained","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Veeam Backup & Replication","year":"2023","cvss_score":7.5,"severity":"high","kev":true,"impact":"Encrypted credentials in the configuration database can be obtained -> access to backup infrastructure hosts","attack_vector":"Network (remote)","remediation":"Control-plane: patch + rotate every credential the backup server held","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-27532"],"status":"curated","published":"2023-03-10"},{"id":"CVE-2023-28432","cve":"CVE-2023-28432","aliases":[],"title":"MinIO: Cluster returns all env vars incl. MINIO_SECRET_KEY and MINIO_ROOT_PASSWORD","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"MinIO","year":"2023","cvss_score":7.5,"severity":"high","kev":true,"impact":"Cluster returns all env vars incl. MINIO_SECRET_KEY and MINIO_ROOT_PASSWORD","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade + ROTATE root password and every derived tenant key","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-28432"],"status":"curated","published":"2023-03-22"},{"id":"CVE-2023-28840","cve":"CVE-2023-28840","aliases":[],"title":"Docker / moby: Swarm overlay-network encryption silently not applied","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"Swarm overlay-network encryption silently not applied; cluster traffic sent in the clear","attack_vector":"Anyone with access to the underlay network between nodes","remediation":"Upgrade moby; if using Swarm overlay encryption assume traffic was exposed","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-28840"],"status":"curated","published":"2023-04-04"},{"id":"CVE-2023-29552","cve":"CVE-2023-29552","aliases":["SLP reflective amplification","AMI-SA-2023004"],"title":"AMI MegaRAC SPx 12 (Service Location Protocol service): Every BMC running SLP is a free DDoS cannon pointed","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx 12 (Service Location Protocol service)","year":"2023","cvss_score":7.5,"severity":"high","kev":true,"impact":"Every BMC running SLP is a free DDoS cannon pointed at the internet, with an amplification factor reported above 2000x. The damage is usually reputational and contractual before it is technical: your management subnet starts sourcing multi-gigabit reflected floods, upstream transit providers null-route you, and the abuse tickets land on the operator. Secondary effect is that the BMC's own control plane saturates, so you lose out-of-band management of the affected nodes exactly when you need it. CISA has this in the Known Exploited Vulnerabilities catalog - it is being used in the wild, not theoretically.","attack_vector":"Unauthenticated UDP to the SLP service on the BMC. Dangerous specifically where BMC management interfaces have been given routable or internet-reachable addresses, which happens more often than operators admit in leased colo, in early-stage neocloud builds, and on remote edge racks reached over a public IP.","remediation":"Config-only, no flash, no reboot - and it is one of the few in this cluster you can fix today. Disable SLP on the BMC (AMI ships it disabled on SPx_12 fixed builds and SPx_13 does not support it at all), and block UDP/427 inbound at the edge. Then do the structural fix: get BMC interfaces off any publicly routable address and behind a VPN or bastion. Verify by scanning your own management ranges from outside, because operators routinely discover BMCs they did not know were exposed.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2023004.pdf","https://www.bitsight.com/blog/new-high-severity-vulnerability-cve-2023-29552-discovered-service-location-protocol-slp","https://www.cisa.gov/known-exploited-vulnerabilities-catalog","https://nvd.nist.gov/vuln/detail/CVE-2023-29552"],"status":"curated","published":"2023-04-25"},{"id":"CVE-2023-31032","cve":"CVE-2023-31032","aliases":[],"title":"DGX A100 SBIOS: Improper crypto config / UI-layer restriction in SBIOS","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX A100 SBIOS","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"Improper crypto config / UI-layer restriction in SBIOS","attack_vector":"Local operator","remediation":"Flash SBIOS 1.25+; node power cycle","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31032","https://github.com/NVIDIA/product-security/tree/main/2024/5513"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-627"],"fleet":{"pain_class":"firmware-flash"},"published":"2024-01-12"},{"id":"CVE-2023-31035","cve":"CVE-2023-31035","aliases":[],"title":"DGX A100 SBIOS: Improper input validation in SBIOS","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX A100 SBIOS","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"Improper input validation in SBIOS","attack_vector":"Local operator","remediation":"Flash SBIOS 1.25+","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31035","https://github.com/NVIDIA/product-security/tree/main/2024/5510"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-20"],"fleet":{"pain_class":"firmware-flash"},"published":"2024-01-12"},{"id":"CVE-2023-31036","cve":"CVE-2023-31036","aliases":[],"title":"Triton Inference Server: RCE / privesc / data tampering via model-load path traversal (`--model-control explicit`)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"RCE / privesc / data tampering via model-load path traversal (`--model-control explicit`)","attack_vector":"Authenticated user of the inference endpoint; malicious model repo","remediation":"Upgrade Triton to 2.40+; rebuild inference serving images; disable explicit model control on multi-tenant endpoints","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5509/5509.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-23"],"published":"2024-01-12"},{"id":"CVE-2023-3127","cve":"CVE-2023-3127","aliases":["ICSA-23-192-02"],"title":"Software House iSTAR Ultra, Ultra LT, Ultra G2 and Edge G2 door controllers: An unauthenticated user can log","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Software House iSTAR Ultra, Ultra LT, Ultra G2 and Edge G2 door controllers","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"An unauthenticated user can log into the controller with administrator rights. No exploitation, no chain - just log in as admin on the panel that governs the doors into the hall and the cages inside it. Administrator on an iSTAR panel means unlocking doors, enrolling credentials, changing door schedules, and editing or clearing the local event log so the entry leaves no record. Whoever walks in can pull drives containing model weights and customer data, attach a console to a running node, plug into the out-of-band management switch and reach every BMC in the row, or leave a hardware implant behind. For a bare-metal GPU provider this is a direct breach of the physical isolation guarantee sold to tenants, and because the log can be cleared from the same session, you may never be able to prove it did or did not happen. This is the fourth distinct critical or high finding on the iSTAR platform in this database's window, which is itself the finding: treat the platform as needing continuous advisory tracking rather than set-and-forget.","attack_vector":"Unauthenticated network access to the controller on the physical-security VLAN. Nothing else is required. That VLAN typically also carries CCTV, intercom and the security integrator's remote-support path, any of which is a route in from a wider network.","remediation":"Firmware update per Johnson Controls' advisory for each affected iSTAR model - a security-integrator engagement with doors in local fallback during the flash, so it needs staff at affected doors for the window. Because the flaw grants administrator without authentication, assume any panel reachable during the exposure window may have had credentials added or the event log cleared: audit the panel's credential list and the head-end's cardholder database against a known-good baseline after patching. Structurally, put the physical-security VLAN behind a firewall with an explicit allow-list from the head-end only, and stop treating that VLAN as trusted because it is 'the security network'.","references":["https://www.cisa.gov/news-events/ics-advisories/icsa-23-192-02","https://nvd.nist.gov/vuln/detail/CVE-2023-3127","https://www.johnsoncontrols.com/cyber-solutions/security-advisories"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-31315","cve":"CVE-2023-31315","aliases":[],"title":"AMD CPU (Sinkclose): Sinkclose: SMM lock bypass","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"AMD CPU (Sinkclose)","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"Sinkclose: SMM lock bypass - ring-0 attacker gains SMM execution, enabling firmware-persistent implants that survive reimaging","attack_vector":"Local ring-0 (post-kernel-escape), i.e. the second stage after any kernel privesc above","remediation":"AGESA/BIOS firmware update + reboot; requires a full firmware rollout across the fleet, not a package update. Some older EPYC SKUs never got fixes","references":["https://access.redhat.com/security/cve/CVE-2023-31315"],"status":"curated","fleet":{"ubiquity":"very common - AMD EPYC is a standard GPU-server host CPU and the host side of MI300 platforms; the flaw reaches back to 2006-era silicon","remediation_pain":"microcode+reboot / firmware-flash via an AGESA/BIOS update per node; AMD initially declined to patch some older Zen parts, leaving unpatchable-mitigate-only nodes in mixed fleets","pain_class":"unpatchable / mitigate-only","why_fleet_wide":"A tenant (or escaped container) with kernel access reaches System Management Mode, the most privileged mode on the box, and can plant an SMM implant invisible to OS and hypervisor that survives a disk wipe."},"published":"2024-08-12"},{"id":"CVE-2023-31320","cve":"CVE-2023-31320","aliases":[],"title":"AMD Radeon Graphics display driver - input validation: Improper input validation in the Radeon display driver lets","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD Radeon Graphics display driver - input validation","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"Improper input validation in the Radeon display driver lets an attacker corrupt the display and deny service. On headless datacenter Instinct nodes the display pipeline is barely exercised, so real exposure is low - included for completeness of the AMD GPU driver picture rather than because it should move up your queue.","attack_vector":"Local, via the display driver path. Effectively inert on headless compute nodes.","remediation":"Update the AMD graphics driver and reload or reboot. Low priority on headless GPU fleets.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31320","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2023-11-14"},{"id":"CVE-2023-33411","cve":"CVE-2023-33411","aliases":[],"title":"Supermicro BMC web server on X11 and M11 based boards with firmware up to 3.17.02: An unauthenticated attacker reads","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC web server on X11 and M11 based boards with firmware up to 3.17.02","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"An unauthenticated attacker reads files out of the BMC's filesystem. In practice that is the credential store, configuration, and keys - which is how this becomes the first link in a chain rather than an information-disclosure footnote: the traversal hands over the credentials needed to exploit CVE-2023-33412 and CVE-2023-33413 on the same box, and because BMC credentials are typically identical across a fleet, one node's disclosure unlocks all of them. Directory traversal in the HTTP server, reachable with no authentication whatsoever.","attack_vector":"Anything routable to the BMC's HTTP interface, unauthenticated. No credential, no host foothold, no tenant access needed - only a network path to the out-of-band management VLAN.","remediation":"Firmware flash to BMC 3.17.02 or later per board, from Supermicro's December 2023 advisory. Because this one needs no credentials, it should be at the front of the queue for any X11 fleet, and any node that was ever exposed to an untrusted network should have its BMC credentials treated as compromised and rotated - to unique per-node values, not another shared password. Network isolation buys time but does not help if your management VLAN is flat and reachable from tenant hosts.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-33411","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2023/33xxx/CVE-2023-33411.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-36835","cve":"CVE-2023-36835","aliases":[],"title":"Juniper Junos OS PFE on QFX10000 Series (VXLAN tunnel routing): A specific *valid* IP packet that needs to be routed","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Juniper Junos OS PFE on QFX10000 Series (VXLAN tunnel routing)","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"A specific *valid* IP packet that needs to be routed over a VXLAN tunnel wedges the Packet Forwarding Engine on a QFX10000. Because the trigger is a legitimate packet rather than a malformed one, no input filter catches it, and QFX10000 sits in the spine or super-spine role where a wedge partitions the fabric rather than dropping one rack.","attack_vector":"A network-based attacker — or, given the trigger is valid traffic, an unlucky workload — sending the specific packet into a VXLAN-routed path.","remediation":"Junos upgrade plus reboot on affected QFX10000 devices, staged so redundant spines are never both down. No filtering workaround, since the triggering packet is valid.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-36835"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["fabric-dos"],"published":"2023-07-14"},{"id":"CVE-2023-38950","cve":"CVE-2023-38950","aliases":[],"title":"ZKTeco BioTime v8.5.5 (iclock API path traversal): Unauthenticated arbitrary file read on the BioTime server","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"ZKTeco BioTime v8.5.5 (iclock API path traversal)","year":"2023","cvss_score":7.5,"severity":"high","kev":true,"impact":"Unauthenticated arbitrary file read on the BioTime server via the iclock API - and this one is in CISA's Known Exploited Vulnerabilities catalog, meaning it is being used against real targets, not theorised about. BioTime is the time-and-attendance and access management server that pairs with ZKTeco terminals, so the files worth reading include its configuration, database credentials and the stored data on who badges in where and when. Combined with the companion issues in the same disclosure set (an unauthenticated administrator password reset via a hidden API, and authenticated arbitrary file write) an attacker moves from file read to full control of the access-management server, and from there to door control. The KEV listing should drive urgency: if you have BioTime anywhere in the facility, treat it as a live target. Personnel movement data is also a targeting asset in its own right - it tells an attacker when the hall is unstaffed.","attack_vector":"Unauthenticated HTTP to the BioTime server's iclock API. BioTime is commonly published to the corporate network for HR and facilities use, and internet exposure is not rare because the product is sold on remote attendance management. Active exploitation is confirmed, so assume internet-reachable instances are already being scanned.","remediation":"Patch immediately to BioTime 9.0.1 (build 20240617.19506) or later per the vendor - and because this is KEV-listed with a known-exploited history, patching is not the end of the work: assume compromise on any instance that was internet-reachable, rotate the database and administrator credentials, audit the cardholder and access-rule data against a known-good baseline, and review door-open events for the exposure window. Then remove BioTime from internet and general corporate reachability entirely. If it is running on a server that also holds anything else, isolate it.","references":["https://www.cisa.gov/known-exploited-vulnerabilities-catalog","https://nvd.nist.gov/vuln/detail/CVE-2023-38950","https://nvd.nist.gov/vuln/detail/CVE-2023-38949"],"status":"curated"},{"id":"CVE-2023-39538","cve":"CVE-2023-39538","aliases":["LogoFAIL"],"title":"UEFI image parsers, AMI AptioV: Unrestricted upload of a crafted BMP logo parsed by the BIOS at boot","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"UEFI image parsers, AMI AptioV","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"Unrestricted upload of a crafted BMP logo parsed by the BIOS at boot; leads to code execution in DXE, before Secure Boot is enforced. Persistent, invisible to the OS","attack_vector":"Local write access to the ESP","remediation":"BIOS/UEFI firmware update per platform — the slowest update in the stack, gated on the ODM shipping an AMI rebase. No dbx-style shortcut exists because the flaw is in the firmware, not a signed binary","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-39538"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-12-06"},{"id":"CVE-2023-39539","cve":"CVE-2023-39539","aliases":["LogoFAIL"],"title":"UEFI image parsers, AMI AptioV: Second LogoFAIL image-parser flaw in AMI AptioV BIOS","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"UEFI image parsers, AMI AptioV","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"Second LogoFAIL image-parser flaw in AMI AptioV BIOS; same pre-Secure-Boot code execution primitive","attack_vector":"Local, ESP write","remediation":"BIOS firmware update; on GPU nodes this means a full BIOS flash cycle with the node drained","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-39539"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-12-06"},{"id":"CVE-2023-4346","cve":"CVE-2023-4346","aliases":[],"title":"KNX devices using KNX Connection Authorization Option 1 (BCU key): An attacker sets the BCU key on KNX devices","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"KNX devices using KNX Connection Authorization Option 1 (BCU key)","year":"2023","cvss_score":7.5,"severity":"high","kev":true,"impact":"An attacker sets the BCU key on KNX devices and permanently locks legitimate operators out - the device cannot be reset or reprogrammed without vendor intervention or physical replacement. This is in CISA's Known Exploited Vulnerabilities catalog, and it has been used in the wild to brick building-automation installations. Where KNX drives lighting, HVAC zone control or blinds in a facility, this is a denial of service you cannot recover from with a reboot or a firmware push: the devices are electronically bricked and the fix is a truck roll with replacement hardware, potentially across every affected device on the bus. In a datacenter context the exposure is usually peripheral rather than core cooling - KNX is more common in European commercial buildings than in purpose-built halls - but it appears in mixed-use and converted buildings, and in office and ancillary space attached to a datacenter. The operator-facing point is the recovery profile: unlike almost everything else on this list, there is no software remedy after the fact.","attack_vector":"Any device that can send telegrams on the KNX bus, or on a KNXnet/IP segment that routes onto it. KNX has no meaningful authentication in its classic form, so this requires only reachability. That means the building network, a KNX/IP router bridging segments, or physical access to the twisted-pair bus in an accessible space.","remediation":"Prevention only - once devices are locked, recovery means vendor unlock procedures where they exist or hardware replacement. Set the BCU key yourself to a known value on every device during commissioning so an attacker cannot claim it, and record it securely. Isolate KNXnet/IP strictly: no route from tenant, corporate or internet networks to the KNX segment, and audit every KNX/IP router for whether it is bridging more than it should. Where the installation supports it, migrate to KNX Secure (KNX Data Secure / IP Secure), which adds authentication and is the actual fix - a device-by-device upgrade project. Given the KEV listing, treat any internet-reachable KNXnet/IP interface as an emergency.","references":["https://www.cisa.gov/known-exploited-vulnerabilities-catalog","https://nvd.nist.gov/vuln/detail/CVE-2023-4346"],"status":"curated"},{"id":"CVE-2023-44199","cve":"CVE-2023-44199","aliases":[],"title":"Juniper Junos OS Packet Forwarding Engine (MX Series): Improper handling of unusual conditions in the Packet Forwarding","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Juniper Junos OS Packet Forwarding Engine (MX Series)","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"Improper handling of unusual conditions in the Packet Forwarding Engine lets an unauthenticated network attacker deny service. A PFE-level failure is worse than a control-plane one — it stops the data plane, so traffic stops even if the routing engine stays up.","attack_vector":"Unauthenticated, network-based, against the PFE on Junos MX platforms.","remediation":"Junos upgrade plus reboot. MX platforms are usually the cluster's edge/border routers, so plan around a redundant pair — patch one side, fail over, patch the other.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-44199"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["fabric-dos"],"published":"2023-10-13"},{"id":"CVE-2023-44487","cve":"CVE-2023-44487","aliases":[],"title":"Envoy: \"HTTP/2 Rapid Reset\": stream-cancellation flood exhausts server resources","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"\"HTTP/2 Rapid Reset\": stream-cancellation flood exhausts server resources. Exploited in the wild at massive scale in late 2023","attack_vector":"Unauthenticated network","remediation":"Upgrade Envoy, Istio, Traefik, ingress-nginx and the Go runtime in every Go-based controller. Broad, multi-component rollout; no GPU drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-44487"],"status":"curated","published":"2023-10-10"},{"id":"CVE-2023-4486","cve":"CVE-2023-4486","aliases":["ICSA-23-341-03"],"title":"Johnson Controls Metasys NAE55 / SNE / SNC network engines and Facility Explorer F4-SNC (before 11.0.6 / 12.0.4)","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Johnson Controls Metasys NAE55 / SNE / SNC network engines and Facility Explorer F4-SNC (before 11.0.6 / 12.0.4)","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"Sending invalid credentials to the login endpoint knocks the engine over. These engines are not dashboards - they are the supervisory controllers that execute the control logic for air handlers, CRAHs and chilled-water plant, and coordinate the field controllers under them. Denial of service on an NAE or SNE means the hall loses coordinated thermal control and falls back to whatever the individual field controllers do standalone, which is typically last-known-value or a crude failsafe. With a 40-140 kW rack density and no supervisory loop closing on actual return-air temperature, you are running blind during exactly the period a competent attacker would also be driving load or blocking alarms. It is trivially repeatable, needs no valid credentials, and can be held indefinitely - so it is a sustained availability attack on the hall rather than a one-shot.","attack_vector":"Unauthenticated TCP to the engine's login endpoint on the facility network. Metasys engines are field-mounted in mechanical rooms and IDF closets and sit on the building VLAN; they are frequently reachable from anywhere on that VLAN with no ACL, because the assumption is that only the ADS talks to them. In practice the mechanical contractor's laptop, the fire-alarm integrator's gateway and the landlord's network all live there too.","remediation":"Firmware update on each engine to 11.0.6 / 12.0.4 or later. That is a controller flash, done by a Johnson Controls technician or certified integrator, engine by engine, with each engine offline during the flash - meaning a real maintenance window on live cooling that many operators will not schedule during peak season. Realistic interim control is an ACL that permits the engine's login port only from the ADS, which removes the attack surface without touching firmware. In a leased colo you cannot flash the landlord's engines: require the firmware version in writing and make DoS resilience of the mechanical control layer an explicit item in the SLA conversation, because a cooling outage caused by the landlord's unpatched engine still kills your training run.","references":["https://www.cisa.gov/news-events/ics-advisories/icsa-23-341-03","https://nvd.nist.gov/vuln/detail/CVE-2023-4486","https://www.johnsoncontrols.com/cyber-solutions/security-advisories"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-45232","cve":"CVE-2023-45232","aliases":["PixieFail","AMI-SA-2024001"],"title":"AMI AptioV UEFI BIOS (EDK II network stack, IPv6): An infinite loop when the firmware parses unknown options in an IPv6","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI AptioV UEFI BIOS (EDK II network stack, IPv6)","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"An infinite loop when the firmware parses unknown options in an IPv6 Destination Options header. A single crafted packet hangs the node in pre-boot firmware - it never reaches the OS, never reports in, and cannot be recovered by a normal reboot because it will hang again on the next boot as long as the attacker keeps sending. On a GPU cluster this is a targeted capacity-denial tool: hold a set of nodes out of the scheduler indefinitely, and because the node is stuck below the OS your host-level monitoring shows nothing but silence.","attack_vector":"Network-reachable, unauthenticated, no interaction, low complexity - one packet during the node's network boot window. Anything that can put IPv6 traffic on the provisioning or boot segment qualifies, including a compromised neighbouring node.","remediation":"BIOS update with the patched EDK II network package - firmware flash plus reboot per node, vendor-rebase-gated. The cheap and immediately available mitigation is the same as for the rest of the PixieFail family: disable network/PXE boot in BIOS where it is not needed, and where it is, put the provisioning network behind strict segmentation so no untrusted host can send packets into the boot window. BIOS setup change plus one reboot, versus a firmware flash campaign.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/2024/AMI-SA-2024001.pdf","https://github.com/tianocore/edk2/security/advisories/GHSA-hc6x-cw6p-gj7h","https://nvd.nist.gov/vuln/detail/CVE-2023-45232"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-01-16"},{"id":"CVE-2023-45233","cve":"CVE-2023-45233","aliases":["PixieFail","VU#132380"],"title":"EDK II NetworkPkg (IPv6 Destination Options header, PadN option parsing): Same shape as the unknown-option hang but","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"EDK II NetworkPkg (IPv6 Destination Options header, PadN option parsing)","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"Same shape as the unknown-option hang but reached through the PadN option, which is trivially craftable. One packet parks a node in firmware forever. In a netboot-driven cluster this is a cheap way to deny an operator their fleet during a reprovisioning window, and the failure looks like a hardware fault rather than an attack.","attack_vector":"Unauthenticated, on-link attacker sending crafted IPv6 packets to nodes during network boot.","remediation":"Firmware flash from the server OEM, one reboot per node. No runtime fix. Practical interim control is to disable IPv6 network boot in the UEFI setup (a config change, deployable via the OEM's remote BIOS-settings tooling without a flash) and to keep the provisioning VLAN reachable only from the deployment controllers.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-45233","https://kb.cert.org/vuls/id/132380","https://github.com/tianocore/edk2/security/advisories/GHSA-hc6x-cw6p-gj7h"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-01-16"},{"id":"CVE-2023-45236","cve":"CVE-2023-45236","aliases":["PixieFail","VU#132380"],"title":"EDK II NetworkPkg (TCP initial sequence number generation): The firmware's TCP initial sequence numbers","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"EDK II NetworkPkg (TCP initial sequence number generation)","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"The firmware's TCP initial sequence numbers are predictable, so an off-path attacker can inject into or hijack the boot-time TCP session. In practice that means substituting the payload the node is downloading - the HTTP-boot image, the kernel, the initrd - without ever being on the wire. For a bare-metal GPU cloud that HTTP-boots tenant images, this is a supply-chain swap at provisioning time that no post-boot integrity check will notice if the swapped image is what gets measured.","attack_vector":"Off-path attacker who can guess the ISN - no need to sit on the provisioning segment at all, which makes this materially worse than the on-link PixieFail bugs. Unauthenticated, pre-OS.","remediation":"OEM BIOS update; the fix replaces the ISN generator, so there is no configuration toggle that helps. Flash + reboot per node. Compensating control while you wait: use HTTPS boot with proper certificate validation rather than plain HTTP/TFTP, and verify signatures on the downloaded image inside the boot flow rather than relying on transport integrity.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-45236","https://kb.cert.org/vuls/id/132380","https://github.com/tianocore/edk2/security/advisories/GHSA-hc6x-cw6p-gj7h"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-01-16"},{"id":"CVE-2023-45237","cve":"CVE-2023-45237","aliases":["PixieFail","VU#132380"],"title":"EDK II NetworkPkg (PseudoRandom number generation used by the network stack): The weak PRNG behind the previous issue","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"EDK II NetworkPkg (PseudoRandom number generation used by the network stack)","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"The weak PRNG behind the previous issue - the firmware's randomness source is not random enough for anything security-relevant it feeds, including sequence numbers and transaction identifiers used during netboot. The operator-visible consequence is that boot-time network exchanges are spoofable by an attacker who does not need to see them.","attack_vector":"Off-path or on-path attacker predicting firmware-generated values during network boot. Unauthenticated, pre-OS.","remediation":"Firmware flash via the server OEM. Reboot per node. There is nothing to configure - the entropy source is compiled in. Until patched, assume boot-time network exchanges are forgeable and lean on cryptographic verification of the boot payload (signed images, Secure Boot with your own keys) rather than on network trust.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-45237","https://kb.cert.org/vuls/id/132380","https://github.com/tianocore/edk2/security/advisories/GHSA-hc6x-cw6p-gj7h"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-01-16"},{"id":"CVE-2023-4607","cve":"CVE-2023-4607","aliases":["LEN-140960"],"title":"Lenovo XClarity Controller (XCC) - permission API: An authenticated XCC user can change the permissions of any user","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lenovo XClarity Controller (XCC) - permission API","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"An authenticated XCC user can change the permissions of any user through a crafted API command - including their own. Fixed in the same advisory as the password-change flaw and with the same practical result: whatever limited BMC account you issued becomes an administrative one, and the attacker inherits out-of-band power control, Virtual Media boot, console access and firmware update rights on the node. The privilege model in this XCC generation should be treated as advisory rather than enforcing until patched.","attack_vector":"Any authenticated XCC account, at any privilege level, reaching the XCC over the out-of-band management VLAN.","remediation":"Flash XCC to the per-model version in LEN-140960 - out-of-band, per-node, no host reboot, no job drain. Because the fix is per-SKU, treat it as one campaign covering both this and CVE-2023-4606. Until patched, the only real control is reducing the number of XCC accounts that exist at all, since privilege tiers are not a boundary here.","references":["https://support.lenovo.com/us/en/product_security/LEN-140960","https://nvd.nist.gov/vuln/detail/CVE-2023-4607"],"status":"curated","published":"2023-10-25"},{"id":"CVE-2023-49298","cve":"CVE-2023-49298","aliases":[],"title":"OpenZFS: Block-cloning path can replace file contents with zero bytes, potentially disabling security mechanisms","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"OpenZFS","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"Block-cloning path can replace file contents with zero bytes, potentially disabling security mechanisms","attack_vector":"Network (remote)","remediation":"Data-plane: module upgrade, node drain + reboot; scrub and verify datasets","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-49298"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2023-11-24"},{"id":"CVE-2023-49933","cve":"CVE-2023-49933","aliases":[],"title":"Slurm: Improper message-integrity enforcement allows RPC traffic modification","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Slurm","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"Improper message-integrity enforcement allows RPC traffic modification","attack_vector":"Anyone on the cluster management network","remediation":"Upgrade Slurm; isolate the Slurm control network","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-49933"],"status":"curated","published":"2023-12-14"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2023-49936","cve":"CVE-2023-49936","aliases":[],"title":"Slurm (NULL pointer dereference in RPC handling): A crafted message crashes the Slurm daemon. On slurmctld that stalls","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Slurm (NULL pointer dereference in RPC handling)","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"A crafted message crashes the Slurm daemon. On slurmctld that stalls every scheduling decision on the cluster - no new job starts, no allocations, and GPUs sit idle until the controller is back. This is the sixth of the December 2023 batch and is the only one of that batch not already in the database.","attack_vector":"Network reach to a Slurm daemon. Same exposure surface as the rest of the 2023-12 batch, so anything that can send RPCs to slurmctld or slurmd.","remediation":"Upgrade to Slurm 22.05.11, 23.02.7 or 23.11.1 and restart the daemons. slurmctld restarts preserve running jobs, so this is a low-drama upgrade - do it in the same window as the rest of the 2023-12 fixes if you have not already.","references":["https://lists.schedmd.com/pipermail/slurm-announce/2023/000103.html","https://nvd.nist.gov/vuln/detail/CVE-2023-49936"],"status":"curated"},{"id":"CVE-2023-50272","cve":"CVE-2023-50272","aliases":["HPESBHF04584"],"title":"HPE iLO 5 / iLO 6 (authentication bypass): Authentication bypass on the iLO itself, remotely, with no credentials","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HPE iLO 5 / iLO 6 (authentication bypass)","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"Authentication bypass on the iLO itself, remotely, with no credentials. That is the whole out-of-band plane on a ProLiant or Apollo node: power control, Virtual Media to boot an attacker-supplied image, remote console into the tenant's session, and a firmware-level foothold that persists through host reimaging. The CVSS vector marks scope as changed, which reflects exactly that - getting the BMC gets you more than the BMC. Affects iLO 5 from v2.63 up to (not including) v3.00, and iLO 6 from v1.05 up to v1.55, so it spans both the Gen10/Gen10 Plus and Gen11 fleets.","attack_vector":"Anything routable to the iLO address on the out-of-band management VLAN, unauthenticated. Attack complexity is rated high, so it is not a trivial one-shot, but it requires no account and no host access - the exposure is defined purely by who can reach the iLO.","remediation":"Flash iLO 5 to v3.00 or later, iLO 6 to v1.55 or later. Out-of-band, per-node, via the iLO web UI, iLOrest, Redfish or OneView - no host reboot and no drain of running jobs; the iLO resets itself and OOB access is unavailable for a couple of minutes. Interim config-only control: restrict the iLO management network to an explicit allowlist of jump hosts, since there is no per-feature toggle that closes an authentication bypass.","references":["https://support.hpe.com/hpesc/public/docDisplay?docLocale=en_US&docId=hpesbhf04584en_us","https://nvd.nist.gov/vuln/detail/CVE-2023-50272"],"status":"curated","published":"2023-12-19"},{"id":"CVE-2023-50275","cve":"CVE-2023-50275","aliases":[],"title":"HPE OneView (clusterService authentication bypass to DoS): Authentication bypass against the OneView cluster service","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HPE OneView (clusterService authentication bypass to DoS)","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"Authentication bypass against the OneView cluster service causing denial of service - loss of the management plane for the HPE estate.","attack_vector":"Unauthenticated network access to the OneView appliance.","remediation":"Apply the OneView update per HPESBGN04586.","references":["https://support.hpe.com/hpesc/public/docDisplay?docId=hpesbgn04586en_us&docLocale=en_US"],"status":"curated"},{"id":"CVE-2023-51232","cve":"CVE-2023-51232","aliases":[],"title":"Dagster (webserver): Directory traversal","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Dagster (webserver)","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"Directory traversal → sensitive file disclosure","attack_vector":"Unauthenticated network to the Dagster webserver","remediation":"Upgrade past 1.5.11","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-51232"],"status":"curated","published":"2025-07-07"},{"id":"CVE-2023-52454","cve":"CVE-2023-52454","aliases":["nvmet-tcp invalid H2C PDU length panic","NVMe/TCP DATAL kernel NULL deref"],"title":"Linux kernel - NVMe-oF TCP target, drivers/nvme/target/tcp.c: A host sending an H2CData command with a DATAL","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel - NVMe-oF TCP target, drivers/nvme/target/tcp.c","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"A host sending an H2CData command with a DATAL inconsistent with the packet size drives a NULL pointer dereference in nvmet_tcp_build_pdu_iovec() and panics the target kernel. The PDU length was also never checked against the MAXH2CDATA value the target itself advertised during connection setup. On a storage node serving many GPU tenants, one malformed PDU from one tenant takes the node down and every attached volume with it - and since NVMe/TCP is unauthenticated by default, the attacker does not need to be a tenant at all, just reachable.","attack_vector":"Connect to the NVMe/TCP target and send an H2CData PDU whose DATAL does not match the actual packet size, or which exceeds the negotiated MAXH2CDATA. Trivial to construct, instantly fatal to the target, and repeatable after every reboot until patched.","remediation":"Host reboot / kernel upgrade on nvmet-tcp targets. Interim: restrict port 4420 to known initiator addresses and enable in-band authentication (on a patched kernel) so an arbitrary peer cannot reach the PDU parser. If a storage node is serving production tenants and cannot be rebooted immediately, the firewall restriction is the meaningful control - this is remotely triggerable with no state.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2023/CVE-2023-52454.json","https://nvd.nist.gov/vuln/detail/CVE-2023-52454"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["fabric-dos"],"published":"2024-02-23"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-52513","cve":"CVE-2023-52513","aliases":[],"title":"Linux kernel (drivers/infiniband/sw/siw): When soft-iWARP fails to process an inbound MPA connection request","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/sw/siw)","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"When soft-iWARP fails to process an inbound MPA connection request immediately, the half-built endpoint unlinks its listener but the later socket-close path still dereferences it, crashing the node on a NULL listener. A peer that opens and breaks connections against a siw listener panics the host and takes out every tenant on it.","attack_vector":"Pre-authentication and remote: MPA request handling is the very first stage of iWARP connection setup, so any host on the IP fabric that can reach a siw listening endpoint drives it. Conditional on the siw (soft-iWARP) module being loaded and a listener bound - common where software RDMA is used for testing or for NVMe-oF/iSER over plain Ethernet.","remediation":"No fixed version is recorded in this entry; boot a stable kernel carrying the siw connection-failure fix (commits 6e26812e289b / 0d520cdb0cd0). Interim: blacklist the siw module where soft-iWARP is not required, or firewall the iWARP listener ports to trusted peers only.","references":["https://git.kernel.org/stable/c/6e26812e289b374c17677d238164a5a8f5770594","https://git.kernel.org/stable/c/0d520cdb0cd095eac5d00078dfd318408c9b5eed","https://nvd.nist.gov/vuln/detail/CVE-2023-52513"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2023-52883","cve":"CVE-2023-52883","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A NULL pointer dereference in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"A NULL pointer dereference in the amdgpu GEM/VM/command-submission ioctl surface. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: Fix possible null pointer dereference","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52883","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-06-20"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-833","CWE-400"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-54306","cve":"CVE-2023-54306","aliases":[],"title":"Linux kernel (net/tls): A receiver that holds its TCP window at zero keeps the kTLS sender blocked inside tx_lock","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/tls)","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"A receiver that holds its TCP window at zero keeps the kTLS sender blocked inside tx_lock indefinitely, so the thread holding the lock never releases it and the TLS transmit work item wedges. Hung-task reports follow, and the blocked work item ties up the shared workqueue that other sockets on the node depend on.","attack_vector":"Remote and entirely under the peer's control: any client or server the node speaks kTLS to can advertise a zero receive window and hold it. This applies to tenant-facing endpoints and to any node service that opens kTLS connections to addresses a tenant influences. No local access needed.","remediation":"Boot a kernel carrying the linked stable commits (which use interruptible sleep and reschedule the work rather than blocking). Interim: set send timeouts on kTLS sockets and cap per-connection lifetime for peer-facing endpoints.","references":["https://git.kernel.org/stable/c/bde541a57b4204d0a800afbbd3d1c06c9cdb133f","https://git.kernel.org/stable/c/7123a4337bf73132bbfb5437e4dc83ba864a9a1e","https://nvd.nist.gov/vuln/detail/CVE-2023-54306"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2023-6021","cve":"CVE-2023-6021","aliases":[],"title":"Ray (log API): LFI — read any file on the head node, unauthenticated","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ray (log API)","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"LFI — read any file on the head node, unauthenticated","attack_vector":"Unauthenticated network to the dashboard","remediation":"Upgrade to 2.8.1+; head-node secrets and cloud credentials are in scope","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-6021"],"status":"curated","published":"2023-11-16"},{"id":"CVE-2024-0096","cve":"CVE-2024-0096","aliases":[],"title":"ChatRTX: Local privesc","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"ChatRTX","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Local privesc","attack_vector":"Local Windows user","remediation":"Consumer app; no DC action","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0096","https://github.com/NVIDIA/product-security/tree/main/2024/5533"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:C/C:H/I:H/A:N","cwe":["CWE-269"],"published":"2024-05-14"},{"id":"CVE-2024-0097","cve":"CVE-2024-0097","aliases":[],"title":"ChatRTX: Local privesc","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"ChatRTX","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Local privesc","attack_vector":"Local Windows user","remediation":"Consumer app; no DC action","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0097","https://github.com/NVIDIA/product-security/tree/main/2024/5533"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:C/C:H/I:H/A:N","cwe":["CWE-269"],"published":"2024-05-14"},{"id":"CVE-2024-0101","cve":"CVE-2024-0101","aliases":[],"title":"Mellanox OS / MetroX / Onyx / Skyway: Switch/gateway DoS (improper input validation)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Mellanox OS / MetroX / Onyx / Skyway","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Switch/gateway DoS (improper input validation)","attack_vector":"Network-adjacent unauthenticated on the fabric","remediation":"Upgrade switch OS image; rolling switch reload","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0101","https://github.com/NVIDIA/product-security/tree/main/2024/5559"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-693"],"published":"2024-08-08"},{"id":"CVE-2024-0112","cve":"CVE-2024-0112","aliases":[],"title":"IGX Orin / Jetson AGX Orin: Privesc / DoS","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"IGX Orin / Jetson AGX Orin","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Privesc / DoS","attack_vector":"Local attacker on the device","remediation":"Flash IGX/Jetson firmware","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0112","https://github.com/NVIDIA/product-security/tree/main/2025/5611"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-20"],"published":"2025-02-12"},{"id":"CVE-2024-0113","cve":"CVE-2024-0113","aliases":[],"title":"NVIDIA Mellanox OS, ONYX, Skyway, MetroX-2/MetroX-3 XC: A crafted URI causes CGI path traversal in the switch web","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Mellanox OS, ONYX, Skyway, MetroX-2/MetroX-3 XC","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"A crafted URI causes CGI path traversal in the switch web interface, reaching privilege escalation and information disclosure on the switch itself. These are the InfiniBand and Ethernet switches carrying your east-west training traffic and, on MetroX, your inter-site links.","attack_vector":"Network, requires the switch web management interface to be reachable and a user to follow a crafted URI. Switch management planes are usually far more reachable inside the datacenter than operators assume.","remediation":"Upgrade the switch OS per bulletin 5563. Cost: a switch OS upgrade means a reboot and a link flap - on a fat-tree fabric plan it leaf-by-leaf with ECMP draining, or you take collective jobs down. Disable the web interface entirely and manage via CLI/gNMI if you can.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0113","https://github.com/NVIDIA/product-security/tree/main/2024/5563"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-35"],"fleet":{"pain_class":"node-reboot"},"published":"2024-08-12"},{"id":"CVE-2024-0762","cve":"CVE-2024-0762","aliases":["UEFIcanhazbufferoverflow","CVE-2024-1598"],"title":"Phoenix SecureCore (TPM configuration / SetupUtility, unsafe UEFI variable handling in SMM): A buffer overflow in how","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Phoenix SecureCore (TPM configuration / SetupUtility, unsafe UEFI variable handling in SMM)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"A buffer overflow in how SecureCore handles a TPM-configuration UEFI variable inside SMM. An attacker overwrites adjacent SMM memory, escalates to ring -2, and installs a bootkit that persists through OS reinstall and disk replacement. Eclypsium's point in naming it was breadth: the same SecureCore code ships across Alder Lake, Coffee Lake, Comet Lake, Ice Lake, Jasper Lake, Kaby Lake, Meteor Lake, Raptor Lake, Rocket Lake and Tiger Lake, so hundreds of PC and server models from multiple OEMs inherit it from one IBV defect. That inheritance pattern is the real lesson for a fleet operator - your firmware exposure is set by an IBV you have no contract with.","attack_vector":"Local admin/root on the host OS writing the vulnerable UEFI variable, then triggering the SMM path. NVD scores it AV:L/PR:L; Phoenix's own CNA scoring assumes higher privilege and higher complexity, which is why the two scores differ (7.8 NVD vs 7.5 Phoenix).","remediation":"BIOS update from your server or system OEM built on the fixed Phoenix SecureCore version - Phoenix lists per-platform fixed versions and OEMs shipped on their own schedules through mid-to-late 2024. Firmware flash plus one reboot per node. No config workaround: the TPM configuration variable is part of normal platform setup and cannot be disabled. Audit by silicon generation rather than by this CVE alone - Phoenix filed the Gemini Lake instance of the same defect as a separate advisory (CVE-2024-1598, fixed in SecureCore for Gemini Lake 4.1.0.567), and OEM release notes commonly cite only one of the two, so a fleet can be patched for the headline case and still exposed on Gemini Lake-based management, edge or storage nodes in the same racks. Where a node cannot be patched promptly, restrict who gets administrative access to the host OS, since that is the entry condition.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0762","https://eclypsium.com/blog/ueficanhazbufferoverflow-widespread-impact-from-vulnerability-in-popular-pc-and-server-firmware/"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"published":"2024-05-14"},{"id":"CVE-2024-10188","cve":"CVE-2024-10188","aliases":[],"title":"LiteLLM: Unauthenticated DoS via `ast.literal_eval` on user input","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LiteLLM","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Unauthenticated DoS via `ast.literal_eval` on user input","attack_vector":"Unauthenticated network","remediation":"Upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-10188"],"status":"curated","published":"2025-03-20"},{"id":"CVE-2024-10403","cve":"CVE-2024-10403","aliases":[],"title":"Brocade Fabric OS (firmware download credential capture): Fabric OS captures the SFTP/FTP server password used","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Brocade Fabric OS (firmware download credential capture)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Fabric OS captures the SFTP/FTP server password used for a firmware download. The credential to your firmware distribution server ends up recoverable from the switch — and that server holds the images you install on every switch in the fabric, so it is a step from 'read a password' to 'supply the next firmware image'. Companion CVE-2023-3489 logs the same password in clear text into SupportSave bundles, which then get emailed to vendor support.","attack_vector":"An attacker with access to the switch or to a SupportSave bundle taken from it.","remediation":"Upgrade Fabric OS past 8.2.3e2 / 9.2.0c / 9.2.1a as applicable — firmware install plus reboot. Immediately: rotate the firmware-server credential, use a single-purpose account with read-only access to the image share, and scrub existing SupportSave archives before sharing them with support.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-10403","https://nvd.nist.gov/vuln/detail/CVE-2023-3489"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-11-21"},{"id":"CVE-2024-11322","cve":"CVE-2024-11322","aliases":[],"title":"CyberPower PowerPanel Business 4.11.0 - Service Watchdog on TCP/2003: An unauthenticated attacker can repeatedly","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"CyberPower PowerPanel Business 4.11.0 - Service Watchdog on TCP/2003","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"An unauthenticated attacker can repeatedly restart the ppbd.exe process via the watchdog service, keeping the power-management daemon permanently down. The consequence is not a crash you notice - it is that the software which would have gracefully shut down your fleet during a utility event is not running when the event happens. This is a denial of the safety mechanism, and its cost only materialises during the incident it was supposed to soften.","attack_vector":"Unauthenticated, to TCP/2003 on the host running PowerPanel Business.","remediation":"Upgrade PowerPanel Business, and firewall TCP/2003 to only the hosts that legitimately need it. Also worth building: an alert on the power-management daemon being down, since the failure mode here is silence.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-11322"],"status":"curated","published":"2025-01-15"},{"id":"CVE-2024-1558","cve":"CVE-2024-1558","aliases":[],"title":"MLflow (`_create_model_version`): Path traversal in model-version creation","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"MLflow (`_create_model_version`)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Path traversal in model-version creation","attack_vector":"Authenticated or unauthenticated model registration","remediation":"Upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-1558"],"status":"curated","published":"2024-04-16"},{"id":"CVE-2024-1561","cve":"CVE-2024-1561","aliases":[],"title":"Gradio (`/component_server`): Arbitrary method invocation on components","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Gradio (`/component_server`)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Arbitrary method invocation on components → local file read","attack_vector":"Unauthenticated network to any exposed Gradio demo","remediation":"Upgrade to 4.19.2+. Tenant-launched Gradio demos on GPU nodes routinely get public share links — provider should block or gate `share=True` egress","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-1561"],"status":"curated","fleet":{"ubiquity":"Very common - Gradio is the default demo/eval UI shipped in AI containers and on Hugging Face Spaces; frequently exposed with `share=True` from a GPU node","remediation_pain":"**Image rebuild** (`daemon-restart` per app) - upgrade to Gradio 4.13.0+ in every image that bundles it, which in practice is most inference/demo images","pain_class":"daemon-restart","why_fleet_wide":"`/component_server` invokes arbitrary `Component` methods, so `move_resource_to_block_cache()` reads any file on the host - API keys and cloud credentials in env/files - from an internet-exposed demo running on a GPU node"},"published":"2024-04-16"},{"id":"CVE-2024-1728","cve":"CVE-2024-1728","aliases":[],"title":"Gradio: Local file inclusion via improper input validation","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Gradio","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Local file inclusion via improper input validation","attack_vector":"Unauthenticated network","remediation":"Upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-1728"],"status":"curated","published":"2024-04-10"},{"id":"CVE-2024-21595","cve":"CVE-2024-21595","aliases":[],"title":"Juniper Junos OS Packet Forwarding Engine (VXLAN + ICMP): A high rate of specific ICMP traffic to a device with VXLAN","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Juniper Junos OS Packet Forwarding Engine (VXLAN + ICMP)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"A high rate of specific ICMP traffic to a device with VXLAN configured deadlocks the Packet Forwarding Engine and leaves the switch unresponsive — recovery requires a manual restart, so it does not self-heal. On a VXLAN/EVPN GPU fabric this is a tenant-reachable way to require hands-on intervention on a leaf, and ICMP is not something most operators filter inside the fabric.","attack_vector":"Unauthenticated, network-based — an attacker able to send ICMP at rate toward a VXLAN-configured Junos device. Any tenant workload qualifies.","remediation":"Junos upgrade plus reboot. Immediate mitigation is a control-plane policer rate-limiting ICMP toward the device — a live config change, no downtime, and it converts a manual-restart outage into a throttled nuisance. Related Junos VXLAN PFE issues: CVE-2023-36835 (QFX10000, PFE wedge on a valid IP packet routed over a VXLAN tunnel) and CVE-2022-22171.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21595","https://nvd.nist.gov/vuln/detail/CVE-2023-36835","https://nvd.nist.gov/vuln/detail/CVE-2022-22171"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["fabric-dos"],"published":"2024-01-12"},{"id":"CVE-2024-24974","cve":"CVE-2024-24974","aliases":[],"title":"OpenVPN: The interactive service pipe is reachable remotely","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"OpenVPN","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"The interactive service pipe is reachable remotely -> interact with the privileged OpenVPN service","attack_vector":"Network (remote)","remediation":"Control-plane: patch; restrict the service named pipe","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-24974"],"status":"curated","published":"2024-07-08"},{"id":"CVE-2024-26147","cve":"CVE-2024-26147","aliases":[],"title":"Helm: Uninitialized variable panic parsing index and plugin YAML","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Helm","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Uninitialized variable panic parsing index and plugin YAML","attack_vector":"Malicious chart repo index","remediation":"Upgrade Helm to 3.14.2+","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26147"],"status":"curated","published":"2024-02-21"},{"id":"CVE-2024-27318","cve":"CVE-2024-27318","aliases":[],"title":"ONNX: Directory traversal in `external_data` — bypass of the 1.13 fix","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"ONNX","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Directory traversal in `external_data` — bypass of the 1.13 fix","attack_vector":"Customer-supplied ONNX model","remediation":"Upgrade past 1.15.0","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-27318"],"status":"curated","published":"2024-02-23"},{"id":"CVE-2024-28028","cve":"CVE-2024-28028","aliases":[],"title":"Intel Neural Compressor: Unauthenticated input-validation failure leading to escalation of privilege in Neural","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel Neural Compressor","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Unauthenticated input-validation failure leading to escalation of privilege in Neural Compressor. Another reason not to leave the optimisation service on an open cluster network.","attack_vector":"Unauthenticated, network-reachable where the service is exposed.","remediation":"Upgrade to v3.0 or later and gate the service behind auth and network policy.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-28028","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01219.html"],"status":"curated","published":"2024-11-13"},{"id":"CVE-2024-28869","cve":"CVE-2024-28869","aliases":[],"title":"Traefik: GET with a Content-Length header hangs the endpoint indefinitely","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Traefik","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"GET with a Content-Length header hangs the endpoint indefinitely; ingress DoS","attack_vector":"Unauthenticated network","remediation":"Rolling Traefik upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-28869"],"status":"curated","published":"2024-04-12"},{"id":"CVE-2024-29966","cve":"CVE-2024-29966","aliases":["CVE-2024-29960","CVE-2024-29965"],"title":"Brocade SANnav OVA appliance image, before v2.3.1 and v2.3.0a: Three defects that together mean every SANnav OVA","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Brocade SANnav OVA appliance image, before v2.3.1 and v2.3.0a","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Three defects that together mean every SANnav OVA deployment shares the same secrets. The documentation ships what is effectively the appliance's root password; every VM built from the official OVA carries identical SSH host keys, so SSH to SANnav is trivially machine-in-the-middle-able; and appliance backups are created world-readable, so a local user can copy a backup, restore it onto their own appliance and recover the passwords of every switch in the fabric. For an operator this is the cheapest possible path from 'someone got a shell somewhere near the management network' to 'attacker holds admin on every FC switch', and therefore to zoning changes that expose one tenant's LUNs to another.","attack_vector":"The root password and SSH keys are usable by anyone with network access to the appliance; the world-readable backup requires only an unprivileged local account on the SANnav host.","remediation":"Upgrade to SANnav 2.3.1 / 2.3.0a, then do the cleanup the upgrade does not do for you: change the appliance root password, regenerate the SSH host keys so your instance is no longer keyed identically to every other deployment, fix permissions on existing backup files and move them off the appliance, and rotate every switch credential that a stolen backup would have exposed. Management-plane work only - no switch firmware flash, no fabric outage.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-29966","https://nvd.nist.gov/vuln/detail/CVE-2024-29960","https://nvd.nist.gov/vuln/detail/CVE-2024-29965"],"status":"curated","published":"2024-04-19"},{"id":"CVE-2024-31142","cve":"CVE-2024-31142","aliases":["XSA-455"],"title":"Xen (x86 speculation): Incorrect logic for BTC/SRSO mitigations","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen (x86 speculation)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Incorrect logic for BTC/SRSO mitigations - guests are unprotected against branch-type confusion despite mitigations appearing enabled","attack_vector":"Tenant VM guest","remediation":"Hypervisor patch + host reboot. Worth flagging: \"mitigation reported as on\" was false, so audit rather than trust the sysfs status","references":["https://xenbits.xen.org/xsa/advisory-455.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-16"},{"id":"CVE-2024-31916","cve":"CVE-2024-31916","aliases":["IBM X-Force 290026"],"title":"IBM OpenBMC bmcweb HTTPS server (FW1050.00 - FW1050.10): Certain URIs on IBM's OpenBMC-derived bmcweb return","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"IBM OpenBMC bmcweb HTTPS server (FW1050.00 - FW1050.10)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Certain URIs on IBM's OpenBMC-derived bmcweb return their content to callers who never authenticated. The Redfish tree is where BMC-side inventory lives - serial numbers, firmware versions, sensor and account metadata - so an unauthenticated reader on the management VLAN gets a precise map of the fleet: which nodes run which firmware, and therefore which nodes are still vulnerable to everything else in this cluster. It is reconnaissance rather than control, but it is the reconnaissance that makes a targeted BMC campaign cheap.","attack_vector":"Unauthenticated HTTPS to the BMC's Redfish/web endpoint. Any host that can route to the management network.","remediation":"Fixed in IBM firmware after FW1050.10; delivery is an OpenPower/Power system firmware update, which on IBM hardware is a supported in-band update path rather than a raw SPI flash, but still a per-node reboot-class operation with a maintenance window. Config-only first move: confirm no BMC in the fleet answers HTTPS from outside your management VLAN, and treat any Redfish data reachable pre-auth as public.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-31916","https://www.ibm.com/support/pages/node/7158679","https://exchange.xforce.ibmcloud.com/vulnerabilities/290026"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-06-27"},{"id":"CVE-2024-34997","cve":"CVE-2024-34997","aliases":[],"title":"joblib (`NumpyArrayWrapper.read_array`): Deserialization vulnerability in joblib 1.4.2","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"joblib (`NumpyArrayWrapper.read_array`)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Deserialization vulnerability in joblib 1.4.2","attack_vector":"Customer-supplied joblib artifact","remediation":"Contested as intended pickle behavior; the real control is format policy, not a version bump","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-34997"],"status":"curated","published":"2024-05-17"},{"id":"CVE-2024-35124","cve":"CVE-2024-35124","aliases":["IBM X-Force 290674"],"title":"IBM OpenBMC default password and session management (FW1020, FW1030, FW1050): The combination of a shipped default","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"IBM OpenBMC default password and session management (FW1020, FW1030, FW1050)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"The combination of a shipped default password and how sessions are managed lets an attacker reach BMC administrator. Full administrative control of the BMC means power state, virtual media, console, firmware update - i.e. the ability to install a persistent implant and to reimage or brick nodes. Default-credential findings are unglamorous and they are also how BMC fleets actually get owned: an operator who racked 400 nodes and changed the OS credentials but not the BMC's is the modal case.","attack_vector":"Network access to the BMC plus, per the CVSS vector, some user interaction and higher attack complexity. Practically: an attacker on the management network against a BMC whose default credential was never rotated or where a stale session can be reused.","remediation":"Fixed in IBM firmware past FW1050.10 / FW1030.50 / FW1020.60 - a per-node system firmware update with a maintenance window. The action that matters more and costs nothing: audit every BMC in the fleet for the vendor default credential right now, rotate to per-node unique passwords, and make credential rotation part of node provisioning rather than a one-time sweep. On a rented bare-metal fleet, also rotate BMC credentials between tenants.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-35124","https://www.ibm.com/support/pages/node/7163195","https://exchange.xforce.ibmcloud.com/vulnerabilities/290674"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-08-13"},{"id":"CVE-2024-35178","cve":"CVE-2024-35178","aliases":[],"title":"Jupyter Server (Windows): Unauthenticated attackers can leak the NTLM hash of the host","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Jupyter Server (Windows)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Unauthenticated attackers can leak the NTLM hash of the host","attack_vector":"Unauthenticated network to the notebook server","remediation":"Upgrade; Windows GPU hosts only","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-35178"],"status":"curated","published":"2024-06-06"},{"id":"CVE-2024-36347","cve":"CVE-2024-36347","aliases":[],"title":"AMD CPU (EntrySign): Improper signature verification in the AMD CPU microcode patch loader","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"AMD CPU (EntrySign)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Improper signature verification in the AMD CPU microcode patch loader - a ring-0 attacker can load arbitrary microcode, defeating SEV-SNP attestation","attack_vector":"Local ring-0 / compromised host; breaks confidential-VM guarantees","remediation":"AGESA/BIOS firmware update + reboot across the fleet; SEV-SNP attestation reports from unpatched hosts cannot be trusted","references":["https://access.redhat.com/security/cve/CVE-2024-36347"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2025-06-27"},{"id":"CVE-2024-36354","cve":"CVE-2024-36354","aliases":["BadRAM (SMM variant)"],"title":"AMD - DIMM SPD address aliasing bypassing SMM isolation (AMD-SB-3014): The BadRAM SPD-aliasing technique aimed at","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD - DIMM SPD address aliasing bypassing SMM isolation (AMD-SB-3014)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"The BadRAM SPD-aliasing technique aimed at System Management Mode rather than at SEV. By lying about a DIMM's size in its serial-presence-detect chip, an attacker creates physical address aliases that let ring-0 code reach into SMRAM and execute at SMM - the level above the hypervisor. Where the SEV-facing BadRAM breaks confidential VMs, this one breaks the platform outright, and it lands beneath every detection tool you run.","attack_vector":"Either brief physical access to the DIMM's SPD chip (the published rig is a ~$10 microcontroller), or - crucially - **ring-0 on a host that has non-compliant DIMMs with unlocked SPD, which needs no physical access at all**. That second path is what makes this a real datacenter concern rather than an evil-maid curiosity.","remediation":"Fixed in AMD PI/AGESA firmware and delivered only as an OEM SBIOS package - AMD ships the PI drop to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before releasing BIOS. **Budget one to six months of OEM lag**, and note that several CVEs in this batch are marked 'no fix planned' on Naples (EPYC 7001) - for those the only remediation is retiring the hardware. Applying it means cordon, drain and a full power cycle per node; there is no driver reload, no live patch and no VBIOS step. The fix adds a boot-time alias-detection scan. Beyond patching, change procurement: specify SPD-lockable DIMMs and verify the lock is actually set, because on non-compliant modules a remote ring-0 attacker reaches this without ever entering your building. Patch this together with the SEV-facing BadRAM CVE - they ship in the same firmware wave.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36354","https://badram.eu/","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3014.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"tags":["tenant-isolation"],"published":"2025-09-06"},{"id":"CVE-2024-36432","cve":"CVE-2024-36432","aliases":[],"title":"BIOS firmware on Supermicro X11DPG-HGX2, X11PDG-QT, X11PDG-OT and X11PDG-SN before version 4.4: An arbitrary memory","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"BIOS firmware on Supermicro X11DPG-HGX2, X11PDG-QT, X11PDG-OT and X11PDG-SN before version 4.4","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"An arbitrary memory write primitive inside platform firmware, on the boards that carry HGX GPU baseboards. What an attacker gets is the ability to corrupt or take over firmware-privileged execution - which on x86 means SMM and the pre-boot environment, below the hypervisor, below the host kernel, and outside anything the operator's security tooling can see. Scope is marked changed in the CVSS vector, meaning the compromise crosses a security boundary. On a GPU host this is the persistence layer under a very expensive, very heavily shared machine. The X11DPG-HGX2 is the head node board for NVIDIA HGX-2 baseboards, so this is literally GPU-platform BIOS rather than generic server BIOS.","attack_vector":"Local, high-privilege access to the host - root or equivalent on the node's operating system, with high attack complexity. The realistic actor is a tenant on leased bare metal, or an attacker who already has host root and wants to convert it into something that survives the node being wiped and re-let.","remediation":"BIOS flash to version 4.4 or later from Supermicro's July 2024 BIOS advisory. BIOS updates on these boards require a host reboot and, on some SKUs, a BMC-mediated update, so this is a per-node maintenance window on machines that are usually running multi-day training jobs - schedule it into a drain cycle rather than expecting an ad-hoc window. There is no config-only mitigation: the write primitive is in the firmware itself.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36432","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2024/36xxx/CVE-2024-36432.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-36433","cve":"CVE-2024-36433","aliases":[],"title":"Supermicro BIOS (arbitrary memory write, X11DPH series): Arbitrary memory write from firmware context on X11DPH boards","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BIOS (arbitrary memory write, X11DPH series)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Arbitrary memory write from firmware context on X11DPH boards, giving a high-privileged local attacker control below the operating system.","attack_vector":"Local high-privilege access.","remediation":"Flash Supermicro BIOS 4.4 or later. Cold reboot required.","references":["https://www.supermicro.com/en/support/security_BIOS_Jul_2024"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-36434","cve":"CVE-2024-36434","aliases":[],"title":"Supermicro BIOS SMM callout (X11DPH-T / X11DPH-Tq): Execution in System Management Mode, the most privileged execution","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BIOS SMM callout (X11DPH-T / X11DPH-Tq)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Execution in System Management Mode, the most privileged execution context on the machine - above the hypervisor, invisible to the OS, and able to write the platform's own firmware. An SMM implant is the deepest persistent foothold available on an x86 node: it survives OS reinstall, disk replacement, and hypervisor redeployment, and no host-based tooling can enumerate it. For a bare-metal operator this is the finding that breaks the tenant handoff guarantee outright, because you cannot prove a returned node is clean. And X11DPH-i before version 4.4 - SMM code calling out to memory the attacker controls, which is the classic route from ring 0 into ring -2.","attack_vector":"Local, host-side, high privilege - root on the node's OS, triggering the SMI that reaches the vulnerable callout. Attack complexity is rated high, so it is a targeted attack rather than an opportunistic one, but a bare-metal tenant has unlimited time and full access to attempt it.","remediation":"BIOS flash to 4.4 or later from Supermicro's July 2024 BIOS advisory, per board. Same image covers the arbitrary-write issues on the neighbouring X11DPH SKUs, so batch them. There is no configuration change that mitigates an SMM callout. If you rent bare metal on these boards, the additional operational control worth adding is a firmware measurement taken at node return and compared against a known-good baseline, since a patched BIOS does not tell you whether the node was implanted before you patched it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36434","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2024/36xxx/CVE-2024-36434.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-37125","cve":"CVE-2024-37125","aliases":[],"title":"Dell SmartFabric OS10 (uncontrolled resource consumption): A remote unauthenticated host can exhaust resources on an","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell SmartFabric OS10 (uncontrolled resource consumption)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"A remote unauthenticated host can exhaust resources on an OS10 switch across 10.5.3.x through 10.5.6.x. No credentials, no adjacency requirement beyond IP reachability — so any tenant workload that can address the switch can attempt it.","attack_vector":"Unauthenticated remote host with IP reachability to the switch.","remediation":"OS10 upgrade plus reload. Interim: control-plane policing and management ACLs restricting who can address the switch at all — live config changes that are worth having permanently.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-37125"],"status":"curated","tags":["fabric-dos"],"published":"2024-09-26"},{"id":"CVE-2024-37775","cve":"CVE-2024-37775","aliases":[],"title":"Sunbird DCIM dcTrack v9.1.2 - ticket location RBAC: Incorrect access control lets an attacker create or update tickets","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Sunbird DCIM dcTrack v9.1.2 - ticket location RBAC","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Incorrect access control lets an attacker create or update tickets against locations they should not have access to, bypassing the RBAC check. In a multi-tenant colo or a shared cage environment, that is a tenant-boundary problem inside the facility workflow system - work orders touching another customer's rack.","attack_vector":"Any authenticated dcTrack user.","remediation":"Upgrade past 9.1.2. If you run dcTrack with per-customer location scoping as a tenant-isolation control, treat that control as having been ineffective for the affected period and review the ticket history.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-37775"],"status":"curated","published":"2024-12-16"},{"id":"CVE-2024-38486","cve":"CVE-2024-38486","aliases":[],"title":"Dell SmartFabric OS10 (command injection): Command injection in SmartFabric OS10 10.5.5.4-10.5.5.10 and 10.5.6.x","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell SmartFabric OS10 (command injection)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Command injection in SmartFabric OS10 10.5.5.4-10.5.5.10 and 10.5.6.x. Listed as a distinct entry because its affected-version window is narrower than the later OS10 command-injection batch, so a fleet on 10.5.5.x needs this specific fix even if it has applied a 10.6.x-targeted advisory elsewhere.","attack_vector":"An attacker able to supply input to the affected OS10 command path.","remediation":"OS10 upgrade plus switch reload. Verify the fixed-version table against your exact running build — Dell's OS10 advisories have heavily overlapping but non-identical version ranges and it is easy to conclude you are patched when you are not.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-38486","https://nvd.nist.gov/vuln/detail/CVE-2024-39577"],"status":"curated","published":"2024-09-06"},{"id":"CVE-2024-38813","cve":"CVE-2024-38813","aliases":[],"title":"VMware vCenter: Privilege escalation to root on vCenter via a crafted network packet","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"VMware vCenter","year":"2024","cvss_score":7.5,"severity":"high","kev":true,"impact":"Privilege escalation to root on vCenter via a crafted network packet [KEV]","attack_vector":"Network access to the management plane","remediation":"vCenter patch + service restart","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-38813"],"status":"curated","fleet":{"pain_class":"daemon-restart"},"published":"2024-09-17"},{"id":"CVE-2024-39321","cve":"CVE-2024-39321","aliases":[],"title":"Traefik: IP allow-lists bypassed via HTTP/3 early data in QUIC 0-RTT with spoofed addresses","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Traefik","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"IP allow-lists bypassed via HTTP/3 early data in QUIC 0-RTT with spoofed addresses","attack_vector":"Unauthenticated network","remediation":"Rolling Traefik upgrade; disable 0-RTT","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-39321"],"status":"curated","published":"2024-07-05"},{"id":"CVE-2024-39719","cve":"CVE-2024-39719","aliases":[],"title":"Ollama: File-existence disclosure via `api/create`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ollama","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"File-existence disclosure via `api/create`","attack_vector":"Unauthenticated network","remediation":"Upgrade; enumerates provider host paths","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-39719"],"status":"curated","published":"2024-10-31"},{"id":"CVE-2024-39722","cve":"CVE-2024-39722","aliases":[],"title":"Ollama: Path traversal in `api/push` discloses server filesystem layout","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ollama","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Path traversal in `api/push` discloses server filesystem layout","attack_vector":"Unauthenticated network","remediation":"Upgrade past 0.1.46","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-39722"],"status":"curated","published":"2024-10-31"},{"id":"CVE-2024-40634","cve":"CVE-2024-40634","aliases":[],"title":"Argo CD: Large JSON payload to /api/webhook DoSes the API server from an unauthenticated position","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Large JSON payload to /api/webhook DoSes the API server from an unauthenticated position","attack_vector":"Unauthenticated network","remediation":"Rolling Argo CD upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-40634"],"status":"curated","published":"2024-07-22"},{"id":"CVE-2024-40992","cve":"CVE-2024-40992","aliases":["RDMA/rxe UD responder length check regression"],"title":"Linux kernel - RDMA/rxe unreliable datagram responder, drivers/infiniband/sw/rxe/rxe_resp.c: The IB architecture says a","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel - RDMA/rxe unreliable datagram responder, drivers/infiniband/sw/rxe/rxe_resp.c","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"The IB architecture says a UD request packet with an invalid length must be silently dropped, but a regression made rxe return a malformed-WQE error that pushes the UD queue pair into the ERROR state instead. A remote sender only has to transmit one oversized UD packet to permanently break the victim's UD queue pair. UD queue pairs carry management traffic and are used by MPI and by connection establishment, so killing them takes the node out of collective communication - the job stalls rather than fails cleanly, which is the expensive failure mode on a large training run.","attack_vector":"Send a UD packet whose payload is larger than the receiver's posted receive buffer. Unauthenticated, connectionless by definition (UD), reachable from anywhere on the fabric that can address the victim's QP. No exploit primitive needed - just a packet that is too big.","remediation":"Host reboot / kernel upgrade. Where rxe is not needed - the common case on clusters with real RNICs - unload and blacklist rdma_rxe instead (config change, no downtime). This one is a stability item as much as a security item: it also fires accidentally under MTU mismatches, so fixing it removes a class of unexplained job hangs.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2024/CVE-2024-40992.json","https://nvd.nist.gov/vuln/detail/CVE-2024-40992"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["fabric-dos"],"published":"2024-07-12"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-908","CWE-200"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-41079","cve":"CVE-2024-41079","aliases":[],"title":"Linux kernel NVMe-oF RDMA target (nvmet, uninitialised completion-entry result field): This is a straight kernel-stack","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel NVMe-oF RDMA target (nvmet, uninitialised completion-entry result field)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"This is a straight kernel-stack disclosure to a remote party. The first two dwords of every NVMe completion entry were left uninitialised on the RDMA transport when the command did not define them - TCP and FC zeroed them, RDMA did not - so the target returned leftover kernel stack contents to whichever initiator issued the command. A remote tenant with an NVMe-oF connection harvests kernel stack bytes at whatever rate it can submit commands, which is exactly the primitive you need to defeat KASLR before using one of the corruption bugs above.","attack_vector":"Remote. Any initiator connected over NVMe-oF/RDMA; the leak arrives in the ordinary completion path, no malformed input required.","remediation":"Kernel update explicitly initialising cqe.result on the RDMA path. Nothing configurable helps - the leak is in normal, well-formed traffic, which also means it produces no anomalous-traffic signal to detect on.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=0990e8a863645496b9e3f91cfcfd63cd95c80319","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2024/CVE-2024-41079.json"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-42145","cve":"CVE-2024-42145","aliases":["IB/core implement a limit on UMAD receive list"],"title":"Linux kernel InfiniBand core (ib_umad): ib_umad kept received management datagrams on an unbounded list","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel InfiniBand core (ib_umad)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"ib_umad kept received management datagrams on an unbounded list. Any node on the fabric that sends MADs faster than userspace drains them exhausts host memory and takes the node down. This is a fabric-wide, unauthenticated availability attack against every host running a subnet manager, an IB diagnostic agent, or UFM's host-side collectors - one malicious or misbehaving endpoint degrades the management plane for the entire cluster.","attack_vector":"Any unauthenticated node attached to the InfiniBand subnet, sending a flood of management datagrams. No credentials anywhere.","remediation":"Upgrade the host kernel to 6.10 or a stable backport (4.19.318, 5.4.280, 5.10.222, 5.15.163, 6.1.98, 6.6.39, 6.9.9) - the fix caps the list at 200k entries. Rolling reboot, prioritizing the nodes that run OpenSM/UFM agents. No firmware flash.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-42145","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2024/CVE-2024-42145.json"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-07-30"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-362","CWE-401"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-42152","cve":"CVE-2024-42152","aliases":[],"title":"Linux kernel NVMe target core (controller teardown racing queue-pair establishment): An initiator that disconnects","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel NVMe target core (controller teardown racing queue-pair establishment)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"An initiator that disconnects while its admin CONNECT is still in flight opens a window where nvmet_sq_destroy() runs concurrently with controller allocation, leaking the controller and its pending async event requests. A remote party controls both halves - connect, then abandon - so the leak is repeatable on demand until the target exhausts memory. Connect-and-drop is also indistinguishable from an unstable client, so it is quiet.","attack_vector":"Remote, unauthenticated. Repeated connect/abort cycles against the target.","remediation":"Kernel update fixing the ordering in nvmet_sq_destroy(). Watch target memory and connection-churn metrics as a detection proxy; allow-list initiators to limit who can churn.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=2f3c22b1d3d7e86712253244797a651998c141fa","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2024/CVE-2024-42152.json"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-42444","cve":"CVE-2024-42444","aliases":[],"title":"AMI AptioV BIOS (TOCTOU race condition): Firmware TOCTOU race allowing execution of arbitrary code on the target device","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI AptioV BIOS (TOCTOU race condition)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Firmware TOCTOU race allowing execution of arbitrary code on the target device.","attack_vector":"Local low-privilege access with user interaction.","remediation":"AMI ships the fix to OEMs, not to you - obtain the updated BIOS from your board/server vendor (Supermicro, Gigabyte, ASRock Rack, Quanta, Tyan etc.) and flash it. Expect a lag of weeks to months between the AMI advisory and an OEM image for your exact SKU, and expect some SKUs never to get one. Cold reboot per node.","references":["https://go.ami.com/hubfs/Security%20Advisories/2025/AMI-SA-2025001.pdf"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-42446","cve":"CVE-2024-42446","aliases":[],"title":"AMI AptioV BIOS (TOCTOU race condition): Second firmware TOCTOU race reaching arbitrary code execution with scope change","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI AptioV BIOS (TOCTOU race condition)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Second firmware TOCTOU race reaching arbitrary code execution with scope change.","attack_vector":"Local high-privilege access; high attack complexity.","remediation":"AMI ships the fix to OEMs, not to you - obtain the updated BIOS from your board/server vendor (Supermicro, Gigabyte, ASRock Rack, Quanta, Tyan etc.) and flash it. Expect a lag of weeks to months between the AMI advisory and an OEM image for your exact SKU, and expect some SKUs never to get one. Cold reboot per node.","references":["https://go.ami.com/hubfs/Security%20Advisories/2025/AMI-SA-2025004.pdf"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-45436","cve":"CVE-2024-45436","aliases":[],"title":"Ollama (`extractFromZipFile`): Zip-slip: archive members extracted outside the parent directory","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ollama (`extractFromZipFile`)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Zip-slip: archive members extracted outside the parent directory","attack_vector":"Customer-supplied model archive","remediation":"Upgrade past 0.1.47","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-45436"],"status":"curated","published":"2024-08-29"},{"id":"CVE-2024-45776","cve":"CVE-2024-45776","aliases":["GRUB2 2025 batch"],"title":"GRUB2 (gettext / message catalogue): Integer overflow reading a crafted translation catalogue gives both","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (gettext / message catalogue)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Integer overflow reading a crafted translation catalogue gives both an out-of-bounds read and write. Language files are unsigned data on the boot partition, which makes them an easy carrier for a bootkit on a node an attacker has held once.","attack_vector":"Attacker-supplied .mo file in GRUB's locale directory - local root or previous tenant.","remediation":"grub2 package update + reboot. Removing locale files from a server image is a cheap surface reduction.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-45776","https://www.openwall.com/lists/oss-security/2025/02/18/3"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-02-18"},{"id":"CVE-2024-45777","cve":"CVE-2024-45777","aliases":["GRUB2 2025 batch"],"title":"GRUB2 (gettext / message catalogue): Second integer overflow in the same translation path, producing a heap","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (gettext / message catalogue)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Second integer overflow in the same translation path, producing a heap out-of-bounds write and pre-boot code execution.","attack_vector":"Attacker-supplied locale catalogue on the boot partition.","remediation":"grub2 package update + reboot; covered by the same distro update as the rest of the batch.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-45777","https://www.openwall.com/lists/oss-security/2025/02/18/3"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-02-19"},{"id":"CVE-2024-45782","cve":"CVE-2024-45782","aliases":["GRUB2 2025 batch"],"title":"GRUB2 (HFS filesystem parser): An unbounded strcpy of the HFS volume name overflows a fixed buffer","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (HFS filesystem parser)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"An unbounded strcpy of the HFS volume name overflows a fixed buffer. About as direct a memory-corruption primitive as this codebase contains, and it fires on nothing more than attaching a crafted volume.","attack_vector":"Attacker-supplied HFS volume, including one presented over BMC virtual media.","remediation":"grub2 package update + reboot. Strip the HFS module if you never boot Apple-formatted media, which on a GPU fleet is always.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-45782","https://www.openwall.com/lists/oss-security/2025/02/18/3"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-03-03"},{"id":"CVE-2024-45807","cve":"CVE-2024-45807","aliases":[],"title":"Envoy: Stream-management bugs in the default oghttp HTTP/2 codec","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Stream-management bugs in the default oghttp HTTP/2 codec","attack_vector":"Unauthenticated network","remediation":"Upgrade Envoy","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-45807"],"status":"curated","published":"2024-09-20"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-46737","cve":"CVE-2024-46737","aliases":[],"title":"Linux kernel NVMe-oF TCP target (nvmet-tcp queue command allocation failure): When command allocation for a new queue","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel NVMe-oF TCP target (nvmet-tcp queue command allocation failure)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"When command allocation for a new queue fails, nr_cmds is left non-zero and the release path walks a NULL array, oopsing the target. A remote initiator induces the allocation failure by opening queues faster than the target can back them - an unauthenticated remote party choosing when the storage target for the whole cluster goes down.","attack_vector":"Remote, unauthenticated, by driving queue creation until allocation fails.","remediation":"Kernel update zeroing nr_cmds on the allocation failure path. Rate-limit and allow-list initiator connections at the network layer in the meantime.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=03e1fd0327fa5e2174567f5fe9290fe21d21b8f4","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2024/CVE-2024-46737.json"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-20"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2024-47866","cve":"CVE-2024-47866","aliases":[],"title":"Ceph RADOS Gateway (RGW): One malformed PUT kills the radosgw process. Sending an object copy with an empty","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ceph RADOS Gateway (RGW)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"One malformed PUT kills the radosgw process. Sending an object copy with an empty x-amz-copy-source header crashes the daemon, so a single tenant can take the shared S3 endpoint offline for everyone and stall every training job that streams checkpoints or datasets through it.","attack_vector":"Any client that can send an HTTP request to the RGW S3 endpoint. Reachable without credentials, so a tenant compute node with network access to the gateway is enough.","remediation":"Upgrade RGW past 19.2.3 and restart the radosgw daemons. Run more than one RGW behind a load balancer with health checks and process supervision so a crash of one gateway does not take the endpoint down.","references":["https://github.com/ceph/ceph/security/advisories/GHSA-mgrm-g92q-f8h8","https://nvd.nist.gov/vuln/detail/CVE-2024-47866"],"status":"curated"},{"id":"CVE-2024-4956","cve":"CVE-2024-4956","aliases":[],"title":"Sonatype Nexus Repository 3: Unauthenticated path traversal","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Sonatype Nexus Repository 3","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Unauthenticated path traversal -> read arbitrary system files from the artifact host","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade to 3.68.1+; rotate anything readable on that host","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-4956"],"status":"curated","published":"2024-05-16"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476","CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-50062","cve":"CVE-2024-50062","aliases":[],"title":"Linux kernel (drivers/infiniband/ulp/rtrs): The RTRS server trusts a connecting client to send its session-info message","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/ulp/rtrs)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"The RTRS server trusts a connecting client to send its session-info message only after every connection of the path is up and the path is CONNECTED. A client that sends it early walks the server into a NULL dereference, so any host on the fabric can panic a storage-serving node and take out every tenant it serves.","attack_vector":"Pre-authentication and remote: the info_req exchange is part of RTRS path establishment, so a peer that can reach the server's RDMA listener drives it with a malformed connection sequence. Conditional on rtrs-srv being loaded and listening (RNBD storage backend). No tenant device node required - this is fabric-side.","remediation":"No fixed version is recorded in this entry; boot a stable kernel carrying the path-establishment sanity checks (commits 394b2f4d5e01 / b5d407666446). Interim: firewall or fabric-ACL the RTRS server port so only trusted initiators can reach it, or stop exporting RTRS targets from shared nodes.","references":["https://git.kernel.org/stable/c/394b2f4d5e014820455af3eb5859eb328eaafcfd","https://git.kernel.org/stable/c/b5d4076664465487a9a3d226756995b12fb73d71","https://nvd.nist.gov/vuln/detail/CVE-2024-50062"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-667","CWE-400"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-50095","cve":"CVE-2024-50095","aliases":[],"title":"Linux kernel (drivers/infiniband/core): A peer that drives enough connection churn across a node's IB port pushes the","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/core)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"A peer that drives enough connection churn across a node's IB port pushes the MAD agent timeout handler into per-request lock thrashing, parking a CPU in the ib_mad workqueue for tens of seconds. Reported as a hard soft-lockup on production RDMA nodes, this is a whole-node stall that every tenant sharing the box feels, not a single-process hang.","attack_vector":"Reachable from the fabric with no authentication: MAD traffic and rdma_cm connection setup between peer nodes is what generates the timed-out work requests. A tenant container holding /dev/infiniband/rdma_cm or uverbs can generate the same connection churn locally. No special device state needed beyond a live IB/RoCE port with the ib_cm/rdma_cm path in use.","remediation":"No fixed version is recorded in this entry; boot a stable kernel that carries the ib_mad timeout batching fix (commits 713adaf0ecfc / 7022a517bf1c). There is no meaningful interim control other than limiting who can open RDMA connections to the node's ports (fabric ACLs, partition keys) and dropping /dev/infiniband from tenants that do not need it.","references":["https://git.kernel.org/stable/c/713adaf0ecfc49405f6e5d9e409d984f628de818","https://git.kernel.org/stable/c/7022a517bf1ca37ef5a474365bcc5eafd345a13a","https://nvd.nist.gov/vuln/detail/CVE-2024-50095"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-50608","cve":"CVE-2024-50608","aliases":[],"title":"Fluent Bit: Prometheus Remote Write input crashes on a Content-Length: 0 packet","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Fluent Bit","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Prometheus Remote Write input crashes on a Content-Length: 0 packet","attack_vector":"Network (remote)","remediation":"Data-plane: DaemonSet image bump across the fleet","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-50608"],"status":"curated","published":"2025-02-18"},{"id":"CVE-2024-50609","cve":"CVE-2024-50609","aliases":[],"title":"Fluent Bit: OpenTelemetry input plugin crashes on a Content-Length: 0 packet","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Fluent Bit","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"OpenTelemetry input plugin crashes on a Content-Length: 0 packet -> log-pipeline outage","attack_vector":"Network (remote)","remediation":"Data-plane: DaemonSet image bump across the fleet","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-50609"],"status":"curated","published":"2025-02-18"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-190"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-53151","cve":"CVE-2024-53151","aliases":[],"title":"Linux NFS-over-RDMA server (svcrdma, xdr_check_write_chunk): An untrusted segcount from the client is multiplied","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux NFS-over-RDMA server (svcrdma, xdr_check_write_chunk)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"An untrusted segcount from the client is multiplied without overflow checking when validating an RDMA Write chunk, so the buffer-overflow check itself is defeated and the server reads past the receive buffer. On an RDMA-attached storage server this is reachable directly from a tenant's HCA.","attack_vector":"Any NFS/RDMA client on the fabric - a tenant compute node with an RDMA NIC and the export mounted over rdma.","remediation":"Update the storage server kernel to one with the svcrdma overflow fix and reboot. If you cannot patch quickly, fall back to NFS over TCP (proto=tcp) on the affected exports, accepting the throughput loss.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git","https://nvd.nist.gov/vuln/detail/CVE-2024-53151"],"status":"curated"},{"id":"CVE-2024-53270","cve":"CVE-2024-53270","aliases":[],"title":"Envoy: Load-shed path assumes an active request exists","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Load-shed path assumes an active request exists; proxy crash","attack_vector":"Unauthenticated network","remediation":"Upgrade Envoy","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53270"],"status":"curated","published":"2024-12-18"},{"id":"CVE-2024-54084","cve":"CVE-2024-54084","aliases":["AMI-SA-2025003"],"title":"AMI AptioV UEFI BIOS: A time-of-check-to-time-of-use race in the BIOS leading to arbitrary code execution","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI AptioV UEFI BIOS","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"A time-of-check-to-time-of-use race in the BIOS leading to arbitrary code execution with a changed scope. This is the quiet half of the March 2025 AMI advisory - it shipped alongside the CVSS 10.0 MegaRAC authentication bypass that got all the attention, and most operators patched the BMC and forgot the BIOS. Successful exploitation puts attacker code in the firmware boot path, below the OS and below any EDR you run, on a node that will keep passing every host-level integrity check you have.","attack_vector":"Local access with high privileges, high attack complexity. Needs root or kernel code on the host plus the ability to win a timing window during a firmware operation. On bare-metal GPU rentals the tenant holds that privilege by contract; on managed nodes it requires a prior host compromise.","remediation":"BIOS update to BKC_5.38 or later - firmware flash plus a full host reboot, per node, gated on your server vendor rebasing. Check specifically whether your March 2025 remediation covered the BIOS: many fleets flashed only the MegaRAC fix for CVE-2024-54085 from the same advisory and left this one open. No config-only mitigation exists for a TOCTOU in firmware.","references":["https://go.ami.com/hubfs/Security%20Advisories/2025/AMI-SA-2025003.pdf","https://nvd.nist.gov/vuln/detail/CVE-2024-54084"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2025-03-11"},{"cvss_vector":"CVSS:3.0/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-1333"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2024-5552","cve":"CVE-2024-5552","aliases":[],"title":"Kubeflow (centraldashboard-angular backend, email validation regex): A catastrophically backtracking regex in the","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Kubeflow (centraldashboard-angular backend, email validation regex)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"A catastrophically backtracking regex in the dashboard's email validation lets an unauthenticated caller pin the backend's CPU with one crafted string. Repeat it and the Kubeflow entry point becomes unusable, so no tenant can reach notebooks, pipelines or the GPU workloads behind them.","attack_vector":"Any unauthenticated client that can reach the centraldashboard-angular backend. Single request, no session.","remediation":"Upgrade the Kubeflow central dashboard to a release with the corrected validation and redeploy. Put a rate limit and request-size cap in front of the dashboard, and set CPU limits on the pod so one abusive request cannot starve the node.","references":["https://huntr.com/bounties/0c1d6432-f385-4c54-beea-9f8c677def5b","https://nvd.nist.gov/vuln/detail/CVE-2024-5552"],"status":"curated"},{"id":"CVE-2024-55567","cve":"CVE-2024-55567","aliases":["INSYDE-SA-2024018"],"title":"Insyde InsydeH2O (UsbCoreDxe SMM module): Another SMM callout in the USB core driver","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (UsbCoreDxe SMM module)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Another SMM callout in the USB core driver - improper input validation lets SMM be redirected into attacker-controlled code outside SMRAM, giving ring -2 execution. Notable mainly because it is the same driver family Insyde has now patched repeatedly (2021 through 2024), which tells an operator something useful: assume the USB stack in your BIOS will need patching again, and build the flash cadence to match rather than treating each one as a one-off.","attack_vector":"Local admin/root on the host OS triggering the vulnerable SMI.","remediation":"OEM BIOS update on Insyde kernel 5.4 / 05.47.01, 5.5 / 05.55.01, 5.6 / 05.62.01, 5.7 / 05.71.01 or later. Firmware flash, one reboot per node. Partial config workaround: disable USB legacy/emulation support in BIOS on headless GPU nodes, which shrinks the reachable surface without a flash - but confirm on your platform that it actually unloads the SMM module.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-55567","https://www.insyde.com/security-pledge/sa-2024018/"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2025-06-12"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-835","CWE-1284"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-57791","cve":"CVE-2024-57791","aliases":[],"title":"Linux kernel SMC (CLC message drain loop, unchecked sock_recvmsg return): The length field in the CLC header is","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel SMC (CLC message drain loop, unchecked sock_recvmsg return)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"The length field in the CLC header is attacker-supplied, and when it exceeds the local buffer the code drains the remainder without checking the receive return value - so a peer that declares a huge length and then stops sending puts the kernel in an unbounded drain loop. One connection from an unauthenticated peer wedges a kernel thread; a handful wedge the node.","attack_vector":"Remote, unauthenticated. Send a CLC header with an oversized length and withhold the rest.","remediation":"Kernel update checking the sock_recvmsg return during the drain. Keep AF_SMC unreachable from tenant networks where SMC is not in use.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=6b80924af6216277892d5f091f5bfc7d1265fa28","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2024/CVE-2024-57791.json"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-7409","cve":"CVE-2024-7409","aliases":[],"title":"QEMU (NBD server): Improper synchronisation during socket closure - DoS of the QEMU NBD server","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"QEMU (NBD server)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Improper synchronisation during socket closure - DoS of the QEMU NBD server","attack_vector":"Unauthenticated network to the NBD port; tenant VM guest","remediation":"QEMU update + restart; do not expose NBD to tenant networks","references":["https://access.redhat.com/security/cve/CVE-2024-7409"],"status":"curated","published":"2024-08-05"},{"cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:C/C:H/I:L/A:N","cwe":["CWE-290"],"fleet":{"pain_class":"unpatchable / mitigate-only"},"id":"CVE-2024-8901","cve":"CVE-2024-8901","aliases":["GHSA-789x-wph8-m68r"],"title":"Kubeflow (AWS ALB Route Directive Adapter for Istio, OIDC JWT validation): The OIDC adapter that Kubeflow adopted for","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubeflow (AWS ALB Route Directive Adapter for Istio, OIDC JWT validation)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"The OIDC adapter that Kubeflow adopted for ALB-fronted deployments validates the JWT but never checks who signed or issued it. An attacker presents a token signed by their own key and is accepted as any federated user, which on Kubeflow means walking into another tenant's notebooks, pipelines and GPU quota. The upstream repo is deprecated and end of life, so no patch is coming.","attack_vector":"Any internet host that can reach an ALB target directly - that is, deployments where ALB targets are exposed rather than only reachable through the load balancer. No valid credentials required.","remediation":"Stop using this adapter; it is end of life and unmaintained. Move Kubeflow authentication to a supported OIDC proxy (oauth2-proxy or the Istio authorization policies with a proper JWKS issuer pin), and make sure ALB target groups are only reachable from the load balancer, never directly.","references":["https://aws.amazon.com/security/security-bulletins/AWS-2024-011/","https://nvd.nist.gov/vuln/detail/CVE-2024-8901"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.0/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-400","CWE-770"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2024-9056","cve":"CVE-2024-9056","aliases":["GHSA-hw8j-hw49-752c"],"title":"BentoML (bundled Gradio app, multipart boundary handling): Appending a long run of characters to a multipart boundary","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"BentoML (bundled Gradio app, multipart boundary handling)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Appending a long run of characters to a multipart boundary makes the server chew through each one, burning CPU until the model endpoint stops answering. An unauthenticated request takes a served model offline, and on a GPU node the pod keeps holding its device allocation while it is useless.","attack_vector":"Any unauthenticated client that can reach the BentoML serving endpoint. Single crafted HTTP request, no session.","remediation":"Upgrade BentoML past 1.4.5 and restart the serving pods. Put a request-size and rate limit in front of the model endpoint, and set CPU limits on the serving container so the abuse cannot spread to co-located pods.","references":["https://github.com/advisories/GHSA-hw8j-hw49-752c","https://nvd.nist.gov/vuln/detail/CVE-2024-9056"],"status":"curated"},{"id":"CVE-2025-0312","cve":"CVE-2025-0312","aliases":[],"title":"Ollama (GGUF import): Crafted GGUF causes DoS on model create","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ollama (GGUF import)","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Crafted GGUF causes DoS on model create","attack_vector":"Customer-supplied GGUF model file","remediation":"Upgrade past 0.3.14; GGUF parsing is unhardened C-adjacent code reachable by any model upload","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0312"],"status":"curated","published":"2025-03-20"},{"id":"CVE-2025-0624","cve":"CVE-2025-0624","aliases":["GRUB2 2025 batch"],"title":"GRUB2 (network config file search): grub_net_search_config_file copies a network-controlled variable with strcpy","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (network config file search)","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"grub_net_search_config_file copies a network-controlled variable with strcpy into a fixed buffer. This is the highest-priority entry in the 2025 batch for anyone running netboot: an attacker on the provisioning segment corrupts GRUB's memory during PXE and takes the node before the OS exists.","attack_vector":"Anyone who can respond on the network boot path - rogue DHCP server, compromised provisioning host, or a tenant that has been given L2 access to the provisioning VLAN by mistake.","remediation":"grub2 package update + reboot, AND rebuild/replace the netboot GRUB binary served over TFTP/HTTP - the served image is the actual attack surface here and patching running nodes does not touch it. Segment the provisioning network away from tenant traffic.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0624","https://access.redhat.com/security/cve/CVE-2025-0624"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-02-19"},{"id":"CVE-2025-1137","cve":"CVE-2025-1137","aliases":[],"title":"IBM Storage Scale (command input neutralization): An authenticated user can execute privileged commands due to improper","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"IBM Storage Scale (command input neutralization)","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"An authenticated user can execute privileged commands due to improper input neutralization. On a shared storage cluster the set of 'authenticated users' is much wider than the set of storage admins — it typically includes every tenant service account with a filesystem role.","attack_vector":"Authenticated user on Storage Scale 5.2.2.0 or 5.2.2.1 under certain configurations.","remediation":"Upgrade Storage Scale past 5.2.2.1 — rolling node upgrade. Review which accounts hold command-line access to the storage cluster in the meantime; most fleets have accumulated more than they intended.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-1137"],"status":"curated","published":"2025-05-10"},{"id":"CVE-2025-14847","cve":"CVE-2025-14847","aliases":[],"title":"MongoDB Server: Mismatched Zlib compressed header lengths","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"MongoDB Server","year":"2025","cvss_score":7.5,"severity":"high","kev":true,"impact":"Mismatched Zlib compressed header lengths -> unauthenticated read of uninitialized heap memory","attack_vector":"Network (remote)","remediation":"Control-plane: patch the metadata store; block unauthenticated wire-protocol reach","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-14847"],"status":"curated","published":"2025-12-19"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-369"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-21885","cve":"CVE-2025-21885","aliases":[],"title":"Linux kernel (drivers/infiniband/hw/bnxt_re): The NVMe-oF target host panics the moment a client connects.","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/hw/bnxt_re)","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"The NVMe-oF target host panics the moment a client connects. Shared-receive-queue page geometry is only filled in for userspace consumers, so a kernel consumer such as nvmet-rdma creates an SRQ with a zero page size and the driver takes a divide-by-zero oops - an immediate whole-node crash triggered from the storage fabric.","attack_vector":"Target-side and reachable from any connecting client: the crash happens in the queue-connect handler, before the NVMe association is established, so no NVMe-level authentication stands in the way. Requires Broadcom bnxt_re HCAs, an exported nvmet-rdma subsystem, and use_srq enabled on the target. If tenants can reach the NVMe-oF listener, any of them can take the storage node down.","remediation":"No fixed release is published in this record - apply the listed stable fix commits or run a current stable kernel. Interim: disable use_srq on nvmet-rdma ports (the target works without it), or restrict the NVMe-oF listener to trusted initiator addresses until the node is patched.","references":["https://git.kernel.org/stable/c/722c3db62bf60cd23acbdc8c4f445bfedae4498e","https://git.kernel.org/stable/c/2cf8e6b52aecb8fbb71c41fe5add3212814031a2","https://nvd.nist.gov/vuln/detail/CVE-2025-21885"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-23320","cve":"CVE-2025-23320","aliases":[],"title":"NVIDIA Triton (Python backend): Information disclosure from the Python backend","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"NVIDIA Triton (Python backend)","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Information disclosure from the Python backend","attack_vector":"Unauthenticated network","remediation":"Patch; leaks the shared-memory key that enables the write primitive","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23320"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-209"],"published":"2025-08-06"},{"id":"CVE-2025-23321","cve":"CVE-2025-23321","aliases":[],"title":"NVIDIA Triton Inference Server: An invalid request triggers a divide-by-zero and kills the server process","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"An invalid request triggers a divide-by-zero and kills the server process. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5687. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23321","https://github.com/NVIDIA/product-security/tree/main/2025/5687"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-369"],"published":"2025-08-06"},{"id":"CVE-2025-23322","cve":"CVE-2025-23322","aliases":[],"title":"NVIDIA Triton Inference Server: Cancelling a stream before it is processed causes a double free and crashes the server","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Cancelling a stream before it is processed causes a double free and crashes the server. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5687. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23322","https://github.com/NVIDIA/product-security/tree/main/2025/5687"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-415"],"published":"2025-08-06"},{"id":"CVE-2025-23323","cve":"CVE-2025-23323","aliases":[],"title":"NVIDIA Triton Inference Server: An integer overflow on an invalid request leads to a segmentation fault","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"An integer overflow on an invalid request leads to a segmentation fault. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5687. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23323","https://github.com/NVIDIA/product-security/tree/main/2025/5687"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-190"],"published":"2025-08-06"},{"id":"CVE-2025-23324","cve":"CVE-2025-23324","aliases":[],"title":"NVIDIA Triton Inference Server: A second integer-overflow-to-segfault path on invalid requests","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"A second integer-overflow-to-segfault path on invalid requests. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5687. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23324","https://github.com/NVIDIA/product-security/tree/main/2025/5687"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-190"],"published":"2025-08-06"},{"id":"CVE-2025-23325","cve":"CVE-2025-23325","aliases":[],"title":"NVIDIA Triton Inference Server: Crafted input drives uncontrolled recursion and exhausts the stack","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Crafted input drives uncontrolled recursion and exhausts the stack. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5687. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23325","https://github.com/NVIDIA/product-security/tree/main/2025/5687"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-674"],"published":"2025-08-06"},{"id":"CVE-2025-23326","cve":"CVE-2025-23326","aliases":[],"title":"NVIDIA Triton Inference Server: Crafted input causes an integer overflow and crashes the server","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Crafted input causes an integer overflow and crashes the server. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5687. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23326","https://github.com/NVIDIA/product-security/tree/main/2025/5687"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-680"],"published":"2025-08-06"},{"id":"CVE-2025-23327","cve":"CVE-2025-23327","aliases":[],"title":"NVIDIA Triton Inference Server: Crafted input causes an integer overflow that reaches data tampering as well as denial","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Crafted input causes an integer overflow that reaches data tampering as well as denial of service - one of the few in this set with an integrity impact, so served results can be corrupted rather than just interrupted. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5687. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23327","https://github.com/NVIDIA/product-security/tree/main/2025/5687"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-190"],"published":"2025-08-06"},{"id":"CVE-2025-23328","cve":"CVE-2025-23328","aliases":[],"title":"NVIDIA Triton Inference Server: Crafted input causes an out-of-bounds write in the server","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Crafted input causes an out-of-bounds write in the server. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5691. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23328","https://github.com/NVIDIA/product-security/tree/main/2025/5691"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-787"],"published":"2025-09-17"},{"id":"CVE-2025-23329","cve":"CVE-2025-23329","aliases":[],"title":"NVIDIA Triton Inference Server: An attacker who can locate and reach the Python backend's shared memory region corrupts","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"An attacker who can locate and reach the Python backend's shared memory region corrupts it directly. Anything co-located in that region belongs to other inference requests. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5691. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23329","https://github.com/NVIDIA/product-security/tree/main/2025/5691"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-284"],"tags":["tenant-isolation"],"published":"2025-09-17"},{"id":"CVE-2025-23331","cve":"CVE-2025-23331","aliases":[],"title":"NVIDIA Triton Inference Server: An invalid request drives an excessive memory allocation and a segmentation fault","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"An invalid request drives an excessive memory allocation and a segmentation fault. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5687. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23331","https://github.com/NVIDIA/product-security/tree/main/2025/5687"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-789"],"published":"2025-08-06"},{"id":"CVE-2025-24357","cve":"CVE-2025-24357","aliases":[],"title":"vLLM (weight loading): `hf_model_weights_iterator` uses `torch.load` without `weights_only`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (weight loading)","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"`hf_model_weights_iterator` uses `torch.load` without `weights_only` → RCE","attack_vector":"Customer-supplied Hub checkpoint","remediation":"Upgrade; a serving engine loading pickle weights inherits every PyTorch pickle CVE","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-24357"],"status":"curated","published":"2025-01-27"},{"id":"CVE-2025-2704","cve":"CVE-2025-2704","aliases":[],"title":"OpenVPN: Corrupting and replaying early-handshake packets against a tls-crypt-v2 server","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"OpenVPN","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Corrupting and replaying early-handshake packets against a tls-crypt-v2 server -> denial of service","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade the VPN concentrator","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-2704"],"status":"curated","published":"2025-04-02"},{"id":"CVE-2025-27817","cve":"CVE-2025-27817","aliases":[],"title":"Apache Kafka (client): SASL/OAUTHBEARER endpoint URLs accept file://","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Apache Kafka (client)","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"SASL/OAUTHBEARER endpoint URLs accept file:// -> arbitrary file read and SSRF","attack_vector":"Network (remote)","remediation":"Control-plane: client library bump across all internal producers/consumers","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-27817"],"status":"curated","published":"2025-06-10"},{"id":"CVE-2025-30202","cve":"CVE-2025-30202","aliases":[],"title":"vLLM (ZeroMQ): DoS and data exposure over ZeroMQ","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (ZeroMQ)","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"DoS and data exposure over ZeroMQ","attack_vector":"Network to the internal socket","remediation":"Upgrade to 0.8.5+","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-30202"],"status":"curated","published":"2025-04-30"},{"id":"CVE-2025-33201","cve":"CVE-2025-33201","aliases":[],"title":"NVIDIA Triton Inference Server: Extra-large payloads trigger an improper-exceptional-condition check and crash","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Extra-large payloads trigger an improper-exceptional-condition check and crash the server. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5734. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33201","https://github.com/NVIDIA/product-security/tree/main/2025/5734"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-754"],"published":"2025-12-03"},{"id":"CVE-2025-33202","cve":"CVE-2025-33202","aliases":[],"title":"NVIDIA Triton Inference Server: Extra-large payloads cause a stack overflow and kill the server","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Extra-large payloads cause a stack overflow and kill the server. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5723. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33202","https://github.com/NVIDIA/product-security/tree/main/2025/5723"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-121"],"published":"2025-11-11"},{"id":"CVE-2025-33211","cve":"CVE-2025-33211","aliases":[],"title":"NVIDIA Triton Inference Server: Improper validation of a specified quantity in input crashes the server","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Improper validation of a specified quantity in input crashes the server. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5734. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33211","https://github.com/NVIDIA/product-security/tree/main/2025/5734"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-1284"],"published":"2025-12-03"},{"id":"CVE-2025-33238","cve":"CVE-2025-33238","aliases":[],"title":"Triton Inference Server: Remote DoS via race condition","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Remote DoS via race condition","attack_vector":"Network-adjacent unauthenticated client","remediation":"Upgrade Triton; redeploy serving images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33238","https://github.com/NVIDIA/product-security/tree/main/2026/5790"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-362"],"published":"2026-03-24"},{"id":"CVE-2025-33254","cve":"CVE-2025-33254","aliases":[],"title":"Triton Inference Server: Remote DoS via race condition in network processing","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Remote DoS via race condition in network processing","attack_vector":"Network-adjacent unauthenticated client","remediation":"Upgrade Triton; redeploy serving images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33254","https://github.com/NVIDIA/product-security/tree/main/2026/5790"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-362"],"published":"2026-03-24"},{"id":"CVE-2025-33255","cve":"CVE-2025-33255","aliases":[],"title":"TensorRT-LLM: Privesc via unsafe deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT-LLM","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Privesc via unsafe deserialization","attack_vector":"Malicious engine/model artifact","remediation":"Bump TensorRT-LLM; rebuild serving images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33255","https://github.com/NVIDIA/product-security/tree/main/2026/5805"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2026-05-20"},{"id":"CVE-2025-37097","cve":"CVE-2025-37097","aliases":[],"title":"HPE Insight Remote Support (unauthenticated denial of service): An unauthenticated attacker takes Insight RS down","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HPE Insight Remote Support (unauthenticated denial of service)","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"An unauthenticated attacker takes Insight RS down, blinding your support and monitoring path for the HPE estate.","attack_vector":"Unauthenticated network access below v7.15.0.646.","remediation":"Upgrade Insight RS to 7.15.0.646.","references":["https://support.hpe.com/hpesc/public/docDisplay?docId=hpesbgn04878en_us&docLocale=en_US"],"status":"curated"},{"id":"CVE-2025-37098","cve":"CVE-2025-37098","aliases":[],"title":"HPE Insight Remote Support (path traversal): Unauthenticated path traversal disclosing files from the IRS server","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HPE Insight Remote Support (path traversal)","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Unauthenticated path traversal disclosing files from the IRS server - including configuration holding device credentials.","attack_vector":"Unauthenticated network access below v7.15.0.646.","remediation":"Upgrade Insight RS to 7.15.0.646 and rotate any device credentials stored in its configuration.","references":["https://support.hpe.com/hpesc/public/docDisplay?docId=hpesbgn04878en_us&docLocale=en_US"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-38018","cve":"CVE-2025-38018","aliases":[],"title":"Linux kernel (net/tls): If a page allocation fails while the TLS strparser is copying a partial record, the receive","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/tls)","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"If a page allocation fails while the TLS strparser is copying a partial record, the receive queue's frag_list is left NULL while full_len still says a record is in flight. The next data_ready dereferences NULL inside the TCP receive softirq - a kernel panic in interrupt context, taking the whole node and every tenant on it.","attack_vector":"Any kTLS RX socket on the node plus memory pressure. A co-tenant can create the pressure (that is a normal condition on a packed GPU node), and the peer keeps feeding partial records; the crash lands in tcp_data_queue -> tls_data_ready -> tls_strp_check_rcv, not in a task context that can be killed cleanly.","remediation":"Boot a kernel carrying the linked stable commits. Interim: keep hard memory limits and reserves on tenant cgroups so the node does not enter page-allocation failure while kTLS connections are live.","references":["https://git.kernel.org/stable/c/8f7f96549bc55e4ef3a6b499bc5011e5de2f46c4","https://git.kernel.org/stable/c/406d05da26835943568e61bb751c569efae071d4","https://nvd.nist.gov/vuln/detail/CVE-2025-38018"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-38035","cve":"CVE-2025-38035","aliases":[],"title":"Linux kernel (drivers/nvme/target): A connecting client that abandons the TCP connection at the right moment during","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/nvme/target)","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"A connecting client that abandons the TCP connection at the right moment during queue setup makes the target call through a NULL socket callback pointer - the crash dump literally shows the instruction pointer at address 0. The storage node panics, taking down every tenant volume it was serving.","attack_vector":"Any peer on the fabric that can reach the nvmet-tcp listening port. No credentials, no NVMe authentication, no tenant device node required - the race is in connection setup itself, before the queue is ever associated with a controller. The attacker only has to open and tear down connections quickly enough that the socket is not in an established state when nvmet_tcp_set_queue_sock runs, which is trivially repeatable from a script. Conditional on the node running nvmet with a TCP port enabled.","remediation":"No fixed version is listed on this record - boot a kernel carrying the linked stable commits. Interim: firewall the nvmet-tcp port so only the storage network can reach it, or disable the TCP port on the target until patched.","references":["https://git.kernel.org/stable/c/6265538446e2426f4bf3b57e91d7680b2047ddd9","https://git.kernel.org/stable/c/17e58be5b49f58bf17799a504f55c2d05ab2ecdc","https://nvd.nist.gov/vuln/detail/CVE-2025-38035"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-401"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-38057","cve":"CVE-2025-38057","aliases":[],"title":"Linux kernel (net/xfrm): Several error paths in the ESP-in-TCP receive code return without freeing the skb, so","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Several error paths in the ESP-in-TCP receive code return without freeing the skb, so malformed or failing frames leak socket buffers. A peer that keeps feeding the error path drains kernel memory on the node until the OOM killer starts taking tenant workloads with it.","attack_vector":"Remote, driven by whatever can reach the espintcp encapsulation socket (TCP port 4500 by default) on a node using ESP-in-TCP encapsulation for IPsec through NAT/middleboxes. Conditional: only nodes with espintcp configured are affected; plain UDP-encapsulated or raw ESP setups are not.","remediation":"Boot a kernel carrying the linked stable commits. Interim: drop ESP-in-TCP encapsulation if it is not required, or firewall the espintcp port to known IKE peers only.","references":["https://git.kernel.org/stable/c/05db2b850a2b8b17f3d1799f563ea1d550e05ed5","https://git.kernel.org/stable/c/e2e1f50fc5ebd2826c4e8c558dc65434382d0c0b","https://nvd.nist.gov/vuln/detail/CVE-2025-38057"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-401"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-38405","cve":"CVE-2025-38405","aliases":[],"title":"Linux kernel (drivers/nvme/target): Every command a client sends to the target carrying metadata (protection","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/nvme/target)","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Every command a client sends to the target carrying metadata (protection information) leaks the bio integrity payload permanently. A tenant or peer issuing metadata-bearing I/O in a loop grows kernel slab without bound until the shared storage node runs out of memory.","attack_vector":"Driven entirely by a connected NVMe-oF client's command stream against an exported namespace - the leak is on the normal inline-bio path, not an error path, so no crafted failure is needed. Any peer allowed to connect to the subsystem can drive it. Conditional on the exported namespace supporting metadata/PI, which is the case for formatted-with-PI backing devices.","remediation":"Update to 6.11 or later (or a stable branch carrying the linked commits). Interim: export namespaces without protection information where the workload permits, and alert on unexplained kmalloc-128 slab growth on target nodes.","references":["https://git.kernel.org/stable/c/431e58d56fcb5ff1f9eb630724a922e0d2a941df","https://git.kernel.org/stable/c/2e2028fcf924d1c6df017033c8d6e28b735a0508","https://nvd.nist.gov/vuln/detail/CVE-2025-38405"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-38590","cve":"CVE-2025-38590","aliases":["net/mlx5e remove skb secpath if xfrm state is not found"],"title":"Linux kernel mlx5_core IPsec RX offload: When hardware reports an xfrm state ID for a decrypted packet whose state","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_core IPsec RX offload","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"When hardware reports an xfrm state ID for a decrypted packet whose state has already been freed, the secpath extension is left attached with length zero and the policy check reads sp->xvec[-1], faulting the kernel. Any remote IPsec peer can crash a node doing hardware IPsec offload on ConnectX - relevant if you encrypt tenant traffic in flight across the fabric.","attack_vector":"Remote IPsec peer, unauthenticated with respect to this bug - the peer just needs to be in an SA that gets torn down while packets are in flight.","remediation":"Upgrade the host kernel to 6.17 or a stable backport (6.6.102, 6.12.42, 6.15.10, 6.16.1). Rolling reboot of nodes doing IPsec offload. Interim: move IPsec off hardware offload to software xfrm (config change, CPU cost, no reboot).","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38590","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2025/CVE-2025-38590.json"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-08-19"},{"id":"CVE-2025-38741","cve":"CVE-2025-38741","aliases":[],"title":"Dell Enterprise SONiC OS 4.5.0 (SSH cryptographic key): The SSH cryptographic-key weakness recurring in Enterprise","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell Enterprise SONiC OS 4.5.0 (SSH cryptographic key)","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"The SSH cryptographic-key weakness recurring in Enterprise SONiC 4.5.0, three years after the same class was fixed in 4.0.x. For an operator the lesson is that SONiC host-key uniqueness is not something to assume from a version number — check it directly on every switch you deploy.","attack_vector":"Unauthenticated, remote against the switch's SSH service.","remediation":"NOS image upgrade plus reboot, then regenerate host keys and refresh your automation's trust store. Consider adding a fleet-wide host-key uniqueness assertion to your provisioning tests.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38741"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-08-04"},{"cwe":["CWE-863","CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-39797","cve":"CVE-2025-39797","aliases":[],"title":"Linux kernel (net/xfrm): Xfrm_alloc_spi could hand out an SPI that is already in use by another inbound SA, because the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Xfrm_alloc_spi could hand out an SPI that is already in use by another inbound SA, because the uniqueness check hashed on destination address as well as SPI. Two live inbound SAs then share an SPI and inbound lookup returns an arbitrary one of them, so ESP packets are matched against the wrong security association - wrong keys, wrong policy, wrong peer attribution. This is an encryption-attribution break with no memory corruption involved: traffic is decrypted (or dropped) under an SA that belongs to someone else.","attack_vector":"Reached through XFRM_MSG_ALLOCSPI, issued by the node's IKE daemon (strongSwan/libreswan) or by any process with CAP_NET_ADMIN in its network namespace, which includes containers granted NET_ADMIN. It is deterministic when the configured SPI range is narrow (charon spi_min/spi_max) and probabilistic but real on any node carrying many child SAs - the shape of a per-tenant IPsec overlay where each tenant pair gets its own SA.","remediation":"Boot a kernel carrying the linked stable commits, and pair it with the follow-up fix for SPI 0 (CVE-2025-39965) which the same change introduced. Interim: widen the IKE daemon's SPI range so collisions are far less likely, and audit for duplicate inbound SPIs with 'ip xfrm state' on nodes running many SAs.","references":["https://git.kernel.org/stable/c/3d8090bb53424432fa788fe9a49e8ceca74f0544","https://git.kernel.org/stable/c/2fc5b54368a1bf1d2d74b4d3b8eea5309a653e38","https://nvd.nist.gov/vuln/detail/CVE-2025-39797"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-39857","cve":"CVE-2025-39857","aliases":[],"title":"Linux kernel (net/smc): On hosts using soft-RoCE, the IB device has no DMA device, and the SMC buffer-mapping path","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/smc)","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"On hosts using soft-RoCE, the IB device has no DMA device, and the SMC buffer-mapping path dereferences that NULL pointer while setting up the receive/send buffers for an inbound connection. An unauthenticated peer connecting to the node crashes the SMC handshake worker and takes the node down - a remote availability kill with no credentials at all.","attack_vector":"Remote and pre-authentication, conditional on soft-RoCE: the crash is in smc_listen_work -> smc_buf_create -> smcr_buf_map_link, so any peer that opens a connection to an SMC-capable listener triggers it when the selected device is the software RoCE driver (rxe) rather than real hardware. Nodes that run rxe for testing, for CPU-only fallback, or inside VMs are exposed; hardware RoCE paths are not.","remediation":"Boot a kernel carrying the fix commits (NULL-checks ibdev->dma_device). Interim: unload/blacklist the rdma_rxe module so soft-RoCE devices are not offered to SMC, or blacklist the smc module on those nodes.","references":["https://git.kernel.org/stable/c/0cdf1fd8fc59d44a48c694324611136910301ef9","https://git.kernel.org/stable/c/eb929910bd4b4165920fa06a87b22cc6cae92e0e","https://nvd.nist.gov/vuln/detail/CVE-2025-39857"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-48956","cve":"CVE-2025-48956","aliases":[],"title":"vLLM (HTTP GET): Single HTTP GET crashes the server","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (HTTP GET)","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Single HTTP GET crashes the server","attack_vector":"Unauthenticated network to the serving port","remediation":"Upgrade to 0.10.1.1+. Trivially exploitable against any internet-exposed tenant endpoint","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-48956"],"status":"curated","published":"2025-08-21"},{"id":"CVE-2025-54502","cve":"CVE-2025-54502","aliases":["APCB SMM driver LocateProtocol misuse"],"title":"AMD Platform Configuration Blob (APCB) SMM driver, EPYC and Instinct MI300A/MI300C: Incorrect use of the UEFI","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"AMD Platform Configuration Blob (APCB) SMM driver, EPYC and Instinct MI300A/MI300C","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Incorrect use of the UEFI LocateProtocol boot service in the APCB SMM driver lets ring 0 escalate to SMM and execute arbitrary code. What makes this one worth flagging for AI operators specifically is the affected list: alongside EPYC 7002 through 9005 it explicitly includes AMD Instinct MI300A and the MI300C-based EPYC 9V64H. Those are accelerator nodes, frequently sold bare metal or with device passthrough, where the customer legitimately has ring 0. Ring -2 code on an MI300A node persists across every tenant that follows.","attack_vector":"Privileged local attacker at ring 0 - on a bare-metal MI300A rental that is the customer by design.","remediation":"Platform Initialization firmware per AMD-SB-7054: MI300A 1.0.0.C (OEM release 2025-12-11), MI300C 1.0.0.3 (2025-12-10), TurinPI 1.0.0.9 for EPYC 9005 (2025-12-31), GenoaPI 1.0.0.H, MilanPI 1.0.0.J, RomePI 1.0.0.P. BIOS flash and reboot per node - on an Instinct fleet that means evicting whatever is training on it. Given the bare-metal exposure, prioritize Instinct nodes over general-purpose EPYC and add a firmware measurement to the between-tenants checklist.","references":["https://www.amd.com/en/resources/product-security/bulletin/AMD-SB-7054.html","https://nvd.nist.gov/vuln/detail/CVE-2025-54502"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2026-04-16"},{"id":"CVE-2025-55551","cve":"CVE-2025-55551","aliases":[],"title":"PyTorch (`torch.linalg.lu`): DoS on slice operation","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"PyTorch (`torch.linalg.lu`)","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"DoS on slice operation","attack_vector":"Tenant-controlled tensor shapes; matters for shared-GPU multi-tenant serving","remediation":"Tenant-owned code path. Provider action is per-tenant GPU/process isolation so a crash does not take out co-tenants","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-55551"],"status":"curated","published":"2025-09-25"},{"id":"CVE-2025-55558","cve":"CVE-2025-55558","aliases":[],"title":"PyTorch (KV/conv path buffer overflow): Buffer overflow when a model combines Conv2d + hardshrink + view","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"PyTorch (KV/conv path buffer overflow)","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Buffer overflow when a model combines Conv2d + hardshrink + view","attack_vector":"Customer-supplied model graph","remediation":"Tenant-owned; provider isolates blast radius per container","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-55558"],"status":"curated","published":"2025-09-25"},{"id":"CVE-2025-5777","cve":"CVE-2025-5777","aliases":[],"title":"Citrix NetScaler ADC/Gateway: \"CitrixBleed 2\" - insufficient input validation","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Citrix NetScaler ADC/Gateway","year":"2025","cvss_score":7.5,"severity":"high","kev":true,"impact":"\"CitrixBleed 2\" - insufficient input validation -> memory overread of session material","attack_vector":"Network (remote)","remediation":"Control-plane: patch and kill all sessions","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-5777"],"status":"curated","published":"2025-06-17"},{"id":"CVE-2025-58187","cve":"CVE-2025-58187","aliases":[],"title":"Go crypto/x509 (Tailscale, Go infra): Name-constraint checking scales non-linearly with certificate size","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Go crypto/x509 (Tailscale, Go infra)","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Name-constraint checking scales non-linearly with certificate size -> CPU exhaustion when validating chains","attack_vector":"Network (remote)","remediation":"Control-plane: Go toolchain bump and rebuild of every TLS-validating service","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-58187"],"status":"curated","published":"2025-10-29"},{"id":"CVE-2025-59425","cve":"CVE-2025-59425","aliases":[],"title":"vLLM (API key comparison): Timing attack recovers the API key","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (API key comparison)","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Timing attack recovers the API key","attack_vector":"Unauthenticated network to the serving port","remediation":"Upgrade to 0.11.0+; do not rely on vLLM's own API key as the tenant auth boundary — front it with a gateway","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-59425"],"status":"curated","published":"2025-10-07"},{"id":"CVE-2025-59538","cve":"CVE-2025-59538","aliases":[],"title":"Argo CD: Azure DevOps webhook credentials mishandled, allowing unauthorised webhook use","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Azure DevOps webhook credentials mishandled, allowing unauthorised webhook use","attack_vector":"Unauthenticated network","remediation":"Rolling Argo CD upgrade; rotate webhook credentials","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-59538"],"status":"curated","published":"2025-10-01"},{"id":"CVE-2025-65002","cve":"CVE-2025-65002","aliases":[],"title":"Fujitsu / Fsas Technologies iRMC S6 BMC (M5-generation servers): A length-boundary bug in BMC authentication","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Fujitsu / Fsas Technologies iRMC S6 BMC (M5-generation servers)","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"A length-boundary bug in BMC authentication - a username of exactly the wrong length gets handled incorrectly and access control does not apply as intended. What the operator loses is the guarantee that their Redfish and web management surface enforces the roles they configured. This is worth carrying in a GPU-fleet database less for Fujitsu's market share than for the pattern: BMC authentication logic keeps failing on boundary conditions in string handling, and operators who assume Redfish role enforcement is sound are relying on code with a long history of exactly this class of defect. Servers before firmware 1.37S - Redfish and web UI access control mishandles the case where a username is exactly 16 characters long.","attack_vector":"Network reachability to the iRMC's Redfish or web interface. Exploitation depends on the username in play hitting the 16-character boundary, which an attacker can arrange when they control account creation or can guess an existing account name of that length.","remediation":"Firmware flash of the iRMC to 1.37S or later - a per-node out-of-band BMC update. A config-only interim step that genuinely helps: audit BMC account names and eliminate any that are exactly 16 characters, which removes the triggering condition without waiting for a flash window. Fujitsu's PSIRT publishes a readable PDF advisory, which puts them ahead of most of the ODM vendors in this database on advisory accessibility.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-65002","https://security.ts.fujitsu.com/ProductSecurity/content/FsasTech-PSIRT-FTI-ISS-2025-082610-Security-Notice.pdf"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-7342","cve":"CVE-2025-7342","aliases":[],"title":"Kubernetes Image Builder: Nutanix/OVA Windows images use default credentials unless overridden","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes Image Builder","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Nutanix/OVA Windows images use default credentials unless overridden","attack_vector":"Unauthenticated network reaching an affected node","remediation":"Rebuild affected Windows node images","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated","published":"2025-08-17"},{"id":"CVE-2026-13460","cve":"CVE-2026-13460","aliases":[],"title":"IBM Storage Scale GUI (hardcoded inter-node token): A hardcoded token in the Storage Scale GUI source, used","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"IBM Storage Scale GUI (hardcoded inter-node token)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"A hardcoded token in the Storage Scale GUI source, used for inter-node cluster communication and REST access. A hardcoded credential in a storage cluster's management path is the same shipped-secret problem as default BMC passwords: it is identical on every deployment, it is in a source tree anyone can read, and rotating it is not something the product expects you to do.","attack_vector":"Anyone who can reach the Storage Scale GUI/REST endpoint and knows the token — which, once published, is everyone.","remediation":"Upgrade Storage Scale past the affected 5.2.3.x / 6.0.x levels. GUI-layer upgrade plus service restart; the filesystem stays up. Immediately restrict the GUI/REST endpoint to a management network — a firewall change, applied live, that matters more than the patch timing.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-13460"],"status":"curated","fleet":{"pain_class":"daemon-restart"},"published":"2026-08-13"},{"id":"CVE-2026-15977","cve":"CVE-2026-15977","aliases":[],"title":"SGLang (`/server_info`): Endpoint returns API keys and SSL keyfile paths","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"SGLang (`/server_info`)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Endpoint returns API keys and SSL keyfile paths","attack_vector":"Network to the serving port with only `--admin-*` partly configured","remediation":"Upgrade; rotate any keys exposed on a previously reachable instance","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-15977"],"status":"curated","published":"2026-07-30"},{"id":"CVE-2026-15978","cve":"CVE-2026-15978","aliases":[],"title":"SGLang (weight exfiltration): Two endpoints allow a remote attacker to pull model weights when no API key is set","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"SGLang (weight exfiltration)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Two endpoints allow a remote attacker to pull model weights when no API key is set","attack_vector":"Unauthenticated network to the serving port","remediation":"Upgrade; for a neocloud hosting customer models, this is direct tenant-IP exfiltration","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-15978"],"status":"curated","published":"2026-07-30"},{"id":"CVE-2026-1669","cve":"CVE-2026-1669","aliases":[],"title":"Keras (HDF5 external links): Arbitrary local file read during model load","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Keras (HDF5 external links)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Arbitrary local file read during model load","attack_vector":"Customer-supplied HDF5 model — reads provider or co-tenant files visible to the process","remediation":"Upgrade; better, disallow HDF5. Note the process' service-account tokens and mounted secrets are in scope","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-1669"],"status":"curated","published":"2026-02-11"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476","CWE-306"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-22998","cve":"CVE-2026-22998","aliases":[],"title":"Linux kernel NVMe-oF TCP target (nvmet-tcp, H2C_DATA PDU before CONNECT): Nvmet_tcp_build_pdu_iovec() dereferences cmd","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel NVMe-oF TCP target (nvmet-tcp, H2C_DATA PDU before CONNECT)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Nvmet_tcp_build_pdu_iovec() dereferences cmd->req.sg and cmd->iov without checking they were ever initialised. Sending an H2C_DATA PDU straight after the ICREQ/ICRESP handshake - before any CONNECT command, before any NVMe-level identification of the initiator - reaches that dereference and crashes the target. This is genuinely pre-authentication: the only thing the attacker completes is the transport handshake, and the payoff is taking down the storage target for every tenant it serves.","attack_vector":"Remote, fully unauthenticated, immediately after TCP connect and the NVMe/TCP ICREQ exchange. No host NQN, no DH-HMAC-CHAP, nothing.","remediation":"Kernel update adding the NULL checks before processing H2C_DATA. Because it is pre-auth, in-band authentication does not help you here - the compensating control is network reachability: the nvmet listener must not be reachable from tenant-routable networks, only from the storage fabric.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=32b63acd78f577b332d976aa06b56e70d054cbba","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-22998.json"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-23148","cve":"CVE-2026-23148","aliases":[],"title":"Linux kernel (drivers/nvme/target): Ordinary client I/O to an nvmet block-device namespace can hit a completion race","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/nvme/target)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Ordinary client I/O to an nvmet block-device namespace can hit a completion race where the target's inline bio is torn down while the same bio is being re-submitted, dereferencing a NULL cgroup pointer in the block layer. The shared storage node crashes on the normal read/write path, so a single tenant's I/O pattern can knock out the target for all of them.","attack_vector":"Driven by a connected NVMe-oF client's regular I/O against an exported namespace backed by a block device - no special opcode and no privilege on the target side. Any tenant or peer that has been allowed to connect to the subsystem can push the target into the window; it is a timing race, so it favours high-rate I/O rather than a crafted packet. Requires nvmet configured with a bdev-backed namespace.","remediation":"Update to 6.12.69 / 6.16 or later. Interim: no clean workaround short of stopping the nvmet subsystem export or moving affected namespaces to a patched node - the path is the normal I/O path, so it cannot be gated by configuration.","references":["https://git.kernel.org/stable/c/ee10b06980acca1d46e0fa36d6fb4a9578eab6e4","https://git.kernel.org/stable/c/68207ceefd71cc74ce4e983fa9bd10c3122e349b","https://nvd.nist.gov/vuln/detail/CVE-2026-23148"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-23242","cve":"CVE-2026-23242","aliases":[],"title":"Linux kernel SoftiWARP receive path (siw_qp_rx, siw_tcp_rx_data header processing): When siw_get_hdr() rejects a header","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel SoftiWARP receive path (siw_qp_rx, siw_tcp_rx_data header processing)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"When siw_get_hdr() rejects a header with -EINVAL before the receive context has been established, the error path still dereferences qp->rx_fpdu->more_ddp_segs on a NULL rx_fpdu. A remote peer sending a deliberately invalid DDP/MPA header crashes the node. On a shared training node that is one unauthenticated packet evicting every co-resident job.","attack_vector":"Remote, unauthenticated. Any peer that can reach the siw TCP listener.","remediation":"Kernel update guarding the more_ddp_segs check on rx_fpdu being present. Same immediate mitigation as the other siw issues: unload the siw module where SoftiWARP is not in use.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=14ab3da122bd18920ad57428f6cf4fade8385142","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-23242.json"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-401","CWE-772"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-23414","cve":"CVE-2026-23414","aliases":[],"title":"Linux kernel (net/tls): The queue that pins encrypted input buffers while the AEAD engine still references them was","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/tls)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"The queue that pins encrypted input buffers while the AEAD engine still references them was only drained from one caller. If the hold operation fails part-way - after some cloned buffers are already queued - those clones are never released, because the drain only runs when a record completed in fully-async mode. Sustained TLS receive traffic under memory pressure therefore leaks kernel socket buffers steadily, and on a shared node that ends in host memory exhaustion rather than a per-connection problem.","attack_vector":"A remote TLS peer feeding records to a kTLS receive socket, with no credentials required - the failure that starts the leak is an allocation failure inside the hold path, which the peer can encourage by driving high record rates and which a co-tenant can supply directly by putting the node under memory pressure. Every tenant container can attach kTLS receive with setsockopt(TLS_RX), so the reachable surface is any TLS connection terminated in-kernel on the node.","remediation":"Update to 6.1.168 / 6.6.131 / 6.12.80 / 6.18 or later. Interim control: monitor skbuff slab growth on nodes terminating kTLS, keep memory headroom so the hold path does not fail, and blacklist the tls ULP if a node cannot be rebooted and is already showing unexplained skb growth.","references":["https://git.kernel.org/stable/c/ac435be7c7613eb13a5a8ceb5182e10b50c9ce87","https://git.kernel.org/stable/c/2dcf324855c34e7f934ce978aa19b645a8f3ee71","https://nvd.nist.gov/vuln/detail/CVE-2026-23414"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-23440","cve":"CVE-2026-23440","aliases":["net/mlx5e race condition during IPsec ESN update"],"title":"Linux kernel mlx5_core IPsec full offload (ESN handling): The extended-sequence-number wrap event can be processed","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_core IPsec full offload (ESN handling)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"The extended-sequence-number wrap event can be processed twice because the arm flag is re-set too late while the xfrm state lock is dropped and retaken. The driver then programs invalid ESN state, anti-replay fails, and all IPsec traffic on that SA halts. Operationally this is a stall of encrypted east-west traffic, not a memory-safety bug - but on an encrypted fabric it looks like a hard partition.","attack_vector":"Remote and unauthenticated in effect: an IPsec peer driving enough traffic to wrap the sequence number reaches the race. No credentials on the host.","remediation":"Upgrade the host kernel to 7.0 or a stable backport (6.6.130, 6.12.78, 6.18.20, 6.19.10). Rolling reboot of IPsec-offload nodes. Interim: shorten SA rekey intervals so ESN wrap is not reached, a config change on the IKE daemon.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-23440","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-23440.json"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-04-03"},{"id":"CVE-2026-24146","cve":"CVE-2026-24146","aliases":[],"title":"Triton Inference Server: DoS via memory exhaustion on malformed input","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"DoS via memory exhaustion on malformed input","attack_vector":"Any client of the inference endpoint","remediation":"Upgrade Triton; redeploy serving images; add request limits","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24146","https://github.com/NVIDIA/product-security/tree/main/2026/5816"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-789"],"published":"2026-04-07"},{"id":"CVE-2026-24158","cve":"CVE-2026-24158","aliases":[],"title":"NVIDIA Triton Inference Server: A large compressed payload - a decompression bomb against the HTTP endpoint","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"A large compressed payload - a decompression bomb against the HTTP endpoint - exhausts the server. Trivially weaponised by any client. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5790. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24158","https://github.com/NVIDIA/product-security/tree/main/2026/5790"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-789"],"published":"2026-03-24"},{"id":"CVE-2026-24163","cve":"CVE-2026-24163","aliases":[],"title":"TensorRT-LLM: RCE via insecure config deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT-LLM","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"RCE via insecure config deserialization","attack_vector":"Malicious model config","remediation":"Bump TensorRT-LLM; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24163","https://github.com/NVIDIA/product-security/tree/main/2026/5805"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2026-05-20"},{"id":"CVE-2026-24173","cve":"CVE-2026-24173","aliases":[],"title":"NVIDIA Triton Inference Server: A malformed request crashes the server outright","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"A malformed request crashes the server outright. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5816. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24173","https://github.com/NVIDIA/product-security/tree/main/2026/5816"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-190"],"published":"2026-04-07"},{"id":"CVE-2026-24174","cve":"CVE-2026-24174","aliases":[],"title":"NVIDIA Triton Inference Server: A second malformed-request crash path","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"A second malformed-request crash path. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5816. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24174","https://github.com/NVIDIA/product-security/tree/main/2026/5816"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-681"],"published":"2026-04-07"},{"id":"CVE-2026-24175","cve":"CVE-2026-24175","aliases":[],"title":"NVIDIA Triton Inference Server: A malformed request header crashes the server before the body is even parsed","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"A malformed request header crashes the server before the body is even parsed. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5816. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24175","https://github.com/NVIDIA/product-security/tree/main/2026/5816"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-248"],"published":"2026-04-07"},{"id":"CVE-2026-24184","cve":"CVE-2026-24184","aliases":[],"title":"NVIDIA Cumulus Linux - LLDP daemon: Crafted LLDP frames overflow a buffer in the LLDP daemon, reaching code execution","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Cumulus Linux - LLDP daemon","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Crafted LLDP frames overflow a buffer in the LLDP daemon, reaching code execution on the switch from an unauthenticated attacker on an adjacent network. LLDP is processed from every connected port by default, so any compromised host in the rack can reach it.","attack_vector":"Adjacent network, unauthenticated. A single compromised server NIC sends LLDP frames to the leaf it is plugged into. This is a host-to-fabric escalation path.","remediation":"Upgrade Cumulus Linux per bulletin 5817. Cost: switch reboot and link flap, sequenced leaf-by-leaf. Interim control: disable LLDP receive on host-facing ports if your tooling does not depend on it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24184","https://github.com/NVIDIA/product-security/tree/main/2026/5817"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:H/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-120"],"fleet":{"pain_class":"node-reboot"},"published":"2026-08-18"},{"id":"CVE-2026-24209","cve":"CVE-2026-24209","aliases":[],"title":"Triton Inference Server: Arbitrary file access via unsafe path operations","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Arbitrary file access via unsafe path operations","attack_vector":"Network client","remediation":"Upgrade Triton; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24209","https://github.com/NVIDIA/product-security/tree/main/2026/5828"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-22"],"published":"2026-05-20"},{"id":"CVE-2026-24210","cve":"CVE-2026-24210","aliases":[],"title":"Triton Inference Server: DoS / memory corruption (integer overflow in buffer allocation)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"DoS / memory corruption (integer overflow in buffer allocation)","attack_vector":"Network client","remediation":"Upgrade Triton; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24210","https://github.com/NVIDIA/product-security/tree/main/2026/5828"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-190"],"published":"2026-05-20"},{"id":"CVE-2026-24212","cve":"CVE-2026-24212","aliases":[],"title":"Isaac Launchable: Info disclosure (unencrypted sensitive data in transit)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Isaac Launchable","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Info disclosure (unencrypted sensitive data in transit)","attack_vector":"Network observer","remediation":"Upgrade the package; enforce TLS","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24212","https://github.com/NVIDIA/product-security/tree/main/2026/5830"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:H/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-319"],"published":"2026-05-26"},{"id":"CVE-2026-24255","cve":"CVE-2026-24255","aliases":[],"title":"NVIDIA Dynamo: MITM via improper TLS certificate verification","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Dynamo","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"MITM via improper TLS certificate verification","attack_vector":"Network attacker on the inter-node path","remediation":"Bump Dynamo; redeploy; verify TLS config","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24255","https://github.com/NVIDIA/product-security/tree/main/2026/5842"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:H/A:N","cwe":["CWE-1023"],"published":"2026-08-04"},{"id":"CVE-2026-24264","cve":"CVE-2026-24264","aliases":[],"title":"Triton Inference Server: DoS via improper exception handling","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"DoS via improper exception handling","attack_vector":"Any inference client","remediation":"Upgrade Triton; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24264","https://github.com/NVIDIA/product-security/tree/main/2026/5848"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-409"],"published":"2026-07-01"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-863","CWE-200"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2026-28229","cve":"CVE-2026-28229","aliases":["GHSA-56px-hm34-xqj5"],"title":"Argo Workflows (Argo Server, WorkflowTemplate / ClusterWorkflowTemplate endpoints): The template endpoints serve","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo Workflows (Argo Server, WorkflowTemplate / ClusterWorkflowTemplate endpoints)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"The template endpoints serve WorkflowTemplates and ClusterWorkflowTemplates to any caller sending an arbitrary Authorization header - literally 'Bearer nothing' works, because the handlers read from a shared informer instead of the caller's identity. Templates routinely embed Secret manifests, registry credentials and internal endpoints, so this hands an outsider the platform team's operational secrets and a full map of every tenant's pipeline.","attack_vector":"Anyone with network reach to the Argo Server API. Any garbage token satisfies the check.","remediation":"Upgrade Argo Server to 3.7.11 or 4.0.2 and restart. Then rotate every credential that appears inside a WorkflowTemplate or ClusterWorkflowTemplate, since exposure leaves no distinguishing log entry.","references":["https://github.com/argoproj/argo-workflows/security/advisories/GHSA-56px-hm34-xqj5","https://nvd.nist.gov/vuln/detail/CVE-2026-28229"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-20","CWE-835"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-31472","cve":"CVE-2026-31472","aliases":[],"title":"Linux kernel (net/xfrm): One crafted inner IPv4 header (tot_len = 0) inside an IPTFS payload puts the receive path into","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"One crafted inner IPv4 header (tot_len = 0) inside an IPTFS payload puts the receive path into an infinite loop in softirq context. The CPU never leaves the loop, so a single packet takes a core out permanently and wedges packet processing on the node - a fabric-wide stall and an outage for every tenant sharing that host, from one packet, repeatable at will.","attack_vector":"The malformed header is inside the decrypted IPTFS payload, so the sender must be a valid peer on the SA - a peer node on the cluster fabric, a compromised node, or a tenant endpoint terminating an overlay tunnel whose key the tenant holds. Conditional on IPTFS mode being configured. No local access to the victim node is required; the inner header never reaches ip_rcv_core, which is where this validation normally happens.","remediation":"Boot a kernel carrying the fix commits below (no fixed stable version published in the record). Interim control: stop terminating tenant-controlled or untrusted IPTFS tunnels on shared nodes, and drop IP-TFS mode in favour of plain ESP tunnel mode until patched.","references":["https://git.kernel.org/stable/c/de6d8e8ce5187f7402c9859b443355e7120c5f09","https://git.kernel.org/stable/c/3db7d4f777a00164582061ccaa99569cd85011a3","https://nvd.nist.gov/vuln/detail/CVE-2026-31472"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-617","CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-31517","cve":"CVE-2026-31517","aliases":[],"title":"Linux kernel (net/xfrm): A peer that mixes zero-copy-eligible and copy-path IPTFS fragments in one datagram makes","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"A peer that mixes zero-copy-eligible and copy-path IPTFS fragments in one datagram makes reassembly call skb_put() on an already non-linear buffer, hitting SKB_LINEAR_ASSERT and taking an invalid-opcode fault in NAPI softirq. That is a hard kernel panic driven by packet shape - every tenant on the node loses its GPUs, and the sender can repeat it after each reboot.","attack_vector":"Inbound ESP on an IPTFS SA; the fragment sequence is chosen by the sender, so any peer holding the SA reaches it - a peer node, a compromised node, or the far side of a tenant overlay tunnel. The published call trace runs straight from the NIC driver's NAPI poll through xfrm4_esp_rcv into iptfs_reassem_cont, so no local access or tenant device node is involved. Conditional on IPTFS mode being configured.","remediation":"Boot a kernel carrying the fix commits below (no fixed stable version published). Interim control: disable IP-TFS mode on SAs that terminate traffic from tenants or from nodes you do not fully trust.","references":["https://git.kernel.org/stable/c/33a7b36268933c75bdc355e5531951e0ea9f1951","https://git.kernel.org/stable/c/7fdfe8f6efeb0e1200e22a903f2471539f54522b","https://nvd.nist.gov/vuln/detail/CVE-2026-31517"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-32666","cve":"CVE-2026-32666","aliases":["CVE-2026-24060","CVE-2026-25086","ICSA-26-078-08"],"title":"Automated Logic WebCTRL / i-Vu server and controllers, BACnet transport trust: This is the vendor formally conceding","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Automated Logic WebCTRL / i-Vu server and controllers, BACnet transport trust","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"This is the vendor formally conceding the structural problem: WebCTRL inherits BACnet's total absence of network-layer authentication and adds no validation of its own, so an attacker on the BACnet segment can spoof packets to the WebCTRL server or to any Automated Logic controller and have them processed as legitimate. The companion issues are just as bad in practice - service traffic including file contents crosses the wire unencrypted and is trivially readable with Wireshark's BACnet dissector, and under some conditions an attacker can bind the WebCTRL service port and impersonate the server without ever injecting code. Operationally this means write commands to setpoints, fan speeds and schedules can be forged, and the operator has no cryptographic way to tell a real command from a fake one. In a GPU hall the practical consequence is that thermal control is only as trustworthy as the physical and VLAN boundary around the BACnet network, which for most operators is much weaker than they assume.","attack_vector":"Any host that can put packets on the BACnet/IP segment. No credentials exist to steal because none are used. This includes the mechanical contractor's laptop, a compromised BMS workstation, a rogue device in an unlocked mechanical room, and - in a leased colo - anything the landlord has on the shared building network. Also reachable through a BACnet router that bridges IP to MS/TP.","remediation":"Partly unpatchable by design. The plaintext and port-binding issues have fixes in current WebCTRL releases and you should take them, but the underlying spoofing exposure is a protocol property: BACnet/IP has no authentication and Automated Logic explicitly says it does not add validation. The only real control is segmentation and physical security of the BACnet segment - dedicated VLAN, no routing to tenant/corporate/internet, port security or 802.1X on the switch ports that carry it, and locked mechanical rooms. Where the vendor supports BACnet Secure Connect (BACnet/SC), moving to it is the actual fix and it is a controller-by-controller project with a contractor, so budget it as a capital line rather than a patch. Leased site: name this in the contract - require BACnet segment isolation with evidence, because you cannot fix someone else's protocol.","references":["https://www.cisa.gov/news-events/ics-advisories/icsa-26-078-08","https://nvd.nist.gov/vuln/detail/CVE-2026-32666","https://nvd.nist.gov/vuln/detail/CVE-2026-24060","https://www.automatedlogic.com/en/company/security-commitment/"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2026-32981","cve":"CVE-2026-32981","aliases":[],"title":"Ray Dashboard: Path traversal in the dashboard static-file handler (port 8265)","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ray Dashboard","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Path traversal in the dashboard static-file handler (port 8265)","attack_vector":"Unauthenticated network","remediation":"Upgrade past 2.8.1","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-32981"],"status":"curated","published":"2026-03-17"},{"id":"CVE-2026-33554","cve":"CVE-2026-33554","aliases":[],"title":"GNU FreeIPMI's ipmi-oem tool before version 1.6.17: The direction of trust is what makes this operator-relevant","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GNU FreeIPMI's ipmi-oem tool before version 1.6.17","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"The direction of trust is what makes this operator-relevant: the vulnerability is in the management station, not the managed node. A malicious or compromised BMC sends a crafted response and corrupts memory in the tool running on your management host. In a fleet that means one compromised BMC can attack the machine that polls every other BMC - so the blast radius of a single bad node extends to your entire out-of-band control plane, including whatever credentials that host holds for the rest of the fleet. Exploitable buffer overflows in the parsing of IPMI response messages, i.e. the client trusts what the BMC sends back.","attack_vector":"Requires the operator's own tooling to talk to a hostile IPMI responder. That happens when a BMC has already been compromised, when a node of unknown provenance is brought into the fleet, or when an attacker on the management VLAN can spoof or intercept IPMI responses - which the weak IPMI session security elsewhere in this list makes plausible.","remediation":"Package update of FreeIPMI to 1.6.17 or later on every management host, monitoring collector and provisioning box that runs ipmi-oem. This is an ordinary distribution package update, so rollout is cheap - no firmware flash, no node reboot, no maintenance window - which makes it one of the few items here you can just fix. The architectural follow-up worth doing: run fleet IPMI polling from a host that holds no other credentials, so a compromise of the poller does not hand over the whole management plane.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-33554","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2026/33xxx/CVE-2026-33554.json"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2026-33697","cve":"CVE-2026-33697","aliases":["CoCoS aTLS relay"],"title":"Cocos AI - attested TLS (aTLS) on AMD SEV-SNP and Intel TDX: The attested-TLS implementation is vulnerable to a relay","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cocos AI - attested TLS (aTLS) on AMD SEV-SNP and Intel TDX","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"The attested-TLS implementation is vulnerable to a relay attack in which an attacker extracts the ephemeral TLS private key used during the intra-handshake attestation, letting them relay a genuine attestation report from a real confidential VM while terminating the session themselves. Affects both the SEV-SNP and TDX deployment targets. The point of confidential AI is that the client can prove it is talking to a specific attested enclave; a relay attack means that proof is worth nothing while the transport still looks correct.","attack_vector":"A network attacker positioned between the client and the confidential workload. No credentials needed - the attack is against the binding between the attestation and the TLS session, not against either one alone.","remediation":"Upgrade Cocos past v0.8.2. The broader operator lesson is worth more than the patch: any attested-TLS design that does not cryptographically bind the attestation report to the exact TLS key in use is relayable, so if you build or buy a confidential-AI stack, make binding an explicit acceptance criterion. Cost: application upgrade and redeploy; no firmware or driver change.","references":["https://github.com/ultravioletrs/cocos/security/advisories/GHSA-vfgg-mvxx-mgg7","https://nvd.nist.gov/vuln/detail/CVE-2026-33697"],"status":"curated","tags":["tenant-isolation"],"published":"2026-03-27"},{"id":"CVE-2026-34486","cve":"CVE-2026-34486","aliases":[],"title":"Apache Tomcat: Missing encryption of sensitive data introduced by the CVE-2026-29146 fix","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Apache Tomcat","year":"2026","cvss_score":7.5,"severity":"high","kev":true,"impact":"Missing encryption of sensitive data introduced by the CVE-2026-29146 fix","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade Tomcat under every Java control-plane/console service","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-34486"],"status":"curated","published":"2026-04-09"},{"id":"CVE-2026-41523","cve":"CVE-2026-41523","aliases":[],"title":"vLLM (activation function loading): Assert-based security check bypass, unauthenticated","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (activation function loading)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Assert-based security check bypass, unauthenticated","attack_vector":"Unauthenticated network","remediation":"Upgrade to 0.22.0+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-41523"],"status":"curated","published":"2026-06-22"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-770"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2026-42294","cve":"CVE-2026-42294","aliases":["GHSA-jcc8-g2q4-9fxq"],"title":"Argo Workflows (Argo Server, webhook interceptor /api/v1/events/): The webhook interceptor buffers the entire request","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo Workflows (Argo Server, webhook interceptor /api/v1/events/)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"The webhook interceptor buffers the entire request body before it authenticates or verifies the signature, so a multi-gigabyte POST to the publicly reachable /api/v1/events/ endpoint drives Argo Server into OOM. Losing Argo Server takes down the UI and API that every tenant uses to submit and inspect workflows.","attack_vector":"Any unauthenticated host that can reach the Argo Server event endpoint. That is the internet wherever the webhook path is exposed for CI or Git integrations.","remediation":"Upgrade Argo Server to 3.7.14 or 4.0.5 and restart. As defense in depth, cap request body size at the ingress or service mesh in front of /api/v1/events/ and set a memory limit on the Argo Server pod so an OOM restarts one pod instead of pressuring the node.","references":["https://github.com/argoproj/argo-workflows/security/advisories/GHSA-jcc8-g2q4-9fxq","https://nvd.nist.gov/vuln/detail/CVE-2026-42294"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-833","CWE-400"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-43253","cve":"CVE-2026-43253","aliases":[],"title":"Linux kernel (drivers/iommu/amd): The AMD IOMMU busy-waits for command completion while holding its spinlock with","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/amd)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"The AMD IOMMU busy-waits for command completion while holding its spinlock with interrupts disabled, so under heavy DMA map/unmap pressure the whole node soft-locks. One tenant generating aggressive IOMMU traffic stalls every other tenant's device on that IOMMU - a shared-node availability failure, and the kernel CNA rated it network-reachable because fabric traffic is what drives the mapping churn.","attack_vector":"Any workload that drives high-rate DMA mapping on an AMD-Vi host with iommu.strict=1: a tenant hammering unmap through a passed-through NIC or GPU, or high packet rates through an SR-IOV VF. Conditional on strict (non-deferred) IOMMU invalidation mode; no host privilege.","remediation":"Update to a stable kernel carrying commits f2f65b28 / 715c2631. Interim: on AMD hosts, avoid iommu.strict=1 where your threat model tolerates deferred invalidation (note that lazy mode itself widens the stale-mapping window, so this is a genuine trade-off), and rate-limit tenant DMA mapping churn if the stack allows.","references":["https://git.kernel.org/stable/c/f2f65b28d802a667119147444ec2ae33eebf9a58","https://git.kernel.org/stable/c/715c263119fd1b918a9fcbd8a36ea5b604a46324","https://nvd.nist.gov/vuln/detail/CVE-2026-43253"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-911","CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-43464","cve":"CVE-2026-43464","aliases":[],"title":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core): When an XDP program shrinks a multi-fragment receive buffer","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"When an XDP program shrinks a multi-fragment receive buffer, the driver loses track of the fragments the XDP core already released, so the page-pool reference count on those pages goes negative. A negative refcount means pages are handed back to the allocator while the receive ring still owns them - page-pool accounting corruption on the RX path that ends in reuse of live DMA pages and node instability.","attack_vector":"Driven from the wire: a fabric peer sends frames large enough to span multiple receive fragments (jumbo / scattered payloads), and any XDP program on the interface that calls bpf_xdp_pull_data() or bpf_xdp_adjust_tail() triggers the miscount. Conditional on legacy (non-striding) RQ plus an attached XDP multi-buffer program - common on nodes running an XDP-based dataplane or eBPF firewall in front of tenants.","remediation":"Update to 6.7.x / 6.13.x / 6.18 or later per your stream. Interim: detach XDP multi-buffer programs from mlx5 interfaces, or switch the interfaces to striding RQ so the legacy-RQ fragment path is not used.","references":["https://git.kernel.org/stable/c/c74557495efb4bd0adefdfc8678ecdbc82a06da3","https://git.kernel.org/stable/c/03cb50e5b74fce8bf6d92b860371b66253cf0f8d","https://nvd.nist.gov/vuln/detail/CVE-2026-43464"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-362","CWE-665"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-45944","cve":"CVE-2026-45944","aliases":[],"title":"Linux kernel (drivers/iommu/intel): The 128-bit VT-d context entry is zeroed with multiple writes while its Present bit","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/intel)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"The 128-bit VT-d context entry is zeroed with multiple writes while its Present bit is still set, so the IOMMU can fetch a torn entry - some fields already cleared, still marked present. The hardware then translates a device's DMA through a half-demolished context, which is undefined behaviour on the structure that binds a device to its tenant's address space.","attack_vector":"Runs on the device-context teardown path: unbinding a device from its domain, which happens when a tenant releases a passthrough device or the operator rebinds a card. Requires VT-d and a race between the CPU zeroing the entry and a hardware fetch, so it is timing-dependent - but the timing is driven by the tenant's own DMA traffic during teardown.","remediation":"Update to a stable kernel carrying commits c716a59e / d2138abc. Interim: quiesce device DMA before releasing a passthrough device (stop the tenant workload, then unbind) rather than tearing down under active traffic.","references":["https://git.kernel.org/stable/c/c716a59e9977d751e5eb54bcfa6a80124cb5067b","https://git.kernel.org/stable/c/d2138abc8f0a7fce4101b7229b43b06811ed083d","https://nvd.nist.gov/vuln/detail/CVE-2026-45944"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-46027","cve":"CVE-2026-46027","aliases":[],"title":"Linux kernel SMC (early link-group access on CLC decline in smc_clc_wait_msg): A peer can send a CLC decline before the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel SMC (early link-group access on CLC decline in smc_clc_wait_msg)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"A peer can send a CLC decline before the connection has been attached to a link group, and the decline handler updates link-group-level state that does not exist yet. The remote side chooses the timing, so an unauthenticated peer declines early and dereferences the unset link group on the node it is talking to.","attack_vector":"Remote, unauthenticated, during SMC handshake - send a decline before link-group setup completes.","remediation":"Kernel update guarding the link-group update on the link group existing. Keep SMC off tenant-reachable interfaces if it is not deliberately used.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=22546729b96fc873b23065dc49e3d73c45cfb874","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-46027.json"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-46114","cve":"CVE-2026-46114","aliases":["RDMA/rxe ATOMIC_WRITE zero-length disclosure","Soft-RoCE skb tailroom leak"],"title":"Linux kernel - RDMA/rxe (Soft-RoCE) responder, drivers/infiniband/sw/rxe/rxe_resp.c: Atomic_write_reply() dereferences","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel - RDMA/rxe (Soft-RoCE) responder, drivers/infiniband/sw/rxe/rxe_resp.c","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Atomic_write_reply() dereferences 8 bytes at the payload address unconditionally, while the rkey check accepted an ATOMIC_WRITE with a RETH length of zero. A remote initiator sending a zero-length ATOMIC_WRITE makes the responder read 8 bytes past the logical end of the packet into the socket buffer's tailroom and then write those bytes into the attacker's own memory region. That is a clean remote read primitive: four bytes of uninitialised kernel memory disclosed to the attacker per probe, repeatable at will, ideal for defeating KASLR or harvesting kernel pointers before a heavier exploit. The IB specification defines ATOMIC_WRITE as exactly 8 bytes, so anything else was always protocol-invalid.","attack_vector":"Remote initiator on an rxe connection sets the RETH length to 0 on an ATOMIC_WRITE and reads back the responder's reply, which now contains kernel tailroom bytes. Repeat to accumulate a memory-disclosure oracle. No local privilege and no authentication beyond reaching the Soft-RoCE endpoint.","remediation":"Host reboot / kernel upgrade. Same family control as the other rxe findings: blacklist and unload rdma_rxe where Soft-RoCE is not intentionally deployed, which is a config change with no downtime and removes this along with the rest. Where rxe is in use, upgrade the kernel and reboot on rolling drain, and firewall UDP/4791 to known peers in the meantime.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2026/CVE-2026-46114.json","https://nvd.nist.gov/vuln/detail/CVE-2026-46114"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-05-28"},{"id":"CVE-2026-46133","cve":"CVE-2026-46133","aliases":["RDMA/rxe unknown opcode ICRC out-of-bounds","Soft-RoCE rxe_opcode table gap"],"title":"Linux kernel - RDMA/rxe (Soft-RoCE) ICRC processing, drivers/infiniband/sw/rxe: The follow-up to CVE-2026-46043, and","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel - RDMA/rxe (Soft-RoCE) ICRC processing, drivers/infiniband/sw/rxe","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"The follow-up to CVE-2026-46043, and the reason to check you have both. The rxe_opcode[] table has 256 entries but only defined IB opcodes are populated; an undefined opcode such as 0xff reads a zero-initialised entry, so the length check added by the previous fix degenerates to a comparison against zero and stops constraining the packet length. rxe_icrc_hdr() then computes length minus the BTH size, which underflows, producing an out-of-bounds read. One unauthenticated UDP packet still panics the node. The defect predates the earlier fix and reaches back to the original Soft-RoCE driver, so any kernel with rxe loaded has carried it for years.","attack_vector":"A single UDP datagram to port 4791 carrying an opcode not defined in the IB specification. No connection state, no authentication, no prior contact with the target. Trivially scriptable and trivially fleet-wide.","remediation":"Host reboot / kernel upgrade to a kernel carrying this fix specifically - patching only CVE-2026-46043 leaves you exposed. As with the rest of the rxe family, the zero-cost control is to blacklist and unload rdma_rxe where Soft-RoCE is not in use (config change, no downtime), which is the right answer on essentially every production GPU node with real RDMA hardware. Where rxe must stay, restrict UDP/4791 at the host firewall and switch ACLs to known peers while the kernel rollout proceeds.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2026/CVE-2026-46133.json","https://nvd.nist.gov/vuln/detail/CVE-2026-46133"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["fabric-dos"],"published":"2026-05-28"},{"id":"CVE-2026-47471","cve":"CVE-2026-47471","aliases":[],"title":"TensorRT-LLM: RCE potential (OOB write in tensor manipulation)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT-LLM","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"RCE potential (OOB write in tensor manipulation)","attack_vector":"Any inference client","remediation":"Bump TensorRT-LLM; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47471","https://github.com/NVIDIA/product-security/tree/main/2026/5840"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:H/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-122"],"published":"2026-07-14"},{"id":"CVE-2026-47476","cve":"CVE-2026-47476","aliases":[],"title":"Triton Inference Server: DoS via resource exhaustion","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"DoS via resource exhaustion","attack_vector":"Any inference client","remediation":"Upgrade Triton; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47476","https://github.com/NVIDIA/product-security/tree/main/2026/5853"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-400"],"published":"2026-07-14"},{"id":"CVE-2026-47477","cve":"CVE-2026-47477","aliases":[],"title":"Triton Inference Server: DoS via stack overflow in recursive processing","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"DoS via stack overflow in recursive processing","attack_vector":"Any inference client","remediation":"Upgrade Triton; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47477","https://github.com/NVIDIA/product-security/tree/main/2026/5853"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-121"],"published":"2026-07-14"},{"id":"CVE-2026-47478","cve":"CVE-2026-47478","aliases":[],"title":"Triton Inference Server: DoS via memory leak (improper resource cleanup)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"DoS via memory leak (improper resource cleanup)","attack_vector":"Any inference client","remediation":"Upgrade Triton; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47478","https://github.com/NVIDIA/product-security/tree/main/2026/5853"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-910"],"published":"2026-07-14"},{"id":"CVE-2026-47479","cve":"CVE-2026-47479","aliases":[],"title":"Triton Inference Server: DoS via allocation-request resource exhaustion","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"DoS via allocation-request resource exhaustion","attack_vector":"Any inference client","remediation":"Upgrade Triton; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47479","https://github.com/NVIDIA/product-security/tree/main/2026/5853"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-400"],"published":"2026-07-14"},{"id":"CVE-2026-47480","cve":"CVE-2026-47480","aliases":[],"title":"Triton Inference Server: DoS via unchecked exception","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"DoS via unchecked exception","attack_vector":"Any inference client","remediation":"Upgrade Triton; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47480","https://github.com/NVIDIA/product-security/tree/main/2026/5853"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-248"],"published":"2026-07-14"},{"id":"CVE-2026-47482","cve":"CVE-2026-47482","aliases":[],"title":"Triton Inference Server: DoS via resource leak in error paths","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"DoS via resource leak in error paths","attack_vector":"Any inference client","remediation":"Upgrade Triton; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47482","https://github.com/NVIDIA/product-security/tree/main/2026/5853"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-401"],"published":"2026-07-14"},{"id":"CVE-2026-47612","cve":"CVE-2026-47612","aliases":[],"title":"NVIDIA Dynamo: Arbitrary file access via path traversal","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Dynamo","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Arbitrary file access via path traversal","attack_vector":"Network client","remediation":"Bump Dynamo; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47612","https://github.com/NVIDIA/product-security/tree/main/2026/5842"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-22"],"published":"2026-08-04"},{"id":"CVE-2026-47613","cve":"CVE-2026-47613","aliases":[],"title":"NVIDIA Dynamo: SSRF via URL handling","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Dynamo","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"SSRF via URL handling","attack_vector":"Tenant-supplied URL","remediation":"Bump Dynamo; add egress network policy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47613","https://github.com/NVIDIA/product-security/tree/main/2026/5842"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-918"],"published":"2026-08-04"},{"id":"CVE-2026-47614","cve":"CVE-2026-47614","aliases":[],"title":"NVIDIA Dynamo: SSRF in remote resource loading","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Dynamo","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"SSRF in remote resource loading","attack_vector":"Tenant-supplied model URI","remediation":"Bump Dynamo; add egress policy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47614","https://github.com/NVIDIA/product-security/tree/main/2026/5842"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-918"],"published":"2026-08-04"},{"id":"CVE-2026-47615","cve":"CVE-2026-47615","aliases":[],"title":"NVIDIA Dynamo: SSRF in model fetching","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Dynamo","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"SSRF in model fetching","attack_vector":"Tenant-supplied model URI","remediation":"Bump Dynamo; add egress policy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47615","https://github.com/NVIDIA/product-security/tree/main/2026/5842"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-918"],"published":"2026-08-04"},{"id":"CVE-2026-47616","cve":"CVE-2026-47616","aliases":[],"title":"NVIDIA Dynamo: SSRF in data retrieval","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Dynamo","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"SSRF in data retrieval","attack_vector":"Tenant-supplied URI","remediation":"Bump Dynamo; add egress policy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47616","https://github.com/NVIDIA/product-security/tree/main/2026/5842"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-918"],"published":"2026-08-04"},{"id":"CVE-2026-47617","cve":"CVE-2026-47617","aliases":[],"title":"NVIDIA Dynamo: SSRF in API communication","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Dynamo","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"SSRF in API communication","attack_vector":"Tenant-supplied URI","remediation":"Bump Dynamo; add egress policy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47617","https://github.com/NVIDIA/product-security/tree/main/2026/5842"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-918"],"published":"2026-08-04"},{"id":"CVE-2026-47618","cve":"CVE-2026-47618","aliases":[],"title":"NVIDIA Dynamo: SSRF in external service communication","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Dynamo","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"SSRF in external service communication","attack_vector":"Tenant-supplied URI","remediation":"Bump Dynamo; add egress policy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47618","https://github.com/NVIDIA/product-security/tree/main/2026/5842"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-918"],"published":"2026-08-04"},{"id":"CVE-2026-47628","cve":"CVE-2026-47628","aliases":[],"title":"NVIDIA Triton Inference Server: Allocation of resources without limits lets an unauthenticated caller exhaust","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Allocation of resources without limits lets an unauthenticated caller exhaust the server. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5865. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47628","https://github.com/NVIDIA/product-security/tree/main/2026/5865"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-770"],"published":"2026-08-18"},{"id":"CVE-2026-47629","cve":"CVE-2026-47629","aliases":[],"title":"NVIDIA Triton Inference Server: Improper input validation crashes the server","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Improper input validation crashes the server. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5865. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47629","https://github.com/NVIDIA/product-security/tree/main/2026/5865"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-20"],"published":"2026-08-18"},{"id":"CVE-2026-50031","cve":"CVE-2026-50031","aliases":[],"title":"GNU FreeIPMI ipmi-oem before 1.6.18: Same shape as its predecessor and the same fleet consequence: a hostile BMC","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GNU FreeIPMI ipmi-oem before 1.6.18","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Same shape as its predecessor and the same fleet consequence: a hostile BMC response corrupts memory in the tool on your management host. The operationally important detail is that this is the second round - an operator who updated to 1.6.17 believing FreeIPMI's OEM response parsing was fixed is still exposed. Treat the ipmi-oem response parser as untrusted code handling untrusted input until proven otherwise, rather than assuming a given release closed the class. A further set of response-message buffer overflows found after the 1.6.17 fix, meaning the first round of hardening did not cover the whole parser.","attack_vector":"Your management tooling querying a BMC that returns crafted responses - a compromised controller, a node brought in from an untrusted source, or an attacker able to interpose on IPMI traffic across the management VLAN.","remediation":"Package update to FreeIPMI 1.6.18 or later everywhere ipmi-oem runs. Cheap: distribution package update, no reboot, no firmware. Given that OEM extension parsing has now produced two rounds of overflows, the stronger move for most GPU operators is to stop using ipmi-oem for routine fleet polling at all - vendor-specific OEM IPMI commands are rarely load-bearing, and dropping them removes this parser from your control plane entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-50031","https://lists.gnu.org/archive/html/info-gnu/2026-06/msg00000.html","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2026/50xxx/CVE-2026-50031.json"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-401"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-52974","cve":"CVE-2026-52974","aliases":[],"title":"Linux kernel (net/tls): When kTLS RX offload fails at tls_dev_add, the rollback frees the software context but never","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/tls)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"When kTLS RX offload fails at tls_dev_add, the rollback frees the software context but never frees the strparser anchor skb allocated during init. Every failed offload attempt leaks an skb, so a tenant that loops failing setsockopt calls bleeds kernel memory until the node starts OOM-killing other tenants' workloads.","attack_vector":"Local and unprivileged, no device node: setsockopt(SOL_TLS, TLS_RX) on a socket bound to a NIC that advertises kTLS RX offload but rejects the add - which happens once the NIC's offload contexts are exhausted, a state the same tenant can create. Only affects nodes with kTLS RX offload-capable NICs (ConnectX-class), which is exactly the storage-path configuration in these fleets.","remediation":"Boot a kernel carrying the linked stable commits. Interim: disable kTLS RX hardware offload (ethtool -K <dev> tls-hw-rx-offload off) so the failing add path is never taken, and cap per-tenant memory.","references":["https://git.kernel.org/stable/c/0c9f399b37ce22a5ed94cc51f03ed07ac7f38e32","https://git.kernel.org/stable/c/688f12aa44511dd57e448eb670075c6302ad1dc1","https://nvd.nist.gov/vuln/detail/CVE-2026-52974"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-401"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-53229","cve":"CVE-2026-53229","aliases":[],"title":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core/en): Every time an XDP_TX transmit fails because the XDP send","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core/en)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Every time an XDP_TX transmit fails because the XDP send queue is full, the driver leaks both the frame and its DMA mapping. Under sustained load that is thousands of device-writable IOMMU mappings left standing for buffers the kernel no longer tracks, plus unbounded memory growth - the node runs out of memory and the stale mappings remain addressable by the NIC.","attack_vector":"A peer on the fabric drives it directly: flood the node hard enough that the XDP transmit queue backs up and every dropped XDP_TX frame leaks. Conditional on AF_XDP zero-copy being in use with an XDP program that returns XDP_TX on an mlx5 interface. No tenant device node required - the leak is in host memory shared by all tenants on the node.","remediation":"Update to a kernel carrying the fix on your stream. Interim: stop using AF_XDP zero-copy with XDP_TX on mlx5 interfaces, or rate-limit untrusted ingress so the XDP send queue does not saturate.","references":["https://git.kernel.org/stable/c/7b3eeba50fbc3b45f279037c29a87a90e8bac1e1","https://git.kernel.org/stable/c/2789b74ae1f4b68333c9d5eec2f3354d07b16e61","https://nvd.nist.gov/vuln/detail/CVE-2026-53229"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:R/S:C/C:L/I:H/A:L","cwe":["CWE-59"],"fleet":{"pain_class":"hot-patch"},"id":"CVE-2026-54572","cve":"CVE-2026-54572","aliases":[],"title":"rclone (local backend, --links): When rclone copies from an untrusted remote with --links, it recreates symlinks","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"rclone (local backend, --links)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"When rclone copies from an untrusted remote with --links, it recreates symlinks without validating the target, so a malicious remote plants a link that makes rclone write outside the destination directory. On a data-ingest node that is arbitrary file write as the mover's user - enough to drop a file into a systemd unit or an authorized_keys path.","attack_vector":"Anyone who controls content on a remote that your rclone job syncs down with --links enabled. In a GPU cluster that includes tenant-writable buckets used as ingest sources.","remediation":"Upgrade rclone to the fixed release. Drop --links from jobs that pull from any source a tenant can write to, and run ingest jobs as an unprivileged user in a directory that contains nothing security-relevant.","references":["https://github.com/rclone/rclone/security/advisories/GHSA-cf44-9pgv-m4xc","https://nvd.nist.gov/vuln/detail/CVE-2026-54572"],"status":"curated"},{"id":"CVE-2026-57231","cve":"CVE-2026-57231","aliases":[],"title":"Podman: Image env var with a key and no value causes Podman to pass the host's value of that variable","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Podman","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Image env var with a key and no value causes Podman to pass the host's value of that variable into the container; host secret leakage","attack_vector":"Malicious image","remediation":"Upgrade Podman to 5.8.4+; sanitise the environment of any host that runs untrusted images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-57231"],"status":"curated","published":"2026-06-26"},{"id":"CVE-2026-5757","cve":"CVE-2026-5757","aliases":[],"title":"Ollama (quantization engine): Unauthenticated remote information disclosure — reads and exfiltrates model data","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ollama (quantization engine)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Unauthenticated remote information disclosure — reads and exfiltrates model data","attack_vector":"Unauthenticated network to the quantization endpoint","remediation":"Upgrade; direct cross-tenant model-weight exposure if Ollama is shared","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-5757"],"status":"curated","published":"2026-06-26"},{"id":"CVE-2026-62430","cve":"CVE-2026-62430","aliases":["XSA-503"],"title":"Xen (vRTC): Out-of-bounds read in vRTC emulation - hypervisor memory disclosure to a guest","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen (vRTC)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Out-of-bounds read in vRTC emulation - hypervisor memory disclosure to a guest","attack_vector":"Tenant VM guest","remediation":"Hypervisor patch + reboot/evacuation","references":["https://xenbits.xen.org/xsa/advisory-503.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-07-28"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476","CWE-1284"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-64048","cve":"CVE-2026-64048","aliases":[],"title":"Linux kernel SMC-D client (CHID matching against unpopulated ism_dev slot): Slot 0 of the client's ISM device array is","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel SMC-D client (CHID matching against unpopulated ism_dev slot)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Slot 0 of the client's ISM device array is reserved for a V1 device and left zeroed when only V2 devices are found. The accepted-CHID matcher compares from index 0 using the CHID alone, so a malicious server replying with CHID 0 matches the empty slot, the client selects a NULL device, and the following lgr_lock dereference faults. The client is the victim here: a hostile SMC server crashes every node that connects to it.","attack_vector":"Remote, from the server side. A malicious or compromised SMC peer answers a V2-only proposal with CHID 0.","remediation":"Kernel update rejecting a CHID-0 match against an empty slot. Do not let tenant workloads act as SMC servers for host-level clients, and keep SMC disabled where it is not intentional.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=277740023def559a4a2ddc3e8e784ee37a0f16a9","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-64048.json"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-64210","cve":"CVE-2026-64210","aliases":[],"title":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core): Two CPUs write to the internal control send queue without","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Two CPUs write to the internal control send queue without holding a lock, so one overwrites the other's work-queue entries and advances the producer index over them. The queue that feeds receive-buffer posting ends up with garbage descriptors ('Bad OP in ICOSQ CQE'), which stalls receive on that channel and takes packet reception away from every workload sharing the interface.","attack_vector":"Racy path is in NAPI poll, so sustained receive load plus an interrupt-affinity change is enough to hit it - a busy fabric peer supplies the load. Conditional on AF_XDP zero-copy being active on the mlx5 interface, which puts the XSK UMR posting path and the IRQ-trigger path on the same queue. No tenant device node needed; it is host-shared driver state.","remediation":"Update to a kernel carrying the fix on your stream. Interim: stop running AF_XDP zero-copy sockets on mlx5 interfaces, and pin IRQ affinity so channels do not migrate between CPUs under load.","references":["https://git.kernel.org/stable/c/8d3b91e7d81000d295cd914d4d9d6f860252e2bf","https://git.kernel.org/stable/c/c326f9c68921e2f14dfcecb2f6b4216313d50248","https://nvd.nist.gov/vuln/detail/CVE-2026-64210"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-65315","cve":"CVE-2026-65315","aliases":[],"title":"Ollama (GGUF metadata parser): Uncontrolled memory allocation","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ollama (GGUF metadata parser)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Uncontrolled memory allocation → remote crash","attack_vector":"Customer-supplied GGUF","remediation":"Upgrade; GGUF parser hardening is still incomplete three years on","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-65315"],"status":"curated","published":"2026-07-21"},{"id":"CVE-2026-69111","cve":"CVE-2026-69111","aliases":[],"title":"Milvus: Unauthenticated DoS terminating service components","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Milvus","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Unauthenticated DoS terminating service components","attack_vector":"Unauthenticated network","remediation":"Upgrade past 2.6.22","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-69111"],"status":"curated","published":"2026-08-05"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-20","CWE-400"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-72330","cve":"CVE-2026-72330","aliases":[],"title":"Linux kernel (net/tls): A remote peer sends a zero-length TLS 1.3 application_data record - which the RFC explicitly","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/tls)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"A remote peer sends a zero-length TLS 1.3 application_data record - which the RFC explicitly permits as a traffic-analysis countermeasure, so this is legal traffic, not an attack payload - and the kTLS read_sock consumer treats the zero bytes consumed as backpressure. The empty record is requeued at the head of the receive list and retried forever, so every subsequent record on that connection is blocked behind it. One legal record from the far end permanently wedges a kTLS connection carrying storage or control traffic, with no way to recover short of tearing the connection down.","attack_vector":"Any remote TLS peer, with no credentials and no local access - it only has to be the other end of an established kTLS session. Conditional on the consumer using the read_sock path rather than recvmsg: that means sockmap/BPF splicing datapaths and in-kernel consumers, which is exactly how a service mesh or storage proxy terminates kTLS. Plain recvmsg users are unaffected because that path already consumes empty records correctly.","remediation":"Boot a kernel carrying the fix commits below (no fixed stable version published). Interim control: route kTLS traffic through recvmsg-based consumers rather than sockmap/read_sock splicing until patched, and add connection-level liveness timeouts so a wedged connection is torn down instead of stalling a storage path indefinitely.","references":["https://git.kernel.org/stable/c/0867b0f2513ebc1c475af9898c97f4772a68d964","https://git.kernel.org/stable/c/c6b440cf766a557b08d25f1b571b3d57d039686e","https://nvd.nist.gov/vuln/detail/CVE-2026-72330"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-772"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-74385","cve":"CVE-2026-74385","aliases":[],"title":"Linux kernel (drivers/nvme/target): A client that completes the TLS handshake against the NVMe-oF TCP target and then","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/nvme/target)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"A client that completes the TLS handshake against the NVMe-oF TCP target and then lets the socket fall out of the established state leaks the whole queue structure and its socket - the target never notices the failure and never cleans up. Repeating it exhausts memory and socket resources on the shared storage node.","attack_vector":"Reachable from any peer on the fabric that can complete a TLS handshake with the nvmet-tcp listener; the leak is in the handshake-completion callback, so it lands before the controller association is established. Conditional on nvmet-tcp being configured with TLS - a plain-text target is not affected by this path.","remediation":"No fixed version is listed on this record - boot a kernel carrying the linked stable commits. Interim: restrict fabric reachability of the TLS-enabled nvmet-tcp port to the storage network, and watch for queue/socket counts that climb without a matching client population.","references":["https://git.kernel.org/stable/c/cba2ee57fd302727aea7d41e9d9cd0969f5df0fb","https://git.kernel.org/stable/c/22aa70f9a0544643ec37d442b6fcb1833d804462","https://nvd.nist.gov/vuln/detail/CVE-2026-74385"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-401","CWE-772"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-74386","cve":"CVE-2026-74386","aliases":[],"title":"Linux kernel (drivers/nvme/target): Every connection that dies partway through queue allocation on the NVMe-oF TCP","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/nvme/target)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Every connection that dies partway through queue allocation on the NVMe-oF TCP target leaks a page-fragment cache reference that is never reclaimed. An unauthenticated peer that loops connect-then-drop drives unbounded kernel memory growth on the storage node until it OOMs, which is a full outage for every tenant using that target.","attack_vector":"Any peer with network reach to the nvmet-tcp listening port. The leak happens in nvmet_tcp_alloc_queue's error path, before the connection is authenticated or associated with a subsystem, so no credentials and no tenant device node are needed - just repeated half-open connections. Requires the node to be running nvmet with a TCP port enabled.","remediation":"No fixed version is listed on this record - boot a kernel carrying the linked stable commits. Interim: restrict who can reach the nvmet-tcp port (storage VLAN only), rate-limit new connections to it at the firewall, and monitor slab growth on target nodes.","references":["https://git.kernel.org/stable/c/a43a9abc1ebf663f0aa56a729106f68dd9c77da6","https://git.kernel.org/stable/c/ba3209704b3cd46961e4e081af5c52a780785648","https://nvd.nist.gov/vuln/detail/CVE-2026-74386"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-667","CWE-401"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-74396","cve":"CVE-2026-74396","aliases":[],"title":"Linux kernel (drivers/infiniband/hw/mlx5): When on-demand-paging translation-table population fails, the UMR path","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/hw/mlx5)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"When on-demand-paging translation-table population fails, the UMR path returns without running its cleanup: the DMA mapping and buffer leak, and if the shared emergency translation page was in use its global mutex is left permanently locked. Every subsequent memory-region update on that node then blocks forever - the RDMA stack wedges for all tenants on the box, not just the one that hit the failure.","attack_vector":"Requires Mellanox/NVIDIA mlx5 with ODP in use - the standard NIC under most GPU clusters. Triggered from a tenant container holding /dev/infiniband/uverbs* that registers on-demand-paging memory regions and drives the populate failure. The CNA scored it network-reachable with no privileges, reflecting that page-fault population is also driven by remote RDMA traffic against the tenant's ODP regions.","remediation":"No fixed version is listed in the record - take the stable kernel carrying ffa85a2c1979 (or 9619909d4869 / 1eae35b37923) and reboot; once the mutex is stuck only a reboot clears it. Interim: disable ODP (do not advertise implicit ODP to tenants) on nodes that do not need it.","references":["https://git.kernel.org/stable/c/ffa85a2c197935ace6f1634ad9eb0a44bc615670","https://git.kernel.org/stable/c/9619909d4869afe720904c6888a289b9ac3055b8","https://nvd.nist.gov/vuln/detail/CVE-2026-74396"],"status":"curated","tags":["tenant-isolation"]},{"id":"NCVD-2018-001-rocev2-lossless-ethernet-fabric","cve":null,"aliases":["PFC deadlock","PFC storm","head-of-line blocking","congestion spreading","Revisiting Network Support for RDMA (SIGCOMM 2018)"],"title":"RoCEv2 lossless Ethernet fabric - IEEE 802.1Qbb Priority Flow Control: RoCE requires a lossless network, which in","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"RoCEv2 lossless Ethernet fabric - IEEE 802.1Qbb Priority Flow Control","year":"2018","cvss_score":7.5,"severity":"high","kev":false,"impact":"RoCE requires a lossless network, which in practice means Priority Flow Control, and PFC brings head-of-line blocking, congestion spreading, and occasional deadlocks as documented properties rather than as bugs. PFC pause frames propagate backwards hop by hop, so a single misbehaving or malicious endpoint that refuses to drain its receive queue can pause its upstream switch port, which pauses its upstream, and so on until a large fraction of an AI training fabric stops forwarding. A cyclic buffer dependency produces a deadlock that does not clear on its own. This is the highest-blast-radius item in this slice: one tenant NIC can wedge a whole rail of a GPU cluster and take every job on it down at once.","attack_vector":"An endpoint on the lossless priority class stops consuming traffic - by design, or through a driver hang, or via a hardware fault - and its NIC emits sustained PFC pause frames. Congestion spreads to unrelated flows sharing the same priority queue on intermediate switches, including flows between tenants who have nothing to do with the source. Deadlock arises when routing plus link failures create a cycle of buffer dependencies, and it persists until an operator intervenes. Triggering it requires no protocol violation whatsoever, which is why it also happens accidentally.","remediation":"Config change plus, on many platforms, a switch reload. Enable PFC watchdog on every switch (Cisco NX-OS, Arista EOS, NVIDIA Cumulus, SONiC all ship one) so a stuck queue is drained and the port error-disabled instead of pausing the fabric - this is a config push, but on some platforms the underlying buffer/QoS profile change needs a switch reload, so schedule per-rail. Keep the lossless class narrow (one priority, storage and RDMA separated), tune ECN/DCQCN thresholds to mark before PFC ever fires, and prefer deadlock-free routing (up-down or edge-disjoint) on Clos fabrics. Longer term, move to RoCE implementations that tolerate loss (the IRN design line, and DCQCN-plus-selective-retransmit NIC firmware) so PFC can be disabled entirely - firmware flash plus a driver upgrade fleet-wide, and a redesign of the QoS profile, so treat as a project.","references":["https://arxiv.org/abs/1806.08159","https://arxiv.org/abs/2207.10898"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["fabric-dos"]},{"id":"NCVD-2018-006-rocev2-lossless-ethernet-fabric","cve":null,"aliases":["PFC deadlock","PFC storm","head-of-line blocking","congestion spreading","Revisiting Network Support for RDMA (SIGCOMM 2018)"],"title":"RoCEv2 lossless Ethernet fabric - IEEE 802.1Qbb Priority Flow Control: RoCE requires a lossless network, which in","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"RoCEv2 lossless Ethernet fabric - IEEE 802.1Qbb Priority Flow Control","year":"2018","cvss_score":7.5,"severity":"high","kev":false,"impact":"RoCE requires a lossless network, which in practice means Priority Flow Control, and PFC brings head-of-line blocking, congestion spreading, and occasional deadlocks as documented properties rather than as bugs. PFC pause frames propagate backwards hop by hop, so a single misbehaving or malicious endpoint that refuses to drain its receive queue can pause its upstream switch port, which pauses its upstream, and so on until a large fraction of an AI training fabric stops forwarding. A cyclic buffer dependency produces a deadlock that does not clear on its own. This is the highest-blast-radius item in this slice: one tenant NIC can wedge a whole rail of a GPU cluster and take every job on it down at once.","attack_vector":"An endpoint on the lossless priority class stops consuming traffic - by design, or through a driver hang, or via a hardware fault - and its NIC emits sustained PFC pause frames. Congestion spreads to unrelated flows sharing the same priority queue on intermediate switches, including flows between tenants who have nothing to do with the source. Deadlock arises when routing plus link failures create a cycle of buffer dependencies, and it persists until an operator intervenes. Triggering it requires no protocol violation whatsoever, which is why it also happens accidentally.","remediation":"Config change plus, on many platforms, a switch reload. Enable PFC watchdog on every switch (Cisco NX-OS, Arista EOS, NVIDIA Cumulus, SONiC all ship one) so a stuck queue is drained and the port error-disabled instead of pausing the fabric - this is a config push, but on some platforms the underlying buffer/QoS profile change needs a switch reload, so schedule per-rail. Keep the lossless class narrow (one priority, storage and RDMA separated), tune ECN/DCQCN thresholds to mark before PFC ever fires, and prefer deadlock-free routing (up-down or edge-disjoint) on Clos fabrics. Longer term, move to RoCE implementations that tolerate loss (the IRN design line, and DCQCN-plus-selective-retransmit NIC firmware) so PFC can be disabled entirely - firmware flash plus a driver upgrade fleet-wide, and a redesign of the QoS profile, so treat as a project.","references":["https://arxiv.org/abs/1806.08159","https://arxiv.org/abs/2207.10898"],"status":"curated","tags":["fabric-dos"]},{"id":"NCVD-2021-005-infiniband-roce-communication-ma","cve":null,"aliases":["ReDMArk DoS","RDMA connection-manager resource exhaustion","RNIC queue-pair exhaustion"],"title":"InfiniBand/RoCE Communication Manager (CM) and RNIC connection-state resources: RNICs hold per-connection state in a","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"InfiniBand/RoCE Communication Manager (CM) and RNIC connection-state resources","year":"2021","cvss_score":7.5,"severity":"high","kev":false,"impact":"RNICs hold per-connection state in a fixed, small on-chip resource pool. ReDMArk demonstrated that an unauthenticated peer can drive a victim RNIC to exhaustion simply by opening connections through the RDMA Connection Manager, or by sending malformed CM datagrams that leave half-open state behind. Once the pool is full the victim node can no longer accept new queue pairs, which on a GPU cluster means NCCL rendezvous fails, jobs hang at the allreduce, and the node effectively drops out of the scheduler while still appearing healthy to a ping-based health check.","attack_vector":"Any node that can reach the victim's CM service (UD QP1 on InfiniBand, UDP/4791 plus the CM port on RoCE) opens connections in a loop, or sends CM REQ messages and never completes the handshake. Because CM traffic is unauthenticated and is processed before any application-level identity exists, no tenant credentials are needed. In a shared cluster a single misbehaving or malicious container with RDMA device access is enough.","remediation":"No patch. Config change: rate-limit CM traffic at the switch or in the host's RDMA CM (per-source connection caps), enforce per-tenant P_Key partitions so a tenant can only reach nodes in its own job, and use SR-IOV VF resource limits so one VF cannot consume the PF's whole QP pool. Vendor firmware on newer ConnectX/BlueField parts adds per-VF resource quotas - that is a firmware flash plus a host reboot, rolling. Cheapest immediate mitigation is scheduler-side: do not co-schedule untrusted tenants onto nodes sharing an RNIC.","references":["https://www.usenix.org/conference/usenixsecurity21/presentation/rothenberger","https://netsec.ethz.ch/publications/papers/sec21summer-redmark.pdf"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"tags":["fabric-dos"]},{"id":"NCVD-2021-011-infiniband-roce-communication-ma","cve":null,"aliases":["ReDMArk DoS","RDMA connection-manager resource exhaustion","RNIC queue-pair exhaustion"],"title":"InfiniBand/RoCE Communication Manager (CM) and RNIC connection-state resources: RNICs hold per-connection state in a","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"InfiniBand/RoCE Communication Manager (CM) and RNIC connection-state resources","year":"2021","cvss_score":7.5,"severity":"high","kev":false,"impact":"RNICs hold per-connection state in a fixed, small on-chip resource pool. ReDMArk demonstrated that an unauthenticated peer can drive a victim RNIC to exhaustion simply by opening connections through the RDMA Connection Manager, or by sending malformed CM datagrams that leave half-open state behind. Once the pool is full the victim node can no longer accept new queue pairs, which on a GPU cluster means NCCL rendezvous fails, jobs hang at the allreduce, and the node effectively drops out of the scheduler while still appearing healthy to a ping-based health check.","attack_vector":"Any node that can reach the victim's CM service (UD QP1 on InfiniBand, UDP/4791 plus the CM port on RoCE) opens connections in a loop, or sends CM REQ messages and never completes the handshake. Because CM traffic is unauthenticated and is processed before any application-level identity exists, no tenant credentials are needed. In a shared cluster a single misbehaving or malicious container with RDMA device access is enough.","remediation":"No patch. Config change: rate-limit CM traffic at the switch or in the host's RDMA CM (per-source connection caps), enforce per-tenant P_Key partitions so a tenant can only reach nodes in its own job, and use SR-IOV VF resource limits so one VF cannot consume the PF's whole QP pool. Vendor firmware on newer ConnectX/BlueField parts adds per-VF resource quotas - that is a firmware flash plus a host reboot, rolling. Cheapest immediate mitigation is scheduler-side: do not co-schedule untrusted tenants onto nodes sharing an RNIC.","references":["https://www.usenix.org/conference/usenixsecurity21/presentation/rothenberger","https://netsec.ethz.ch/publications/papers/sec21summer-redmark.pdf"],"status":"curated","tags":["fabric-dos"]},{"id":"NCVD-2024-006-nvidia-gpudirect-rdma-gpu-bar1-w","cve":null,"aliases":["GPUDirect RDMA BAR1 exposure","nvidia-peermem","nvidia_p2p_get_pages","GPU memory registered as an RDMA memory region"],"title":"NVIDIA GPUDirect RDMA - GPU BAR1 window peer-mapped to the RNIC, nvidia-peermem / nvidia_p2p_get_pages: GPUDirect RDMA","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPUDirect RDMA - GPU BAR1 window peer-mapped to the RNIC, nvidia-peermem / nvidia_p2p_get_pages","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"GPUDirect RDMA works by exposing GPU memory through a PCIe BAR window (BAR1) so a third-party device - the NIC - can DMA into it directly, with nvidia_p2p_get_pages() pinning the range and handing the peer device physical addresses. The practical consequence for a cluster operator is that every weakness in RDMA memory protection now applies to HBM, not just host DRAM: an rkey guessed or injected per the ReDMArk findings resolves to a region backed by GPU memory, and a successful remote read returns model weights or KV cache straight out of the GPU with no CPU involvement and no host-side trace. The GPU's own MMU is not in this path - the RNIC's rkey check is the entire access control. NVIDIA's documentation notes the 64-bit p2pToken is randomised specifically to keep an adversary from guessing it, which is an acknowledgement that guessability is the threat model here.","attack_vector":"Requires an attacker able to reach the victim's RNIC on the fabric and to guess or inject against the memory region that covers GPU memory - the ReDMArk and NeVerMore primitives. A second, local path: any process that can obtain a peer-mapping token or that shares the RDMA device can register GPU memory it should not reach, since the peer-memory client trusts the calling context. Stale mappings are a third: the nvidia_p2p callback must free the page table on deallocation, and a mapping that outlives its buffer leaves the NIC pointed at memory that has been handed to another context.","remediation":"Driver upgrade plus config change. Keep the NVIDIA GPU driver, nvidia-peermem (or the in-tree dma-buf peer path on recent kernels), and the RDMA stack on current versions - a driver upgrade requiring a host reboot on GPU nodes, so batch it with a scheduled drain. Config: never share an RNIC or an IB device node between tenants when GPUDirect is enabled, register the narrowest possible GPU regions with the least permission, use a per-tenant protection domain, and prefer dma-buf-based registration with explicit lifetime over legacy peermem pinning. Where GPUDirect is not actually needed for a workload, disabling it removes the exposure entirely at a bandwidth cost. Combine with the fabric partitioning in the InfiniBand and RoCE entries - the RDMA-layer fixes are what actually protect the GPU memory.","references":["https://docs.nvidia.com/cuda/gpudirect-rdma/","https://www.usenix.org/conference/usenixsecurity21/presentation/rothenberger","https://arxiv.org/abs/2202.08080"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"]},{"id":"NCVD-2024-009-nvidia-gpudirect-rdma-gpu-bar1-w","cve":null,"aliases":["GPUDirect RDMA BAR1 exposure","nvidia-peermem","nvidia_p2p_get_pages","GPU memory registered as an RDMA memory region"],"title":"NVIDIA GPUDirect RDMA - GPU BAR1 window peer-mapped to the RNIC, nvidia-peermem / nvidia_p2p_get_pages: GPUDirect RDMA","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPUDirect RDMA - GPU BAR1 window peer-mapped to the RNIC, nvidia-peermem / nvidia_p2p_get_pages","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"GPUDirect RDMA works by exposing GPU memory through a PCIe BAR window (BAR1) so a third-party device - the NIC - can DMA into it directly, with nvidia_p2p_get_pages() pinning the range and handing the peer device physical addresses. The practical consequence for a cluster operator is that every weakness in RDMA memory protection now applies to HBM, not just host DRAM: an rkey guessed or injected per the ReDMArk findings resolves to a region backed by GPU memory, and a successful remote read returns model weights or KV cache straight out of the GPU with no CPU involvement and no host-side trace. The GPU's own MMU is not in this path - the RNIC's rkey check is the entire access control. NVIDIA's documentation notes the 64-bit p2pToken is randomised specifically to keep an adversary from guessing it, which is an acknowledgement that guessability is the threat model here.","attack_vector":"Requires an attacker able to reach the victim's RNIC on the fabric and to guess or inject against the memory region that covers GPU memory - the ReDMArk and NeVerMore primitives. A second, local path: any process that can obtain a peer-mapping token or that shares the RDMA device can register GPU memory it should not reach, since the peer-memory client trusts the calling context. Stale mappings are a third: the nvidia_p2p callback must free the page table on deallocation, and a mapping that outlives its buffer leaves the NIC pointed at memory that has been handed to another context.","remediation":"Driver upgrade plus config change. Keep the NVIDIA GPU driver, nvidia-peermem (or the in-tree dma-buf peer path on recent kernels), and the RDMA stack on current versions - a driver upgrade requiring a host reboot on GPU nodes, so batch it with a scheduled drain. Config: never share an RNIC or an IB device node between tenants when GPUDirect is enabled, register the narrowest possible GPU regions with the least permission, use a per-tenant protection domain, and prefer dma-buf-based registration with explicit lifetime over legacy peermem pinning. Where GPUDirect is not actually needed for a workload, disabling it removes the exposure entirely at a bandwidth cost. Combine with the fabric partitioning in the InfiniBand and RoCE entries - the RDMA-layer fixes are what actually protect the GPU memory.","references":["https://docs.nvidia.com/cuda/gpudirect-rdma/","https://www.usenix.org/conference/usenixsecurity21/presentation/rothenberger","https://arxiv.org/abs/2202.08080"],"status":"curated","tags":["tenant-isolation"]},{"id":"NCVD-2025-006-nvidia-bluefield-3-rnic-dpu-micr","cve":null,"aliases":["Noisy Neighbor","RDMA state saturation attack","RDMA pipeline saturation attack","BlueField-3 resource exhaustion","arXiv:2510.12629"],"title":"NVIDIA BlueField-3 RNIC/DPU - microarchitectural state and pipeline resources under containerized multi-tenancy","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA BlueField-3 RNIC/DPU - microarchitectural state and pipeline resources under containerized multi-tenancy","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Measured on current NVIDIA BlueField-3 hardware, two families of resource-exhaustion attack from a co-resident container cost the victim container up to 93.9% of its bandwidth, a 1,117x latency increase, and a 115% rise in NIC cache misses. Pipeline saturation additionally produces severe link-level congestion with strong amplification - small verb requests translating into disproportionately large resource consumption, so the attacker spends almost nothing. For a GPU cloud this is a direct, reproducible way for one tenant to make another tenant's distributed training run at a fraction of its purchased throughput, on the exact hardware generation being deployed for AI fabrics today.","attack_vector":"A container with ordinary RDMA access issues verb patterns designed either to saturate RNIC connection/translation state (state saturation) or to congest the NIC's processing pipeline (pipeline saturation). No kernel exploit, no privilege escalation, no fabric spoofing - just legitimate verbs at an adversarial shape. Because the amplification is high, an attacker constrained by its own bandwidth quota can still consume the shared NIC.","remediation":"No patch available; NVIDIA has not published an advisory for this class. Config change: enforce per-container caps on queue pairs, memory regions, and completion queues via the RDMA cgroup controller and per-VF resource limits (applied at container/VF creation, no reboot), and monitor per-container verb rates so abuse is at least attributable. The paper proposes HT-Verbs, a telemetry-driven throttling framework that needs no hardware change but is research code, not a product. Structural fix is a dedicated VF or NIC per tenant. Treat RDMA bandwidth SLAs on shared BlueField-3 ports as unenforceable until you have measured your own configuration.","references":["https://arxiv.org/abs/2510.12629","https://www.usenix.org/conference/nsdi23/presentation/kong"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"tags":["tenant-isolation"]},{"id":"NCVD-2025-015-nvidia-bluefield-3-rnic-dpu-micr","cve":null,"aliases":["Noisy Neighbor","RDMA state saturation attack","RDMA pipeline saturation attack","BlueField-3 resource exhaustion","arXiv:2510.12629"],"title":"NVIDIA BlueField-3 RNIC/DPU - microarchitectural state and pipeline resources under containerized multi-tenancy","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA BlueField-3 RNIC/DPU - microarchitectural state and pipeline resources under containerized multi-tenancy","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Measured on current NVIDIA BlueField-3 hardware, two families of resource-exhaustion attack from a co-resident container cost the victim container up to 93.9% of its bandwidth, a 1,117x latency increase, and a 115% rise in NIC cache misses. Pipeline saturation additionally produces severe link-level congestion with strong amplification - small verb requests translating into disproportionately large resource consumption, so the attacker spends almost nothing. For a GPU cloud this is a direct, reproducible way for one tenant to make another tenant's distributed training run at a fraction of its purchased throughput, on the exact hardware generation being deployed for AI fabrics today.","attack_vector":"A container with ordinary RDMA access issues verb patterns designed either to saturate RNIC connection/translation state (state saturation) or to congest the NIC's processing pipeline (pipeline saturation). No kernel exploit, no privilege escalation, no fabric spoofing - just legitimate verbs at an adversarial shape. Because the amplification is high, an attacker constrained by its own bandwidth quota can still consume the shared NIC.","remediation":"No patch available; NVIDIA has not published an advisory for this class. Config change: enforce per-container caps on queue pairs, memory regions, and completion queues via the RDMA cgroup controller and per-VF resource limits (applied at container/VF creation, no reboot), and monitor per-container verb rates so abuse is at least attributable. The paper proposes HT-Verbs, a telemetry-driven throttling framework that needs no hardware change but is research code, not a product. Structural fix is a dedicated VF or NIC per tenant. Treat RDMA bandwidth SLAs on shared BlueField-3 ports as unenforceable until you have measured your own configuration.","references":["https://arxiv.org/abs/2510.12629","https://www.usenix.org/conference/nsdi23/presentation/kong"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.0/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-400"],"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2025-016-bentoml-1-3-9-bundled-gradio-app","cve":null,"aliases":["GHSA-hh3j-9m59-p8vc"],"title":"BentoML 1.3.9 (bundled Gradio app, /login endpoint): The /login endpoint of the integrated Gradio app processes each","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"BentoML 1.3.9 (bundled Gradio app, /login endpoint)","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"The /login endpoint of the integrated Gradio app processes each appended character in a malformed multipart boundary, consuming resources until the service is unavailable. No CVE was assigned to the BentoML advisory itself; the underlying Gradio issue is tracked as CVE-2024-8966. Effect on an operator is a model server that stops serving while still holding its GPU.","attack_vector":"Unauthenticated, no user interaction. Anything that can reach the BentoML serving port.","remediation":"Upgrade BentoML past 1.3.9 and restart serving pods. Cap request size at the ingress in front of the endpoint as a durable mitigation for this whole class.","references":["https://github.com/advisories/GHSA-hh3j-9m59-p8vc","https://huntr.com/bounties/e467ec92-0ad1-4461-8468-1beabf701b9f"],"status":"curated"},{"id":"NCVD-2026-001-openbmc-bmcweb-http-1-1-expect-1","cve":null,"aliases":["GHSA-p3gc-68x5-g9w3","CONN-F2"],"title":"OpenBMC bmcweb HTTP/1.1 Expect: 100-continue handling: bmcweb applies a 4 KB body limit to unauthenticated requests","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"OpenBMC bmcweb HTTP/1.1 Expect: 100-continue handling","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"bmcweb applies a 4 KB body limit to unauthenticated requests, and the Expect: 100-continue path returns before that limit is applied. An unauthenticated client streams roughly 10 MB into BMC memory per connection - a 2,560x amplification over the intended cap - and drives the BMC out of memory. Redfish, the web UI, KVM and serial-over-LAN all die with the process. What makes this one worth tracking specifically: it has a published GitHub advisory but no CVE was ever assigned, so it does not appear in NVD, OSV or the vendor scanners most operators rely on. If your firmware risk process is 'match SBOM against NVD', you will not see it. Intel firmware shipped in January 2026 still contained it.","attack_vector":"Unauthenticated HTTP(S) to bmcweb on the BMC management interface, using a stock HTTP client. No credentials, no host access.","remediation":"Fixed on 2026-04-21 in bmcweb commit 0b2049b0, released in bmcweb 3.0.0; affected versions are 2.18.0 and earlier. Reaching a deployed fleet needs a BMC firmware flash - per node, out-of-band, and dependent on your ODM rebasing to bmcweb 3.0.0, which for most server vendors will lag by quarters. Config-only mitigation and the realistic near-term answer: put the BMC HTTPS port behind an ACL so only management jump hosts can reach it, and do not expose the web interface beyond the management VLAN. Because there is no CVE, add it to your firmware acceptance checklist by bmcweb version rather than expecting a scanner to flag it.","references":["https://github.com/openbmc/bmcweb/security/advisories/GHSA-p3gc-68x5-g9w3","https://seclists.org/fulldisclosure/2026/May/24","https://binreaper.pages.dev/posts/2026-05-27-bmcweb-disclosure/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"NCVD-2026-002-openbmc-bmcweb-http-2-content-le","cve":null,"aliases":["H2-F2"],"title":"OpenBMC bmcweb HTTP/2 Content-Length handling: bmcweb passes the client-supplied Content-Length straight","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"OpenBMC bmcweb HTTP/2 Content-Length handling","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"bmcweb passes the client-supplied Content-Length straight into a reserve() call. One HEADERS frame declaring a nonsense length like 999999999999 throws bad_alloc and takes down the single-threaded event loop - the entire management interface, from one frame, with no body sent and no credentials. Notable for the disclosure hygiene as much as the bug: it was patched the same day as the Expect-header issue but received no advisory of its own and no CVE, so nothing downstream records that affected firmware is vulnerable. An operator auditing by advisory count will undercount bmcweb's exposure.","attack_vector":"Unauthenticated HTTP/2 to bmcweb on the management interface. A single frame is enough; no payload, no session.","remediation":"Fixed 2026-04-21 in bmcweb commit 62526bb0, same release train as the Expect fix (bmcweb 3.0.0). Since it shares a release with the advisory-carrying bug, one BMC firmware flash covers both - per node, out-of-band, ODM-rebase-lagged. Until then, config-only: restrict who can reach the BMC HTTPS port, and disable h2 if your tooling permits. Verify by bmcweb version, not by scanner output, since no CVE exists to match against.","references":["https://seclists.org/fulldisclosure/2026/May/24","https://binreaper.pages.dev/posts/2026-05-27-bmcweb-disclosure/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"NCVD-2026-003-openbmc-bmcweb-http-2-body-buffe","cve":null,"aliases":["H2-F1"],"title":"OpenBMC bmcweb HTTP/2 body buffering (HttpBody::reader, nghttp2 flow control): The HTTP/2 code path in bmcweb appends","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"OpenBMC bmcweb HTTP/2 body buffering (HttpBody::reader, nghttp2 flow control)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"The HTTP/2 code path in bmcweb appends incoming DATA frames with no size check at all, and nghttp2's default window auto-replenishment means the data flows before authentication happens. Same outcome as the Expect bug but worse, because it bypasses the body limit entirely rather than merely leaking past it, and because h2 is negotiated by default over ALPN so it is the path a modern client takes automatically. The researcher rates it the most severe of the four findings. Reported as unfixed on master at time of disclosure, with no advisory and no CVE - so this is a known, public, unpatched pre-auth DoS in the daemon that owns every out-of-band control path on your fleet.","attack_vector":"Unauthenticated HTTP/2 over TLS to bmcweb on the management interface. ALPN negotiates h2 by default, so no unusual client is needed.","remediation":"No fix as of disclosure; proposed Gerrit patches were posted to the OpenBMC mailing list. This is a config-only situation until upstream lands a fix and your ODM rebases - so ACL the BMC HTTPS port to management jump hosts only, and if your tooling can live on HTTP/1.1, disabling h2 in ALPN on the BMC removes the path. Assume no firmware you can buy today contains a fix. When one exists it arrives as a per-node out-of-band BMC firmware flash with the usual ODM lag.","references":["https://seclists.org/fulldisclosure/2026/May/24","https://binreaper.pages.dev/posts/2026-05-27-bmcweb-disclosure/"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2026-015-nvme-over-fabrics-discovery-cont","cve":null,"aliases":["NVMe-oF discovery controller unauthenticated by default","NVMe/TCP no auth by default","nvmet_host_allowed discovery bypass"],"title":"NVMe-over-Fabrics discovery controller - Linux kernel nvmet (drivers/nvme/target/discovery.c), NVMe/TCP and NVMe/RDMA","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVMe-over-Fabrics discovery controller - Linux kernel nvmet (drivers/nvme/target/discovery.c), NVMe/TCP and NVMe/RDMA","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"The NVMe-oF discovery controller answers any peer that can reach the port, by design. In Linux nvmet this is explicit in the source: nvmet_host_allowed() returns true unconditionally for the discovery subsystem, so every Get Log Page against the discovery controller is served pre-authentication over TCP, RDMA, or FC. The response enumerates every subsystem NQN, transport type, and address on that target. For a neocloud running disaggregated NVMe behind a tenant-reachable fabric, that hands an attacker the complete map of which customer's storage lives where and what NQN to spoof to reach it - the reconnaissance step that makes NQN spoofing practical. NVMe/TCP additionally ships with no authentication and no transport encryption unless in-band DH-HMAC-CHAP and TLS are explicitly configured, which is not the default in most deployments.","attack_vector":"The attacker points nvme discover at any reachable target IP on port 4420 (or the RDMA equivalent) with an arbitrary Host NQN and receives the full discovery log page. No credentials, no prior connection, no exploit. From there they connect to a named subsystem asserting a permitted Host NQN. The same unauthenticated surface is what makes memory-disclosure bugs in the discovery path (see CVE-2026-64320) remotely reachable pre-auth.","remediation":"Config change, no reboot, do it now: enable in-band DH-HMAC-CHAP on the target and require it for I/O subsystems; enable TLS for NVMe/TCP where the kernel version supports it; bind the discovery controller to a management interface unreachable from tenant networks rather than to the tenant storage fabric; and use per-subsystem allowed_hosts lists rather than the permissive default. All of this is nvmet configfs or SPDK RPC and applies at runtime, but initiators must be reconfigured in lockstep, so plan a rolling reattach per tenant. Firewall NVMe/TCP 4420 to known initiator addresses as an immediate stopgap.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2026/CVE-2026-64320.json","https://arxiv.org/abs/2202.08080","https://arxiv.org/abs/1903.09355"],"status":"curated","fleet":{"pain_class":"hot-patch"},"tags":["tenant-isolation"]},{"id":"NCVD-2026-038-nvme-over-fabrics-discovery-cont","cve":null,"aliases":["NVMe-oF discovery controller unauthenticated by default","NVMe/TCP no auth by default","nvmet_host_allowed discovery bypass"],"title":"NVMe-over-Fabrics discovery controller - Linux kernel nvmet (drivers/nvme/target/discovery.c), NVMe/TCP and NVMe/RDMA","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVMe-over-Fabrics discovery controller - Linux kernel nvmet (drivers/nvme/target/discovery.c), NVMe/TCP and NVMe/RDMA","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"The NVMe-oF discovery controller answers any peer that can reach the port, by design. In Linux nvmet this is explicit in the source: nvmet_host_allowed() returns true unconditionally for the discovery subsystem, so every Get Log Page against the discovery controller is served pre-authentication over TCP, RDMA, or FC. The response enumerates every subsystem NQN, transport type, and address on that target. For a neocloud running disaggregated NVMe behind a tenant-reachable fabric, that hands an attacker the complete map of which customer's storage lives where and what NQN to spoof to reach it - the reconnaissance step that makes NQN spoofing practical. NVMe/TCP additionally ships with no authentication and no transport encryption unless in-band DH-HMAC-CHAP and TLS are explicitly configured, which is not the default in most deployments.","attack_vector":"The attacker points nvme discover at any reachable target IP on port 4420 (or the RDMA equivalent) with an arbitrary Host NQN and receives the full discovery log page. No credentials, no prior connection, no exploit. From there they connect to a named subsystem asserting a permitted Host NQN. The same unauthenticated surface is what makes memory-disclosure bugs in the discovery path (see CVE-2026-64320) remotely reachable pre-auth.","remediation":"Config change, no reboot, do it now: enable in-band DH-HMAC-CHAP on the target and require it for I/O subsystems; enable TLS for NVMe/TCP where the kernel version supports it; bind the discovery controller to a management interface unreachable from tenant networks rather than to the tenant storage fabric; and use per-subsystem allowed_hosts lists rather than the permissive default. All of this is nvmet configfs or SPDK RPC and applies at runtime, but initiators must be reconfigured in lockstep, so plan a rolling reattach per tenant. Firewall NVMe/TCP 4420 to known initiator addresses as an immediate stopgap.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2026/CVE-2026-64320.json","https://arxiv.org/abs/2202.08080","https://arxiv.org/abs/1903.09355"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2020-13817","cve":"CVE-2020-13817","aliases":[],"title":"ntpd (transmit timestamp prediction): A remote attacker who can predict transmit timestamps can crash ntpd or, worse","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"ntpd (transmit timestamp prediction)","year":"2020","cvss_score":7.4,"severity":"high","kev":false,"impact":"A remote attacker who can predict transmit timestamps can crash ntpd or, worse, change the system time. Changing the time on a cluster node is a more interesting attack than crashing it: certificates become valid or invalid, log timestamps stop lining up with reality, and scheduled jobs fire at the wrong moment.","attack_vector":"Remote attacker able to predict the client's transmit timestamps for outgoing packets.","remediation":"Upgrade ntp to 4.2.8p14 or later, or move to chrony/NTS. Package upgrade and service restart. Monitor for step changes in system time as a detection control — most fleets do not alert on this and should.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-13817"],"status":"curated","fleet":{"pain_class":"daemon-restart"},"published":"2020-06-04"},{"id":"CVE-2021-1228","cve":"CVE-2021-1228","aliases":[],"title":"Cisco Nexus 9000 in ACI mode (fabric infrastructure VLAN): A device plugged into a normal front-panel port can talk its","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Cisco Nexus 9000 in ACI mode (fabric infrastructure VLAN)","year":"2021","cvss_score":7.4,"severity":"high","kev":false,"impact":"A device plugged into a normal front-panel port can talk its way onto the ACI infrastructure VLAN — the fabric's own control plane. From there an attacker sees and can influence the fabric's internal signalling rather than one tenant's EPG. In a multi-tenant ACI build this is the boundary that separates 'a tenant' from 'the fabric operator'.","attack_vector":"Unauthenticated, adjacent — physical or logical access to a leaf front-panel port. Any tenant with a bare-metal node, or anyone who can plug into a rack, is in position.","remediation":"ACI software upgrade across the APIC cluster and the leaf/spine switches — a staged fabric upgrade, not a single reload, and Cisco's recommended sequence takes hours on a large pod. Interim mitigation is strict port-level admission control and disabling unused ports, both live config changes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1228","https://nvd.nist.gov/vuln/detail/CVE-2019-1890"],"status":"curated","tags":["tenant-isolation"],"published":"2021-02-24"},{"id":"CVE-2021-3493","cve":"CVE-2021-3493","aliases":[],"title":"Linux kernel (overlayfs, Ubuntu patch): OverlayFS file-capability privilege escalation","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (overlayfs, Ubuntu patch)","year":"2021","cvss_score":7.4,"severity":"high","kev":true,"impact":"OverlayFS file-capability privilege escalation; trivially weaponised inside unprivileged user namespaces [KEV]","attack_vector":"Any tenant process in a container","remediation":"Ubuntu-specific patch delta. Livepatchable on Ubuntu Pro Livepatch; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2021-3493"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-04-17"},{"cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:H/I:H/A:N","cwe":["CWE-319"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2021-45104","cve":"CVE-2021-45104","aliases":["HTCONDOR-2022-0002"],"title":"HTCondor (daemon-to-daemon channel, negotiator/startd/schedd): Secret material crosses the network in the clear when","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HTCondor (daemon-to-daemon channel, negotiator/startd/schedd)","year":"2021","cvss_score":7.4,"severity":"high","kev":false,"impact":"Secret material crosses the network in the clear when weak encryption is configured or when any 8.8-or-older daemon is still in the pool. An attacker who captures it can take over another user's slot and run code as that user, and can impersonate the negotiator and startd well enough to make the schedd hand over other users' jobs.","attack_vector":"Passive capture of HTCondor traffic between daemons - so anyone on the cluster network path, including a tenant on a compute node with a promiscuous interface or a compromised switch.","remediation":"Upgrade to HTCondor 9.0.10 or 9.5.1, restart all daemons, and remove every pre-9.0 daemon from the pool - the mixed-version case is what reintroduces the cleartext path. Enable strong daemon-to-daemon encryption explicitly rather than relying on defaults.","references":["https://htcondor.org/security/vulnerabilities/HTCONDOR-2022-0002.html","https://nvd.nist.gov/vuln/detail/CVE-2021-45104"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2022-21656","cve":"CVE-2022-21656","aliases":[],"title":"Envoy: Type-confusion in default certificate validation","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy","year":"2022","cvss_score":7.4,"severity":"high","kev":false,"impact":"Type-confusion in default certificate validation","attack_vector":"Unauthenticated network","remediation":"Upgrade Envoy","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-21656"],"status":"curated","published":"2022-02-22"},{"id":"CVE-2022-21953","cve":"CVE-2022-21953","aliases":[],"title":"Rancher: Missing authorization allows an authenticated user to create a shell pod with kubectl access","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Rancher","year":"2022","cvss_score":7.4,"severity":"high","kev":false,"impact":"Missing authorization allows an authenticated user to create a shell pod with kubectl access in the local cluster; management-plane takeover","attack_vector":"Any authenticated Rancher user","remediation":"Upgrade Rancher","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-21953"],"status":"curated","published":"2023-02-07"},{"id":"CVE-2022-28734","cve":"CVE-2022-28734","aliases":[],"title":"GRUB2 (HTTP chunked transfer): Out-of-bounds write handling chunked HTTP responses during HTTP boot","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (HTTP chunked transfer)","year":"2022","cvss_score":7.4,"severity":"high","kev":false,"impact":"Out-of-bounds write handling chunked HTTP responses during HTTP boot. An attacker who can answer or MITM the HTTP boot request executes code in the bootloader on a node that has not booted an OS yet - the cleanest possible foothold on a bare-metal fleet.","attack_vector":"Network position on the provisioning path during HTTP boot: rogue DHCP, DNS spoofing, or a compromised provisioning server.","remediation":"grub2 package update + reboot, and replace the served netboot binary. If you HTTP-boot, move to HTTPS boot with a pinned CA and isolate the provisioning VLAN - the transport was the actual weakness here.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28734","https://access.redhat.com/security/cve/CVE-2022-28734"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2023-07-20"},{"id":"CVE-2022-31671","cve":"CVE-2022-31671","aliases":[],"title":"Harbor registry: P2P preheat execution logs readable/updatable by any authenticated user via job ID enumeration","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Harbor registry","year":"2022","cvss_score":7.4,"severity":"high","kev":false,"impact":"P2P preheat execution logs readable/updatable by any authenticated user via job ID enumeration","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; job logs often contain registry credentials","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31671"],"status":"curated","published":"2024-11-14"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:C/C:L/I:L/A:L","cwe":["CWE-22"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2022-35919","cve":"CVE-2022-35919","aliases":[],"title":"MinIO (admin server-update API): An authenticated request to the server-update admin API traverses out of the intended","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"MinIO (admin server-update API)","year":"2022","cvss_score":7.4,"severity":"high","kev":false,"impact":"An authenticated request to the server-update admin API traverses out of the intended directory, letting the caller read files on the MinIO host outside the object store - config, keys, whatever the service account can open. It also crosses into the host filesystem, so the blast radius is the node, not just a bucket.","attack_vector":"Any authenticated MinIO user that can reach the admin API port.","remediation":"Upgrade to RELEASE.2022-07-30T05-21-40Z or later and restart. Keep the admin API on a management-only listener that tenant workloads cannot route to, and run MinIO as an unprivileged user with a minimal filesystem view.","references":["https://github.com/minio/minio/security/advisories/GHSA-gr9v-6pcm-rqvg","https://nvd.nist.gov/vuln/detail/CVE-2022-35919"],"status":"curated"},{"id":"CVE-2022-47630","cve":"CVE-2022-47630","aliases":["TFV-10","TF-A X.509 parser out-of-bounds read"],"title":"Arm Trusted Firmware-A through v2.8, X.509 certificate parser used by Trusted Board Boot (get_ext, auth_nvctr)","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arm Trusted Firmware-A through v2.8, X.509 certificate parser used by Trusted Board Boot (get_ext, auth_nvctr)","year":"2022","cvss_score":7.4,"severity":"high","kev":false,"impact":"The code that validates the secure boot certificate chain reads out of bounds on a malformed certificate. So the component whose entire job is to decide whether firmware is trustworthy can be driven off the rails by the untrusted input it is inspecting - dangerous read side effects and leakage of microarchitectural state. It undermines the chain of trust at the exact point where a bare-metal operator is trying to prove to the next tenant that the box is clean.","attack_vector":"An attacker able to place a crafted certificate in the boot chain: control of the firmware image or the firmware-update path. On bare-metal GPU rental, a prior tenant with flash write access. Not remote.","remediation":"Upgrade TF-A past v2.8 with the TFV-10 fix and have the OEM re-issue the platform firmware. Flash + reboot + drain per node. There is no runtime mitigation - the parser runs before anything you control. Pair it with the operational control that actually helps: measure and attest boot firmware between tenants instead of trusting the parser.","references":["https://trustedfirmware-a.readthedocs.io/en/latest/security_advisories/security-advisory-tfv-10.html","https://nvd.nist.gov/vuln/detail/CVE-2022-47630"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2023-01-16"},{"id":"CVE-2023-0286","cve":"CVE-2023-0286","aliases":[],"title":"OpenSSL: X.400 address type confusion in X.509 GeneralName","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"OpenSSL","year":"2023","cvss_score":7.4,"severity":"high","kev":false,"impact":"X.400 address type confusion in X.509 GeneralName - crash or possible memory disclosure during cert verification","attack_vector":"Unauthenticated network","remediation":"Package update + service restart","references":["https://access.redhat.com/security/cve/CVE-2023-0286"],"status":"curated","fleet":{"pain_class":"daemon-restart"},"published":"2023-02-08"},{"id":"CVE-2023-1829","cve":"CVE-2023-1829","aliases":[],"title":"Linux kernel (net/sched tcindex): Use-after-free in the tcindex traffic-control filter - local root","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/sched tcindex)","year":"2023","cvss_score":7.4,"severity":"high","kev":false,"impact":"Use-after-free in the tcindex traffic-control filter - local root","attack_vector":"Any tenant process in a container with CAP_NET_ADMIN in a userns","remediation":"Livepatchable; otherwise drain + reboot. Blacklist `cls_tcindex`","references":["https://access.redhat.com/security/cve/CVE-2023-1829"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2023-04-12"},{"id":"CVE-2023-40548","cve":"CVE-2023-40548","aliases":["shim 15.8 batch"],"title":"shim (verify_sbat_section): Integer overflow leading to heap overflow while verifying the SBAT section on 32-bit","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"shim (verify_sbat_section)","year":"2023","cvss_score":7.4,"severity":"high","kev":false,"impact":"Integer overflow leading to heap overflow while verifying the SBAT section on 32-bit systems. The irony matters operationally: SBAT is the mechanism meant to revoke vulnerable bootloaders, and parsing it is itself exploitable.","attack_vector":"Crafted binary presented to shim on a 32-bit EFI implementation. Rare on server-class GPU hardware, common on 32-bit-EFI edge and embedded boxes.","remediation":"shim package update + reboot. Low priority on x86-64 server fleets, real on any 32-bit-EFI hardware you still operate.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-40548","https://www.openwall.com/lists/oss-security/2024/01/26/1"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-01-29"},{"id":"CVE-2024-46817","cve":"CVE-2024-46817","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.4,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Stop amdgpu_dm initialize when stream nums greater than 6","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46817","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-27"},{"id":"CVE-2025-20111","cve":"CVE-2025-20111","aliases":[],"title":"Cisco Nexus 3000/9000 (health monitoring diagnostics): The health monitoring diagnostics subsystem on Nexus 3000 and","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Cisco Nexus 3000/9000 (health monitoring diagnostics)","year":"2025","cvss_score":7.4,"severity":"high","kev":false,"impact":"The health monitoring diagnostics subsystem on Nexus 3000 and 9000 switches in standalone NX-OS mode can be driven into a denial of service by an unauthenticated attacker. The irony is worth noting for operators — the subsystem whose job is to tell you the switch is healthy is itself the thing that takes it down.","attack_vector":"Unauthenticated, remote — traffic reaching the affected diagnostics handling on the switch.","remediation":"NX-OS upgrade plus reload per switch, staged across MLAG/ECMP pairs. Restrict management-plane reachability with CoPP in the meantime — a live config change.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-20111"],"status":"curated","tags":["fabric-dos"],"published":"2025-02-26"},{"id":"CVE-2026-18556","cve":"CVE-2026-18556","aliases":[],"title":"N-able N-central: Authentication bypass using an alternate path or channel on the RMM server","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"N-able N-central","year":"2026","cvss_score":7.4,"severity":"high","kev":true,"impact":"Authentication bypass using an alternate path or channel on the RMM server","attack_vector":"Network (remote)","remediation":"Control-plane: patch; the RMM has agent reach into every managed host","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-18556"],"status":"curated","published":"2026-08-01"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:R/S:C/C:H/I:N/A:N","cwe":["CWE-22"],"fleet":{"pain_class":"hot-patch"},"id":"CVE-2026-24123","cve":"CVE-2026-24123","aliases":["GHSA-6r62-w2q3-48hf"],"title":"BentoML (bentofile.yaml path fields: description, docker.setup_script, docker.dockerfile_template","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"BentoML","year":"2026","cvss_score":7.4,"severity":"high","kev":false,"impact":"Several bentofile path fields accept traversal, so building an attacker's bento copies arbitrary files from the builder's filesystem into the bento archive. SSH keys, cloud credential files and environment secrets ride along into an artifact that is then pushed to a registry the attacker can read - a clean supply-chain exfiltration path off a shared build host.","attack_vector":"Anyone who can get a victim to run bentoml build on their bentofile. The exfiltration completes when the resulting bento is pushed or shared.","remediation":"Upgrade BentoML to 1.4.34 or later. Build untrusted bentos in an isolated container with no credential files present, and scan any bento built from an outside source before pushing it to a shared registry.","references":["https://github.com/bentoml/BentoML/security/advisories/GHSA-6r62-w2q3-48hf","https://nvd.nist.gov/vuln/detail/CVE-2026-24123"],"status":"curated"},{"id":"CVE-2026-47473","cve":"CVE-2026-47473","aliases":[],"title":"TensorRT-LLM: Memory corruption (improper array bounds checking)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT-LLM","year":"2026","cvss_score":7.4,"severity":"high","kev":false,"impact":"Memory corruption (improper array bounds checking)","attack_vector":"Any inference client","remediation":"Bump TensorRT-LLM; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47473","https://github.com/NVIDIA/product-security/tree/main/2026/5840"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-123"],"published":"2026-07-14"},{"cwe":["CWE-362","CWE-200"],"fleet":{"pain_class":"node-drain"},"id":"CVE-2026-68093","cve":"CVE-2026-68093","aliases":[],"title":"Linux kernel (arch/x86/kvm/svm): After a CPU offline/online cycle, KVM's ASID generation counter is reset in a way that","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (arch/x86/kvm/svm)","year":"2026","cvss_score":7.4,"severity":"high","kev":false,"impact":"After a CPU offline/online cycle, KVM's ASID generation counter is reset in a way that lets two vCPUs belonging to DIFFERENT VMs run on the same physical CPU with the same ASID. They then share nested-page-table TLB entries, so one tenant's guest resolves addresses through another tenant's translations. This is direct cross-tenant memory read and write exposure, with no attacker skill required once the condition exists.","attack_vector":"No tenant action needed to create the hazard - it is created by the operator's own CPU hotplug (maintenance, core parking, power management) on AMD nodes running kvm_amd. Any two co-resident guests pinned to the recycled pCPU can end up sharing an ASID; a tenant that can influence its own vCPU placement and probe memory will notice foreign data. AMD SVM only.","remediation":"Update to a kernel with the referenced stable commits. Interim, and this one is actionable today: stop doing CPU offline/online cycles on AMD hypervisor nodes carrying live guests - drain the node before any hotplug operation, and reboot rather than hot-cycling CPUs.","references":["https://git.kernel.org/stable/c/2028b81321dc757b6875b99c10d908e349c141e2","https://git.kernel.org/stable/c/60283726f2845bd78b95efbd0e50b93944780477","https://nvd.nist.gov/vuln/detail/CVE-2026-68093"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:R/S:C/C:N/I:H/A:N","cwe":["CWE-295","CWE-347"],"fleet":{"pain_class":"hot-patch"},"id":"NCVD-2026-056-sigstore-cosign-verify-blob-veri","cve":null,"aliases":["GHSA-fx35-mq7g-6g98"],"title":"Sigstore cosign (verify-blob / verify-blob-attestation, legacy JSON bundle): SUPPLY CHAIN, VERIFICATION BYPASS: keyless","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Sigstore cosign (verify-blob / verify-blob-attestation, legacy JSON bundle)","year":"2026","cvss_score":7.4,"severity":"high","kev":false,"impact":"SUPPLY CHAIN, VERIFICATION BYPASS: keyless identity checks silently stop applying. When cosign verify-blob reads the cert field of a legacy JSON bundle and it fails to parse as an X.509 certificate, cosign quietly falls back to treating the input as a raw public key and sets co.SigVerifier. Downstream, whenever co.SigVerifier is set, certificate chain validation and CheckCertificatePolicy are skipped — so --certificate-identity and --certificate-oidc-issuer are ignored and verification passes for any key the attacker chooses, as long as the signature is valid for it. Worse for anyone using explicit keys: the fallback assigns unconditionally, so a public key embedded in a legacy bundle overwrites a key the operator passed via --key. For a GPU cluster this sits on the path where model weights, datasets and release blobs get validated before being pulled onto nodes; the verification step returns success while checking nothing meaningful. OCI image verification (cosign verify) and the modern protobuf bundle format are not affected.","attack_vector":"Network / artifact delivery, unauthenticated from the attacker's side, with user interaction in that a victim runs verify-blob against an attacker-supplied legacy JSON bundle. Affects cosign v3 up to 3.1.2 and v2 up to 2.6.4.","remediation":"Upgrade cosign past 3.1.2 (v3) or 2.6.4 (v2) wherever it runs — CI runners, admission controllers, node bootstrap scripts, developer machines. Move off legacy JSON bundles to the standardized Sigstore bundle format (--new-bundle-format), which is unaffected and is the default in cosign v3. Re-verify any blob whose only assurance came from a legacy-bundle verify-blob run on an affected version.","references":["https://github.com/sigstore/cosign/security/advisories/GHSA-fx35-mq7g-6g98"],"status":"curated"},{"id":"CVE-2018-3615","cve":"CVE-2018-3615","aliases":["Foreshadow","L1TF-SGX"],"title":"Intel SGX (L1 terminal fault on enclave pages): Speculative execution lets code outside an enclave read the enclave's","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SGX (L1 terminal fault on enclave pages)","year":"2018","cvss_score":7.3,"severity":"high","kev":false,"impact":"Speculative execution lets code outside an enclave read the enclave's data out of the L1 data cache, breaking the confidentiality guarantee SGX exists to provide. On a confidential-compute offering this is the whole product failing: the host, or a co-resident tenant with the right sibling-thread placement, reads enclave secrets including sealing and attestation material.","attack_vector":"Local code execution on the same physical core as the enclave. In a cloud that includes a co-tenant on a sibling hyperthread, which is why the mitigation story involves disabling SMT.","remediation":"Microcode update plus OS/hypervisor mitigations. Intel microcode for this can be late-loaded at boot by the OS without waiting for an OEM BIOS release, so the practical path is: update the microcode package, reboot, and disable SMT (or enforce core scheduling) on nodes that run untrusted co-tenants. Also re-attest and re-provision any enclave secrets that existed on unpatched hardware - the microcode fix does not un-leak them.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-3615","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00161.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"tags":["tenant-isolation"],"published":"2018-08-14"},{"id":"CVE-2020-8595","cve":"CVE-2020-8595","aliases":[],"title":"Istio: Authentication Policy exact-path matching allows unauthorized access to HTTP paths","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Istio","year":"2020","cvss_score":7.3,"severity":"high","kev":false,"impact":"Authentication Policy exact-path matching allows unauthorized access to HTTP paths","attack_vector":"Unauthenticated network","remediation":"Rolling istiod upgrade plus sidecar restart; sidecar restart means restarting tenant pods","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-8595"],"status":"curated","published":"2020-02-12"},{"id":"CVE-2021-1074","cve":"CVE-2021-1074","aliases":[],"title":"NVIDIA GPU Display Driver for Windows, installer: A race between the installer's integrity check and execution lets","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver for Windows, installer","year":"2021","cvss_score":7.3,"severity":"high","kev":false,"impact":"A race between the installer's integrity check and execution lets an unprivileged local user swap in a malicious file and get it run by an administrator's install. The window is short, which is a mitigation and not a fix - driver rollouts are scheduled and repeatable, so an attacker sitting on the box knows exactly when to try.","attack_vector":"An unprivileged local user present on the host while an administrator runs the driver installer.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1074"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-04-21"},{"id":"CVE-2021-1075","cve":"CVE-2021-1075","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): Use of a dangling pointer in the escape handler, rated code-execution","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2021","cvss_score":7.3,"severity":"high","kev":false,"impact":"Use of a dangling pointer in the escape handler, rated code-execution capable. Guest-agnostic local kernel bug - unprivileged code to SYSTEM on the GPU host.","attack_vector":"Any local user with GPU device access on a Windows host.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1075"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-04-21"},{"id":"CVE-2021-1085","cve":"CVE-2021-1085","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): The guest can write to a shared memory location after the host has validated its","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":7.3,"severity":"high","kev":false,"impact":"The guest can write to a shared memory location after the host has validated its contents - a double-fetch that NVIDIA rates as escalation-capable. The tenant passes the host's checks with clean data, then substitutes their own. vGPU 12.x before 12.2, 11.x before 11.4, 8.x before 8.7.","attack_vector":"Any unprivileged user inside a guest VM able to race the host's validation window.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1085"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2021-04-29"},{"id":"CVE-2022-21797","cve":"CVE-2022-21797","aliases":[],"title":"joblib: Arbitrary code execution via `eval` on the `pre_dispatch` flag in `Parallel()`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"joblib","year":"2022","cvss_score":7.3,"severity":"high","kev":false,"impact":"Arbitrary code execution via `eval` on the `pre_dispatch` flag in `Parallel()`","attack_vector":"Tenant code passing attacker-influenced strings","remediation":"Upgrade joblib >= 1.2.0 in base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-21797"],"status":"curated","published":"2022-09-26"},{"id":"CVE-2022-23817","cve":"CVE-2022-23817","aliases":[],"title":"AMD Secure Processor Secure OS - memory buffer checking: A malicious trusted application can read and write the ASP","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor Secure OS - memory buffer checking","year":"2022","cvss_score":7.3,"severity":"high","kev":false,"impact":"A malicious trusted application can read and write the ASP Secure OS's kernel virtual address space because buffer bounds are not checked at the TA boundary. That is full privilege escalation inside the secure processor - the attacker is now the security engine, not a client of it.","attack_vector":"Local, requires the ability to load a malicious trusted application into the ASP (signing-key compromise or a legitimately signed but attacker-controlled TA).","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-23817","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2024-08-13"},{"id":"CVE-2022-31097","cve":"CVE-2022-31097","aliases":[],"title":"Grafana: Stored XSS via Unified Alerting","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Grafana","year":"2022","cvss_score":7.3,"severity":"high","kev":false,"impact":"Stored XSS via Unified Alerting -> privilege escalation to Grafana admin","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31097"],"status":"curated","published":"2022-07-15"},{"id":"CVE-2023-29057","cve":"CVE-2023-29057","aliases":["LEN-118321"],"title":"Lenovo XClarity Controller (XCC) - LDAP/AD authorization: When XCC is configured to authenticate against Active","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lenovo XClarity Controller (XCC) - LDAP/AD authorization","year":"2023","cvss_score":7.3,"severity":"high","kev":false,"impact":"When XCC is configured to authenticate against Active Directory, a user's local XCC account permissions silently override the permissions their directory group grants - and the local ones can be higher. The result is privilege escalation that your identity provider cannot see: you revoke someone's admin rights in AD, XCC keeps honouring the stale local grant, and they retain out-of-band control of the node. For an operator this is a offboarding and least-privilege failure more than an exploit, which makes it easy to miss - nothing looks broken, and the directory tells you the access is gone.","attack_vector":"A valid XCC user in a deployment where LDAP/AD is configured for authentication and authorisation and the user also has a local XCC account. No exploit code required - the misbehaviour is in the authorisation logic itself.","remediation":"Flash XCC to the version listed for your model in LEN-118321 - out-of-band, per-node, no host reboot and no drain. Alongside the flash, do the config work that actually closes the gap: enumerate local XCC accounts on every node and delete the ones that shadow directory identities, because patching the precedence logic does not remove local accounts that are already there.","references":["https://support.lenovo.com/us/en/product_security/LEN-118321","https://nvd.nist.gov/vuln/detail/CVE-2023-29057"],"status":"curated","published":"2023-04-28"},{"id":"CVE-2023-29164","cve":"CVE-2023-29164","aliases":["INTEL-SA-00990","CVE-2023-49144","CVE-2023-35123","INTEL-SA-01078","CVE-2025-20097"],"title":"BMC firmware for Intel Server Boards S2600WF / S2600ST / S2600BP before 02.01.0017 and M50CYP, and OpenBMC firmware","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"BMC firmware for Intel Server Boards S2600WF / S2600ST / S2600BP before 02.01.0017 and M50CYP, and OpenBMC firmware…","year":"2023","cvss_score":7.3,"severity":"high","kev":false,"impact":"Improper access control in the BMC firmware on Intel's own server boards, with companion out-of-bounds read and uncaught-exception issues in the OpenBMC builds Intel ships for current Eagle Stream and Birch Stream platforms. BMC compromise is node ownership below the OS: virtual media, power control, KVM, and the firmware update path into BIOS and ME. It survives host reimage by construction and carries into the next tenant, and mass BMC access is a hall-level power-off capability. The OpenBMC entries matter because many neocloud whitebox designs run Intel's OpenBMC tree rather than a commercial BMC stack, and those builds are updated far less often.","attack_vector":"Access to the BMC - network access on the out-of-band management path for the network-facing issues, privileged local access for the information-disclosure path, and the host KCS interface where it has not been disabled.","remediation":"BMC firmware update to 02.01.0017 or later on the S2600 boards, and to the fixed OpenBMC branch (egs-1.15-0 / bhs-0.27 or later) on the current platforms - from Intel or from the ODM that built the board (Quanta, Wiwynn, Supermicro, Inventec). BMC updates do not need a host reboot, so this is cheap to roll relative to BIOS work. If you run whitebox nodes on Intel's OpenBMC, you own the update cadence yourself - there is no OEM pushing it to you, and that is the real finding here. Keep the BMC network isolated, credentials unique per node, and KCS disabled.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-29164","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00990.html","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01078.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-31008","cve":"CVE-2023-31008","aliases":[],"title":"DGX H100 BMC (IPMI): Privesc","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX H100 BMC (IPMI)","year":"2023","cvss_score":7.3,"severity":"high","kev":false,"impact":"Privesc","attack_vector":"Network-adjacent IPMI client","remediation":"Flash BMC 23.08.18 out-of-band","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5473/5473.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:H/A:H","cwe":["CWE-20"],"published":"2023-09-20"},{"id":"CVE-2023-31016","cve":"CVE-2023-31016","aliases":[],"title":"GPU Display Driver (Windows): Arbitrary code exec via uncontrolled search path","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver (Windows)","year":"2023","cvss_score":7.3,"severity":"high","kev":false,"impact":"Arbitrary code exec via uncontrolled search path","attack_vector":"Local low-priv user","remediation":"Upgrade Oct-2023 driver branch","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5491/5491.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-427"],"published":"2023-11-02"},{"id":"CVE-2023-34239","cve":"CVE-2023-34239","aliases":[],"title":"Gradio: Lack of path filtering","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Gradio","year":"2023","cvss_score":7.3,"severity":"high","kev":false,"impact":"Lack of path filtering → arbitrary file read from the host","attack_vector":"Unauthenticated network to the demo","remediation":"Upgrade; a Gradio demo runs with the tenant's full container filesystem access","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-34239"],"status":"curated","published":"2023-06-08"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:L/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-53500","cve":"CVE-2023-53500","aliases":[],"title":"Linux kernel (net/xfrm): A qdisc that reuses skb","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2023","cvss_score":7.3,"severity":"high","kev":false,"impact":"A qdisc that reuses skb->cb during enqueue clobbers the state decode_session6 relies on, so transmitting IPv6 through an xfrm interface reads freed slab memory while deciding which policy and SA the packet belongs to. Memory-safety break in the path that classifies traffic for encryption.","attack_vector":"Requires an xfrm interface with a cb-clobbering qdisc such as sfb attached, then any IPv6 transmit through it - neighbour discovery traffic is enough. A container with CAP_NET_ADMIN in its own netns can attach the qdisc itself and then send; otherwise it depends on the node's qdisc configuration on the encrypted overlay device.","remediation":"Boot a kernel carrying the linked stable commits. Interim: do not attach sfb (or other cb-using qdiscs) to xfrm interfaces, and drop CAP_NET_ADMIN from tenant containers so they cannot attach one.","references":["https://git.kernel.org/stable/c/da4cbaa75ed088b6d70db77b9103a27e2359e243","https://git.kernel.org/stable/c/db0e50741f0387f388e9ec824ea7ae8456554d5b","https://nvd.nist.gov/vuln/detail/CVE-2023-53500"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2023-54098","cve":"CVE-2023-54098","aliases":[],"title":"Linux i915 GVT-g mediated GPU virtualisation (debugfs teardown): Companion to the vGPU debugfs cleanup bug: GVT-g","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux i915 GVT-g mediated GPU virtualisation (debugfs teardown)","year":"2023","cvss_score":7.3,"severity":"high","kev":false,"impact":"Companion to the vGPU debugfs cleanup bug: GVT-g destroys debugfs state without checking the DRM minor's debugfs root is still valid, crashing the host on teardown. Again on the vGPU lifecycle path, which is the tenant boundary on GVT-g deployments.","attack_vector":"Whoever can drive vGPU create/destroy on the host.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-54098","https://git.kernel.org/stable/c/ae9a61511736cc71a99f01e8b7b90f6fb6128ed8"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-12-24"},{"id":"CVE-2024-11622","cve":"CVE-2024-11622","aliases":[],"title":"HPE Insight Remote Support (XXE): Third XXE variant in Insight RS enabling remote information disclosure","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HPE Insight Remote Support (XXE)","year":"2024","cvss_score":7.3,"severity":"high","kev":false,"impact":"Third XXE variant in Insight RS enabling remote information disclosure.","attack_vector":"Unauthenticated network access.","remediation":"Patch per HPESBGN04731 - all three XXE issues ship in the same update.","references":["https://support.hpe.com/hpesc/public/docDisplay?docId=hpesbgn04731en_us&docLocale=en_US"],"status":"curated"},{"id":"CVE-2024-21960","cve":"CVE-2024-21960","aliases":[],"title":"AMD Optimizing CPU Libraries (AOCL) - installation directory permissions: AOCL installs with permissive directory","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD Optimizing CPU Libraries (AOCL) - installation directory permissions","year":"2024","cvss_score":7.3,"severity":"high","kev":false,"impact":"AOCL installs with permissive directory permissions, so a low-privileged user can replace library files that privileged processes later load - straightforward privilege escalation to arbitrary code execution. Worth flagging for AI operators specifically: AOCL (BLIS, libFLAME, AOCL-LibM) is exactly what gets installed on AMD nodes to accelerate the CPU side of an ML pipeline, so it is likely present on your hosts and likely loaded by jobs running as someone else.","attack_vector":"Local, low-privileged user who can write into the AOCL installation directory. If tenants share a node and AOCL lives somewhere world-writable, one tenant poisons the next tenant's math library.","remediation":"Update AOCL and correct the directory permissions - this is a filesystem ACL fix plus a package update, no reboot and no firmware. Audit the permissions on every math and ML library directory on shared nodes while you are there; the same mistake recurs across vendor-supplied HPC packages.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21960","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2025-05-13"},{"id":"CVE-2024-22415","cve":"CVE-2024-22415","aliases":[],"title":"jupyter-lsp: Unauthenticated file read/write through the LSP extension","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"jupyter-lsp","year":"2024","cvss_score":7.3,"severity":"high","kev":false,"impact":"Unauthenticated file read/write through the LSP extension","attack_vector":"Network attacker reaching JupyterLab","remediation":"Upgrade; the extension bypasses notebook auth","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-22415"],"status":"curated","published":"2024-01-18"},{"id":"CVE-2024-25621","cve":"CVE-2024-25621","aliases":[],"title":"containerd: Overly broad default permissions on containerd-managed directories","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2024","cvss_score":7.3,"severity":"high","kev":false,"impact":"Overly broad default permissions on containerd-managed directories","attack_vector":"Any local process or partially-escaped container on the node","remediation":"Rolling containerd upgrade with node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-25621"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2025-11-06"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:L/A:N","cwe":["CWE-200"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-26992","cve":"CVE-2024-26992","aliases":[],"title":"Linux kernel (arch/x86/kvm/vmx): With adaptive PEBS exposed, KVM never guaranteed that LBR MSRs held guest values","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (arch/x86/kvm/vmx)","year":"2024","cvss_score":7.3,"severity":"high","kev":false,"impact":"With adaptive PEBS exposed, KVM never guaranteed that LBR MSRs held guest values before VM entry, so a guest can request LBR entries in its PEBS records and read back host RIPs - host kernel addresses and host execution paths - straight into the VM. Adaptive records also carry memory-access info that can sidestep the host's PMU event filter, and KVM generated adaptive records even when the guest asked for basic ones.","attack_vector":"Guest-driven and unprivileged inside the VM: the tenant programs PEBS and LBR MSRs from its own vCPU. Requires an Intel host with vPMU enabled and PEBS enumerated in the guest's CPUID - the case whenever performance counters are handed to tenants.","remediation":"Update to a stable kernel with the fix (adaptive PEBS support is dropped; no fixed release is enumerated, take the branch with commit 7a7650b3ac23). Interim control: do not expose the vPMU to tenant guests - run with kvm.enable_pmu=0 or strip PEBS/LBR from guest CPUID.","references":["https://git.kernel.org/stable/c/7a7650b3ac23e5fc8c990f00e94f787dc84e3175","https://git.kernel.org/stable/c/037e48ceccf163899374b601afb6ae8d0bf1d2ac","https://nvd.nist.gov/vuln/detail/CVE-2024-26992"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-36292","cve":"CVE-2024-36292","aliases":[],"title":"Intel Data Center GPU Flex Series - Windows driver: Improper buffer restrictions in the Flex Series Windows driver let","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel Data Center GPU Flex Series - Windows driver","year":"2024","cvss_score":7.3,"severity":"high","kev":false,"impact":"Improper buffer restrictions in the Flex Series Windows driver let an authenticated user knock the GPU out. Flex Series is the media/VDI-oriented datacenter part, so the affected hosts are typically multi-session - meaning any logged-in user, not just an administrator.","attack_vector":"Local, authenticated, on a Windows host running the Flex Series driver before 31.0.101.4314.","remediation":"Update the Flex Series Windows driver to 31.0.101.4314 or later. Cost: Windows driver replacement means a node reboot; drain sessions first.","references":["https://www.intel.com/content/www/us/en/security-center/default.html","https://nvd.nist.gov/vuln/detail/CVE-2024-36292"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-05-13"},{"id":"CVE-2024-36339","cve":"CVE-2024-36339","aliases":[],"title":"AMD Optimizing CPU Libraries (AOCL) - DLL hijacking: A DLL search-order hijack in AOCL lets an attacker get","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD Optimizing CPU Libraries (AOCL) - DLL hijacking","year":"2024","cvss_score":7.3,"severity":"high","kev":false,"impact":"A DLL search-order hijack in AOCL lets an attacker get their library loaded by a privileged process, reaching arbitrary code execution. Same practical exposure as the AOCL permissions issue: the CPU math libraries underneath your AMD-node ML stack become a privilege-escalation vector.","attack_vector":"Local, requires the ability to place a library where the search order will find it first. Windows-oriented, though the search-path pattern has Linux analogues via LD_LIBRARY_PATH on badly configured hosts.","remediation":"Update AOCL. Check that no shared-node job can influence library search paths for privileged processes - on Linux that means auditing LD_LIBRARY_PATH handling in your job launcher and any setuid tooling. Package update, no reboot, no firmware.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36339","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2025-05-13"},{"id":"CVE-2024-37130","cve":"CVE-2024-37130","aliases":[],"title":"Dell OpenManage Server Administrator (XSL hijacking local privilege escalation): A local low-privileged user hijacks","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Dell OpenManage Server Administrator (XSL hijacking local privilege escalation)","year":"2024","cvss_score":7.3,"severity":"high","kev":false,"impact":"A local low-privileged user hijacks an XSL load path and escalates to admin, gaining full control of the machine.","attack_vector":"Local low-privilege user on a host running OMSA 11.0.1.0 or earlier.","remediation":"Upgrade OMSA past 11.0.1.0. Agent update on every host that runs it.","references":["https://www.dell.com/support/kbdoc/en-us/000225914/dsa-2024-264-dell-openmanage-server-administrator-omsa-security-update-for-local-privilege-escalation-via-xsl-hijacking-vulnerability"],"status":"curated"},{"id":"CVE-2024-38552","cve":"CVE-2024-38552","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.3,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Fix potential index out of bounds in color transformation function","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-38552","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-06-19"},{"id":"CVE-2024-44933","cve":"CVE-2024-44933","aliases":[],"title":"Linux bnxt_en driver (bnxt_fill_hw_rss_tbl): Memory out-of-bounds in the RSS indirection-table path of the Broadcom NIC","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux bnxt_en driver (bnxt_fill_hw_rss_tbl)","year":"2024","cvss_score":7.3,"severity":"high","kev":false,"impact":"Memory out-of-bounds in the RSS indirection-table path of the Broadcom NIC driver, giving kernel panic, denial of service, or an information leak out of kernel memory. The information-leak arm is the one that matters for multi-tenancy — kernel memory on a shared host contains other workloads' data.","attack_vector":"Reachable through ring-reservation and RSS configuration paths in the driver; local host privilege or specific traffic conditions depending on the trigger.","remediation":"Kernel/driver upgrade and host reboot. Rolling across the fleet during normal maintenance; no firmware flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-44933"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-08-26"},{"id":"CVE-2024-45333","cve":"CVE-2024-45333","aliases":[],"title":"Intel Data Center GPU Flex Series - Windows driver: A further improper access control in the Flex Series Windows driver","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel Data Center GPU Flex Series - Windows driver","year":"2024","cvss_score":7.3,"severity":"high","kev":false,"impact":"A further improper access control in the Flex Series Windows driver allowing an authenticated local user to cause denial of service.","attack_vector":"Local, authenticated, on a Windows host running the driver before 31.0.101.4314.","remediation":"Update to 31.0.101.4314 or later - the same release that fixes CVE-2024-36292, so patch both in one window. Cost: node reboot.","references":["https://www.intel.com/content/www/us/en/security-center/default.html","https://nvd.nist.gov/vuln/detail/CVE-2024-45333"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-05-13"},{"id":"CVE-2024-46811","cve":"CVE-2024-46811","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.3,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Fix index may exceed array range within fpu_update_bw_bounding_box","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46811","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-27"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:L","cwe":["CWE-200","CWE-908"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-50096","cve":"CVE-2024-50096","aliases":[],"title":"Linux kernel (drivers/gpu/drm/nouveau): When the device-to-host copy behind a page fault silently fails, the fault","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/nouveau)","year":"2024","cvss_score":7.3,"severity":"high","kev":false,"impact":"When the device-to-host copy behind a page fault silently fails, the fault handler still hands the process a HIGH_USER page that was never written. The tenant reads whatever was previously in that page - residual data from other workloads on the node. This is the cross-tenant information-disclosure case, and the upstream fix calls it a security vulnerability in those words.","attack_vector":"Tenant holding /dev/dri/renderD* on nouveau using SVM/HMM device memory: fault a migrated page back to host RAM while the copy engine fails (a hung or erroring GPU, itself tenant-inducible). No capabilities required. nouveau device-memory (dmem) path only.","remediation":"Update to a kernel carrying the fix (stable commits below; no fixed_in published). Interim: disable SVM/HMM device memory on nouveau nodes, or do not share nouveau GPUs across tenants.","references":["https://git.kernel.org/stable/c/fd9bb7e996bab9b9049fffe3f3d3b50dee191d27","https://git.kernel.org/stable/c/73f75d2b5aee5a735cf64b8ab4543d5c20dbbdd9","https://nvd.nist.gov/vuln/detail/CVE-2024-50096"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-50117","cve":"CVE-2024-50117","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amd): A NULL pointer dereference in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amd)","year":"2024","cvss_score":7.3,"severity":"high","kev":false,"impact":"A NULL pointer dereference in the amdgpu firmware, ACPI and IP-block initialisation. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd: Guard against bad data for ATIF ACPI method","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-50117","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-11-05"},{"id":"CVE-2024-53104","cve":"CVE-2024-53104","aliases":[],"title":"Linux kernel (uvcvideo): Out-of-bounds write parsing UVC_VS_UNDEFINED frames - exploited in the wild","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (uvcvideo)","year":"2024","cvss_score":7.3,"severity":"high","kev":true,"impact":"Out-of-bounds write parsing UVC_VS_UNDEFINED frames - exploited in the wild [KEV]","attack_vector":"Local user with USB/device access","remediation":"Livepatchable; otherwise drain + reboot. Near-zero exposure on headless GPU servers - deprioritise behind the netfilter set, but it will still show on every compliance scan","references":["https://access.redhat.com/security/cve/CVE-2024-53104"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-12-02"},{"id":"CVE-2024-53108","cve":"CVE-2024-53108","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.3,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Adjust VSDB parser for replay feature","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53108","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-12-02"},{"id":"CVE-2024-53674","cve":"CVE-2024-53674","aliases":[],"title":"HPE Insight Remote Support (XXE): XML external entity injection allowing remote users to disclose information","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HPE Insight Remote Support (XXE)","year":"2024","cvss_score":7.3,"severity":"high","kev":false,"impact":"XML external entity injection allowing remote users to disclose information from the IRS server.","attack_vector":"Unauthenticated network access.","remediation":"Patch per HPESBGN04731.","references":["https://support.hpe.com/hpesc/public/docDisplay?docId=hpesbgn04731en_us&docLocale=en_US"],"status":"curated"},{"id":"CVE-2024-53675","cve":"CVE-2024-53675","aliases":[],"title":"HPE Insight Remote Support (XXE): Second XXE path in Insight RS enabling information disclosure","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HPE Insight Remote Support (XXE)","year":"2024","cvss_score":7.3,"severity":"high","kev":false,"impact":"Second XXE path in Insight RS enabling information disclosure.","attack_vector":"Unauthenticated network access.","remediation":"Patch per HPESBGN04731.","references":["https://support.hpe.com/hpesc/public/docDisplay?docId=hpesbgn04731en_us&docLocale=en_US"],"status":"curated"},{"id":"CVE-2024-57804","cve":"CVE-2024-57804","aliases":["CVE-2024-57807"],"title":"Linux kernel mpi3mr driver (Broadcom tri-mode 9600-series HBA/RAID) and megaraid_sas driver: Rapidly toggling PHY","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel mpi3mr driver (Broadcom tri-mode 9600-series HBA/RAID) and megaraid_sas driver","year":"2024","cvss_score":7.3,"severity":"high","kev":false,"impact":"Rapidly toggling PHY enable/disable through the SAS transport sysfs interface corrupts the controller's persistent and current configuration pages on Broadcom tri-mode 9600-series adapters. The corruption is in configuration state the controller keeps, not just in kernel memory - so a host-side operation writes bad persistent state into the storage controller, which is precisely the boundary a tenant handoff is supposed to reset. The megaraid_sas sibling is a lock-ordering deadlock between reset and scan paths that hangs the SCSI host. Both are availability and integrity problems at the controller layer: a node whose controller config pages are corrupt may enumerate drives differently or fail to bring arrays up after the next reboot, and that failure surfaces on the next tenant, not the one who caused it.","attack_vector":"Local root on the bare-metal host, through the SAS transport sysfs PHY controls. On a bare-metal GPU rental this is the tenant themselves - they legitimately have root and the sysfs interface is not namespaced.","remediation":"Kernel/driver update on the host and a reboot. The deeper point for operators is that this is a class you cannot patch away: a bare-metal tenant with root has a large, mostly unaudited surface of sysfs and ioctl interfaces that write persistent state into the storage controller. Between tenants, do not just reimage - re-read and, if your controller tooling supports it, restore the controller configuration to a known-good baseline (storcli/StorCLI config restore or the equivalent), and verify controller firmware version and config page integrity as part of the handoff checklist.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git","https://nvd.nist.gov/vuln/detail/CVE-2024-57804","https://nvd.nist.gov/vuln/detail/CVE-2024-57807"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-01-11"},{"id":"CVE-2025-10164","cve":"CVE-2025-10164","aliases":[],"title":"SGLang (`/update_weights_from_tensor`): Unsafe deserialization of the `serialized_named_tensors` argument","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"SGLang (`/update_weights_from_tensor`)","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"Unsafe deserialization of the `serialized_named_tensors` argument","attack_vector":"Network to the SGLang HTTP API","remediation":"Upgrade past 0.4.6; the weight-update endpoint must not be tenant-reachable","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-10164"],"status":"curated","fleet":{"ubiquity":"Very common - SGLang is the main vLLM alternative for high-throughput LLM serving on neoclouds; affects 0.4.6 through <0.5.4","remediation_pain":"`daemon-restart` - upgrade to 0.5.4+ and roll every serving replica","pain_class":"daemon-restart","why_fleet_wide":"`update_weights_from_tensor` pickle-deserializes attacker input with no authentication, giving unauthenticated remote code execution on every SGLang serving process reachable on the network"},"published":"2025-09-09"},{"id":"CVE-2025-22830","cve":"CVE-2025-22830","aliases":["AMI-SA-2025006"],"title":"AMI AptioV UEFI BIOS: A race condition in the BIOS that a skilled local attacker can drive to resource exhaustion","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI AptioV UEFI BIOS","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"A race condition in the BIOS that a skilled local attacker can drive to resource exhaustion, with AMI rating full confidentiality, integrity and availability impact and a subsequent-system impact as well. The operator-facing outcome is a node that fails to complete firmware operations or wedges in early boot - and given the CIA rating AMI assigns, a successfully-won race is a path to firmware-level compromise rather than just a hang. Winning a race in firmware is exactly the kind of bug that is unreliable in a lab and reliable at fleet scale, where an attacker gets thousands of boot attempts.","attack_vector":"Local access with high privileges and some user interaction, and the attack requires specific conditions to be present - AMI's CVSS v4 vector marks attack requirements as present, meaning the attacker needs the machine in a particular state. Realistically: root on the host plus the ability to trigger a reboot or a firmware operation, which any tenant of a bare-metal node has.","remediation":"BIOS update to AptioV_5.040 or later - firmware flash plus a host reboot per node, gated on your server vendor rebasing the AMI BKC for your SKU. No config-only workaround. Because the trigger involves reboots and firmware operations, one thing you can do without patching is restrict who can initiate BIOS updates and firmware operations out-of-band, and log every one of them.","references":["https://go.ami.com/hubfs/Security%20Advisories/2025/AMI-SA-2025006.pdf","https://nvd.nist.gov/vuln/detail/CVE-2025-22830"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2025-08-12"},{"id":"CVE-2025-22833","cve":"CVE-2025-22833","aliases":[],"title":"AMI AptioV BIOS (unchecked buffer copy): Buffer copy without size checking in firmware leading to arbitrary code","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI AptioV BIOS (unchecked buffer copy)","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"Buffer copy without size checking in firmware leading to arbitrary code execution.","attack_vector":"Local low-privilege access with user interaction.","remediation":"AMI ships the fix to OEMs, not to you - obtain the updated BIOS from your board/server vendor (Supermicro, Gigabyte, ASRock Rack, Quanta, Tyan etc.) and flash it. Expect a lag of weeks to months between the AMI advisory and an OEM image for your exact SKU, and expect some SKUs never to get one. Cold reboot per node.","references":["https://go.ami.com/hubfs/Security%20Advisories/2025/AMI-SA-2025008.pdf"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-23241","cve":"CVE-2025-23241","aliases":[],"title":"Intel ice driver (Ethernet 800 Series, Linux kernel mode): Kernel-mode flaw in the 800-series Ethernet Linux driver (an","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel ice driver (Ethernet 800 Series, Linux kernel mode)","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"Kernel-mode flaw in the 800-series Ethernet Linux driver (an integer overflow) reachable by an authenticated user for privilege escalation. Part of the same 2025 batch - patch them as one unit rather than individually.","attack_vector":"Authenticated local user; tenants on nodes exposing VFs or RDMA devices.","remediation":"Fixed in the Intel out-of-tree ice driver (or the equivalent in-kernel version). Updating the driver requires unloading and reloading the module, which drops every link on that NIC - on a node whose RDMA fabric carries collective traffic, that is a job-killing event, so drain first. If you take it via a distro kernel update instead, it is a reboot. No firmware flash for the driver-side fixes. Target ice 1.17.2 or later.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23241","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01296.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-08-12"},{"id":"CVE-2025-23242","cve":"CVE-2025-23242","aliases":[],"title":"NVIDIA Riva: Unauthorized access to the speech service (insufficient access control)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Riva","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"Unauthorized access to the speech service (insufficient access control)","attack_vector":"Network client of the Riva endpoint","remediation":"Upgrade Riva NIM containers; redeploy; put auth in front of the endpoint","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23242","https://github.com/NVIDIA/product-security/tree/main/2025/5625"],"status":"curated","fleet":{"ubiquity":"Common - shipped as a NIM-style microservice, deployed by neoclouds offering managed speech endpoints","remediation_pain":"`daemon-restart` (upgrade to Riva 2.19.0 and redeploy the service)","pain_class":"daemon-restart","why_fleet_wide":"Improper access control in the service auth layer, network-reachable with no user interaction: an unauthenticated caller escalates privileges into the hosting cloud environment and can read other tenants' data"},"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:L/I:L/A:L","cwe":["CWE-284"],"published":"2025-03-11"},{"id":"CVE-2025-23257","cve":"CVE-2025-23257","aliases":[],"title":"NVIDIA DOCA: Local privesc via insecure file permissions","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DOCA","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"Local privesc via insecure file permissions","attack_vector":"Local user on the DPU/host","remediation":"Upgrade DOCA packages; restart DPU services","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23257","https://github.com/NVIDIA/product-security/tree/main/2025/5655"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-732"],"published":"2025-09-04"},{"id":"CVE-2025-23258","cve":"CVE-2025-23258","aliases":[],"title":"NVIDIA DOCA: Local privesc via insecure file permissions","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DOCA","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"Local privesc via insecure file permissions","attack_vector":"Local user on the DPU/host","remediation":"Upgrade DOCA packages","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23258","https://github.com/NVIDIA/product-security/tree/main/2025/5655"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-732"],"published":"2025-09-04"},{"id":"CVE-2025-23277","cve":"CVE-2025-23277","aliases":[],"title":"GPU Display Driver: Access-control bypass","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"Access-control bypass -> privesc","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23277","https://github.com/NVIDIA/product-security/tree/main/2025/5670"],"status":"curated","fleet":{"ubiquity":"Universal - Linux and Windows driver","remediation_pain":"`node-reboot`","pain_class":"node-reboot","why_fleet_wide":"Out-of-bounds memory access in kernel mode from a local tenant: cheap denial of service against a whole GPU node"},"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-284"],"published":"2025-08-02"},{"id":"CVE-2025-23344","cve":"CVE-2025-23344","aliases":[],"title":"NVIDIA NVDebug tool: NVDebug allows an actor to run code on the platform host as a non-privileged user, reaching code","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NVDebug tool","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"NVDebug allows an actor to run code on the platform host as a non-privileged user, reaching code execution and privilege escalation.","attack_vector":"Local, low privileges, user interaction. An unprivileged account on the platform host plus an operator running the tool.","remediation":"Update NVDebug per bulletin 5696. Cost: trivial tool replacement, no drain.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23344","https://github.com/NVIDIA/product-security/tree/main/2025/5696"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-78"],"fleet":{"pain_class":"hot-patch"},"published":"2025-09-09"},{"id":"CVE-2025-29951","cve":"CVE-2025-29951","aliases":[],"title":"AMD Secure Processor (ASP) bootloader - buffer overflow: A buffer overflow in the ASP bootloader gives an attacker a","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor (ASP) bootloader - buffer overflow","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"A buffer overflow in the ASP bootloader gives an attacker a memory overwrite in the secure processor's boot path, ending in privilege escalation and arbitrary code execution below the OS. Anything the ASP protects on that node - SEV keys, fTPM, boot measurement - is then attacker-controlled, and the foothold persists across OS reinstalls.","attack_vector":"Local. Needs the ability to influence what the bootloader parses, i.e. SPI ROM write access or a subverted firmware update path.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. Lock down SPI ROM writes and platform BIOS update authentication as the interim control; there is no OS-level mitigation for bootloader code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-29951","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2026-02-10"},{"id":"CVE-2025-30167","cve":"CVE-2025-30167","aliases":[],"title":"Jupyter Core (Windows): Config read from a shared writable path","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Jupyter Core (Windows)","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"Config read from a shared writable path → code execution as another user","attack_vector":"Co-tenant local user","remediation":"Upgrade to 5.8.0+","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-30167"],"status":"curated","published":"2025-06-03"},{"id":"CVE-2025-31133","cve":"CVE-2025-31133","aliases":[],"title":"runc: Insufficient verification of masked-path bind mounts (/dev/null replaced by symlink) enables container","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"runc","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"Insufficient verification of masked-path bind mounts (/dev/null replaced by symlink) enables container escape to host root","attack_vector":"Any tenant workload / malicious image with a custom mount config","remediation":"Replace runc on all nodes; running containers stay vulnerable so drain required","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-31133"],"status":"curated","fleet":{"ubiquity":"Universal - all runc ≤1.2.7 / 1.3.2 / 1.4.0-rc.2","remediation_pain":"`node-drain` - binary replace plus recreation of every container; running containers are not retroactively protected","pain_class":"node-drain","why_fleet_wide":"Replacing `/dev/null` in the container with a symlink into host `/proc` makes critical host procfs entries mount writable, yielding container escape to host root from a tenant-controlled image"},"published":"2025-11-06"},{"id":"CVE-2025-33181","cve":"CVE-2025-33181","aliases":[],"title":"Cumulus Linux / NVOS: Command injection (local)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Cumulus Linux / NVOS","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"Command injection (local)","attack_vector":"Local switch operator","remediation":"Upgrade Cumulus Linux / NVOS","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33181","https://github.com/NVIDIA/product-security/tree/main/2026/5722"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-77"],"published":"2026-02-24"},{"id":"CVE-2025-33205","cve":"CVE-2025-33205","aliases":[],"title":"NVIDIA NeMo Framework: A predefined variable pulls in functionality from an untrusted control sphere, reaching code","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NeMo Framework","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"A predefined variable pulls in functionality from an untrusted control sphere, reaching code execution. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5729 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33205","https://github.com/NVIDIA/product-security/tree/main/2025/5729"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-829"],"published":"2025-11-25"},{"id":"CVE-2025-33212","cve":"CVE-2025-33212","aliases":[],"title":"NVIDIA NeMo Framework: Loading a maliciously crafted model file bypasses the framework's control mechanisms and reaches","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NeMo Framework","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"Loading a maliciously crafted model file bypasses the framework's control mechanisms and reaches code execution. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5736 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33212","https://github.com/NVIDIA/product-security/tree/main/2025/5736"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2025-12-16"},{"id":"CVE-2025-33228","cve":"CVE-2025-33228","aliases":[],"title":"CUDA Toolkit: Local privesc via command injection","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"Local privesc via command injection","attack_vector":"Local user running toolkit binaries","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33228","https://github.com/NVIDIA/product-security/tree/main/2026/5755"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-78"],"published":"2026-01-20"},{"id":"CVE-2025-33229","cve":"CVE-2025-33229","aliases":[],"title":"CUDA Toolkit: Code exec via untrusted library load","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"Code exec via untrusted library load","attack_vector":"Local user / malicious image layer","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33229","https://github.com/NVIDIA/product-security/tree/main/2026/5755"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-427"],"published":"2026-01-20"},{"id":"CVE-2025-33230","cve":"CVE-2025-33230","aliases":[],"title":"CUDA Toolkit: Local privesc via command injection","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"Local privesc via command injection","attack_vector":"Local user","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33230","https://github.com/NVIDIA/product-security/tree/main/2026/5755"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-78"],"published":"2026-01-20"},{"id":"CVE-2025-38508","cve":"CVE-2025-38508","aliases":[],"title":"Linux x86/sev - Secure TSC frequency calculation (TSC_FACTOR): Secure TSC is how an SEV-SNP guest gets a timebase","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux x86/sev - Secure TSC frequency calculation (TSC_FACTOR)","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"Secure TSC is how an SEV-SNP guest gets a timebase it can trust rather than one the host can manipulate. The guest's frequency calculation ignored TSC_FACTOR and so drifted from the real mean TSC frequency. A confidential guest whose clock is wrong makes wrong decisions about timeouts, certificate validity and rate limiting - and clock manipulation is a known lever for defeating the very side-channel defences that confidential computing depends on.","attack_vector":"Affects SEV-SNP guests using Secure TSC. Not an active attack so much as a broken trusted-time guarantee that a host-side adversary can lean on.","remediation":"Fixed in the **guest** kernel, not the host - the hardening lives in the SEV-ES/SNP guest's #VC handler and interrupt entry code. That inverts the usual rollout: you can patch every hypervisor you own and still be exposed, because the protection has to be in the tenant's own VM image. As an operator your job is to ship updated confidential-guest images (or tell tenants which minimum kernel to run) and, where you can, enforce it as an admission requirement. Each guest picks the fix up on its next boot; no host reboot, no firmware update. The fix is in the guest's x86/sev code, so confidential-VM images need updating; host patching alone does not deliver it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38508"],"status":"curated","published":"2025-08-16"},{"id":"CVE-2025-40334","cve":"CVE-2025-40334","aliases":[],"title":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu): Missing or insufficient validation of","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu)","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu user-mode queues (doorbell submission path). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdgpu: validate userq buffer virtual address and size","attack_vector":"Local. Reachable by any process with a render node open that can create user-mode queues - the normal ROCm submission path, reachable from an unprivileged container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-40334","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-12-09"},{"id":"CVE-2025-52881","cve":"CVE-2025-52881","aliases":[],"title":"runc: Attacker misdirects runc writes to /proc via racing symlinks","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"runc","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"Attacker misdirects runc writes to /proc via racing symlinks; can defeat LSM labelling and escape","attack_vector":"Any tenant workload","remediation":"Replace runc on all nodes; drain required","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-52881"],"status":"curated","fleet":{"ubiquity":"Universal - same runc version range","remediation_pain":"`node-drain` - runc upgrade + container recreation across the whole fleet","pain_class":"node-drain","why_fleet_wide":"LSM (AppArmor/SELinux) bypass that makes arbitrary procfs writes easy, turning the other two into reliable host root; AWS, Alibaba and every distro shipped emergency runc rebuilds"},"published":"2025-11-06"},{"id":"CVE-2025-54386","cve":"CVE-2025-54386","aliases":[],"title":"Traefik: Path traversal in the WASM plugin installation mechanism","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Traefik","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"Path traversal in the WASM plugin installation mechanism","attack_vector":"Anyone who can supply a Traefik plugin","remediation":"Rolling Traefik upgrade; restrict plugin sources","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-54386"],"status":"curated","published":"2025-08-02"},{"id":"CVE-2025-7647","cve":"CVE-2025-7647","aliases":[],"title":"llama-index-core: Predictable hardcoded cache directory","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"llama-index-core","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"Predictable hardcoded cache directory → local hijack","attack_vector":"Co-tenant local user on a shared node","remediation":"Upgrade past 0.12.44; matters on shared bare-metal or shared `/tmp`","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-7647"],"status":"curated","published":"2025-09-27"},{"id":"CVE-2025-9905","cve":"CVE-2025-9905","aliases":[],"title":"Keras (HDF5 path): Code execution from crafted `.h5`/`.hdf5` model despite safe mode","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Keras (HDF5 path)","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"Code execution from crafted `.h5`/`.hdf5` model despite safe mode","attack_vector":"Customer-supplied legacy HDF5 model","remediation":"Block `.h5` ingest entirely; the legacy format has no safe loader","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-9905"],"status":"curated","published":"2025-09-19"},{"id":"CVE-2025-9906","cve":"CVE-2025-9906","aliases":[],"title":"Keras: Code execution from crafted `.keras` archive despite safe mode","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Keras","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"Code execution from crafted `.keras` archive despite safe mode","attack_vector":"Customer-supplied model file","remediation":"Upgrade; same class as above","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-9906"],"status":"curated","published":"2025-09-19"},{"id":"CVE-2026-13201","cve":"CVE-2026-13201","aliases":[],"title":"KubeVirt: safepath OpenAtNoFollow resolves via /proc/self/fd, defeating the symlink protection","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"KubeVirt","year":"2026","cvss_score":7.3,"severity":"high","kev":false,"impact":"safepath OpenAtNoFollow resolves via /proc/self/fd, defeating the symlink protection","attack_vector":"Cluster user with virt-launcher access","remediation":"Upgrade KubeVirt","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-13201"],"status":"curated","published":"2026-06-24"},{"id":"CVE-2026-15793","cve":"CVE-2026-15793","aliases":[],"title":"BuildKit: git.checkoutbundle=true on a malicious git source yields crafted command invocation on the build host","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"BuildKit","year":"2026","cvss_score":7.3,"severity":"high","kev":false,"impact":"git.checkoutbundle=true on a malicious git source yields crafted command invocation on the build host","attack_vector":"Anyone using the raw low-level build API","remediation":"Upgrade BuildKit; block raw LLB API for tenants","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-15793"],"status":"curated","published":"2026-07-21"},{"id":"CVE-2026-2033","cve":"CVE-2026-2033","aliases":[],"title":"MLflow (artifact handler): Directory traversal","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"MLflow (artifact handler)","year":"2026","cvss_score":7.3,"severity":"high","kev":false,"impact":"Directory traversal → RCE via the artifact handler","attack_vector":"Remote attacker with artifact-write access","remediation":"Upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-2033"],"status":"curated","published":"2026-02-20"},{"cvss_vector":"CVSS:4.0/AV:N/AC:H/AT:N/PR:L/UI:A/VC:H/VI:H/VA:H/SC:N/SI:N/SA:N","cwe":["CWE-79"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2026-23960","cve":"CVE-2026-23960","aliases":["GHSA-cv78-6m8q-ph82"],"title":"Argo Workflows (Argo Server, artifact directory listing renderer): Object names are printed into the artifact directory","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo Workflows (Argo Server, artifact directory listing renderer)","year":"2026","cvss_score":7.3,"severity":"high","kev":false,"impact":"Object names are printed into the artifact directory listing HTML without escaping, and those names come from whatever files a workflow wrote. Any tenant who can run a workflow plants stored JavaScript that executes under the Argo Server origin in another user's browser, then drives the Argo API with that victim's privileges - submitting workflows, reading templates, deleting other tenants' runs.","attack_vector":"A tenant with workflow-submit rights plants the payload; an admin or another tenant triggers it by browsing the artifact listing in the Argo UI.","remediation":"Upgrade to 3.6.17 or 3.7.8 and restart Argo Server. Until then, tell operators not to browse artifact directory listings for workflows submitted by untrusted tenants, and consider a CSP at the ingress in front of the UI.","references":["https://github.com/argoproj/argo-workflows/security/advisories/GHSA-cv78-6m8q-ph82","https://nvd.nist.gov/vuln/detail/CVE-2026-23960"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-24156","cve":"CVE-2026-24156","aliases":[],"title":"NVIDIA DALI: RCE via unsafe deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DALI","year":"2026","cvss_score":7.3,"severity":"high","kev":false,"impact":"RCE via unsafe deserialization","attack_vector":"Malicious dataset/pipeline","remediation":"Bump DALI; rebuild data-loading images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24156","https://github.com/NVIDIA/product-security/tree/main/2026/5811"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2026-04-07"},{"id":"CVE-2026-24180","cve":"CVE-2026-24180","aliases":[],"title":"NVIDIA DALI: Code exec via OOB write on malformed tensor dims","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DALI","year":"2026","cvss_score":7.3,"severity":"high","kev":false,"impact":"Code exec via OOB write on malformed tensor dims","attack_vector":"Malicious dataset","remediation":"Bump DALI; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24180","https://github.com/NVIDIA/product-security/tree/main/2026/5814"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-122"],"published":"2026-06-09"},{"id":"CVE-2026-24181","cve":"CVE-2026-24181","aliases":[],"title":"NVIDIA DALI: Code exec via buffer overflow","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DALI","year":"2026","cvss_score":7.3,"severity":"high","kev":false,"impact":"Code exec via buffer overflow","attack_vector":"Malicious dataset","remediation":"Bump DALI; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24181","https://github.com/NVIDIA/product-security/tree/main/2026/5814"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-129"],"published":"2026-06-09"},{"id":"CVE-2026-24206","cve":"CVE-2026-24206","aliases":[],"title":"Triton Inference Server: Weak authentication","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2026","cvss_score":7.3,"severity":"high","kev":false,"impact":"Weak authentication -> unauthorized model access","attack_vector":"Network client of the endpoint","remediation":"Upgrade Triton; redeploy serving images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24206","https://github.com/NVIDIA/product-security/tree/main/2026/5828"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:L/I:L/A:L","cwe":["CWE-288"],"published":"2026-05-20"},{"id":"CVE-2026-24229","cve":"CVE-2026-24229","aliases":[],"title":"TensorRT-LLM: Missing authentication in authorization checks","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT-LLM","year":"2026","cvss_score":7.3,"severity":"high","kev":false,"impact":"Missing authentication in authorization checks","attack_vector":"Network client","remediation":"Bump TensorRT-LLM; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24229","https://github.com/NVIDIA/product-security/tree/main/2026/5840"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:U/C:H/I:L/A:L","cwe":["CWE-306"],"published":"2026-07-14"},{"cwe":["CWE-190","CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-31491","cve":"CVE-2026-31491","aliases":[],"title":"Linux kernel (drivers/infiniband/hw/irdma): Queue-depth arithmetic was done in 32 bits, so a tenant passing a huge","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/hw/irdma)","year":"2026","cvss_score":7.3,"severity":"high","kev":false,"impact":"Queue-depth arithmetic was done in 32 bits, so a tenant passing a huge SQ/RQ/SRQ size overflows the calculation and the driver reports success with a silently truncated queue. The queue the hardware and the userspace library then disagree about is a direct route to out-of-bounds work-queue-entry writes.","attack_vector":"A tenant container holding /dev/infiniband/uverbs* on an Intel irdma node calls create_qp/create_srq with a size near U32_MAX. No fabric peer or host root required.","remediation":"No fixed version is listed in the record - take the stable kernel carrying 3f08351de5ca (or cbd852f5700e / e37afcb56ae0) and reboot. Interim: remove /dev/infiniband/* from untrusted containers on irdma nodes.","references":["https://git.kernel.org/stable/c/3f08351de5ca4f2f724b86ad252fbc21289467e1","https://git.kernel.org/stable/c/cbd852f5700eb3f64392452faf693ac45cae8281","https://nvd.nist.gov/vuln/detail/CVE-2026-31491"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-40110","cve":"CVE-2026-40110","aliases":[],"title":"Jupyter Server (Origin validation): `re.match` used for Origin validation","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Jupyter Server (Origin validation)","year":"2026","cvss_score":7.3,"severity":"high","kev":false,"impact":"`re.match` used for Origin validation → prefix-match bypass","attack_vector":"Malicious page visited by the notebook user","remediation":"Upgrade past 2.17.0","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-40110"],"status":"curated","published":"2026-05-05"},{"id":"CVE-2026-46680","cve":"CVE-2026-46680","aliases":[],"title":"containerd: Numeric User directive that fails 32-bit parsing is treated as a username, changing the effective UID","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2026","cvss_score":7.3,"severity":"high","kev":false,"impact":"Numeric User directive that fails 32-bit parsing is treated as a username, changing the effective UID","attack_vector":"Malicious image","remediation":"Rolling containerd upgrade with node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-46680"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2026-07-01"},{"id":"CVE-2026-62290","cve":"CVE-2026-62290","aliases":[],"title":"cert-manager: Challenge resource handling flaw in cert-manager 1.18.0-1.19.5 and 1.20.x","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"cert-manager","year":"2026","cvss_score":7.3,"severity":"high","kev":false,"impact":"Challenge resource handling flaw in cert-manager 1.18.0-1.19.5 and 1.20.x","attack_vector":"Cluster user with namespace access who can create Certificates","remediation":"Rolling cert-manager upgrade to 1.19.6/1.20.3+; no GPU drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-62290"],"status":"curated","published":"2026-07-16"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:N/A:L","cwe":["CWE-668"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-63914","cve":"CVE-2026-63914","aliases":[],"title":"Linux kernel (net/xfrm, net/key): No memory corruption here, but a clean namespace boundary break. SA migration","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm, net/key)","year":"2026","cvss_score":7.3,"severity":"high","kev":false,"impact":"No memory corruption here, but a clean namespace boundary break. SA migration notifications were multicast to the init namespace regardless of which namespace issued them, so a migration triggered inside a tenant namespace is delivered to the host's IKE daemon as if it were the host's own event - carrying an attacker-chosen selector and new endpoint address. A tenant can therefore feed the host key-management daemon a forged MOBIKE-style address update, and in the other direction a tenant's own daemon never sees its migrations, so address updates inside a namespace silently do not work. Both halves mean fabric encryption ends up pointed at an endpoint someone else chose.","attack_vector":"A tenant container with CAP_NET_ADMIN in its own user+network namespace issues XFRM_MSG_MIGRATE (or the PF_KEY equivalent) on its own namespace's socket; the notification is delivered to XFRMNLGRP_MIGRATE and PF_KEY listeners in the init namespace, i.e. to the host's IKE daemon. Exploitability depends on whether that daemon acts on migrate notifications without re-validating the originating namespace - most do not check, because until this fix the notification could only come from init_net. No fabric access needed.","remediation":"Boot a kernel carrying the fix commits below (no fixed stable version published). Interim control: do not run a host IKE daemon subscribed to XFRMNLGRP_MIGRATE or PF_KEY BROADCAST_ALL on nodes where tenants hold CAP_NET_ADMIN in their own namespaces, and remove that capability from tenant namespaces.","references":["https://git.kernel.org/stable/c/bafc7d0774b9bf52909c70ed990bc5ccf7ec4bad","https://git.kernel.org/stable/c/6df8157547347b5257bf640a0ae3dfc4411e06cd","https://nvd.nist.gov/vuln/detail/CVE-2026-63914"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:A/AC:H/PR:H/UI:N/S:C/C:H/I:H/A:N","cwe":["CWE-706","CWE-863"],"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2026-046-cilium-sds-secret-sync-for-l7-ne","cve":null,"aliases":["GHSA-xqhm-7xhv-6ppj"],"title":"Cilium (SDS secret sync for L7 network policies): A tenant confined to their own namespace reaches across and rewrites","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium (SDS secret sync for L7 network policies)","year":"2026","cvss_score":7.3,"severity":"high","kev":false,"impact":"A tenant confined to their own namespace reaches across and rewrites another tenant's L7 policy enforcement. When SDS is enabled for L7 policies, synced Secrets and ConfigMaps land in a shared namespace where names can collide. An attacker who can create a Secret or ConfigMap in their own namespace plus a CiliumNetworkPolicy referencing it picks a name that collides with a victim's object, and their object overwrites the synchronised SDS object the victim's policy depends on — including policies using headerMatches.secret. The result is that the attacker chooses whether another tenant's traffic is permitted or denied: open a path that should be closed, or close one that should be open and take out that tenant's service. Namespace-scoped create rights on ordinary Kubernetes objects is exactly the permission set a GPU cloud hands every customer.","attack_vector":"Adjacent network / in-cluster. Requires namespace-scoped rights to create a Secret or ConfigMap and a referencing CiliumNetworkPolicy, plus knowledge of the victim object's name, on a cluster with SDS enabled for L7 policies.","remediation":"Upgrade to Cilium 1.17.18, 1.18.12 or 1.19.6 and roll the agents. Workaround without upgrading: disable policy Secret synchronisation and manage referenced policy Secrets directly in the configured policy-secrets namespace (cilium-secrets by default), which removes the shared naming surface. Audit existing synced objects for unexpected owners.","references":["https://github.com/cilium/cilium/security/advisories/GHSA-xqhm-7xhv-6ppj"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2018-7078","cve":"CVE-2018-7078","aliases":[],"title":"HPE iLO4 / iLO5: Remote code execution on the management controller","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HPE iLO4 / iLO5","year":"2018","cvss_score":7.2,"severity":"high","kev":false,"impact":"Remote code execution on the management controller","attack_vector":"Network, authenticated","remediation":"iLO firmware update (iLO4 <2.60, iLO5 <1.30)","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-7078"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2018-08-06"},{"id":"CVE-2018-7105","cve":"CVE-2018-7105","aliases":[],"title":"HPE iLO3/4/5: Arbitrary code execution on the iLO","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HPE iLO3/4/5","year":"2018","cvss_score":7.2,"severity":"high","kev":false,"impact":"Arbitrary code execution on the iLO","attack_vector":"Network","remediation":"iLO firmware update across three generations simultaneously — the version matrix is the operational cost","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-7105"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2018-09-27"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:H/UI:N/S:U/C:H/I:H/A:H","fleet":{"pain_class":"daemon-restart"},"id":"CVE-2019-17272","cve":"CVE-2019-17272","aliases":[],"title":"NetApp ONTAP Select Deploy administration utility (privilege escalation): An administrative user of the Deploy utility","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"NetApp ONTAP Select Deploy administration utility (privilege escalation)","year":"2019","cvss_score":7.2,"severity":"high","kev":false,"impact":"An administrative user of the Deploy utility escalates beyond their assigned role, collapsing any separation between a limited storage operator and full control of the deployment tooling.","attack_vector":"An existing administrative account on any version of ONTAP Select Deploy, reachable over the network.","remediation":"Upgrade Deploy to a fixed release. Do not rely on role separation inside Deploy as a security boundary on affected versions - treat every Deploy admin as fully privileged until patched.","references":["https://security.netapp.com/advisory/ntap-20191121-0002/","https://nvd.nist.gov/vuln/detail/CVE-2019-17272"],"status":"curated"},{"id":"CVE-2019-19029","cve":"CVE-2019-19029","aliases":[],"title":"Harbor: SQL injection via user-groups","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Harbor","year":"2019","cvss_score":7.2,"severity":"high","kev":false,"impact":"SQL injection via user-groups","attack_vector":"Authenticated registry user","remediation":"Upgrade Harbor","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-19029"],"status":"curated","published":"2020-03-20"},{"id":"CVE-2019-19577","cve":"CVE-2019-19577","aliases":[],"title":"Xen on AMD - x86 HVM pagetable height update: AMD HVM guest OS users can trigger a data-structure access during a","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen on AMD - x86 HVM pagetable height update","year":"2019","cvss_score":7.2,"severity":"high","kev":false,"impact":"AMD HVM guest OS users can trigger a data-structure access during a pagetable-height update, causing denial of service or possibly gaining privileges. Privilege escalation out of a guest into the hypervisor is the worst outcome available on a virtualised host - the attacker moves from one tenant's VM to controlling all of them.","attack_vector":"From inside an AMD HVM guest under Xen. Tenant-reachable.","remediation":"Fixed in Xen (XSA-310). Update the hypervisor and reboot the host; no firmware step. Affects Xen through 4.12.x.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-19577","https://xenbits.xen.org/xsa/advisory-310.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2019-12-11"},{"id":"CVE-2020-12967","cve":"CVE-2020-12967","aliases":["SEVerity","SEVurity"],"title":"AMD SEV / SEV-ES - missing nested page table protection: SEV and SEV-ES do not protect the nested page tables, so a","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV / SEV-ES - missing nested page table protection","year":"2020","cvss_score":7.2,"severity":"high","kev":false,"impact":"SEV and SEV-ES do not protect the nested page tables, so a malicious hypervisor can remap the guest's physical address space underneath it. The SEVurity and SEVerity research chains turned this into arbitrary code execution inside the encrypted guest: because encryption is keyed by physical address and the host controls the mapping, the host can move ciphertext blocks around to assemble instructions of its choosing inside the victim VM. The guest's memory is encrypted and the attacker still gets code execution in it.","attack_vector":"Requires a malicious administrator who has compromised the hypervisor. Affects SEV and SEV-ES; SEV-SNP's reverse-map table is the architectural answer to exactly this.","remediation":"**Not fixable on SEV or SEV-ES** - the missing protection is architectural, which is why AMD built SEV-SNP with the RMP. The remediation is a platform migration: run confidential workloads on SEV-SNP-capable EPYC (Milan 7003 and later) with SNP actually enabled, not on SEV or SEV-ES. If you are running SEV or SEV-ES today and telling customers their VMs are protected from you, that claim does not hold. Requires new hardware or at minimum a firmware/BIOS enablement pass plus reconfiguring your VMM to launch SNP guests.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-12967","https://www.usenix.org/conference/woot20/presentation/werner","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"tags":["tenant-isolation"],"published":"2021-05-13"},{"id":"CVE-2020-24637","cve":"CVE-2020-24637","aliases":[],"title":"ArubaOS GRUB2 implementation (secure boot): Two flaws in ArubaOS's GRUB2 implementation allow secure boot","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ArubaOS GRUB2 implementation (secure boot)","year":"2020","cvss_score":7.2,"severity":"high","kev":false,"impact":"Two flaws in ArubaOS's GRUB2 implementation allow secure boot to be bypassed, leading to remote compromise. Secure boot on a network device is the control that stops a compromise from becoming permanent; bypassing it means an attacker's image survives reimaging and firmware updates. Same structural problem as the Cisco NX-OS image-verification bypass, on a different vendor.","attack_vector":"An attacker able to influence the boot chain — via administrative access or the documented remote path.","remediation":"ArubaOS upgrade plus reload; this is a bootloader-level fix so it must be applied per device and cannot be worked around in config. Treat any device suspected of pre-patch compromise as needing replacement or a verified out-of-band reflash rather than an in-place upgrade.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-24637"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2020-12-11"},{"id":"CVE-2020-26122","cve":"CVE-2020-26122","aliases":[],"title":"Inspur NF5266M5 through firmware 3.21.2 and other Inspur M5-generation servers: An attacker with administrative reach","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Inspur NF5266M5 through firmware 3.21.2 and other Inspur M5-generation servers","year":"2020","cvss_score":7.2,"severity":"high","kev":false,"impact":"An attacker with administrative reach installs their own BMC firmware on Inspur server hardware, and from that point the operator has no reliable way to determine what is running on the controller. The outcomes are the full BMC set - power, console, virtual media, boot control - plus an implant that survives host reimaging and node reprovisioning. Inspur M5 hardware shows up in a lot of Asia-Pacific cloud and HPC estates and in secondhand markets that neoclouds buy from, so this is a real inventory question for anyone assembling capacity opportunistically rather than from a single vendor. The BMC's firmware verification is weak and lacks the checks needed to establish that an image is genuine before flashing it.","attack_vector":"Administrative privilege on the BMC, reachable over the management network. Shared or default BMC credentials on secondhand hardware make this a realistic starting position rather than a theoretical one.","remediation":"Firmware update from Inspur - and obtaining it is the problem. Inspur's English security bulletin path returns a hard 404 and the site's homepage carries no PSIRT link, so there is no reachable vendor advisory channel for a non-Chinese-market operator. Given US export and entity-list constraints on Inspur, many operators will also find vendor support unavailable regardless. The practical position: treat Inspur BMC firmware as unverifiable, isolate these BMCs on a management VLAN with no tenant-reachable route, rotate credentials to per-node unique values, and for secondhand M5 hardware assume the BMC may already carry unknown firmware and plan accordingly.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-26122","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2020/26xxx/CVE-2020-26122.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-287"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2021-20288","cve":"CVE-2021-20288","aliases":[],"title":"Ceph MON (CephX authentication): The monitor does not sanitize other_keys when handling CEPHX_GET_AUTH_SESSION_KEY, so","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ceph MON (CephX authentication)","year":"2021","cvss_score":7.2,"severity":"high","kev":false,"impact":"The monitor does not sanitize other_keys when handling CEPHX_GET_AUTH_SESSION_KEY, so an attacker who has any CephX credential (or who can force one to be reissued) can reuse a key and authenticate as a different, higher-privileged Ceph entity. That is escalation from one tenant's cephx identity to another's, including admin-level access to pools they do not own.","attack_vector":"Anyone holding a valid CephX credential for the cluster and able to reach the monitors on the cluster/public network - which includes any tenant compute node that mounts RBD or CephFS natively.","remediation":"Upgrade the monitors to Ceph 14.2.20 or later (and matching Octopus/Pacific builds), then restart ceph-mon. Rotate CephX keys afterwards, since anything issued before the fix could already have been reused. Keep the Ceph public network off tenant-routable paths.","references":["https://access.redhat.com/security/cve/CVE-2021-20288","https://nvd.nist.gov/vuln/detail/CVE-2021-20288"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2021-26311","cve":"CVE-2021-26311","aliases":["undeSErVed","SEV memory remapping"],"title":"AMD SEV / SEV-ES - guest address space rearrangement undetected by attestation: A malicious hypervisor can rearrange","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV / SEV-ES - guest address space rearrangement undetected by attestation","year":"2021","cvss_score":7.2,"severity":"high","kev":false,"impact":"A malicious hypervisor can rearrange memory in the guest's address space and the SEV attestation mechanism does not notice, because attestation measures content rather than placement. That lets the host reorder the guest's own encrypted pages to build a different program out of the same measured bytes - the attestation report still validates while the VM executes something the tenant never wrote. This is the failure that makes 'the attestation passed' an insufficient answer on SEV/SEV-ES.","attack_vector":"Malicious hypervisor. Affects SEV and SEV-ES; SEV-SNP addresses it with the reverse-map table enforcing page ownership and mapping.","remediation":"**Not fixable on SEV or SEV-ES** - migrate confidential workloads to SEV-SNP, where the RMP enforces the guest-physical to system-physical mapping. Practically that means EPYC Milan (7003) or newer with SNP enabled in SBIOS and a VMM that launches SNP guests. If you are attesting SEV/SEV-ES guests today, understand that a passing attestation report does not establish the guest is running the code that was measured.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26311","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2021-05-13"},{"id":"CVE-2021-26344","cve":"CVE-2021-26344","aliases":[],"title":"AMD PSP1 Configuration Block (APCB) parsing: An out-of-bounds memory write while the platform processes the AMD PSP1","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD PSP1 Configuration Block (APCB) parsing","year":"2021","cvss_score":7.2,"severity":"high","kev":false,"impact":"An out-of-bounds memory write while the platform processes the AMD PSP1 Configuration Block. Reaching it requires the ability to modify and re-sign the BIOS image, which is a high bar - but the payoff is memory corruption inside the secure processor's configuration path, i.e. control of the root of trust with a signature that validates.","attack_vector":"Local, and requires both BIOS image modification and the ability to sign the result. That combination points at a signing-key compromise or an insider in the firmware build pipeline rather than a runtime attacker.","remediation":"Fixed in AMD reference firmware (AGESA / PSP / SEV firmware) and delivered only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo and the ODMs each rebuild and requalify AMD's AGESA drop before shipping. **Expect one to six months of OEM lag**, and on end-of-support platforms expect nothing. Applying it is a drain plus full power cycle, not a driver reload. Verify by reading back the PSP/SMU firmware version afterwards rather than trusting the BIOS version string. The signing prerequisite means your firmware supply chain is the real control here - who can sign a BIOS for your fleet, and how is that key held?","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26344","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2024-08-13"},{"id":"CVE-2021-36784","cve":"CVE-2021-36784","aliases":[],"title":"Rancher: restricted-admin role escalates to full admin","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Rancher","year":"2021","cvss_score":7.2,"severity":"high","kev":false,"impact":"restricted-admin role escalates to full admin","attack_vector":"A restricted-admin user","remediation":"Upgrade Rancher","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-36784"],"status":"curated","published":"2022-05-02"},{"id":"CVE-2022-30704","cve":"CVE-2022-30704","aliases":["INTEL-SA-00717","CVE-2021-0187"],"title":"Intel TXT SINIT Authenticated Code Module for some Intel processors: Improper initialization in the SINIT ACM","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel TXT SINIT Authenticated Code Module for some Intel processors","year":"2022","cvss_score":7.2,"severity":"high","kev":false,"impact":"Improper initialization in the SINIT ACM - the Intel-signed module that performs the TXT measured launch. A privileged local attacker can escalate through it, which means the dynamic root of trust that a TXT launch is supposed to establish can be subverted at the moment it is created. If you use TXT (directly, or via a launch control policy that gates whether a node may join a secure pool), an attacker can make a compromised node produce a passing launch. Nodes admitted to a trusted pool on that basis are not trustworthy, and the compromise is at firmware level so it crosses tenant handoff.","attack_vector":"A privileged local user on the node - local root or SMM-capable code.","remediation":"New SINIT ACM delivered inside a BIOS/platform-firmware update from the OEM (Dell, HPE, Supermicro, Lenovo, Gigabyte, Quanta), plus updating any standalone SINIT binary your tboot/launch stack loads. Host reboot and job drain. If you gate scheduling on TXT launch results, update the launch control policy hashes at the same time or the policy will fail closed and strand capacity.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-30704","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00717.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-33196","cve":"CVE-2022-33196","aliases":[],"title":"Intel Xeon memory controller configuration (with SGX): Memory controller configuration registers are left with","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Xeon memory controller configuration (with SGX)","year":"2022","cvss_score":7.2,"severity":"high","kev":false,"impact":"Memory controller configuration registers are left with permissions that let a privileged local user reach a privilege escalation on SGX-enabled Xeon platforms. Memory controller configuration is the layer that enforces where enclave memory lives, so getting it wrong is structurally worse than a normal ring-0 bug on a confidential-compute host.","attack_vector":"Privileged local access on the host.","remediation":"Platform BIOS/firmware update from the OEM, plus TCB recovery and re-attestation. BIOS means a per-node drain, a reboot, and waiting on OEM packaging - budget quarters, not weeks, on server boards.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-33196","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00738.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2023-02-16"},{"id":"CVE-2022-41804","cve":"CVE-2022-41804","aliases":[],"title":"Intel Xeon processors (SGX/TDX error injection): Unauthorised error injection against SGX or TDX on affected Xeon parts","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Xeon processors (SGX/TDX error injection)","year":"2022","cvss_score":7.2,"severity":"high","kev":false,"impact":"Unauthorised error injection against SGX or TDX on affected Xeon parts gives a privileged user an escalation. Fault injection against a TEE is the same family as Plundervolt: the host corrupts the protected computation until it gives up its secrets.","attack_vector":"Privileged local access on the host.","remediation":"Microcode update - late-loadable at boot on most distributions, so this one does not wait on an OEM BIOS release - plus a TCB recovery and re-attestation. Reboot required.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-41804","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00837.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"tags":["tenant-isolation"],"published":"2023-08-11"},{"id":"CVE-2022-42278","cve":"CVE-2022-42278","aliases":[],"title":"NVIDIA DGX BMC (AMI-derived management controller): The BMC's SPX REST API lets an authorised attacker read and write","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVIDIA DGX BMC (AMI-derived management controller)","year":"2022","cvss_score":7.2,"severity":"high","kev":false,"impact":"The BMC's SPX REST API lets an authorised attacker read and write arbitrary locations in the IPMI server process memory, reaching code execution inside the BMC. A BMC compromise on a DGX gives an attacker power control, virtual media, serial console and a persistent foothold under the host OS on a node holding eight GPUs.","attack_vector":"Network access to the BMC management interface holding credentials at some authorised level. Whether that is 'remote' depends entirely on how genuinely isolated your OOB network is - in practice most fleets have a jump host, a DCIM integration or a monitoring collector that bridges it.","remediation":"Update the DGX BMC firmware bundle per bulletin 5435. Cost: BMC firmware usually updates without draining the GPUs, but the BMC resets and out-of-band access drops for a few minutes. Pair the patch with an actual audit of who can route to the BMC subnet - that control is worth more than the patch.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42278","https://github.com/NVIDIA/product-security/tree/main/2022/5435"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-119"],"fleet":{"pain_class":"firmware-flash"},"published":"2023-01-13"},{"id":"CVE-2022-42279","cve":"CVE-2022-42279","aliases":[],"title":"DGX servers BMC: OS command injection on BMC","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX servers BMC","year":"2022","cvss_score":7.2,"severity":"high","kev":false,"impact":"OS command injection on BMC","attack_vector":"Authenticated mgmt-LAN admin","remediation":"Flash BMC 2.09.00+ out-of-band","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42279","https://github.com/NVIDIA/product-security/tree/main/2024/5513"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-78"],"published":"2023-01-13"},{"id":"CVE-2022-42289","cve":"CVE-2022-42289","aliases":[],"title":"DGX-2 SBIOS: OS command injection in firmware","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX-2 SBIOS","year":"2022","cvss_score":7.2,"severity":"high","kev":false,"impact":"OS command injection in firmware","attack_vector":"Network-adjacent authenticated admin","remediation":"Flash SBIOS/BMC out-of-band","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42289","https://github.com/NVIDIA/product-security/tree/main/2024/5513"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-78"],"fleet":{"pain_class":"firmware-flash"},"published":"2023-01-13"},{"id":"CVE-2022-42290","cve":"CVE-2022-42290","aliases":[],"title":"DGX-2 SBIOS: OS command injection in BIOS","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX-2 SBIOS","year":"2022","cvss_score":7.2,"severity":"high","kev":false,"impact":"OS command injection in BIOS","attack_vector":"Network-adjacent authenticated admin","remediation":"Flash SBIOS out-of-band","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42290","https://github.com/NVIDIA/product-security/tree/main/2024/5513"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-78"],"fleet":{"pain_class":"firmware-flash"},"published":"2023-01-13"},{"id":"CVE-2023-25507","cve":"CVE-2023-25507","aliases":[],"title":"NVIDIA DGX BMC (AMI-derived management controller): The DGX-1 BMC's SPX REST API accepts injected shell commands","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVIDIA DGX BMC (AMI-derived management controller)","year":"2023","cvss_score":7.2,"severity":"high","kev":false,"impact":"The DGX-1 BMC's SPX REST API accepts injected shell commands from an authorised caller, giving direct command execution on the management controller. A BMC compromise on a DGX gives an attacker power control, virtual media, serial console and a persistent foothold under the host OS on a node holding eight GPUs.","attack_vector":"Network access to the BMC management interface holding credentials at some authorised level. Whether that is 'remote' depends entirely on how genuinely isolated your OOB network is - in practice most fleets have a jump host, a DCIM integration or a monitoring collector that bridges it.","remediation":"Update the DGX BMC firmware bundle per bulletin 5458. Cost: BMC firmware usually updates without draining the GPUs, but the BMC resets and out-of-band access drops for a few minutes. Pair the patch with an actual audit of who can route to the BMC subnet - that control is worth more than the patch.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25507","https://github.com/NVIDIA/product-security/tree/main/2023/5458"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-77"],"fleet":{"pain_class":"firmware-flash"},"published":"2023-04-22"},{"id":"CVE-2023-25549","cve":"CVE-2023-25549","aliases":["SEVD-2023-101-04"],"title":"Schneider Electric StruxureWare Data Center Expert (V7.9.2 and prior) - network settings endpoint: Code injection","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Schneider Electric StruxureWare Data Center Expert (V7.9.2 and prior) - network settings endpoint","year":"2023","cvss_score":7.2,"severity":"high","kev":false,"impact":"Code injection through a parameter of the DCE network-settings endpoint gives remote code execution on the DCIM appliance. Same outcome as the other DCE RCEs: control of the facility-layer aggregation point.","attack_vector":"Authenticated remote access to the DCE administrative interface.","remediation":"Upgrade past V7.9.2 (SEVD-2023-101-04 covers this whole batch - CVE-2023-25547 through -25555 - so patch once). Rotate stored device credentials afterwards.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25549"],"status":"curated","published":"2023-04-18"},{"id":"CVE-2023-29002","cve":"CVE-2023-29002","aliases":[],"title":"Cilium: Debug mode logs the contents of the cilium-secrets namespace, including TLS private keys","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2023","cvss_score":7.2,"severity":"high","kev":false,"impact":"Debug mode logs the contents of the cilium-secrets namespace, including TLS private keys","attack_vector":"Anyone with log-pipeline read access","remediation":"Disable agent debug mode; rotate the TLS keys in cilium-secrets; scrub logs","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-29002"],"status":"curated","published":"2023-04-18"},{"id":"CVE-2023-31037","cve":"CVE-2023-31037","aliases":[],"title":"BlueField-2 / BlueField-3 DPU BMC: Code injection on DPU BMC","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"BlueField-2 / BlueField-3 DPU BMC","year":"2023","cvss_score":7.2,"severity":"high","kev":false,"impact":"Code injection on DPU BMC -> DPU takeover below the host OS","attack_vector":"Network-adjacent mgmt access to the DPU BMC","remediation":"Flash DPU BMC firmware out-of-band; DPU reset drops tenant networking, schedule drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31037","https://github.com/NVIDIA/product-security/tree/main/2024/5511"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"],"fleet":{"pain_class":"firmware-flash"},"published":"2024-01-24"},{"id":"CVE-2023-31313","cve":"CVE-2023-31313","aliases":[],"title":"AMD Power Management Firmware (PMFW) - unintended proxy to the System Management Unit: The GPU power management","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Power Management Firmware (PMFW) - unintended proxy to the System Management Unit","year":"2023","cvss_score":7.2,"severity":"high","kev":false,"impact":"The GPU power management firmware acts as an unintended proxy, letting a privileged attacker relay malformed messages through to the System Management Unit. The SMU is the always-on microcontroller that governs clocks, voltages and power limits across the platform; getting arbitrary messages to it through the GPU firmware path means the GPU stack becomes a route into platform-level control. It is the confused-deputy pattern applied to firmware: PMFW is trusted by the SMU, so whoever controls PMFW inherits that trust.","attack_vector":"Local, privileged attacker able to send messages to the GPU power management firmware.","remediation":"Fixed in AMD GPU firmware, which on Instinct parts is delivered as a firmware bundle through the ROCm/amdgpu driver package (the PSP loads the signed blobs at driver init) rather than through the server BIOS. Practically: update the AMD GPU driver/firmware package, then **drain the node and reboot** - the firmware is loaded once at driver init, so a reload of the module with no process holding /dev/kfd is the minimum, and a reboot is what you will actually schedule. Some fixes at this layer also require a **GPU VBIOS flash** via AMD's amdvbflash/amdfwtool, which is an offline, per-card operation with real bricking risk - check the AMD bulletin for whether a VBIOS update is called out before assuming a driver package covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31313","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2026-02-12"},{"id":"CVE-2023-3260","cve":"CVE-2023-3260","aliases":[],"title":"Dataprobe iBoot PDU: Authenticated OS command injection on the PDU","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dataprobe iBoot PDU","year":"2023","cvss_score":7.2,"severity":"high","kev":false,"impact":"Authenticated OS command injection on the PDU — an attacker who reaches the power controller can cut power to racks, and can pivot from the PDU into the management network","attack_vector":"Network, authenticated","remediation":"PDU firmware update to 1.44.08042023; PDUs are rarely in the patch pipeline at all, so the real cost is building one","references":["https://thehackernews.com/2023/08/multiple-flaws-in-cyberpower-and.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-08-14"},{"id":"CVE-2023-32666","cve":"CVE-2023-32666","aliases":[],"title":"Intel 4th Gen Xeon on-chip debug and test interface (with SGX or TDX): The on-chip debug and test interface has","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel 4th Gen Xeon on-chip debug and test interface (with SGX or TDX)","year":"2023","cvss_score":7.2,"severity":"high","kev":false,"impact":"The on-chip debug and test interface has improper access control on 4th-generation Xeon when SGX or TDX is enabled, giving a privileged user escalation. Debug interfaces reaching into a TEE is the single worst shape for a confidential-compute claim, because it bypasses the architectural boundary entirely rather than working around it.","attack_vector":"Privileged local access on the host.","remediation":"Microcode/platform firmware update plus TCB recovery. Where the fix lands in microcode it can be late-loaded at boot; where it lands in BIOS, expect the OEM lag. Re-attest afterwards.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-32666","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00986.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2024-03-14"},{"id":"CVE-2023-34341","cve":"CVE-2023-34341","aliases":["AMI-SA-2023005","NVIDIA OSR review"],"title":"AMI MegaRAC SPx (SPX REST API): Arbitrary read and write into the memory of the BMC's IPMI server process via the SPX","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx (SPX REST API)","year":"2023","cvss_score":7.2,"severity":"high","kev":false,"impact":"Arbitrary read and write into the memory of the BMC's IPMI server process via the SPX REST API. That is a full primitive - the attacker reads out credentials, session tokens and keys held in that process, then writes to redirect execution. In fleet terms, one compromised BMC admin credential turns into code execution on the controller and from there into persistent firmware-level control of the node.","attack_vector":"Network-reachable REST API, requires high privileges - an administrative BMC account. The realistic path is credential reuse: fleets provision BMCs from a template and end up with the same admin password across an entire rack or SKU, so a single leaked credential from one node's config, a Redfish scraper, or a decommissioned host escalates to every node sharing it.","remediation":"Firmware flash to SPx_12.7 / SPx_13.5, out-of-band per node, ODM-gated. The higher-leverage work is credential hygiene and is config-only: unique per-node BMC admin passwords generated and stored by your secrets manager, no admin credential embedded in provisioning images or monitoring configs, and Redfish accounts scoped to read-only where they only scrape telemetry.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2023005.pdf","https://nvd.nist.gov/vuln/detail/CVE-2023-34341"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-06-12"},{"id":"CVE-2023-34343","cve":"CVE-2023-34343","aliases":["AMI-SA-2023005","NVIDIA OSR review"],"title":"AMI MegaRAC SPx (SPX REST API): Shell command injection through the BMC's REST API","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx (SPX REST API)","year":"2023","cvss_score":7.2,"severity":"high","kev":false,"impact":"Shell command injection through the BMC's REST API. An administrative BMC user gets a root shell on the management controller's Linux - which is a large step up from what the web UI lets them do, because from a BMC shell the attacker can write firmware, install a persistent implant in the BMC's own flash, and pivot to the host. The gap between 'has a BMC admin password' and 'owns the node forever' closes here.","attack_vector":"Network-reachable REST API with an administrative BMC account. Same shared-credential exposure as the rest of the SPX REST API family: assume any leaked BMC admin password is a fleet-wide credential unless you have proven otherwise.","remediation":"Firmware flash to SPx_12.7 / SPx_13.5, out-of-band per node, ODM-gated. Config-only compensations that actually reduce blast radius: unique BMC credentials per node, an allowlist ACL restricting who can reach the BMC web/REST port at all, and logging of BMC authentication to your SIEM so a credential-spray across the management VLAN is visible.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2023005.pdf","https://nvd.nist.gov/vuln/detail/CVE-2023-34343"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-06-12"},{"id":"CVE-2023-40289","cve":"CVE-2023-40289","aliases":[],"title":"Supermicro BMC (IPMI web interface, command injection): Command injection that turns a BMC administrator account","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC (IPMI web interface, command injection)","year":"2023","cvss_score":7.2,"severity":"high","kev":false,"impact":"Command injection that turns a BMC administrator account into shell on the BMC's own Linux. That matters more than it sounds: BMC admin is a constrained management role, whereas BMC shell means arbitrary firmware modification, access to the host over the internal bridges, and a place to hide that no host-side agent can inspect. Chained after any of the XSS bugs in the same batch, an operator merely visiting a page is enough to reach it.","attack_vector":"An authenticated BMC administrator - or, realistically, an attacker who chained an XSS in the same firmware to ride an admin's session. Requires network reach to the BMC web interface.","remediation":"BMC firmware flash per board, out-of-band. Supermicro fixes ship per-SKU and lag disclosure, so expect a long tail of boards with no image. Interim controls that work today: keep the BMC web UI off any routable network, require a jump host, and stop using shared BMC admin credentials across the fleet so one compromise is not fleet-wide.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-40289","https://www.binarly.io/advisories"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-03-27"},{"id":"CVE-2023-5528","cve":"CVE-2023-5528","aliases":[],"title":"Kubernetes (in-tree storage): Crafted PV/pod on Windows nodes escalates to node admin via in-tree storage plugin","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (in-tree storage)","year":"2023","cvss_score":7.2,"severity":"high","kev":false,"impact":"Crafted PV/pod on Windows nodes escalates to node admin via in-tree storage plugin","attack_vector":"Cluster user able to create pods and PVs","remediation":"Rolling control-plane and kubelet upgrade; Windows node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-5528"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2023-11-14"},{"id":"CVE-2024-0161","cve":"CVE-2024-0161","aliases":["DSA-2024-006"],"title":"Dell PowerEdge Server BIOS (SMM communication buffer): The BIOS fails to properly validate the SMM communication","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell PowerEdge Server BIOS (SMM communication buffer)","year":"2024","cvss_score":7.2,"severity":"high","kev":false,"impact":"The BIOS fails to properly validate the SMM communication buffer, so a low-privilege local attacker can write into SMRAM. System Management Mode is the most privileged execution context on the machine - more privileged than the hypervisor, invisible to it, and reachable regardless of what OS is booted. An attacker who writes to SMRAM owns the node in a way that no reimage, no disk wipe and no hypervisor-level control removes. On a multi-tenant bare-metal GPU fleet this is the canonical persistent-implant primitive.","attack_vector":"A low-privilege local account on the host. Notably this does NOT need root or administrator - so a tenant workload running as an ordinary user is in scope, as is anything that gets modest code execution through an application bug.","remediation":"System BIOS update. Stage it via iDRAC/Lifecycle Controller or OME, but it lands only on the next reboot, so it costs a job drain and a maintenance window per node. Per-platform version floors are in the advisory table. There is no config-only mitigation for an SMM handler bug - you cannot turn SMM off. Prioritise nodes that run untrusted or multi-tenant workloads over internal-only nodes when sequencing the rollout.","references":["https://www.dell.com/support/kbdoc/en-us/000222979/dsa-2024-006-security-update-for-dell-poweredge-server-bios-for-an-improper-smm-communication-buffer-verification-vulnerability","https://nvd.nist.gov/vuln/detail/CVE-2024-0161"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-03-13"},{"id":"CVE-2024-10237","cve":"CVE-2024-10237","aliases":[],"title":"Supermicro BMC firmware validation (MBD-X12DPG-OA6): Root-of-Trust bypass","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC firmware validation (MBD-X12DPG-OA6)","year":"2024","cvss_score":7.2,"severity":"high","kev":false,"impact":"Root-of-Trust bypass — firmware image authentication design flaw lets a modified image pass BMC inspection and signature verification. Persistent below-OS implant","attack_vector":"Network, high-privilege BMC access","remediation":"BMC flash with a fixed Supermicro build; the RoT bypass means prior firmware state cannot be trusted, so treat affected nodes as requiring re-attestation, not just patching","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-10237"],"status":"curated","fleet":{"ubiquity":"very common - Supermicro is a dominant GPU-server ODM for neoclouds","remediation_pain":"firmware-flash + RoT re-provisioning - the flaw is in the update-validation path itself, so a compromised node may need physical-access recovery of the SPI/BMC flash","pain_class":"physical access","why_fleet_wide":"A signed-firmware-validation bypass means the fleet's defense against malicious firmware is itself the bug: an attacker can push a persistent BMC implant that survives host reinstall and is invisible to the OS."},"published":"2025-02-04"},{"id":"CVE-2024-10238","cve":"CVE-2024-10238","aliases":[],"title":"Supermicro BMC firmware image verification routine on MBD-X12DPG-OA6: A crafted update image smashes the stack","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC firmware image verification routine on MBD-X12DPG-OA6","year":"2024","cvss_score":7.2,"severity":"high","kev":false,"impact":"A crafted update image smashes the stack of the code that was supposed to be checking that image, giving the attacker execution inside the BMC's update path. The result is an operator-invisible firmware implant on a GPU node: control of power sequencing, virtual media, and the console, plus the ability to falsify what the BMC reports about itself. On a board specifically sold for GPU workloads, this is the persistence mechanism an attacker wants after gaining any temporary administrative access. A dual-socket GPU-oriented board. The parser fails to bound-check a length field in the image container before copying, so the verification code itself is the memory-safety bug.","attack_vector":"An authenticated high-privilege BMC session that can submit a firmware image. Reachability is whatever your management network allows - typically the OOB VLAN plus whatever provisioning automation holds BMC admin credentials.","remediation":"Firmware flash with the fixed BMC image from Supermicro's January 2025 BMC/IPMI advisory. There is no config-only fix, because the vulnerable code is the update handler. Practical hardening in the meantime: deny BMC firmware upload from anything except a single hardened provisioning host, and treat any node whose BMC accepted an unexpected image as needing an out-of-band SPI reflash rather than a software update.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-10238","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2024/10xxx/CVE-2024-10238.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-10239","cve":"CVE-2024-10239","aliases":[],"title":"Supermicro OpenBMC firmware image verification (MBD-X12DPG-OA6), fat->fsd.max_fld field: The BMC's own firmware-image","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro OpenBMC firmware image verification (MBD-X12DPG-OA6), fat->fsd.max_fld field","year":"2024","cvss_score":7.2,"severity":"high","kev":false,"impact":"The BMC's own firmware-image parser overflows its stack on an unchecked length field in the image header - meaning the code that is supposed to decide whether an update is trustworthy can be taken over by the untrusted update itself. This is the bug class that defeats a hardware root of trust: it does not matter that the SoC verifies signatures if you get control inside the verifier before verification completes. The prize is a permanent BMC implant that survives host reimage, node redeployment and tenant handover, on a platform explicitly marketed as an OpenBMC server board.","attack_vector":"Requires BMC administrator privileges to submit a firmware image - so it is a second-stage bug. But 'admin on the BMC' is exactly what the credential-disclosure, default-password and authentication-bypass entries elsewhere in this cluster hand an attacker, and it is also what a malicious insider or a compromised firmware-management pipeline already has.","remediation":"Fixed in Supermicro's January 2025 BMC/IPMI firmware drop - a per-node out-of-band BMC flash, with the ordinary brick risk if interrupted. Beyond patching, this argues for treating BMC firmware update authority as a privileged operation in its own right: restrict which systems hold BMC admin credentials, require change control on firmware pushes, and verify BMC flash contents out-of-band after updates rather than trusting the BMC's own report of what it is running.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-10239","https://www.supermicro.com/en/support/security_BMC_IPMI_Jan_2025","https://eclypsium.com/wp-content/uploads/OpenBMC-Security-in-Practice.pdf"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2025-02-04"},{"id":"CVE-2024-21820","cve":"CVE-2024-21820","aliases":[],"title":"Intel Xeon memory controller configuration (with SGX): Incorrect default permissions on Xeon memory controller","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Xeon memory controller configuration (with SGX)","year":"2024","cvss_score":7.2,"severity":"high","kev":false,"impact":"Incorrect default permissions on Xeon memory controller configuration when SGX is in use, reachable by a privileged local user for escalation. Same advisory family as the conditions check issue and fixed by the same platform update.","attack_vector":"Privileged local access on the host.","remediation":"OEM platform BIOS update, drain and reboot, then re-attest enclaves.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21820","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01079.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2024-11-13"},{"id":"CVE-2024-22095","cve":"CVE-2024-22095","aliases":[],"title":"Intel Server D50DNP UEFI firmware (PlatformVariableInitDxe): Improper input validation in a UEFI DXE driver on Intel","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Server D50DNP UEFI firmware (PlatformVariableInitDxe)","year":"2024","cvss_score":7.2,"severity":"high","kev":false,"impact":"Improper input validation in a UEFI DXE driver on Intel Server D50DNP boards gives a privileged local user escalation into firmware. D50DNP is a dense datacenter server board, so this is squarely an AI-datacenter platform. Firmware-level escalation means persistence below the OS that survives reimaging.","attack_vector":"Privileged local access on the host.","remediation":"Fixed in platform BIOS/UEFI firmware. That means an OEM release, a per-node drain, a flash and a cold reboot - and OEM availability commonly lags the Intel advisory by quarters on server boards. There is no microcode or OS-level shortcut for this class; budget it as a fleet-wide maintenance campaign, not a patch.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-22095","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01080.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-16"},{"id":"CVE-2024-22453","cve":"CVE-2024-22453","aliases":[],"title":"Dell PowerEdge Server BIOS (heap-based buffer overflow): A high-privileged local attacker writes to memory it should","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell PowerEdge Server BIOS (heap-based buffer overflow)","year":"2024","cvss_score":7.2,"severity":"high","kev":false,"impact":"A high-privileged local attacker writes to memory it should not reach, with scope change - firmware-level memory corruption on the server platform.","attack_vector":"Local high-privilege access to the host.","remediation":"Flash the fixed PowerEdge BIOS. A BIOS update is a cold reboot per node and cannot be done live - on a GPU fleet that means draining jobs and taking the box out of the scheduler, so batch it with other firmware work rather than doing a standalone pass.","references":["https://www.dell.com/support/kbdoc/en-us/000223209/dsa-2024-105-security-update-for-dell-poweredge-server-bios-for-a-heap-based-buffer-overflow-vulnerability"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-2659","cve":"CVE-2024-2659","aliases":[],"title":"Lenovo ThinkSystem SMM / SMM2 and FPC (command injection): An authenticated user with elevated privileges executes","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Lenovo ThinkSystem SMM / SMM2 and FPC (command injection)","year":"2024","cvss_score":7.2,"severity":"high","kev":false,"impact":"An authenticated user with elevated privileges executes system commands on the chassis System Management Module or Fan/Power Controller - the components that own power and thermal control for a whole enclosure of ThinkSystem nodes.","attack_vector":"Authenticated high-privilege access to the SMM/SMM2 or FPC management interface.","remediation":"Apply the Lenovo firmware update per LEN-140420. Chassis-level management firmware flash; nodes keep running but chassis management is interrupted during the update.","references":["https://support.lenovo.com/us/en/product_security/LEN-140420"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-28248","cve":"CVE-2024-28248","aliases":[],"title":"Cilium: HTTP policies not consistently applied to all traffic","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2024","cvss_score":7.2,"severity":"high","kev":false,"impact":"HTTP policies not consistently applied to all traffic; L7 policy bypass","attack_vector":"Any pod on the cluster network","remediation":"Rolling Cilium upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-28248"],"status":"curated","published":"2024-03-18"},{"id":"CVE-2024-38509","cve":"CVE-2024-38509","aliases":["LEN-156781"],"title":"Lenovo XClarity Controller (XCC) - IPMI command handler: A specially crafted IPMI command gives an authenticated XCC","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lenovo XClarity Controller (XCC) - IPMI command handler","year":"2024","cvss_score":7.2,"severity":"high","kev":false,"impact":"A specially crafted IPMI command gives an authenticated XCC user arbitrary code execution on the controller. Code execution on the BMC is the terminal outcome for a node: the attacker holds power control, Virtual Media, console and the firmware write path, and can plant an implant that lives below the hypervisor and survives every reimage. It is one of a cluster of XCC command-injection issues Lenovo fixed across 2024 reachable through IPMI, the SSH captive shell and file upload - the IPMI path is the one that matters most because IPMI is so often left enabled for legacy automation.","attack_vector":"An authenticated XCC user with elevated privileges sending IPMI commands - so the exposure is your administrative BMC credentials plus anything on the management VLAN that can reach the IPMI service. Compromise of an automation host that holds XCC admin creds is the realistic path.","remediation":"Flash XCC to the per-model version in LEN-156781 - out-of-band, per-node, no host reboot and no drain. The high-value config-only mitigation is to disable IPMI over LAN on XCC where your tooling has moved to Redfish, which removes this entire command surface rather than fixing one handler in it. Budget for migrating any remaining ipmitool-based automation first.","references":["https://support.lenovo.com/us/en/product_security/LEN-156781","https://nvd.nist.gov/vuln/detail/CVE-2024-38509"],"status":"curated","published":"2024-07-26"},{"id":"CVE-2024-41942","cve":"CVE-2024-41942","aliases":[],"title":"JupyterHub: A user granted limited access can escalate","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"JupyterHub","year":"2024","cvss_score":7.2,"severity":"high","kev":false,"impact":"A user granted limited access can escalate","attack_vector":"Authenticated notebook user on a shared hub","remediation":"Upgrade to 4.1.6/5.1.0+","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-41942"],"status":"curated","published":"2024-08-08"},{"id":"CVE-2024-42442","cve":"CVE-2024-42442","aliases":["AMI-SA-2024004"],"title":"AMI AptioV UEFI BIOS (SMM): A memory-bounds bug in the BIOS that lets an attacker execute code outside the intended","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI AptioV UEFI BIOS (SMM)","year":"2024","cvss_score":7.2,"severity":"high","kev":false,"impact":"A memory-bounds bug in the BIOS that lets an attacker execute code outside the intended System Management Mode. What sets it apart from the rest of the AptioV set is the attack vector: AMI scores it as network-reachable, not local. A firmware bug you can reach over the wire is a different risk class from one that needs host root first - it means the BIOS attack surface is exposed through a management path rather than only through the host OS, and network segmentation of that path becomes load-bearing.","attack_vector":"Network, with high privileges required - an administrative account on whichever management path exposes the BIOS operation. In practice that means a BMC or out-of-band management credential, which is why this bug chains so naturally with the MegaRAC credential and REST API issues in this same cluster: BMC admin access becomes host firmware code execution.","remediation":"BIOS update to BKC_5.37 or later - firmware flash plus a host reboot, per node, vendor-rebase-gated. Because the vector is network with privileges, there is real config-only mitigation available now: isolate the BMC and management plane so that no untrusted party can reach the privileged management interface, and give every node unique management credentials so one leak does not reach the fleet.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/2024/AMI-SA-2024004.pdf","https://nvd.nist.gov/vuln/detail/CVE-2024-42442"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-11-12"},{"id":"CVE-2024-56161","cve":"CVE-2024-56161","aliases":["EntrySign"],"title":"AMD Zen microcode patch loader (CPU ROM signature verification): The CPU ROM's microcode patch loader verified patch","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Zen microcode patch loader (CPU ROM signature verification)","year":"2024","cvss_score":7.2,"severity":"high","kev":false,"impact":"The CPU ROM's microcode patch loader verified patch signatures using AES-CMAC with a published example key, so an attacker with local administrator privilege can sign and load their own microcode onto Zen 1 through Zen 4 CPUs. Loading arbitrary microcode means redefining what x86 instructions do - the attacker can make RDRAND return a constant, disable checks, or backdoor the CPU beneath every layer of software. For a confidential-computing operator this is the ballgame: SEV-SNP's entire guarantee rests on the CPU behaving as specified, so a host administrator can now read and tamper with any SEV-SNP guest's memory while the attestation report still looks clean. Google's security team demonstrated the full chain.","attack_vector":"Local, requires ring-0 / host administrator privilege. Not remote, and not reachable from a tenant container. But 'host administrator' is exactly the adversary SEV-SNP exists to exclude, which is why a local-admin bug is a confidential-computing catastrophe rather than a routine escalation. Persistence note: microcode does not survive a power cycle, so an attacker must reload it each boot - which also means a cold boot clears an implant.","remediation":"Fixed by an AMD microcode patch. Two delivery routes, and the difference matters: the linux-firmware amd-ucode blobs load early at boot (initramfs) and need only a reboot, while the durable fix is the microcode embedded in the OEM SBIOS/AGESA package, which carries the usual one-to-six-month OEM lag and a full power cycle. **For confidential computing you need the SBIOS route**: microcode late-loaded by the OS is not part of what SEV-SNP attests, so a guest checking the attestation report cannot tell the fix is present. AMD does not support late-loading microcode on a running EPYC host - treat this as reboot-required. After patching, expect the reported TCB version to change and plan the VCEK certificate refresh accordingly. Specifically: AMD shipped fixed microcode plus an updated ASP bootloader in AGESA (December 2024 / released publicly February 2025). Zen 1-4 EPYC and Ryzen are affected; Zen 5 is not. The critical operator step people skip is the attestation side - after the TCB bump, refresh VCEK certs and require tenants to pin the new minimum TCB, otherwise you are patched but still accepting attestations that would have been valid on a backdoored host.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-56161","https://github.com/google/security-research/security/advisories/GHSA-4xq7-4mgh-gp6w","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2025-02-03"},{"id":"CVE-2024-8531","cve":"CVE-2024-8531","aliases":["SEVD-2024-282-01"],"title":"Schneider Electric Data Center Expert - upgrade bundle signature verification: Improper cryptographic signature","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Schneider Electric Data Center Expert - upgrade bundle signature verification","year":"2024","cvss_score":7.2,"severity":"high","kev":false,"impact":"Improper cryptographic signature verification on DCE upgrade bundles: a manipulated bundle can carry arbitrary bash scripts that execute as root. Your patching process becomes the attack. Anyone who can place a bundle in front of the appliance - a compromised mirror, an internal file share, a helpful vendor email - gets root on the system holding the facility's power and cooling credentials.","attack_vector":"Requires getting a crafted upgrade bundle to the appliance. In practice: whoever runs DCE upgrades, or anyone who can tamper with where the bundles are staged.","remediation":"Upgrade DCE per SEVD-2024-282-01 to a build that verifies signatures correctly. Until then, treat upgrade bundles as untrusted code: obtain them only over an authenticated channel from the vendor, verify hashes out of band, and stage them somewhere with restricted write access.","references":["https://download.schneider-electric.com/files?p_Doc_Ref=SEVD-2024-282-01&p_enDocType=Security+and+Safety+Notice&p_File_Name=SEVD-2024-282-01.pdf"],"status":"curated","published":"2024-10-11"},{"id":"CVE-2024-9180","cve":"CVE-2024-9180","aliases":[],"title":"HashiCorp Vault: Operator with write on the root namespace identity endpoint escalates self/others to the root policy","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HashiCorp Vault","year":"2024","cvss_score":7.2,"severity":"high","kev":false,"impact":"Operator with write on the root namespace identity endpoint escalates self/others to the root policy","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; re-scope operator policies","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-9180"],"status":"curated","published":"2024-10-10"},{"id":"CVE-2024-9474","cve":"CVE-2024-9474","aliases":[],"title":"Palo Alto PAN-OS: Admin with mgmt-interface access performs firewall actions as root","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Palo Alto PAN-OS","year":"2024","cvss_score":7.2,"severity":"high","kev":true,"impact":"Admin with mgmt-interface access performs firewall actions as root; chained with CVE-2024-0012","attack_vector":"Network (remote)","remediation":"Control-plane: same patch window; assume compromise if mgmt was exposed","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-9474"],"status":"curated","published":"2024-11-18"},{"id":"CVE-2025-0032","cve":"CVE-2025-0032","aliases":[],"title":"AMD CPU microcode patch loading - improper cleanup: Improper cleanup during microcode patch loading gives a local","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD CPU microcode patch loading - improper cleanup","year":"2025","cvss_score":7.2,"severity":"high","kev":false,"impact":"Improper cleanup during microcode patch loading gives a local administrator another route to load malicious CPU microcode, costing integrity of x86 instruction execution itself. This is the same category of failure as EntrySign and lands in the same place: the CPU can be made to lie, and everything built on top of it - SEV-SNP guest isolation included - inherits the lie.","attack_vector":"Local, administrator privilege. Reboot clears a loaded implant, but the attacker who has admin can simply reload it every boot.","remediation":"Fixed by an AMD microcode patch. Two delivery routes, and the difference matters: the linux-firmware amd-ucode blobs load early at boot (initramfs) and need only a reboot, while the durable fix is the microcode embedded in the OEM SBIOS/AGESA package, which carries the usual one-to-six-month OEM lag and a full power cycle. **For confidential computing you need the SBIOS route**: microcode late-loaded by the OS is not part of what SEV-SNP attests, so a guest checking the attestation report cannot tell the fix is present. AMD does not support late-loading microcode on a running EPYC host - treat this as reboot-required. After patching, expect the reported TCB version to change and plan the VCEK certificate refresh accordingly.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0032","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2025-09-06"},{"id":"CVE-2025-12006","cve":"CVE-2025-12006","aliases":[],"title":"Supermicro BMC firmware validation logic on the X12STW-F motherboard: An attacker with administrative reach to the BMC","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC firmware validation logic on the X12STW-F motherboard","year":"2025","cvss_score":7.2,"severity":"high","kev":false,"impact":"An attacker with administrative reach to the BMC installs a firmware image of their own construction. What they walk away with is a controller that owns node power, console, virtual media and the host's boot path, and that keeps owning it after the operator wipes and reprovisions the machine. Because BMC credentials in most fleets are identical across every node, one admin-level compromise converts directly into a fleet-wide persistent foothold that no host-level EDR or reimage cycle will find. The same class of image-verification weakness as its X13 sibling, showing the flaw spans two board generations rather than one SKU.","attack_vector":"An authenticated BMC session at administrator privilege reachable over the network. That bar is lower than it sounds in real datacenters: shared or default BMC credentials, a leaked provisioning secret, or chaining any of the several authenticated Supermicro BMC command-execution bugs in this same list will get an attacker there.","remediation":"Firmware flash per node using the fixed image from Supermicro's January 2026 BMC/IPMI batch, matched to the exact board SKU. Config-only mitigation is partial but worth doing immediately: rotate BMC administrator credentials so they are unique per node rather than shared fleet-wide, and remove any standing admin accounts used by automation in favour of scoped operator-level accounts. Network isolation of the management VLAN remains the backstop for nodes that cannot be taken down for a flash.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-12006","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2025/12xxx/CVE-2025-12006.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-20037","cve":"CVE-2025-20037","aliases":[],"title":"Intel CSME firmware (TOCTOU): A time-of-check/time-of-use race in CSME firmware lets a privileged local user escalate","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel CSME firmware (TOCTOU)","year":"2025","cvss_score":7.2,"severity":"high","kev":false,"impact":"A time-of-check/time-of-use race in CSME firmware lets a privileged local user escalate into the management engine. Recent, and a reminder that the CSME attack surface is still producing findings on current platforms.","attack_vector":"Privileged local access on the host, plus winning a race.","remediation":"Fixed in Intel CSME/SPS firmware, which reaches you as an OEM BIOS or firmware package - not as a microcode or OS update. That means: wait for your server vendor to ship it, drain the node, flash, and reboot. OEM availability is the long pole and routinely lags the Intel advisory by one or more quarters on server platforms. Track it per platform SKU, because vendors ship these unevenly across their own product lines.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-20037","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01280.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-08-12"},{"id":"CVE-2025-20053","cve":"CVE-2025-20053","aliases":[],"title":"Intel Xeon processor firmware (SGX enabled): Improper buffer restrictions in Xeon firmware on SGX-enabled parts, giving","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Xeon processor firmware (SGX enabled)","year":"2025","cvss_score":7.2,"severity":"high","kev":false,"impact":"Improper buffer restrictions in Xeon firmware on SGX-enabled parts, giving a privileged local user an escalation path. Firmware-level, so it lands underneath anything the OS can defend.","attack_vector":"Privileged local access on the host.","remediation":"Platform firmware/BIOS update from the OEM plus SGX TCB recovery. Full drain and reboot per node; OEM release timing dominates the rollout.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-20053","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01313.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2025-08-12"},{"id":"CVE-2025-26403","cve":"CVE-2025-26403","aliases":[],"title":"Intel Xeon 6 memory subsystem (with SGX or TDX): An out-of-bounds write in the Xeon 6 memory subsystem reachable when","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Xeon 6 memory subsystem (with SGX or TDX)","year":"2025","cvss_score":7.2,"severity":"high","kev":false,"impact":"An out-of-bounds write in the Xeon 6 memory subsystem reachable when SGX or TDX is enabled, escalating privilege for a privileged local user. An OOB write in the memory subsystem of a confidential-compute platform undercuts both the SGX and the TDX guarantee on the same silicon.","attack_vector":"Privileged local access on a Xeon 6 host with SGX or TDX enabled.","remediation":"OEM platform firmware/BIOS update, plus TDX module and SGX TCB recovery as applicable. Drain and reboot; re-attest every TD and enclave afterwards.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-26403","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01367.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2025-08-12"},{"id":"CVE-2025-32086","cve":"CVE-2025-32086","aliases":[],"title":"Intel Xeon 6 DDRIO configuration (with SGX or TDX): An improperly implemented security check in DDRIO configuration on","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Xeon 6 DDRIO configuration (with SGX or TDX)","year":"2025","cvss_score":7.2,"severity":"high","kev":false,"impact":"An improperly implemented security check in DDRIO configuration on Xeon 6 lets a privileged local user escalate on SGX/TDX-enabled platforms. DDRIO sits between the memory controller and DRAM, which is the layer memory-encryption integrity depends on.","attack_vector":"Privileged local access on a Xeon 6 host with SGX or TDX enabled.","remediation":"OEM platform firmware/BIOS update plus TCB recovery for both SGX and TDX. Drain, reboot, re-attest.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-32086","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01367.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2025-08-12"},{"cvss_vector":"CVSS:4.0/AV:N/AC:L/AT:P/PR:L/UI:N/VC:N/VI:H/VA:H/SC:N/SI:H/SA:H","cwe":["CWE-770"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2025-32777","cve":"CVE-2025-32777","aliases":["GHSA-hg79-fw4p-25p8"],"title":"Volcano (scheduler, Elastic service and extender plugin response handling): The scheduler reads unbounded responses","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Volcano (scheduler, Elastic service and extender plugin response handling)","year":"2025","cvss_score":7.2,"severity":"high","kev":false,"impact":"The scheduler reads unbounded responses from the Elastic service and from extender plugins. Since operators commonly run those in separate pods or on separate nodes, compromising one lets an attacker cross the node isolation boundary and take down the shared Volcano scheduler - which stops GPU job placement for every tenant on the cluster.","attack_vector":"An attacker who has already compromised the Elastic service or an extender plugin process, typically running on a different pod or node than the scheduler.","remediation":"Upgrade Volcano to 1.9.1, 1.10.2, 1.11.2, 1.11.0-network-topology-preview.3 or 1.12.0-alpha.2 depending on your track, and restart the scheduler. Also treat extender plugin endpoints as untrusted input and keep them on a restricted network path.","references":["https://github.com/volcano-sh/volcano/security/advisories/GHSA-hg79-fw4p-25p8","https://nvd.nist.gov/vuln/detail/CVE-2025-32777"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-58770","cve":"CVE-2025-58770","aliases":["AMI-SA-2025009"],"title":"AMI AptioV UEFI BIOS: Improper handling of insufficient permissions in the BIOS lets a low-privileged local user","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI AptioV UEFI BIOS","year":"2025","cvss_score":7.2,"severity":"high","kev":false,"impact":"Improper handling of insufficient permissions in the BIOS lets a low-privileged local user escalate their authorization, with integrity and availability impact that AMI scores as reaching the subsequent system too. What makes this one worth prioritising over its neighbours is that AMI's own CVSS vector marks exploit maturity as proof-of-concept - meaning working exploit code exists publicly, not just a theoretical write-up. It is also the newest entry in AMI's published series, so ODM rebased images are the least likely to be available.","attack_vector":"Local access with only low privileges required and no user interaction. That is a notably low bar for a firmware bug - it does not need root, so an unprivileged process or a compromised service account on the host is enough to start escalating toward firmware.","remediation":"BIOS update to AptioV_5.041 or later: firmware flash plus a full host reboot, per node, and expect the longest ODM lag of anything in this cluster because the advisory is recent. No config-only fix. Given the low privilege requirement, the interim control is ordinary host hardening - reduce what unprivileged local code exists on GPU nodes at all, and treat any node where untrusted tenant code runs as already exposed until the BIOS is updated.","references":["https://go.ami.com/hubfs/Security%20Advisories/2025/AMI-SA-2025009.pdf","https://nvd.nist.gov/vuln/detail/CVE-2025-58770"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2025-12-12"},{"id":"CVE-2025-6198","cve":"CVE-2025-6198","aliases":[],"title":"Supermicro BMC firmware validation (MBD-X13SEM-F): Second-generation RoT bypass","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC firmware validation (MBD-X13SEM-F)","year":"2025","cvss_score":7.2,"severity":"high","kev":false,"impact":"Second-generation RoT bypass — crafted image passes BMC firmware validation logic (CWE-347, improper signature verification); survives reimaging","attack_vector":"Network, high-privilege","remediation":"Per-node out-of-band BMC flash to the fixed Supermicro build; no host-side mitigation exists","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-6198"],"status":"curated","fleet":{"ubiquity":"very common - affects an additional Supermicro product set beyond CVE-2024-10237","remediation_pain":"firmware-flash - out-of-band per node; once RoT is defeated, prior attestation evidence is untrustworthy and nodes need re-baselining","pain_class":"firmware-flash","why_fleet_wide":"Bypasses the BMC Root of Trust, so the firmware-signing anchor a neocloud relies on for tenant-isolation claims is defeated below the OS."},"published":"2025-09-19"},{"id":"CVE-2025-62626","cve":"CVE-2025-62626","aliases":[],"title":"AMD CPUs - attacker influence over RDSEED entropy: A local attacker can influence the values RDSEED returns, causing","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD CPUs - attacker influence over RDSEED entropy","year":"2025","cvss_score":7.2,"severity":"high","kev":false,"impact":"A local attacker can influence the values RDSEED returns, causing consumers to draw insufficient entropy. Anything on the node that seeds a key, nonce or token from the hardware RNG - including confidential guests that deliberately chose the hardware source because they do not trust the host - gets attacker-influenced material. The failure is silent: the instruction reports success, and nothing downstream can tell the difference until someone else predicts the key.","attack_vector":"Local. The attacker influences entropy consumed by other software on the same machine, including guests.","remediation":"Mitigated by AMD microcode plus, on most of these, a kernel-side change - and the durable delivery vehicle is the OEM SBIOS/AGESA package, which carries **one to six months of OEM lag** and needs a drained node and a full power cycle. The linux-firmware amd-ucode blobs get you the microcode sooner via initramfs early-load and a reboot, but AMD does not support late-loading microcode on a running EPYC host, so either way this is reboot-required, not a live patch. Also update guest and host kernels so the OS entropy pool does not lean solely on RDSEED. Any long-lived key generated on an affected host before patching should be rotated - the fix protects future output, not keys already derived.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-62626","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2025-11-21"},{"id":"CVE-2025-67818","cve":"CVE-2025-67818","aliases":[],"title":"Weaviate: Crafted entry name with an absolute path","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Weaviate","year":"2025","cvss_score":7.2,"severity":"high","kev":false,"impact":"Crafted entry name with an absolute path → arbitrary file write","attack_vector":"Tenant with data-insert permission","remediation":"Upgrade past 1.33.4; data insertion is a filesystem-write primitive","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-67818"],"status":"curated","published":"2025-12-12"},{"id":"CVE-2025-7937","cve":"CVE-2025-7937","aliases":[],"title":"Supermicro BMC firmware validation (MBD-X12STW): RoT bypass, crafted firmware image accepted","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC firmware validation (MBD-X12STW)","year":"2025","cvss_score":7.2,"severity":"high","kev":false,"impact":"RoT bypass, crafted firmware image accepted; demonstrates the 2024 fix was incomplete","attack_vector":"Network, high-privilege","remediation":"Second BMC flash cycle on nodes already patched for CVE-2024-10237 — the operational cost is that a fleet gets flashed twice for the same class of bug","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-7937"],"status":"curated","fleet":{"ubiquity":"very common - same Supermicro BMC image family","remediation_pain":"firmware-flash - a second flash cycle on nodes already flashed once; the first remediation round did not hold","pain_class":"firmware-flash","why_fleet_wide":"Shows fleet firmware remediation is not one-and-done: the patch was bypassed, forcing a repeat out-of-band flash campaign across the same node population."},"published":"2025-09-19"},{"id":"CVE-2025-8076","cve":"CVE-2025-8076","aliases":[],"title":"Supermicro BMC web server request handling on MBD-X13SEDW-F: Any account that can log into the BMC web interface can","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC web server request handling on MBD-X13SEDW-F","year":"2025","cvss_score":7.2,"severity":"high","kev":false,"impact":"Any account that can log into the BMC web interface can turn that login into code execution on the controller. For an operator this collapses the distinction between 'someone has a BMC password' and 'someone owns the node out of band' - they get power control, console, virtual-media boot of an attacker image, and a persistence point below the hypervisor that reimaging will not clear. A post-authentication stack buffer overflow triggered by a crafted payload to the management web UI.","attack_vector":"An authenticated high-privilege session against the BMC's HTTP interface, reachable from anywhere routable to the out-of-band management network. In fleets that share one BMC password across every node - still the norm - a single credential leak makes this exploitable everywhere at once.","remediation":"Firmware flash from Supermicro's November 2025 BMC/IPMI advisory. Config-only measures that reduce exposure now: put BMC web access behind a bastion so it is not reachable from the general management subnet, rotate to per-node BMC credentials, and disable the BMC web UI on nodes managed purely via Redfish or IPMI. Flashing is per node and out of band; plan it alongside the other X13 BMC fixes in the same advisory so you only take one flash window.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-8076","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2025/8xxx/CVE-2025-8076.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-8727","cve":"CVE-2025-8727","aliases":[],"title":"Supermicro BMC web interface (stack buffer overflow, X13SEDW-F): Second authenticated stack overflow in the BMC web","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC web interface (stack buffer overflow, X13SEDW-F)","year":"2025","cvss_score":7.2,"severity":"high","kev":false,"impact":"Second authenticated stack overflow in the BMC web function, giving full BMC compromise.","attack_vector":"Authenticated BMC web access with high privilege.","remediation":"Flash the fixed Supermicro BMC firmware for the affected board SKU. BMC flash does not require a host reboot, but it does drop out-of-band management for several minutes per node - script it and stagger it so you never lose OOB across a whole rack at once. Check your exact MBD- part number against the advisory; Supermicro scopes these narrowly. Ships in the same November 2025 firmware batch as the other CoreWeave-reported findings.","references":["https://www.supermicro.com/en/support/security_BMC_IPMI_Nov_2025"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2026-23759","cve":"CVE-2026-23759","aliases":[],"title":"Perle IOLAN STS/SCS terminal server (firmware before 6.0): A logged-in user of the restricted admin shell (Telnet","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Perle IOLAN STS/SCS terminal server (firmware before 6.0)","year":"2026","cvss_score":7.2,"severity":"high","kev":false,"impact":"A logged-in user of the restricted admin shell (Telnet or SSH) can break out of that shell and run arbitrary OS commands as root. The 'ps' subcommand doesn't sanitize its arguments before handing them to a shell, so an operator account that was only supposed to have limited diagnostic access ends up with full root on the terminal server.","attack_vector":"Requires an authenticated login to the restricted shell (any account that can reach the 'ps' command), then injects shell metacharacters after the subcommand.","remediation":"Firmware upgrade to 6.0 or later — this is a shell-sanitization bug in the restricted-admin feature, not something a permission change alone fixes. Flash each terminal server one at a time; expect a brief loss of the serial sessions it's terminating during the reboot.","references":["https://www.perle.com/downloads/server_sds_sts_rackmount.shtml","https://www.vulncheck.com/advisories/perle-iolan-sts-scs-authenticated-command-injection-via-shell-ps"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2026-03-17"},{"id":"CVE-2026-3820","cve":"CVE-2026-3820","aliases":[],"title":"Supermicro BMC SMTP service configuration handler on AS-2115HS-TNR and related boards: Crafted characters injected","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC SMTP service configuration handler on AS-2115HS-TNR and related boards","year":"2026","cvss_score":7.2,"severity":"high","kev":false,"impact":"Crafted characters injected into the SMTP configuration are executed by the underlying BMC system when the mail process is invoked. The attacker gets arbitrary code execution or a hard denial of service on the controller, and Supermicro's own wording allows for permanent compromise of the controller - meaning an implant that persists in BMC flash. For an operator this is the same nightmare as any BMC RCE: out-of-band power and console control, virtual-media boot of an attacker image, and a foothold below the reimage boundary. The 2026 recurrence of the same notification-service injection class that produced CVE-2023-35861 three years earlier.","attack_vector":"An attacker who has obtained BMC administrator privileges and can reach the BMC's configuration interface over the network. Shared fleet-wide BMC credentials, leaked provisioning secrets, or a chained lower-privilege bug all put an attacker at this level.","remediation":"Firmware flash from Supermicro's June 2026 BMC/IPMI advisory batch, per board SKU. Config-only interim mitigation: disable BMC SMTP alerting and rotate BMC admin credentials to per-node unique values. The fact that this is the second SMTP-handler injection in this codebase in three years is itself operator-relevant - treat the BMC notification subsystem as untrusted attack surface and disable it on nodes that get their alerting from IPMI polling or Redfish instead.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-3820","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2026/3xxx/CVE-2026-3820.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2026-42306","cve":"CVE-2026-42306","aliases":[],"title":"Docker / moby: Race condition during `docker cp` mount setup allows escape/host access","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2026","cvss_score":7.2,"severity":"high","kev":false,"impact":"Race condition during `docker cp` mount setup allows escape/host access","attack_vector":"Any tenant workload on a node where docker cp is used","remediation":"Upgrade Docker Engine to 29.5.1+; daemon restart","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-42306"],"status":"curated","published":"2026-06-12"},{"id":"CVE-2026-6973","cve":"CVE-2026-6973","aliases":[],"title":"Ivanti Endpoint Manager Mobile: Improper input validation","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ivanti Endpoint Manager Mobile","year":"2026","cvss_score":7.2,"severity":"high","kev":true,"impact":"Improper input validation -> authenticated admin achieves remote code execution","attack_vector":"Network (remote)","remediation":"Control-plane: patch; restrict admin access to the ops network","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-6973"],"status":"curated","published":"2026-05-07"},{"id":"CVE-2026-9777","cve":"CVE-2026-9777","aliases":["ZDI-26-381"],"title":"ATEN Unizon fleet management platform: Unizon is ATEN's centralized manager for its KVM and PDU fleet. The restoreDB","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"ATEN Unizon fleet management platform","year":"2026","cvss_score":7.2,"severity":"high","kev":false,"impact":"Unizon is ATEN's centralized manager for its KVM and PDU fleet. The restoreDB function doesn't validate a user-supplied path before writing to it, letting an authenticated attacker write files anywhere on the host and execute code with SYSTEM privileges — full compromise of the platform that has management-plane reach into every KVM and PDU it administers.","attack_vector":"Requires an authenticated account on Unizon (the advisory doesn't specify a high privilege tier is needed), then sends a crafted path to the restoreDB endpoint.","remediation":"Software upgrade to the patched Unizon release. Since Unizon is the single management server for the whole device fleet rather than per-device firmware, this is one upgrade — but treat it as urgent given the blast radius (SYSTEM-level access to the platform that manages every connected KVM/PDU).","references":["https://www.aten.com/global/en/supportcenter/info/security-advisory/30/","https://www.zerodayinitiative.com/advisories/ZDI-26-381/"],"status":"curated","tags":["tenant-isolation"],"published":"2026-06-24"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:C/C:L/I:L/A:N","fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2026-047-cilium-vlan-interface-datapath-t","cve":null,"aliases":["GHSA-vh48-r624-p8v7"],"title":"Cilium (VLAN interface datapath, TCX attachment): Ingress host policies and L7 network policies silently stop being","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium (VLAN interface datapath, TCX attachment)","year":"2026","cvss_score":7.2,"severity":"high","kev":false,"impact":"Ingress host policies and L7 network policies silently stop being enforced on nodes where Cilium is attached to both a VLAN interface and its parent. The datapath misreads packets arriving on the VLAN interface as having already traversed the proxy and the host firewall, so it skips enforcement at the destination endpoint. An operator reading their policy objects sees deny rules that look correct and are not running. Anyone who can send traffic to a pod on such a node reaches it regardless of policy — from outside the node, unauthenticated. VLAN-per-tenant is a common physical-network design in GPU colocation and bare-metal neocloud builds, and TCX attachment is on by default on kernel 6.2 and newer, so the affected configuration is the modern one rather than an exotic one. Same-node pod traffic and non-host L3/L4 policies are not affected, which narrows the blast radius but also makes the gap easy to miss in testing.","attack_vector":"Network, unauthenticated, from outside the node. Requires the node to have Cilium attached to both a VLAN interface and its parent interface, with TCX attachment enabled.","remediation":"Upgrade to Cilium 1.18.9 or 1.17.15. If you cannot upgrade now, disable TCX with the agent flag --enable-tcx=false or bpf.enableTCX: false in the Helm chart, accepting the performance change. Egress policies on the source side, where configured to deny, may still hold and are worth checking as interim cover. Verify enforcement empirically on VLAN-attached nodes rather than trusting the policy objects.","references":["https://github.com/cilium/cilium/security/advisories/GHSA-vh48-r624-p8v7"],"status":"curated"},{"id":"CVE-2019-0090","cve":"CVE-2019-0090","aliases":["Intel x86 Root of Trust: loss of trust"],"title":"Intel CSME / Converged Security and Management Engine (mask ROM): A flaw in the CSME boot ROM window before memory","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel CSME / Converged Security and Management Engine (mask ROM)","year":"2019","cvss_score":7.1,"severity":"high","kev":false,"impact":"A flaw in the CSME boot ROM window before memory protections engage allows code execution in the engine that is the hardware root of trust for the whole platform - it is what verifies BIOS under Boot Guard, backs Intel PTT/fTPM, and holds the chipset key from which platform-unique keys derive. Researchers demonstrated extraction of that chipset key, which forges the identity of the platform itself: EPID-based attestation, DRM, and any measurement chain rooted in CSME become untrustworthy, and the compromise is not visible from the OS at all.","attack_vector":"Local or physical access to the machine during the early boot window. On bare-metal GPU rental, a tenant with root plus a reboot is the realistic actor; on a colo floor, so is anyone with hands.","remediation":"Cannot be fully fixed. The vulnerable code is in mask ROM, so no firmware update replaces it - Intel's CSME updates only narrow the exploitation window. The durable answer is hardware generations that are not affected, and until then treating CSME-rooted attestation as advisory rather than authoritative. If you sell attestation guarantees, do not root them here; root them in a discrete device you control.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-0090","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00213.html"],"status":"curated","published":"2019-05-17"},{"id":"CVE-2019-5687","cve":"CVE-2019-5687","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): A kernel object created by the escape handler gets default","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2019","cvss_score":7.1,"severity":"high","kev":false,"impact":"A kernel object created by the escape handler gets default permissions that expose it to accounts that should not reach it. Local privilege boundary erosion inside the driver.","attack_vector":"Any local user on the host who can open the over-permissive object.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://support.lenovo.com/us/en/product_security/LEN-28096","https://nvd.nist.gov/vuln/detail/CVE-2019-5687"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2019-08-06"},{"id":"CVE-2019-5697","cve":"CVE-2019-5697","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): The vGPU Manager grants a guest access to memory the guest does not own. That is the","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2019","cvss_score":7.1,"severity":"high","kev":false,"impact":"The vGPU Manager grants a guest access to memory the guest does not own. That is the core vGPU isolation promise failing - a tenant VM reads memory belonging to the host or to another tenant's vGPU.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU on the affected host.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-5697"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2019-11-09"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:N","cwe":["CWE-522"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2020-27781","cve":"CVE-2020-27781","aliases":[],"title":"CephFS (via OpenStack Manila native driver): A Manila user can request access for an existing CephFS identity and get","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"CephFS (via OpenStack Manila native driver)","year":"2020","cvss_score":7.1,"severity":"high","kev":false,"impact":"A Manila user can request access for an existing CephFS identity and get that identity's credentials handed back, which means stealing another tenant's CephFS key. With that key the attacker mounts and reads/writes shares that belong to somebody else.","attack_vector":"Any tenant able to issue Manila share-access requests against a cluster using the native CephFS driver.","remediation":"Apply the Ceph and Manila updates that scope credential creation to the requesting project, restart the manila-share and ceph-mgr volumes module, then rotate every CephFS auth ID that Manila created before the fix.","references":["https://access.redhat.com/security/cve/CVE-2020-27781","https://nvd.nist.gov/vuln/detail/CVE-2020-27781"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:C/C:N/I:N/A:H","cwe":["CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2020-4411","cve":"CVE-2020-4411","aliases":[],"title":"IBM Spectrum Scale kernel module: An unauthenticated local trigger takes down the Spectrum Scale kernel module and with","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"IBM Spectrum Scale kernel module","year":"2020","cvss_score":7.1,"severity":"high","kev":false,"impact":"An unauthenticated local trigger takes down the Spectrum Scale kernel module and with it filesystem access on the node. Scope is changed, so the blast radius is the whole node rather than the calling process.","attack_vector":"Local access to a node running Spectrum Scale 4.2.0.0-4.2.3.21 or 5.0.0.0-5.0.4.3. No credentials needed.","remediation":"Upgrade to the fixed level named in IBM's bulletin, rebuild the portability layer and reboot each node as it is drained.","references":["https://www.ibm.com/support/pages/node/6209002","https://nvd.nist.gov/vuln/detail/CVE-2020-4411"],"status":"curated"},{"id":"CVE-2020-5366","cve":"CVE-2020-5366","aliases":["DSA-2020-128"],"title":"Dell iDRAC9 (web interface, local file inclusion): A path-traversal / local-file-inclusion flaw lets a low-privilege","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC9 (web interface, local file inclusion)","year":"2020","cvss_score":7.1,"severity":"high","kev":false,"impact":"A path-traversal / local-file-inclusion flaw lets a low-privilege iDRAC user read files outside the intended directory on the BMC. The value to an attacker is escalation material - configuration, credentials and session data that upgrade a read-only monitoring account into real control of the service processor. This is the classic privilege-ladder rung: the account you handed to a monitoring agent becomes the account that owns the node's out-of-band plane.","attack_vector":"An authenticated low-privilege iDRAC operator or read-only account reaching the iDRAC web interface over the management VLAN. No host access needed.","remediation":"Flash iDRAC9 to 4.20.20.20 or later - out-of-band, per-node, no host reboot, no drain. There is no clean config-only mitigation for this one beyond tightening who holds iDRAC accounts at all, so treat it as a firmware campaign. Worth pairing with an audit of low-privilege iDRAC service accounts, which are usually shared and rarely rotated.","references":["https://www.dell.com/support/article/en-us/sln322125/dsa-2020-128-idrac-local-file-inclusion-vulnerability?lang=en","https://nvd.nist.gov/vuln/detail/CVE-2020-5366"],"status":"curated","published":"2020-07-09"},{"id":"CVE-2020-5970","cve":"CVE-2020-5970","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): Guest-supplied data size is not validated, letting a tenant tamper with host-side","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2020","cvss_score":7.1,"severity":"high","kev":false,"impact":"Guest-supplied data size is not validated, letting a tenant tamper with host-side state or take the GPU down for co-tenants. vGPU 8.x before 8.4, 9.x before 9.4, 10.x before 10.3.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5970"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2020-06-30"},{"id":"CVE-2020-5972","cve":"CVE-2020-5972","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): Uninitialised local pointers later freed in the vGPU plugin - a guest-triggered free","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2020","cvss_score":7.1,"severity":"high","kev":false,"impact":"Uninitialised local pointers later freed in the vGPU plugin - a guest-triggered free of an arbitrary pointer, which is a host memory-corruption primitive. vGPU 8.x before 8.4, 9.x before 9.4, 10.x before 10.3.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5972"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2020-06-30"},{"id":"CVE-2020-5983","cve":"CVE-2020-5983","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin + host kernel module): The host can be made to write outside the frame-buffer region","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin + host kernel module)","year":"2020","cvss_score":7.1,"severity":"high","kev":false,"impact":"The host can be made to write outside the frame-buffer region allocated to a guest. That is a direct breach of the vGPU frame-buffer partition - one tenant's writes landing in memory belonging to the host or another tenant's vGPU. vGPU 8.x before 8.5, 10.x before 10.4, and 11.0.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU on the affected host.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5983"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2020-10-02"},{"id":"CVE-2020-5985","cve":"CVE-2020-5985","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): Guest-supplied length not validated, allowing host-side data tampering or a","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2020","cvss_score":7.1,"severity":"high","kev":false,"impact":"Guest-supplied length not validated, allowing host-side data tampering or a shared-GPU outage. vGPU 8.x before 8.5, 10.x before 10.4, and 11.0.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5985"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2020-10-02"},{"id":"CVE-2020-5988","cve":"CVE-2020-5988","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): Double free in the vGPU plugin, guest-triggered. Host heap corruption or disclosure.","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2020","cvss_score":7.1,"severity":"high","kev":false,"impact":"Double free in the vGPU plugin, guest-triggered. Host heap corruption or disclosure. vGPU 8.x before 8.5, 10.x before 10.4, and 11.0.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5988"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2020-10-02"},{"id":"CVE-2021-1056","cve":"CVE-2021-1056","aliases":[],"title":"NVIDIA Linux GPU Display Driver (nvidia.ko): Nvidia.ko does not fully honour filesystem permissions when providing GPU","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Linux GPU Display Driver (nvidia.ko)","year":"2021","cvss_score":7.1,"severity":"high","kev":false,"impact":"Nvidia.ko does not fully honour filesystem permissions when providing GPU device-level isolation. This is the one that matters for containerised GPU fleets - it means the device-node permission model your container runtime relies on to give each container only its assigned GPUs is not actually enforced by the driver, so a container can reach GPUs it was not allocated. On a shared Kubernetes GPU node running multiple tenants' pods, that is a straight isolation break between tenants. Debian and Gentoo both shipped it as a security update, so distro-packaged fleets are in scope.","attack_vector":"Any container or local user on a Linux GPU host with access to some subset of the NVIDIA device nodes - which is every GPU workload.","remediation":"Install the fixed Linux GPU Display Driver branch. nvidia.ko / nvidia-uvm.ko cannot be replaced while any process holds a GPU, so plan a node drain: cordon the node, stop every CUDA job and GPU container, unload the modules or reboot, install, reload. Container runtimes that bind-mount the driver libraries (nvidia-container-toolkit) need restarting so running pods pick up the new userspace. No firmware flash. On multi-tenant nodes, do not treat device-node permissions or the container toolkit's device isolation as a security boundary until the driver is patched; one tenant per node or MIG-backed partitioning is the only reliable interim control.","references":["https://lists.debian.org/debian-lts-announce/2022/01/msg00013.html","https://security.gentoo.org/glsa/202310-02","https://nvd.nist.gov/vuln/detail/CVE-2021-1056"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2021-01-08"},{"id":"CVE-2021-1058","cve":"CVE-2021-1058","aliases":[],"title":"NVIDIA vGPU software (guest kernel-mode driver + vGPU plugin): Unvalidated input size across the guest kernel-mode","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software (guest kernel-mode driver + vGPU plugin)","year":"2021","cvss_score":7.1,"severity":"high","kev":false,"impact":"Unvalidated input size across the guest kernel-mode driver and vGPU plugin boundary - a tenant tampers with host-side data or crashes the shared GPU. vGPU 8.x before 8.6, 11.0 before 11.3.","attack_vector":"Any user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the host and the vGPU guest driver inside each tenant VM to the fixed release. Host side is a node drain plus reboot; guest side is a per-VM driver install and reboot. Because the guest driver is inside tenant-controlled VMs, in a multi-tenant estate you cannot fully remediate the guest half yourself - the host-side upgrade is the control you own. No VBIOS flash.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1058"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2021-01-08"},{"id":"CVE-2021-1060","cve":"CVE-2021-1060","aliases":[],"title":"NVIDIA vGPU software (guest kernel-mode driver + vGPU plugin): Unvalidated index across the guest driver and vGPU","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software (guest kernel-mode driver + vGPU plugin)","year":"2021","cvss_score":7.1,"severity":"high","kev":false,"impact":"Unvalidated index across the guest driver and vGPU plugin boundary, giving a tenant host-side data tampering or a shared-GPU outage. vGPU 8.x before 8.6, 11.0 before 11.3.","attack_vector":"Any user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the host and the vGPU guest driver inside each tenant VM to the fixed release. Host side is a node drain plus reboot; guest side is a per-VM driver install and reboot. Because the guest driver is inside tenant-controlled VMs, in a multi-tenant estate you cannot fully remediate the guest half yourself - the host-side upgrade is the control you own. No VBIOS flash.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1060"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2021-01-08"},{"id":"CVE-2021-1062","cve":"CVE-2021-1062","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): Unvalidated guest-supplied length in the vGPU plugin","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":7.1,"severity":"high","kev":false,"impact":"Unvalidated guest-supplied length in the vGPU plugin; tenant-triggered host data tampering or denial of service on the shared GPU. vGPU 8.x before 8.6, 11.0 before 11.3.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1062"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2021-01-08"},{"id":"CVE-2021-1064","cve":"CVE-2021-1064","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): The plugin takes a value from the guest, casts it to a pointer and dereferences it.","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":7.1,"severity":"high","kev":false,"impact":"The plugin takes a value from the guest, casts it to a pointer and dereferences it. A tenant chooses a host address to read - arbitrary host-memory disclosure or a crash. vGPU 8.x before 8.6, 11.0 before 11.3.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1064"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2021-01-08"},{"id":"CVE-2021-1065","cve":"CVE-2021-1065","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): Unvalidated guest input in the vGPU plugin leading to host data tampering or a","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":7.1,"severity":"high","kev":false,"impact":"Unvalidated guest input in the vGPU plugin leading to host data tampering or a shared-GPU outage. vGPU 8.x before 8.6, 11.0 before 11.3.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1065"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2021-01-08"},{"id":"CVE-2021-1086","cve":"CVE-2021-1086","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): The vGPU Manager lets guests control resources they are not entitled to, which","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":7.1,"severity":"high","kev":false,"impact":"The vGPU Manager lets guests control resources they are not entitled to, which NVIDIA describes as integrity and confidentiality loss. A tenant reaching resources outside its partition is the vGPU security model failing. vGPU 12.x before 12.2, 11.x before 11.4, 8.x before 8.7.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1086"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2021-04-29"},{"id":"CVE-2021-1090","cve":"CVE-2021-1090","aliases":[],"title":"NVIDIA GPU Display Driver (Windows nvlddmkm.sys + Linux nvidia.ko): Out-of-bounds read or write in the kernel-mode","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver (Windows nvlddmkm.sys + Linux nvidia.ko)","year":"2021","cvss_score":7.1,"severity":"high","kev":false,"impact":"Out-of-bounds read or write in the kernel-mode layer's control-call handler on both Windows and Linux. Unprivileged local code corrupts kernel memory or crashes the node; Gentoo shipped it as a security update.","attack_vector":"Any local user or GPU container with access to the NVIDIA device nodes.","remediation":"Install the fixed GPU Display Driver branch on both Windows and Linux nodes. The kernel component (nvlddmkm.sys / nvidia.ko) cannot be hot-swapped under load, so this is a node drain and reboot per host; restart the container runtime afterwards so mounted driver libraries match the kernel module. No VBIOS or BMC flash.","references":["https://security.gentoo.org/glsa/202310-02","https://nvd.nist.gov/vuln/detail/CVE-2021-1090"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-07-22"},{"id":"CVE-2021-1091","cve":"CVE-2021-1091","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): Hard-link attack: an unprivileged user makes the driver overwrite","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2021","cvss_score":7.1,"severity":"high","kev":false,"impact":"Hard-link attack: an unprivileged user makes the driver overwrite a file that needs elevated privilege to modify. Arbitrary privileged file overwrite, which is a well-trodden route to SYSTEM.","attack_vector":"Any local unprivileged user on a Windows GPU host.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1091"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-07-22"},{"id":"CVE-2021-1119","cve":"CVE-2021-1119","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): Double free in the vGPU Manager that NVIDIA explicitly describes as a","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":7.1,"severity":"high","kev":false,"impact":"Double free in the vGPU Manager that NVIDIA explicitly describes as a write-what-where condition allowing arbitrary code execution. A tenant VM writing chosen values to chosen host addresses is a full hypervisor-host compromise from inside a guest.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1119"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2021-10-29"},{"id":"CVE-2021-21539","cve":"CVE-2021-21539","aliases":[],"title":"Dell iDRAC9: TOCTOU race during simultaneous web-interface access — state corruption on the BMC","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC9","year":"2021","cvss_score":7.1,"severity":"high","kev":false,"impact":"TOCTOU race during simultaneous web-interface access — state corruption on the BMC","attack_vector":"Network, authenticated","remediation":"iDRAC firmware update","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-21539"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2021-04-30"},{"id":"CVE-2021-26332","cve":"CVE-2021-26332","aliases":[],"title":"AMD SEV-ES firmware - TMR placement in MMIO space: SEV-ES firmware does not verify that the Trusted Memory Region is","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV-ES firmware - TMR placement in MMIO space","year":"2021","cvss_score":7.1,"severity":"high","kev":false,"impact":"SEV-ES firmware does not verify that the Trusted Memory Region is outside MMIO space. Point the TMR at MMIO and the secure firmware's private working memory is suddenly backed by device registers the host controls - a route to observing or influencing what the SEV firmware does, costing integrity or availability of confidential guests.","attack_vector":"Hypervisor-privileged attacker who controls where the TMR is placed.","remediation":"Fixed in AMD reference firmware (AGESA / PSP / SEV firmware) and delivered only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo and the ODMs each rebuild and requalify AMD's AGESA drop before shipping. **Expect one to six months of OEM lag**, and on end-of-support platforms expect nothing. Applying it is a drain plus full power cycle, not a driver reload. Verify by reading back the PSP/SMU firmware version afterwards rather than trusting the BIOS version string. This sits inside the SEV-SNP trust boundary, so the update moves the platform's reported TCB version: refresh VCEK certificates from AMD's KDS and update any attestation policy your tenants pin, or confidential guest launches will start failing right after the BIOS lands.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26332","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2022-05-10"},{"id":"CVE-2021-26402","cve":"CVE-2021-26402","aliases":[],"title":"AMD Secure Processor firmware - BIOS mailbox command bounds checking: Insufficient bounds checking while the ASP","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor firmware - BIOS mailbox command bounds checking","year":"2021","cvss_score":7.1,"severity":"high","kev":false,"impact":"Insufficient bounds checking while the ASP firmware handles BIOS mailbox commands lets an attacker write partially-controlled data out of bounds into SMM or SEV-protected memory. Both destinations are places the OS is explicitly not allowed to reach: SMM is the most privileged execution mode on x86, and SEV memory belongs to confidential guests.","attack_vector":"Local, via the BIOS mailbox interface - requires host privilege.","remediation":"Fixed in AMD reference firmware (AGESA / PSP / SEV firmware) and delivered only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo and the ODMs each rebuild and requalify AMD's AGESA drop before shipping. **Expect one to six months of OEM lag**, and on end-of-support platforms expect nothing. Applying it is a drain plus full power cycle, not a driver reload. Verify by reading back the PSP/SMU firmware version afterwards rather than trusting the BIOS version string. This sits inside the SEV-SNP trust boundary, so the update moves the platform's reported TCB version: refresh VCEK certificates from AMD's KDS and update any attestation policy your tenants pin, or confidential guest launches will start failing right after the BIOS lands.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26402","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2023-01-11"},{"id":"CVE-2021-27364","cve":"CVE-2021-27364","aliases":[],"title":"Linux iSCSI: Unprivileged user can craft Netlink messages to scsi_transport_iscsi","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux iSCSI","year":"2021","cvss_score":7.1,"severity":"high","kev":false,"impact":"Unprivileged user can craft Netlink messages to scsi_transport_iscsi","attack_vector":"Local","remediation":"Data-plane: kernel patch + reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-27364"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-03-07"},{"id":"CVE-2021-28507","cve":"CVE-2021-28507","aliases":[],"title":"Arista EOS (service ACLs): Service ACL bypass for OpenConfig gNOI and RESTCONF","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (service ACLs)","year":"2021","cvss_score":7.1,"severity":"high","kev":false,"impact":"Service ACL bypass for OpenConfig gNOI and RESTCONF — the compensating control for the above does not hold","attack_vector":"Network","remediation":"EOS upgrade; important because it invalidates \"we ACL'd the management API\" as a mitigation","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-28507"],"status":"curated","published":"2022-01-14"},{"id":"CVE-2021-28692","cve":"CVE-2021-28692","aliases":[],"title":"Xen - x86 IOMMU command timeout detection and handling: Xen's IOMMU command timeout handling is inappropriate, so IOMMU","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen - x86 IOMMU command timeout detection and handling","year":"2021","cvss_score":7.1,"severity":"high","kev":false,"impact":"Xen's IOMMU command timeout handling is inappropriate, so IOMMU operations that do not complete in time are mishandled. On a host doing PCI passthrough - which is every GPU cloud - the IOMMU is the component enforcing that an assigned device can only DMA into its owner's memory. Mishandled timeouts mean that enforcement can be left in an indeterminate state while devices keep running.","attack_vector":"Requires a guest able to generate IOMMU load, i.e. a guest with an assigned device doing heavy DMA - normal behaviour for a passed-through GPU or NIC.","remediation":"Fixed in Xen (XSA-373). Hypervisor update plus host reboot. Especially relevant on GPU nodes, where passed-through accelerators and RDMA NICs push the IOMMU hard enough to hit timeout paths that idle hosts never reach.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-28692","https://xenbits.xen.org/xsa/advisory-373.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2021-06-30"},{"id":"CVE-2021-3571","cve":"CVE-2021-3571","aliases":[],"title":"linuxptp / ptp4l (transparent clock on little-endian): A crafted PTP packet against ptp4l running as a transparent","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"linuxptp / ptp4l (transparent clock on little-endian)","year":"2021","cvss_score":7.1,"severity":"high","kev":false,"impact":"A crafted PTP packet against ptp4l running as a transparent clock on a little-endian machine — i.e. every x86 and ARM64 server in your cluster — produces a fault. Transparent-clock mode is exactly the configuration used when PTP is carried across switches inside the cluster, so the affected deployment is the mainstream one, not an edge case.","attack_vector":"Remote, unauthenticated — a crafted PTP message from anything that can reach the node's PTP port.","remediation":"linuxptp package upgrade plus ptp4l restart. Same segmentation advice as the companion issue: PTP traffic should not be sourceable by tenant workloads.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-3571"],"status":"curated","published":"2021-07-09"},{"id":"CVE-2021-36309","cve":"CVE-2021-36309","aliases":[],"title":"Dell Enterprise SONiC OS (information disclosure): An authenticated user can extract sensitive information","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell Enterprise SONiC OS (information disclosure)","year":"2021","cvss_score":7.1,"severity":"high","kev":false,"impact":"An authenticated user can extract sensitive information from Enterprise SONiC 3.3.0 and earlier. On a switch, 'sensitive information' generally means credentials for the things the switch talks to — TACACS/RADIUS secrets, SNMP communities, image-server logins — so the impact propagates outward from the device.","attack_vector":"Authenticated user with system access on the switch.","remediation":"NOS image upgrade plus reboot, then rotate every shared secret configured on the switch. The rotation is the expensive part on a fleet that shares TACACS keys across all devices.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-36309"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-10-01"},{"id":"CVE-2021-4460","cve":"CVE-2021-4460","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): An out-of-bounds access in the amdkfd (KFD compute","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2021","cvss_score":7.1,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdkfd (KFD compute driver, /dev/kfd) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdkfd: Fix UBSAN shift-out-of-bounds warning","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-4460","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-10-01"},{"cwe":["CWE-908","CWE-824"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2021-47481","cve":"CVE-2021-47481","aliases":[],"title":"NVIDIA/Mellanox ConnectX driver (mlx5_ib on-demand paging MR creation): The xarray tracking implicit on-demand-paging","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA/Mellanox ConnectX driver (mlx5_ib on-demand paging MR creation)","year":"2021","cvss_score":7.1,"severity":"high","kev":false,"impact":"The xarray tracking implicit on-demand-paging children was not initialised when an ODP memory region was created, so deregistration walked an uninitialised structure and faulted on a wild address (the report shows a fault at 0x800000000 from ib_write_bw). ODP is the mechanism that lets the NIC page tenant memory in and out on demand; corrupt bookkeeping in that layer is bookkeeping about which pages belong to which memory key.","attack_vector":"Local, unprivileged. A tenant registers an ODP memory region and deregisters it - the ordinary lifecycle of any GPUDirect workload using on-demand paging.","remediation":"Kernel update initialising the ODP xarray in reg_create(). Disabling ODP is possible on some stacks but costs the pinning-free memory model that large-model training relies on, so patching is the realistic route.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=5508546631a0f555d7088203dec2614e41b5106e","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2021/CVE-2021-47481.json"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2022-28184","cve":"CVE-2022-28184","aliases":[],"title":"NVIDIA GPU Display Driver - kernel mode layer (Windows nvlddmkm.sys and Linux nvidia.ko): The DxgkDdiEscape handler","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - kernel mode layer (Windows nvlddmkm.sys and Linux nvidia.ko)","year":"2022","cvss_score":7.1,"severity":"high","kev":false,"impact":"The DxgkDdiEscape handler lets an unprivileged user reach administrator-privileged registers, giving direct read/write to GPU control state that should be off-limits to tenant code. Both the Windows and Linux datacenter drivers are affected, so a mixed fleet needs two separate rollouts.","attack_vector":"Local and unprivileged on either OS. On Linux it is reachable from any GPU container via /dev/nvidia*; on Windows from any session holding a GPU handle.","remediation":"Upgrade both the Linux and the Windows datacenter driver branches listed in bulletin 5353. Cost: Linux needs a drain and nvidia.ko reload per node; Windows needs a reboot per node. Two change windows unless your fleet is homogeneous.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28184","https://github.com/NVIDIA/product-security/tree/main/2022/5353"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:N","cwe":["CWE-284"],"fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2022-05-17"},{"cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:L/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-269"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2022-29164","cve":"CVE-2022-29164","aliases":["GHSA-cmv8-6362-r5w9"],"title":"Argo Workflows (Argo Server, HTML artifact serving): A workflow can emit an HTML artifact containing script that, when","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo Workflows (Argo Server, HTML artifact serving)","year":"2022","cvss_score":7.1,"severity":"high","kev":false,"impact":"A workflow can emit an HTML artifact containing script that, when a victim opens the artifact deep link, runs on the Argo Server origin and drives the API through XHR with the victim's session. A low-privilege tenant escalates to whatever an admin who clicks the link can do - create workflows, read templates and secrets, delete other tenants' work.","attack_vector":"A tenant with workflow-submit rights produces the artifact and gets a higher-privileged user to open its link, typically by email.","remediation":"Upgrade Argo Server to 3.2.11 or 3.3.5 and restart. Serve artifacts from a separate origin or object-store domain rather than the Argo Server domain wherever possible, so a malicious artifact cannot borrow the UI's session.","references":["https://github.com/argoproj/argo-workflows/security/advisories/GHSA-cmv8-6362-r5w9","https://nvd.nist.gov/vuln/detail/CVE-2022-29164"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2022-2989","cve":"CVE-2022-2989","aliases":[],"title":"Podman: Incorrect supplementary group handling","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Podman","year":"2022","cvss_score":7.1,"severity":"high","kev":false,"impact":"Incorrect supplementary group handling; information disclosure or data modification","attack_vector":"Any tenant workload","remediation":"Upgrade Podman","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-2989"],"status":"curated","published":"2022-09-13"},{"id":"CVE-2022-2995","cve":"CVE-2022-2995","aliases":[],"title":"CRI-O: Incorrect supplementary group handling leads to information disclosure between workloads","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"CRI-O","year":"2022","cvss_score":7.1,"severity":"high","kev":false,"impact":"Incorrect supplementary group handling leads to information disclosure between workloads","attack_vector":"Any tenant workload","remediation":"Upgrade CRI-O; drain node","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-2995"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2022-09-19"},{"id":"CVE-2022-31612","cve":"CVE-2022-31612","aliases":[],"title":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys): An out-of-bounds read through DxgkDdiEscape leaks","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys)","year":"2022","cvss_score":7.1,"severity":"high","kev":false,"impact":"An out-of-bounds read through DxgkDdiEscape leaks internal kernel information or crashes the node. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5383. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31612","https://github.com/NVIDIA/product-security/tree/main/2022/5383"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:H","cwe":["CWE-125"],"fleet":{"pain_class":"node-drain"},"published":"2022-11-19"},{"id":"CVE-2022-31613","cve":"CVE-2022-31613","aliases":[],"title":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys): Any local user can null-pointer-dereference","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys)","year":"2022","cvss_score":7.1,"severity":"high","kev":false,"impact":"Any local user can null-pointer-dereference the kernel mode layer and panic the node - no privileges required at all. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5383. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31613","https://github.com/NVIDIA/product-security/tree/main/2022/5383"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:C/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-drain"},"published":"2022-11-19"},{"id":"CVE-2022-32521","cve":"CVE-2022-32521","aliases":["SEVD-2023-010-06"],"title":"Schneider Electric Data Center Expert (versions prior to v7.9.0) - Java deserialization: Unsafe deserialization of data","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Schneider Electric Data Center Expert (versions prior to v7.9.0) - Java deserialization","year":"2022","cvss_score":7.1,"severity":"high","kev":false,"impact":"Unsafe deserialization of data posted to the web server yields remote code execution on the DCIM appliance. Standard deserialization bug, non-standard consequence: the host it lands on controls the power and cooling telemetry and credentials for the building.","attack_vector":"Remote, by posting crafted serialized data to the DCE web server.","remediation":"Upgrade to DCE v7.9.0 or later. Rotate stored credentials. Restrict who can reach the DCE web interface to a management jump host.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-32521"],"status":"curated","published":"2023-01-30"},{"id":"CVE-2022-34676","cve":"CVE-2022-34676","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): An out-of-bounds read in the kernel mode handler leaks","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":7.1,"severity":"high","kev":false,"impact":"An out-of-bounds read in the kernel mode handler leaks kernel memory contents or crashes the node. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5415. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34676","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:H/A:H","cwe":["CWE-197"],"fleet":{"pain_class":"node-drain"},"published":"2022-12-30"},{"id":"CVE-2022-35929","cve":"CVE-2022-35929","aliases":[],"title":"cosign / sigstore: `cosign verify-attestation --type` returns a false positive if any attestation exists","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"cosign / sigstore","year":"2022","cvss_score":7.1,"severity":"high","kev":false,"impact":"`cosign verify-attestation --type` returns a false positive if any attestation exists; unsigned images pass policy","attack_vector":"Malicious image","remediation":"Upgrade cosign; re-verify every image admitted while the vulnerable version was in the gate","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-35929"],"status":"curated","published":"2022-08-04"},{"id":"CVE-2022-41737","cve":"CVE-2022-41737","aliases":[],"title":"IBM Storage Scale Container Native Storage Access (namespace boundary): A local attacker can initiate connections from","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"IBM Storage Scale Container Native Storage Access (namespace boundary)","year":"2022","cvss_score":7.1,"severity":"high","kev":false,"impact":"A local attacker can initiate connections from a container outside its current namespace. Network-namespace escape from a storage-access container is a direct route from one tenant's pod onto networks the pod was never meant to touch — including, on most cluster designs, the storage back-end network where authentication is weak because it is assumed to be private.","attack_vector":"A local attacker inside a container using Storage Scale container-native access, versions 5.1.2.1 through 5.1.7.0.","remediation":"Upgrade Container Native Storage Access past 5.1.7.0 — rolling operator/DaemonSet upgrade. Companion issue CVE-2022-41738 allows connections *into* containers from external networks; both are closed by the same upgrade path. Also treat the storage back-end network as authenticated rather than trusted, which is an architectural change and the durable answer.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-41737","https://nvd.nist.gov/vuln/detail/CVE-2022-41738"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"tags":["tenant-isolation"],"published":"2024-02-17"},{"id":"CVE-2022-42262","cve":"CVE-2022-42262","aliases":[],"title":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko): A second unvalidated index path in the","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko)","year":"2022","cvss_score":7.1,"severity":"high","kev":false,"impact":"A second unvalidated index path in the vGPU plugin with the same guest-driven host buffer overrun. The attacker sits inside a tenant guest VM and the blast radius is the hypervisor host, which is running every other tenant's vGPU on the same physical GPU. This is precisely the boundary a vGPU-based multi-tenant offering is sold on.","attack_vector":"A tenant inside their own guest VM, driving the paravirtualised vGPU control interface. Several of these need only an unprivileged process in the guest; the rest need guest root, which a tenant already has on a VM they rented. No host credentials are involved at any point.","remediation":"Patch the vGPU Manager on the hypervisor host per bulletin 5415. Cost: the highest of any class here. The host driver cannot be reloaded while vGPUs are attached, so every tenant VM on that hypervisor must be live-migrated or powered off - a full host drain. NVIDIA also enforces a supported host/guest driver skew, so budget a matching guest-driver campaign in the same window or tenants lose their vGPU on next boot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42262","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-drain"},"tags":["tenant-isolation"],"published":"2022-12-30"},{"id":"CVE-2022-42263","cve":"CVE-2022-42263","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): An integer overflow in the kernel mode handler leaks","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":7.1,"severity":"high","kev":false,"impact":"An integer overflow in the kernel mode handler leaks kernel information or crashes the node. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5415. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42263","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:H","cwe":["CWE-190"],"fleet":{"pain_class":"node-drain"},"published":"2022-12-30"},{"id":"CVE-2022-42264","cve":"CVE-2022-42264","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): An unprivileged user drives the kernel to use","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":7.1,"severity":"high","kev":false,"impact":"An unprivileged user drives the kernel to use an out-of-range pointer offset, producing data corruption, data loss, kernel information disclosure or a crash. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5415. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42264","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:N","cwe":["CWE-823"],"fleet":{"pain_class":"node-drain"},"published":"2022-12-30"},{"id":"CVE-2022-42280","cve":"CVE-2022-42280","aliases":[],"title":"DGX-2 BMC: Path traversal","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX-2 BMC","year":"2022","cvss_score":7.1,"severity":"high","kev":false,"impact":"Path traversal -> file read/write on BMC","attack_vector":"Network-adjacent authenticated","remediation":"Flash DGX-2 BMC firmware out-of-band","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42280","https://github.com/NVIDIA/product-security/tree/main/2024/5513"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:N","cwe":["CWE-22"],"fleet":{"pain_class":"firmware-flash"},"published":"2023-01-13"},{"id":"CVE-2022-42327","cve":"CVE-2022-42327","aliases":["XSA-412"],"title":"Xen (x86): Unintended memory sharing between guests - cross-tenant data exposure","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen (x86)","year":"2022","cvss_score":7.1,"severity":"high","kev":false,"impact":"Unintended memory sharing between guests - cross-tenant data exposure","attack_vector":"Tenant VM guest","remediation":"Hypervisor patch + evacuation/reboot. Direct multi-tenant isolation break","references":["https://xenbits.xen.org/xsa/advisory-412.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2022-11-01"},{"id":"CVE-2022-43755","cve":"CVE-2022-43755","aliases":[],"title":"Rancher: Insufficient entropy means a leaked cattle-token stays usable after rotation","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Rancher","year":"2022","cvss_score":7.1,"severity":"high","kev":false,"impact":"Insufficient entropy means a leaked cattle-token stays usable after rotation","attack_vector":"An attacker who once observed the token","remediation":"Upgrade Rancher; force full token regeneration","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-43755"],"status":"curated","published":"2023-02-07"},{"id":"CVE-2022-48797","cve":"CVE-2022-48797","aliases":[],"title":"Linux memory management (NUMA balancing) as used by Gaudi accelerators: Automatic NUMA balancing migrated copy-on-write","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux memory management (NUMA balancing) as used by Gaudi accelerators","year":"2022","cvss_score":7.1,"severity":"high","kev":false,"impact":"Automatic NUMA balancing migrated copy-on-write pages that the Gaudi accelerator still had pinned, silently corrupting data in flight to the device. This is a correctness and integrity bug rather than an access-control one, but on a training cluster silent tensor corruption is worse than a crash because it poisons checkpoints before anyone notices.","attack_vector":"No attacker needed - it fires on normal accelerator workloads whenever automatic NUMA balancing is enabled on the node.","remediation":"Take the kernel fix and reboot. As an interim mitigation, disable automatic NUMA balancing (numa_balancing=disable) on accelerator nodes, which is a boot-parameter change and therefore still needs a drain and reboot. No firmware component.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-48797","https://git.kernel.org/stable/c/254090925e16abd914c87b4ad1b489440d89c4c3"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-07-16"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:H","cwe":["CWE-787","CWE-125"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-49883","cve":"CVE-2022-49883","aliases":[],"title":"Linux kernel (arch/x86/kvm): A guest that is not advertised long mode makes the host's SMM emulator walk 16","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (arch/x86/kvm)","year":"2022","cvss_score":7.1,"severity":"high","kev":false,"impact":"A guest that is not advertised long mode makes the host's SMM emulator walk 16 general-purpose register slots through a 32-bit SMRAM image, running off the end of the host-side buffer. The overrun happens in host kernel memory while emulating SMI entry and RSM, so a tenant VM gets an out-of-bounds host kernel read/write instead of a contained guest fault.","attack_vector":"Fully guest-driven on a 64-bit host: a tenant running a VM whose CPUID lacks X86_FEATURE_LM raises an SMI (a guest can direct one at itself through the emulated local APIC) and executes RSM. No VMM cooperation, no ioctl, no device node inside the container or guest is required.","remediation":"Boot a kernel carrying the linked stable fix; the record enumerates no fixed release, so pick the stable point release that contains commit a7ebfbea0f52. Interim control: always expose long mode (X86_FEATURE_LM) in tenant guest CPUID so the 32-bit SMRAM path is never taken, since SMM emulation itself cannot be turned off per VM.","references":["https://git.kernel.org/stable/c/a7ebfbea0f52550d7cdf12c38f3f5eaa7b2b6494","https://git.kernel.org/stable/c/696db303e54f7352623d9f640e6c51d8fa9d5588","https://nvd.nist.gov/vuln/detail/CVE-2022-49883"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2022-50026","cve":"CVE-2022-50026","aliases":[],"title":"habanalabs kernel driver (Gaudi NIC queue validation): A shift-out-of-bounds in Gaudi NIC queue validation: the driver","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"habanalabs kernel driver (Gaudi NIC queue validation)","year":"2022","cvss_score":7.1,"severity":"high","kev":false,"impact":"A shift-out-of-bounds in Gaudi NIC queue validation: the driver computes a queue offset for queue types that are not NIC queues, producing undefined behaviour in the kernel. Practical effect is a kernel oops taking the node out of service; on a hardened kernel it is a panic.","attack_vector":"Local user with the habanalabs device node mapped in - i.e. any Gaudi tenant container.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-50026","https://git.kernel.org/stable/c/01622098aeb05a5efbb727199bbc2a4653393255"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-06-18"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:H","cwe":["CWE-125","CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-50736","cve":"CVE-2022-50736","aliases":[],"title":"Linux kernel (drivers/infiniband/sw/siw): A tenant gets an out-of-bounds kernel array read using values it controls.","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/sw/siw)","year":"2022","cvss_score":7.1,"severity":"high","kev":false,"impact":"A tenant gets an out-of-bounds kernel array read using values it controls. The driver translates completion opcode and status fields through fixed lookup tables without validating them, and those fields sit in a completion queue that is mapped into the tenant's own address space - so the tenant writes the index the kernel then trusts. Usable for kernel memory disclosure or to crash the node.","attack_vector":"Tenant container holding /dev/infiniband/uverbs* on a node with the soft-iWARP driver (siw) loaded. Two ways in: write garbage opcode/status into the mmap'd CQ entries, or push the QP into ERROR state and post work so the flush path builds a completion with an undefined opcode. A fabric peer can force the ERROR state by dropping the connection, so the second path is partly remote-driven.","remediation":"No fixed release is published in this record - apply the listed stable fix commits or run a current stable kernel. Interim: unload/blacklist siw where soft-iWARP is not required, and do not expose /dev/infiniband/* to containers that only need ordinary TCP networking.","references":["https://git.kernel.org/stable/c/6af043089d3f1210776d19b6fdabea610d4c7699","https://git.kernel.org/stable/c/75af03fdf35acf15a3977f7115f6b8d10dff4bc7","https://nvd.nist.gov/vuln/detail/CVE-2022-50736"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2023-0180","cve":"CVE-2023-0180","aliases":[],"title":"GPU Display Driver (GeForce/RTX/Quadro/Tesla/vGPU): Info disclosure + DoS (OOB read in kernel driver)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver (GeForce/RTX/Quadro/Tesla/vGPU)","year":"2023","cvss_score":7.1,"severity":"high","kev":false,"impact":"Info disclosure + DoS (OOB read in kernel driver)","attack_vector":"Any tenant with GPU device access (local, unprivileged)","remediation":"Upgrade driver to the Jan-2023 branch; drain + reboot node (kernel module reload evicts all GPU workloads)","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0180","https://github.com/NVIDIA/product-security/tree/main/2023/5452"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:H","cwe":["CWE-125"],"fleet":{"pain_class":"node-reboot"},"published":"2023-04-01"},{"id":"CVE-2023-0181","cve":"CVE-2023-0181","aliases":[],"title":"GPU Display Driver: Local privesc / data tampering (missing access control)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2023","cvss_score":7.1,"severity":"high","kev":false,"impact":"Local privesc / data tampering (missing access control)","attack_vector":"Any tenant with a container holding /dev/nvidia*","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0181","https://github.com/NVIDIA/product-security/tree/main/2023/5452"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:H/A:H","cwe":["CWE-280"],"fleet":{"pain_class":"node-reboot"},"published":"2023-04-01"},{"id":"CVE-2023-0183","cve":"CVE-2023-0183","aliases":[],"title":"GPU Display Driver: Local privesc (kernel buffer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2023","cvss_score":7.1,"severity":"high","kev":false,"impact":"Local privesc (kernel buffer overflow)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0183","https://github.com/NVIDIA/product-security/tree/main/2023/5452"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-reboot"},"published":"2023-04-01"},{"id":"CVE-2023-0191","cve":"CVE-2023-0191","aliases":[],"title":"GPU Display Driver: Local privesc / data tampering (OOB write)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2023","cvss_score":7.1,"severity":"high","kev":false,"impact":"Local privesc / data tampering (OOB write)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0191","https://github.com/NVIDIA/product-security/tree/main/2023/5452"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:H/A:H","cwe":["CWE-119"],"fleet":{"pain_class":"node-reboot"},"published":"2023-04-01"},{"id":"CVE-2023-25516","cve":"CVE-2023-25516","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): An unprivileged user triggers an integer overflow","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2023","cvss_score":7.1,"severity":"high","kev":false,"impact":"An unprivileged user triggers an integer overflow in the Linux kernel mode layer, yielding information disclosure and denial of service. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5468. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25516","https://github.com/NVIDIA/product-security/tree/main/2023/5468"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:H","cwe":["CWE-190"],"fleet":{"pain_class":"node-drain"},"published":"2023-07-04"},{"id":"CVE-2023-25517","cve":"CVE-2023-25517","aliases":[],"title":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko): The vGPU plugin lets a guest OS control","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko)","year":"2023","cvss_score":7.1,"severity":"high","kev":false,"impact":"The vGPU plugin lets a guest OS control resources it is not authorised for, reaching information disclosure and data tampering on the host. The attacker sits inside a tenant guest VM and the blast radius is the hypervisor host, which is running every other tenant's vGPU on the same physical GPU. This is precisely the boundary a vGPU-based multi-tenant offering is sold on. Same authorisation-failure class as CVE-2022-31609 - the guest gets legitimate-looking access to something belonging to the host or another tenant.","attack_vector":"A tenant inside their own guest VM, driving the paravirtualised vGPU control interface. Several of these need only an unprivileged process in the guest; the rest need guest root, which a tenant already has on a VM they rented. No host credentials are involved at any point.","remediation":"Patch the vGPU Manager on the hypervisor host per bulletin 5468. Cost: the highest of any class here. The host driver cannot be reloaded while vGPUs are attached, so every tenant VM on that hypervisor must be live-migrated or powered off - a full host drain. NVIDIA also enforces a supported host/guest driver skew, so budget a matching guest-driver campaign in the same window or tenants lose their vGPU on next boot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25517","https://github.com/NVIDIA/product-security/tree/main/2023/5468"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:N","cwe":["CWE-285"],"fleet":{"pain_class":"node-drain"},"tags":["tenant-isolation"],"published":"2023-07-04"},{"id":"CVE-2023-25518","cve":"CVE-2023-25518","aliases":[],"title":"Jetson AGX Xavier / Xavier NX: Arbitrary memory R/W (PCIe controller without IOMMU)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Jetson AGX Xavier / Xavier NX","year":"2023","cvss_score":7.1,"severity":"high","kev":false,"impact":"Arbitrary memory R/W (PCIe controller without IOMMU)","attack_vector":"Local attacker with physical access","remediation":"Flash JetPack 32.7.4+; edge fleet only, not core DC","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5466/5466.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:P/AC:H/PR:N/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-923"],"published":"2023-06-23"},{"id":"CVE-2023-31316","cve":"CVE-2023-31316","aliases":[],"title":"AMD Secure Processor - hardware config integrity across power save/restore: Hardware configuration state is not","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor - hardware config integrity across power save/restore","year":"2023","cvss_score":7.1,"severity":"high","kev":false,"impact":"Hardware configuration state is not properly preserved across a power save/restore cycle in the ASP, so an attacker who can write outside the Trusted Memory Range can change security-relevant configuration that comes back wrong after resume. Suspend/resume is a soft spot in every confidential-computing design; here it is the ASP's own configuration that fails to survive intact.","attack_vector":"Local, requires the ability to write outside the TMR - host-privileged code - and a power-state transition to land on.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. Because this sits inside the SEV-SNP trust boundary, the update also moves the platform's reported TCB version: after patching you must refresh VCEK certificates from AMD's KDS and update whatever attestation policy your tenants (or your own confidential-VM control plane) pin against, or every guest launch will start failing validation. Servers rarely suspend, so on a datacenter fleet the practical exposure is lower than the score suggests; still worth catching in the next BIOS wave.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31316","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2026-05-15"},{"id":"CVE-2023-34338","cve":"CVE-2023-34338","aliases":["AMI-SA-2023006","Nozomi Labs BMC audit"],"title":"AMI MegaRAC SPx (BMC TLS certificate / cryptographic keys): A hard-coded certificate and its private key ship inside","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx (BMC TLS certificate / cryptographic keys)","year":"2023","cvss_score":7.1,"severity":"high","kev":false,"impact":"A hard-coded certificate and its private key ship inside the firmware, so the same key is on every BMC built from that image across every customer of that ODM. Anyone who extracts it once - and firmware images are downloadable from vendor support sites - can impersonate any BMC's HTTPS endpoint or decrypt intercepted management traffic. The operator consequence is that TLS on your management plane is decorative: admin passwords, Redfish tokens and KVM sessions are recoverable by an attacker who can interpose.","attack_vector":"Adjacent network with the ability to interpose on BMC traffic and some operator interaction (an admin logging into the BMC). No credentials needed - the attacker supplies the trust. A compromised management jump host, a rogue device on the management VLAN, or an ARP/DHCP position on that segment is enough.","remediation":"Firmware flash to SPx_12.3 / SPx_13.0 or later, but the flash is only half the fix - a fixed image does not retroactively replace a certificate already in place. After flashing you must generate and install a unique per-node BMC certificate signed by your own internal CA, which is a config operation over Redfish and can be automated, no reboot required. Do the certificate rotation even on nodes you cannot yet flash; it is the part that actually removes the shared key.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2023006.pdf","https://nvd.nist.gov/vuln/detail/CVE-2023-34338"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-07-05"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:H/A:H","cwe":["CWE-787","CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-52453","cve":"CVE-2023-52453","aliases":[],"title":"Linux kernel (drivers/vfio/pci/hisilicon): The VFIO migration save and resume paths do not advance the data pointer by","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/vfio/pci/hisilicon)","year":"2023","cvss_score":7.1,"severity":"high","kev":false,"impact":"The VFIO migration save and resume paths do not advance the data pointer by the file offset once PRE_COPY is in use, so device state is written to and read from the wrong place in the migration buffer. On the destination the device is restored from corrupted state and starts issuing bad DMA - the upstream log shows SMMU fault events and device queue timeouts. A passthrough device programmed from mis-indexed state is a device operating outside what the host thinks it authorised.","attack_vector":"The resume side parses the incoming migration stream, so the mis-indexed data comes from the migration source rather than from the tenant directly. Reached whenever a device is migrated with PRE_COPY enabled. Conditional on HiSilicon ACC accelerators (ZIP/SEC/HPRE) passed through with the hisi_acc_vfio_pci variant driver - not present on typical GPU fleets, though the pattern generalises to any variant driver's migration-state handling.","remediation":"The record lists no fixed release; boot a kernel carrying the stable fix commits below. Interim control: disable live migration for hisi_acc passthrough devices, or disable the PRE_COPY phase, until patched.","references":["https://git.kernel.org/stable/c/45f80b2f230df10600e6fa1b83b28bf1c334185e","https://git.kernel.org/stable/c/6bda81e24a35a856f58e6a5786de579b07371603","https://nvd.nist.gov/vuln/detail/CVE-2023-52453"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-0074","cve":"CVE-2024-0074","aliases":[],"title":"GPU Display Driver: Local privesc / data tampering (OOB write)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2024","cvss_score":7.1,"severity":"high","kev":false,"impact":"Local privesc / data tampering (OOB write)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0074","https://github.com/NVIDIA/product-security/tree/main/2024/5520"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:H/A:H","cwe":["CWE-788"],"fleet":{"pain_class":"node-reboot"},"published":"2024-03-27"},{"id":"CVE-2024-0128","cve":"CVE-2024-0128","aliases":[],"title":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko): The host vGPU plugin lets a guest reach","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko)","year":"2024","cvss_score":7.1,"severity":"high","kev":false,"impact":"The host vGPU plugin lets a guest reach global GPU resources, producing information disclosure, data tampering and privilege escalation across the partition. The attacker sits inside a tenant guest VM and the blast radius is the hypervisor host, which is running every other tenant's vGPU on the same physical GPU. This is precisely the boundary a vGPU-based multi-tenant offering is sold on. 'Global resources' on a shared physical GPU means state belonging to other tenants - this is a cross-tenant read/write primitive, not just a host bug.","attack_vector":"A tenant inside their own guest VM, driving the paravirtualised vGPU control interface. Several of these need only an unprivileged process in the guest; the rest need guest root, which a tenant already has on a VM they rented. No host credentials are involved at any point.","remediation":"Patch the vGPU Manager on the hypervisor host per bulletin 5586. Cost: the highest of any class here. The host driver cannot be reloaded while vGPUs are attached, so every tenant VM on that hypervisor must be live-migrated or powered off - a full host drain. NVIDIA also enforces a supported host/guest driver skew, so budget a matching guest-driver campaign in the same window or tenants lose their vGPU on next boot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0128","https://github.com/NVIDIA/product-security/tree/main/2024/5586"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:N","cwe":["CWE-732"],"fleet":{"pain_class":"node-drain"},"tags":["tenant-isolation"],"published":"2024-10-26"},{"id":"CVE-2024-0150","cve":"CVE-2024-0150","aliases":[],"title":"GPU Display Driver: Local privesc (kernel driver buffer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2024","cvss_score":7.1,"severity":"high","kev":false,"impact":"Local privesc (kernel driver buffer overflow)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0150","https://github.com/NVIDIA/product-security/tree/main/2025/5614"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-reboot"},"published":"2025-01-28"},{"id":"CVE-2024-21801","cve":"CVE-2024-21801","aliases":[],"title":"Intel TDX module: Insufficient control-flow management in the TDX module lets a privileged host user deny service","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel TDX module","year":"2024","cvss_score":7.1,"severity":"high","kev":false,"impact":"Insufficient control-flow management in the TDX module lets a privileged host user deny service - i.e. the host can wedge confidential VMs. Availability rather than confidentiality, but on a confidential-compute product the host being able to kill TDs at will is still a boundary the design claims to hold.","attack_vector":"Privileged host user.","remediation":"Update the Intel TDX module. The TDX module is loaded by the SEAM loader at boot, so the practical rollout is: stage the new module, drain every trust domain off the node, and reboot. It is not a live-patchable component and running TDs cannot be migrated through it. After the update, every TD must re-attest because the TDX module SVN is part of the attestation report - so anything that pinned the old measurement will fail until you update your attestation policy too. No OEM BIOS release needed for the module itself, which makes this materially faster than a platform firmware update.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21801","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01070.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-08-14"},{"id":"CVE-2024-25743","cve":"CVE-2024-25743","aliases":["Heckler-class","interrupt injection"],"title":"SEV-ES / SEV-SNP guest kernel - injection of virtual interrupts 0 and 14: An untrusted hypervisor can inject virtual","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"SEV-ES / SEV-SNP guest kernel - injection of virtual interrupts 0 and 14","year":"2024","cvss_score":7.1,"severity":"high","kev":false,"impact":"An untrusted hypervisor can inject virtual interrupts 0 (divide error) and 14 (page fault) into an SEV-SNP or SEV-ES guest at arbitrary points, reaching userspace signal handlers - in particular SIGFPE - inside the confidential VM. The Heckler research showed this class turning into authentication bypass inside the guest: inject an interrupt at the right instruction and a login check returns the wrong answer. The host never touches guest memory, so no memory-integrity mechanism catches it.","attack_vector":"Malicious hypervisor against its own guest. Requires timing precision but no guest vulnerability.","remediation":"Fixed in the **guest** kernel, not the host - the hardening lives in the SEV-ES/SNP guest's #VC handler and interrupt entry code. That inverts the usual rollout: you can patch every hypervisor you own and still be exposed, because the protection has to be in the tenant's own VM image. As an operator your job is to ship updated confidential-guest images (or tell tenants which minimum kernel to run) and, where you can, enforce it as an admission requirement. Each guest picks the fix up on its next boot; no host reboot, no firmware update. Fixed by guest-kernel hardening (Linux 6.9+ restricts which interrupts a guest accepts from the hypervisor). Enable Restricted Injection where the platform and guest support it. Again this is a guest-image problem, so operators should publish a minimum kernel and gate confidential workloads on it rather than assuming host patching covers them.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-25743","https://ahoi-attacks.github.io/heckler/","https://www.amd.com/en/resources/product-security.html"],"status":"curated","tags":["tenant-isolation"],"published":"2024-05-15"},{"id":"CVE-2024-26672","cve":"CVE-2024-26672","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu): A NULL pointer dereference in the amdgpu power","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu)","year":"2024","cvss_score":7.1,"severity":"high","kev":false,"impact":"A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: Fix variable 'mca_funcs' dereferenced before NULL check in 'amdgpu_mca_smu_get_mca_entry()'","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26672","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-04-02"},{"id":"CVE-2024-27029","cve":"CVE-2024-27029","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): An out-of-bounds access in the amdgpu kernel driver core - a","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2024","cvss_score":7.1,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu kernel driver core - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: fix mmhub client id out-of-bounds access","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-27029","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-05-01"},{"id":"CVE-2024-32878","cve":"CVE-2024-32878","aliases":[],"title":"llama.cpp (`gguf_init_from_file`): Use of uninitialized heap variable","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"llama.cpp (`gguf_init_from_file`)","year":"2024","cvss_score":7.1,"severity":"high","kev":false,"impact":"Use of uninitialized heap variable → double free","attack_vector":"Customer-supplied GGUF file","remediation":"Rebuild","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-32878"],"status":"curated","published":"2024-04-26"},{"cwe":["CWE-672","CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-35960","cve":"CVE-2024-35960","aliases":[],"title":"NVIDIA/Mellanox ConnectX flow steering core (mlx5 fs_core rule tree linkage): Flow steering is what decides which","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA/Mellanox ConnectX flow steering core (mlx5 fs_core rule tree linkage)","year":"2024","cvss_score":7.1,"severity":"high","kev":false,"impact":"Flow steering is what decides which packets reach which virtual function, and therefore which tenant. add_rule_fg() only linked newly created rules into the steering tree when their refcount was 1, while create_flow_handle deliberately reuses identical existing rules - so a handle could end up holding a rule with refcount 2 that was never linked, with NULL parent and root. Deleting the flow group then dereferences NULL. Unlinked steering rules are steering state the driver has lost track of, on the structure that enforces tenant traffic separation on a shared adapter.","attack_vector":"Local. Reached by the ordinary create/delete cycle of flow steering rules - which in an SR-IOV or switchdev cluster is driven by the CNI and by tenant network policy changes, i.e. by tenant-visible actions.","remediation":"Kernel update linking rules into the tree correctly regardless of refcount. No configuration workaround; flow steering cannot be turned off on an SR-IOV adapter.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=1263b0b26077b1183c3c45a0a2479573a351d423","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2024/CVE-2024-35960.json"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:H","cwe":["CWE-125"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-36960","cve":"CVE-2024-36960","aliases":[],"title":"Linux kernel (drivers/gpu/drm/vmwgfx): A tenant that asks for a fence event on its DRM fd and then reads the fd back","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/vmwgfx)","year":"2024","cvss_score":7.1,"severity":"high","kev":false,"impact":"A tenant that asks for a fence event on its DRM fd and then reads the fd back gets more bytes than the event actually contains, so adjacent kernel heap memory is copied into its buffer. That is a straight kernel-memory read primitive from an unprivileged process - enough to defeat KASLR and to harvest residual data left by other work on the box.","attack_vector":"An unprivileged process inside a VMware-hosted guest holding /dev/dri/card* or /dev/dri/renderD*: issue the vmwgfx fence-event ioctl, then read() the DRM fd. No capabilities needed. Only applies where vmwgfx is the GPU driver (ESXi/Workstation guests), so the exposure is tenant VMs, not bare-metal GPU nodes.","remediation":"Boot a kernel carrying the fix (the record lists no fixed_in - take the stable backport from your distro tree, commits below). Interim: do not expose /dev/dri to untrusted processes inside vmwgfx guests.","references":["https://git.kernel.org/stable/c/2f527e3efd37c7c5e85e8aa86308856b619fa59f","https://git.kernel.org/stable/c/cef0962f2d3e5fd0660c8efb72321083a1b531a9","https://nvd.nist.gov/vuln/detail/CVE-2024-36960"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-46722","cve":"CVE-2024-46722","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): An out-of-bounds access in the amdgpu kernel driver core - a","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2024","cvss_score":7.1,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu kernel driver core - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: fix mc_data out-of-bounds read warning","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46722","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-09-18"},{"id":"CVE-2024-46723","cve":"CVE-2024-46723","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): An out-of-bounds access in the amdgpu kernel driver core - a","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2024","cvss_score":7.1,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu kernel driver core - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: fix ucode out-of-bounds read warning","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46723","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-09-18"},{"id":"CVE-2024-46724","cve":"CVE-2024-46724","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): An out-of-bounds access in the amdgpu kernel driver core - a","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2024","cvss_score":7.1,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu kernel driver core - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: Fix out-of-bounds read of df_v1_7_channel_number","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46724","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-09-18"},{"id":"CVE-2024-46731","cve":"CVE-2024-46731","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): An out-of-bounds access in the amdgpu power","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2024","cvss_score":7.1,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu power management (SMU/powerplay) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/pm: fix the Out-of-bounds read warning","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46731","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-18"},{"id":"CVE-2024-46815","cve":"CVE-2024-46815","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.1,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Check num_valid_sets before accessing reader_wm_sets[]","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46815","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-27"},{"cwe":["CWE-190","CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-47719","cve":"CVE-2024-47719","aliases":[],"title":"Linux kernel (drivers/iommu/iommufd): A tenant supplies an IOVA and user pointer whose alignment math overflows, so","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/iommufd)","year":"2024","cvss_score":7.1,"severity":"high","kev":false,"impact":"A tenant supplies an IOVA and user pointer whose alignment math overflows, so iommufd allocates the mapping area over a different IOVA range than the one it validated. A DMA window ends up somewhere the kernel did not intend it - the precondition for a passthrough device reaching memory outside its tenant's assignment.","attack_vector":"Any holder of /dev/iommu - a VMM for a passthrough VM, or a container handed the iommufd node - issuing IOMMU_IOAS_MAP with a crafted iova/user-pointer pair. Found by syzkaller, but the overflow is in the production area-allocation path (iopt_alloc_area_pages), not in test-only code; CONFIG_IOMMUFD_TEST only makes it visible as a WARN. No host root required.","remediation":"The record lists no fixed release; boot a kernel carrying the stable fix commits below. Interim control: do not expose /dev/iommu to tenants; mediate DMA mapping through the host VMM rather than handing the fd into the guest's container.","references":["https://git.kernel.org/stable/c/cd6dd564ae7d99967ef50078216929418160b30e","https://git.kernel.org/stable/c/a6e9f9fd14772c0b23c6d1d7002d98f9d27cb1f6","https://nvd.nist.gov/vuln/detail/CVE-2024-47719"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-49894","cve":"CVE-2024-49894","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.1,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Fix index out of bounds in degamma hardware format translation","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49894","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"id":"CVE-2024-53150","cve":"CVE-2024-53150","aliases":[],"title":"Linux kernel (ALSA usb-audio): Out-of-bounds read finding clock sources","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (ALSA usb-audio)","year":"2024","cvss_score":7.1,"severity":"high","kev":true,"impact":"Out-of-bounds read finding clock sources [KEV]","attack_vector":"Local user with USB device access","remediation":"Livepatchable; otherwise drain + reboot. Blacklist snd-usb-audio on servers","references":["https://access.redhat.com/security/cve/CVE-2024-53150"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-12-24"},{"cwe":["CWE-190"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-22091","cve":"CVE-2025-22091","aliases":[],"title":"NVIDIA/Mellanox ConnectX driver (mlx5_ib memory-region page size selection): Mlx5_ib stored the result of its","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA/Mellanox ConnectX driver (mlx5_ib memory-region page size selection)","year":"2025","cvss_score":7.1,"severity":"high","kev":false,"impact":"Mlx5_ib stored the result of its best-page-size search in an unsigned int. Registering a large physically contiguous region - 4GB, which is entirely ordinary for a GPU training buffer pinned for GPUDirect - makes the driver select a 4GB page size, the variable wraps to zero, and the memory key is built from a bogus page size. A memory key whose page size does not describe the memory it maps is precisely the failure mode that lets an RDMA translation land somewhere other than where the registering tenant intended.","attack_vector":"Local, unprivileged. Any tenant registering a very large contiguous buffer through the normal verbs path - this is triggered by legitimate large-model workloads, not only by attack, which is what makes it likely to be latent in the fleet.","remediation":"Kernel update widening the page_size variables to unsigned long across the mlx5_ib MR path. Cannot be worked around by configuration - the trigger is buffer size, which the tenant chooses.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=01fd737776ca0f17a96d83cd7f0840ce130b9a02","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2025/CVE-2025-22091.json"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-22226","cve":"CVE-2025-22226","aliases":[],"title":"VMware ESXi / Workstation / Fusion: Out-of-bounds read in HGFS leaks vmx process memory to the guest","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"VMware ESXi / Workstation / Fusion","year":"2025","cvss_score":7.1,"severity":"high","kev":true,"impact":"Out-of-bounds read in HGFS leaks vmx process memory to the guest [KEV]","attack_vector":"Tenant VM guest (local admin inside the VM)","remediation":"ESXi patch + host reboot with evacuation","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-22226"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-03-04"},{"id":"CVE-2025-23270","cve":"CVE-2025-23270","aliases":[],"title":"IGX Orin / Jetson (power mgmt): DoS / hardware damage via power-management abuse","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"IGX Orin / Jetson (power mgmt)","year":"2025","cvss_score":7.1,"severity":"high","kev":false,"impact":"DoS / hardware damage via power-management abuse","attack_vector":"Local attacker on the device","remediation":"Flash firmware; edge fleet","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23270","https://github.com/NVIDIA/product-security/tree/main/2025/5662"],"status":"curated","cvss_vector":"CVSS:3.1/AV:P/AC:H/PR:N/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-392"],"published":"2025-07-17"},{"id":"CVE-2025-23278","cve":"CVE-2025-23278","aliases":[],"title":"GPU Display Driver: Local privesc (OOB write)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2025","cvss_score":7.1,"severity":"high","kev":false,"impact":"Local privesc (OOB write)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23278","https://github.com/NVIDIA/product-security/tree/main/2025/5670"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:H/A:H","cwe":["CWE-129"],"fleet":{"pain_class":"node-reboot"},"published":"2025-08-02"},{"id":"CVE-2025-23360","cve":"CVE-2025-23360","aliases":[],"title":"NVIDIA NeMo Framework: A relative path traversal gives arbitrary file write, reaching code execution","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NeMo Framework","year":"2025","cvss_score":7.1,"severity":"high","kev":false,"impact":"A relative path traversal gives arbitrary file write, reaching code execution. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5623 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23360","https://github.com/NVIDIA/product-security/tree/main/2025/5623"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:H/A:H","cwe":["CWE-23"],"published":"2025-03-11"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:C/C:N/I:N/A:H","cwe":["CWE-667"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-40038","cve":"CVE-2025-40038","aliases":[],"title":"Linux kernel (arch/x86/kvm/svm): On AMD hosts that cannot report the next RIP, KVM's WRMSR/HLT/INVD fastpath has to","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (arch/x86/kvm/svm)","year":"2025","cvss_score":7.1,"severity":"high","kev":false,"impact":"On AMD hosts that cannot report the next RIP, KVM's WRMSR/HLT/INVD fastpath has to decode the guest instruction, which reads guest memory - and it does so with interrupts disabled. A guest can therefore make the host kernel sleep in atomic context, which is a scheduling-while-atomic condition that can hang or crash the node; the impact is a shared-node DoS, not an escape.","attack_vector":"Guest-driven and unprivileged inside the VM: the tenant executes HLT, WRMSR or INVD from a page KVM must fault in to decode. Conditional on next-RIP being unavailable - older AMD silicon, or a host running with kvm_amd.nrips=0.","remediation":"Update to a stable kernel with the linked fix (no fixed release enumerated; take the branch carrying commit da2a3c231f7f). Interim control: run on hosts with NRIPS support and never set kvm_amd.nrips=0.","references":["https://git.kernel.org/stable/c/da2a3c231f7f2a5ac146d972b8c1d7d84aff6d70","https://git.kernel.org/stable/c/cd3efb93677c4b0cf76348882fb429165fee33fd","https://nvd.nist.gov/vuln/detail/CVE-2025-40038"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-41239","cve":"CVE-2025-41239","aliases":[],"title":"VMware ESXi / Workstation / Tools: Uninitialised memory in vSockets discloses host memory to the guest","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"VMware ESXi / Workstation / Tools","year":"2025","cvss_score":7.1,"severity":"high","kev":false,"impact":"Uninitialised memory in vSockets discloses host memory to the guest","attack_vector":"Tenant VM guest","remediation":"ESXi patch + host reboot; also VMware Tools update in guests","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-41239"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-07-15"},{"id":"CVE-2025-6202","cve":"CVE-2025-6202","aliases":["Phoenix"],"title":"SK Hynix DDR5 DIMMs (manufactured 2021-01 through 2024-12): Rowhammer bit flips on DDR5, which had been assumed out","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"SK Hynix DDR5 DIMMs (manufactured 2021-01 through 2024-12)","year":"2025","cvss_score":7.1,"severity":"high","kev":false,"impact":"Rowhammer bit flips on DDR5, which had been assumed out of reach because of on-die ECC and improved TRR. Affects SK Hynix DDR5 DIMMs produced between January 2021 and December 2024 - a very large share of DDR5 installed in AI host nodes bought in that window. Impact is integrity of host memory, with the usual escalation to privilege via page-table corruption.","attack_vector":"Local attacker on the node. High attack complexity, low privileges - a tenant workload with sustained memory access is the model.","remediation":"Take an inventory of DIMM vendor and date code across the fleet (dmidecode -t memory) before anything else, because the exposure is specific. Mitigation guidance is to raise the DRAM refresh rate - tripling it substantially raises the bar at a measurable memory-bandwidth cost - which is a BIOS-level change requiring drain and reboot per node. There is no microcode or OS patch. Monitor correctable ECC error rates as the detection signal.","references":["https://comsec.ethz.ch/phoenix","https://security.googleblog.com/2025/09/supporting-rowhammer-research-to.html","https://nvd.nist.gov/vuln/detail/CVE-2025-6202"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-09-15"},{"id":"CVE-2025-6242","cve":"CVE-2025-6242","aliases":[],"title":"vLLM (`MediaConnector` SSRF): SSRF via `load_from_url` in multimodal input handling","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (`MediaConnector` SSRF)","year":"2025","cvss_score":7.1,"severity":"high","kev":false,"impact":"SSRF via `load_from_url` in multimodal input handling","attack_vector":"Unauthenticated request supplying an image URL — reaches cloud metadata endpoints and internal control planes","remediation":"Upgrade and block link-local metadata (169.254.169.254) at the pod network layer. Provider-owned: IMDS reachability from tenant pods is an infrastructure decision","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-6242"],"status":"curated","published":"2025-10-07"},{"id":"CVE-2025-66448","cve":"CVE-2025-66448","aliases":[],"title":"vLLM (`Nemotron_Nano_VL_Config`): RCE via a config class evaluated at model load","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (`Nemotron_Nano_VL_Config`)","year":"2025","cvss_score":7.1,"severity":"high","kev":false,"impact":"RCE via a config class evaluated at model load","attack_vector":"Customer-supplied model config","remediation":"Upgrade to 0.11.1+","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-66448"],"status":"curated","published":"2025-12-01"},{"id":"CVE-2025-68313","cve":"CVE-2025-68313","aliases":[],"title":"AMD Zen 5 RDSEED (16-bit and 32-bit variants): On Zen 5, the 16-bit and 32-bit forms of RDSEED return zero far more","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD Zen 5 RDSEED (16-bit and 32-bit variants)","year":"2025","cvss_score":7.1,"severity":"high","kev":false,"impact":"On Zen 5, the 16-bit and 32-bit forms of RDSEED return zero far more often than randomness allows, while still setting the carry flag to signal success. Any software that trusts RDSEED's success indicator - kernel entropy pools, TLS libraries, key generation inside confidential guests - silently consumes zeros as if they were seed material. This is a hardware entropy failure that announces nothing; you find it by auditing, not by observing symptoms.","attack_vector":"Not an attack in the usual sense - it is a silicon defect that any workload on affected Zen 5 parts hits passively. The exposure is that an attacker who knows a target derived keys from RDSEED on affected hardware can search a drastically reduced keyspace.","remediation":"Fixed in the Linux kernel (x86/CPU/AMD) by masking off the broken RDSEED variants on affected Zen 5 parts, so the OS stops trusting them and falls back to working entropy sources. Take the distro kernel update and reboot the node; no firmware, microcode or BIOS step. Guest kernels need the same fix, so confidential-VM images must be updated too. Rotate any long-lived key material generated on affected Zen 5 hosts before the fix - the patch stops the bleeding but does not un-weaken existing keys.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-68313"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-12-16"},{"cwe":["CWE-362","CWE-787"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-23361","cve":"CVE-2026-23361","aliases":[],"title":"Linux kernel (drivers/pci/controller/dwc): Raising an MSI-X interrupt is a posted PCI write, and the endpoint driver","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/pci/controller/dwc)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"Raising an MSI-X interrupt is a posted PCI write, and the endpoint driver tears down the outbound address-translation entry that write depends on before the write has landed. If the write loses the race it is delivered through a stale or reassigned translation - the upstream commit states plainly that it may corrupt host memory, and the reporter's SMMU logs show the resulting unauthorised writes being caught as translation faults. That is a device writing into memory it was never given.","attack_vector":"Applies to a machine running Linux in PCIe *endpoint* mode - a DPU, smart storage appliance or nvmet-pci-epf target built on a DesignWare endpoint controller - not to a normal GPU compute node. The victim is the host on the other side of the link: the endpoint's misdirected DMA lands in host memory. It is driven by ordinary traffic, not by a special request - the reporter hit it with fio at high queue depth against nvmet-pci-epf, so any busy client of that endpoint makes it fire. Whether it is exploitable rather than merely corrupting depends on what the stale translation points at; an IOMMU on the host side turns most instances into faults rather than writes.","remediation":"On any PCIe-endpoint-mode Linux appliance in the fleet, boot a kernel where dw_pcie_ep_raise_msix_irq() reads back the MSI-X address to flush the posted write before unmapping the ATU entry. Interim: keep an IOMMU enabled and enforcing on the host side of every such link so a misdirected endpoint write faults instead of landing, and avoid MSI-X on affected endpoint firmware where MSI is a workable substitute.","references":["https://git.kernel.org/stable/c/a7afb8f810c04845fdfc58c57d9cf0cc5f23ced0","https://git.kernel.org/stable/c/6f60a783860c77b309f7d81003b6a0c73feca49e","https://nvd.nist.gov/vuln/detail/CVE-2026-23361"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-24185","cve":"CVE-2026-24185","aliases":[],"title":"NVIDIA NVOS (network switches): With PKA-only SSH mode enabled, an administrator can inadvertently leave an alternative","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NVOS (network switches)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"With PKA-only SSH mode enabled, an administrator can inadvertently leave an alternative authentication path open. If the default password was never changed, that path grants unauthorised switch access - so the switch looks key-only while still accepting a known password.","attack_vector":"Network, adjacent, low privileges. The exposure only exists where the default password survived deployment, which is exactly the switch nobody revisited after racking.","remediation":"Update NVOS per bulletin 5817 and, more importantly, verify that no switch still holds a default password - the configuration audit matters more than the patch here. Cost: switch reboot for the upgrade; the password audit costs nothing and should happen today.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24185","https://github.com/NVIDIA/product-security/tree/main/2026/5817"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:H/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-288"],"fleet":{"pain_class":"node-reboot"},"published":"2026-08-18"},{"id":"CVE-2026-24195","cve":"CVE-2026-24195","aliases":[],"title":"vGPU Manager: Guest-to-host impact via invalid guest-driver input","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU Manager","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"Guest-to-host impact via invalid guest-driver input","attack_vector":"Tenant VM guest","remediation":"vGPU Manager upgrade; evacuate guest VMs","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24195","https://github.com/NVIDIA/product-security/tree/main/2026/5821"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:C/C:N/I:N/A:H","cwe":["CWE-20"],"published":"2026-05-26"},{"id":"CVE-2026-24196","cve":"CVE-2026-24196","aliases":[],"title":"GPU Display Driver: Info disclosure (OOB read of graphics memory)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"Info disclosure (OOB read of graphics memory)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; rolling reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24196","https://github.com/NVIDIA/product-security/tree/main/2026/5821"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:H","cwe":["CWE-125"],"fleet":{"pain_class":"node-reboot"},"published":"2026-05-26"},{"id":"CVE-2026-24779","cve":"CVE-2026-24779","aliases":[],"title":"vLLM (`MediaConnector`): SSRF, recurrence of CVE-2025-6242","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (`MediaConnector`)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"SSRF, recurrence of CVE-2025-6242","attack_vector":"Unauthenticated request with an attacker-supplied media URL","remediation":"Upgrade to 0.14.1+; enforce IMDSv2 / metadata deny at the network layer","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24779"],"status":"curated","published":"2026-01-27"},{"id":"CVE-2026-25960","cve":"CVE-2026-25960","aliases":[],"title":"vLLM (`load_from_url_async`): Bypass of the CVE-2026-24779 SSRF fix","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (`load_from_url_async`)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"Bypass of the CVE-2026-24779 SSRF fix","attack_vector":"Unauthenticated request","remediation":"Upgrade past 0.15.1. Third SSRF in the same connector — network-layer egress control is the only durable fix","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-25960"],"status":"curated","published":"2026-03-09"},{"id":"CVE-2026-31395","cve":"CVE-2026-31395","aliases":[],"title":"Linux bnxt_en driver (DBG_BUF_PRODUCER async event handler): The async-event handler indexes a fixed array","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux bnxt_en driver (DBG_BUF_PRODUCER async event handler)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"The async-event handler indexes a fixed array with a `type` field supplied by the NIC firmware, without bounds checking — so firmware controls a kernel array index. This is the concrete version of a threat operators often wave at abstractly: if the adapter's firmware is compromised or buggy, it has a direct path into kernel memory corruption on the host. Every argument for verifying NIC firmware provenance at intake rests on bugs of exactly this shape.","attack_vector":"The NIC firmware itself, or anything that can influence what the firmware reports — which includes a firmware image installed at build time or by a previous tenant on bare metal.","remediation":"Kernel/driver upgrade plus host reboot. The durable control is separate: verify and reflash NIC firmware from a known-good image at rack intake and at tenant handoff, so the host is not trusting whatever firmware happens to be on the card.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-31395"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2026-04-03"},{"id":"CVE-2026-31766","cve":"CVE-2026-31766","aliases":[],"title":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu): An out-of-bounds access in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu user-mode queues (doorbell submission path) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: validate doorbell_offset in user queue creation","attack_vector":"Local. Reachable by any process with a render node open that can create user-mode queues - the normal ROCm submission path, reachable from an unprivileged container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-31766","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-05-01"},{"cvss_vector":"CVSS:4.0/AV:N/AC:L/AT:N/PR:L/UI:N/VC:N/VI:L/VA:H/SC:N/SI:N/SA:N","cwe":["CWE-287"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2026-34204","cve":"CVE-2026-34204","aliases":[],"title":"MinIO (server-side encryption / replication): An authenticated tenant can inject SSE metadata through replication","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"MinIO (server-side encryption / replication)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"An authenticated tenant can inject SSE metadata through replication headers, corrupting how objects are recorded as encrypted. The practical outcome is objects that cannot be decrypted afterwards - durable data loss on a replicated bucket, plus confusion about which objects are actually protected.","attack_vector":"Any authenticated MinIO user able to send replication-related headers to the S3 endpoint.","remediation":"Upgrade to the fixed release from GHSA-3rh2-v3gr-35p9 and restart the cluster. Verify readability of objects written during the exposure window on SSE-enabled replicated buckets, and restrict replication header handling to your replication service account.","references":["https://github.com/minio/minio/security/advisories/GHSA-3rh2-v3gr-35p9","https://nvd.nist.gov/vuln/detail/CVE-2026-34204"],"status":"curated"},{"id":"CVE-2026-35155","cve":"CVE-2026-35155","aliases":["DSA-2026-187"],"title":"Dell iDRAC10 (credential handling, race condition): A race in iDRAC10's credential handling leaves secrets","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC10 (credential handling, race condition)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"A race in iDRAC10's credential handling leaves secrets insufficiently protected, letting an authenticated low-privilege user win the race and come back with elevated access to the BMC. Elevated on the BMC means the whole out-of-band toolkit: power, Virtual Media, KVM, firmware. This one matters disproportionately because iDRAC10 ships on 17G PowerEdge - the newest GPU platforms - so the affected fleet is the freshly-racked capacity, not the legacy tier, and it is likely still inside its burn-in window where firmware is whatever shipped from the factory.","attack_vector":"An authenticated low-privilege iDRAC10 account. The exposure is anyone you have handed a non-admin BMC login: remote hands, an integrator, a monitoring service account, or a tenant-facing self-service console that proxies BMC actions.","remediation":"Flash iDRAC10 to 1.30.10.50 or later. Affected builds are 1.20.70.50 and 1.30.05.10 specifically. Out-of-band, per-node, no host reboot and no job drain. Because this hits new deployments, fold the check into rack-acceptance: verify iDRAC10 build before a node ever takes tenant traffic, rather than discovering it in a later sweep. No config-only mitigation - reduce exposure meanwhile by pruning low-privilege iDRAC accounts.","references":["https://www.dell.com/support/kbdoc/en-us/000452298/dsa-2026-187-security-update-for-dell-idrac10-vulnerability","https://nvd.nist.gov/vuln/detail/CVE-2026-35155"],"status":"curated","published":"2026-04-29"},{"cvss_vector":"CVSS:4.0/AV:N/AC:L/AT:N/PR:L/UI:N/VC:N/VI:N/VA:H/SC:N/SI:N/SA:N","cwe":["CWE-770"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2026-39414","cve":"CVE-2026-39414","aliases":[],"title":"MinIO (S3 Select): A crafted S3 Select CSV query makes MinIO allocate memory without bound until the process is","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"MinIO (S3 Select)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"A crafted S3 Select CSV query makes MinIO allocate memory without bound until the process is OOM-killed. One tenant issuing a single query takes down the shared object store for every job on the cluster.","attack_vector":"Any authenticated tenant that can issue S3 Select queries against the endpoint.","remediation":"Upgrade to the release in GHSA-h749-fxx7-pwpg and restart the nodes. If S3 Select is not used by your workloads, deny SelectObjectContent in the bucket policy. Set a memory cgroup limit on the MinIO service so an allocation blowup restarts one node instead of destabilising the host.","references":["https://github.com/minio/minio/security/advisories/GHSA-h749-fxx7-pwpg","https://nvd.nist.gov/vuln/detail/CVE-2026-39414"],"status":"curated"},{"id":"CVE-2026-45856","cve":"CVE-2026-45856","aliases":[],"title":"Linux kernel InfiniBand core (ib_uverbs post_send): ib_uverbs_post_send() takes the work-queue-entry size straight","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel InfiniBand core (ib_uverbs post_send)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"ib_uverbs_post_send() takes the work-queue-entry size straight from userspace with no validation, allocates that size, then reads fields past the allocation - an out-of-bounds read of the kernel heap that leaks kernel memory to an unprivileged process. The receive path validated this; the send path did not. Any tenant process holding an RDMA verbs handle - which on a GPU cluster is every job using NCCL, UCX or MPI - can read host kernel memory.","attack_vector":"A local unprivileged user with an open RDMA verbs device handle. On a shared GPU node that is any tenant running a distributed training job.","remediation":"Upgrade the host kernel to 7.0 or a stable backport (5.10.252, 5.15.202, 6.1.165, 6.6.128, 6.12.75, 6.18.14, 6.19.4). Rolling reboot of every node that exposes /dev/infiniband/uverbs* to workloads - which is all of them on an RDMA cluster.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-45856","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-45856.json"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-05-27"},{"id":"CVE-2026-46199","cve":"CVE-2026-46199","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn4): An out-of-bounds access in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn4)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu firmware, ACPI and IP-block initialisation - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu/vcn4: Prevent OOB reads when parsing dec msg","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-46199","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-05-28"},{"id":"CVE-2026-46204","cve":"CVE-2026-46204","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn4): An out-of-bounds access in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn4)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu firmware, ACPI and IP-block initialisation - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu/vcn4: Prevent OOB reads when parsing IB","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-46204","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-05-28"},{"id":"CVE-2026-46218","cve":"CVE-2026-46218","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): An out-of-bounds access in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu firmware, ACPI and IP-block initialisation - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: Add bounds checking to ib_{get,set}_value","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-46218","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-05-28"},{"id":"CVE-2026-46230","cve":"CVE-2026-46230","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn3): An out-of-bounds access in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn3)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu firmware, ACPI and IP-block initialisation - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu/vcn3: Prevent OOB reads when parsing dec msg","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-46230","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-05-28"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:H/A:H","cwe":["CWE-125","CWE-843"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-52953","cve":"CVE-2026-52953","aliases":[],"title":"Linux kernel (drivers/iommu/intel): Killing a VM that has a device attached through the VT-d nested/PASID path makes","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/intel)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"Killing a VM that has a device attached through the VT-d nested/PASID path makes the host dereference past the end of a static blocked-domain object and take a general protection fault. A tenant kills its own qemu and the host kernel goes down, taking every other tenant sharing that node with it.","attack_vector":"A tenant VM (or its VMM) with a device bound through vfio + iommufd nested domains and a PASID attached simply exits or is killed - releasing the vfio device fd runs the reset path that hits the bug. Conditional on VT-d scalable mode with nested translation in use; no host root and no fabric access needed.","remediation":"Update to a stable kernel carrying commits 88397fad / 1e659db4. Interim: avoid VT-d nested translation (vIOMMU) for tenant VMs on unpatched hosts, and drain co-tenants off nodes that run nested-PASID passthrough until the kernel is updated.","references":["https://git.kernel.org/stable/c/88397fad7914ee74a7880fa5ce01f9eb6bfe0743","https://git.kernel.org/stable/c/1e659db468476733d217c1314c1e0d9244356d6c","https://nvd.nist.gov/vuln/detail/CVE-2026-52953"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-53138","cve":"CVE-2026-53138","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Bound VBIOS record-chain walk loops","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53138","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-06-25"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:H","cwe":["CWE-125","CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-53187","cve":"CVE-2026-53187","aliases":[],"title":"Linux kernel RDMA core (UVERBS_ATTR_ALLOC_DMAH_CPU_ID, DMA handle allocation): The cpu_id a tenant passes when","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel RDMA core (UVERBS_ATTR_ALLOC_DMAH_CPU_ID, DMA handle allocation)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"The cpu_id a tenant passes when allocating a DMA handle goes straight into cpumask_test_cpu() with no range check, so it becomes an unbounded bit index into the cpumask bitmap - an out-of-bounds kernel read at an offset the tenant chooses. On kernels built with CONFIG_DEBUG_PER_CPU_MAPS the same input trips a WARN_ON_ONCE, which on a panic_on_warn fleet (common in hyperscale and neocloud images, because operators want crash dumps rather than silent corruption) converts an unprivileged ioctl into a node reboot - one tenant evicting every other job on the box.","attack_vector":"Local ioctl on /dev/infiniband/uverbs* by any RDMA-capable tenant. Unprivileged.","remediation":"Kernel update rejecting cpu_id values that are not below nr_cpu_ids. If you run panic_on_warn fleet-wide, note that this class of bug turns any WARN reachable from an unprivileged syscall into a tenant-triggered node kill - worth reviewing that policy alongside the patch.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=0efbb6b54ff56300867027d8e0800d0e32226a20","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-53187.json"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-53330","cve":"CVE-2026-53330","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Fix out-of-bounds read in dp_get_eq_aux_rd_interval()","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53330","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-07-01"},{"id":"CVE-2026-53875","cve":"CVE-2026-53875","aliases":[],"title":"picklescan: `scan_pytorch` bypass via forged magic numbers","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"picklescan","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"`scan_pytorch` bypass via forged magic numbers","attack_vector":"Customer-supplied model file","remediation":"Upgrade to 1.0.3+. There is no complete fix for pickle scanning — this is the fifth bypass in the same tool","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53875"],"status":"curated","published":"2026-06-17"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:C/C:N/I:N/A:H","cwe":["CWE-617"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-63806","cve":"CVE-2026-63806","aliases":[],"title":"Linux kernel (virt/kvm): A guest store that splits a page and lands on a datamatch-enabled ioeventfd reaches a BUG_ON","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (virt/kvm)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"A guest store that splits a page and lands on a datamatch-enabled ioeventfd reaches a BUG_ON in KVM's ioeventfd handling because of an alignment assumption that does not hold. Impact is denial of service, not escape - but it is a guest hitting a kernel BUG on the host, so on a node with panic_on_oops set it is a whole-node outage for every co-resident tenant.","attack_vector":"Pure guest-side: emit an unaligned store (e.g. a 16-byte store at page offset 0xffc) where the second page carries a datamatch ioeventfd at offset 0 - a virtio doorbell is exactly such an ioeventfd, so every VM with virtio devices has the target. No host privilege needed.","remediation":"Update to a kernel with the referenced stable commits. No meaningful interim control - ioeventfds are how virtio doorbells work. Set panic_on_oops deliberately: leaving it off keeps the blast radius to the one VM rather than the node.","references":["https://git.kernel.org/stable/c/2426c15c1395b7d5ccf1e5025ca898af7f3decb6","https://git.kernel.org/stable/c/4186c850789906b875a1d263377a4d37c078e317","https://nvd.nist.gov/vuln/detail/CVE-2026-63806"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-64172","cve":"CVE-2026-64172","aliases":["Erratum 1235"],"title":"Linux KVM/SVM - AVIC IPI virtualization on Hygon Family 18h: AVIC inter-processor-interrupt virtualization is unsafe on","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux KVM/SVM - AVIC IPI virtualization on Hygon Family 18h","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"AVIC inter-processor-interrupt virtualization is unsafe on Hygon Family 18h parts, which are derived from AMD Family 17h. With AVIC active, a guest's IPIs can be delivered incorrectly - interrupt delivery reaching the wrong target is a cross-VM correctness failure on the interrupt path, which is the same seam the Heckler-class attacks exploit.","attack_vector":"From inside a guest VM on affected Hygon silicon with AVIC enabled.","remediation":"Fixed in the Linux kernel by disabling AVIC IPI virtualization on affected parts. Distro kernel update plus host reboot. Interim mitigation: disable AVIC in the kvm_amd module parameters (avic=0), which costs some interrupt-heavy performance but needs only a module reload with guests drained. Relevant only if you have Hygon parts in the fleet - worth checking, since they show up in some regional supply chains.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-64172"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-07-19"},{"id":"CVE-2026-65918","cve":"CVE-2026-65918","aliases":[],"title":"torchvision (GIF decoder): Out-of-bounds heap read in `read_from_tensor` GIF decode","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"torchvision (GIF decoder)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"Out-of-bounds heap read in `read_from_tensor` GIF decode","attack_vector":"Customer-supplied image data reaching a vision preprocessing pipeline","remediation":"Rebuild images with torchvision > 0.28.0","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-65918"],"status":"curated","published":"2026-07-23"},{"id":"CVE-2026-68103","cve":"CVE-2026-68103","aliases":[],"title":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu): A correctness defect in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu user-mode queues (doorbell submission path) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu: reject mapping a reserved doorbell to a new queue","attack_vector":"Local. Reachable by any process with a render node open that can create user-mode queues - the normal ROCm submission path, reachable from an unprivileged container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68103","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-10"},{"id":"CVE-2026-68258","cve":"CVE-2026-68258","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A correctness defect in the amdkfd (KFD compute","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"A correctness defect in the amdkfd (KFD compute driver, /dev/kfd) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdkfd: Check bounds on CRIU restore queue type and mqd size","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68258","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-10"},{"id":"CVE-2026-68293","cve":"CVE-2026-68293","aliases":["net/mlx5 MCIA register buffer overflow on 32 dword reads"],"title":"Linux kernel mlx5_core port / transceiver module EEPROM (MCIA register): The MCIA register can return 32 dwords","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_core port / transceiver module EEPROM (MCIA register)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"The MCIA register can return 32 dwords when the device advertises the capability, but the kernel's structure defines only 12, so reading module EEPROM copies past the end of the buffer. In practice this panics the host when an operator or monitoring agent runs ethtool against an optical module - so routine transceiver telemetry becomes a way to take a node down, and any agent that polls optics fleet-wide becomes a fleet-wide outage trigger.","attack_vector":"Local, low-privileged in the sense that the read is triggered from ethtool module-EEPROM queries; the overflowing length comes from what the adapter firmware advertises.","remediation":"Upgrade the host kernel to 7.2 or a stable backport (6.12.101, 6.18.42, 7.1.6). Rolling reboot. Interim: stop automated optics/EEPROM polling on affected kernels - a monitoring config change, no reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68293","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-68293.json"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-10"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:H","cwe":["CWE-125","CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-68420","cve":"CVE-2026-68420","aliases":[],"title":"Linux kernel (net/xfrm): Outbound policies rejected optional tunnel and BEET templates but never got the same check for","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"Outbound policies rejected optional tunnel and BEET templates but never got the same check for IPTFS, so an IPTFS template marked 'level use' walks past the end of the template address array during outbound resolution. The result is a stack out-of-bounds read used to pick a security association - kernel stack contents leaked into SA selection, and a crash when the read lands badly. A tenant able to install policies gets both an information-disclosure primitive and a node crash.","attack_vector":"Two commands and a ping, per the upstream reproducer: install an outbound policy whose first template is 'mode iptfs level use', then send any packet matching the selector. That needs the ability to write xfrm policy - host root, or a tenant container holding CAP_NET_ADMIN in its own user+network namespace. Inbound and forward policies are unaffected. Conditional on IPTFS support being present in the kernel.","remediation":"Boot a kernel carrying the fix commits below (no fixed stable version published). Interim control: do not grant CAP_NET_ADMIN inside tenant user namespaces; if tenants must program IPsec, reject templates with mode iptfs and level use in whatever admission layer sits in front of the xfrm netlink socket.","references":["https://git.kernel.org/stable/c/d7fc6f351c478586980a521d63b0214d9c055e78","https://git.kernel.org/stable/c/9333f4b6f44858fc98eb12bf26b8d2959eb975d5","https://nvd.nist.gov/vuln/detail/CVE-2026-68420"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:N/UI:N/S:U/C:N/I:L/A:H","cwe":["CWE-345","CWE-940"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-68425","cve":"CVE-2026-68425","aliases":[],"title":"Linux kernel InfiniBand MAD layer (kernel RMPP receive reassembly, ib_mad): This is a pre-authentication flaw on the","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel InfiniBand MAD layer (kernel RMPP receive reassembly, ib_mad)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"This is a pre-authentication flaw on the InfiniBand management plane. The kernel started RMPP reassembly for an inbound DATA response purely on the high bits of the transaction ID, before matching the full TID and source address against any outstanding request. An unsolicited response injected onto the fabric can therefore allocate and extend kernel RMPP receive state that no local agent ever asked for - reassembly-state exhaustion and management-agent disruption driven by a peer on the same IB subnet, with no credential of any kind. On a shared fabric the adjacent 'peer' can be another tenant's node, so this is a cross-tenant reach into the layer that carries subnet-manager traffic.","attack_vector":"Adjacent network - any host able to emit MADs onto the same InfiniBand subnet, unauthenticated. That includes every compute node on a shared fabric, and any tenant that has been given umad access on such a node.","remediation":"Kernel update that drops unmatched RMPP DATA responses before reassembly begins. Compensating controls are weak here: IB partitioning (pkeys) does not separate the management class, and M_Key protection only guards SMPs, not the GMP/RMPP path this affects. Patch the fabric-attached nodes.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=45416c87ebcece1e90f3bc5bc172d106b77c6b69","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-68425.json"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-68447","cve":"CVE-2026-68447","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A memory or reference-count leak in the amdkfd (KFD","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"A memory or reference-count leak in the amdkfd (KFD compute driver, /dev/kfd). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amdkfd: clamp v9 CRIU control stack checkpoint copy to BO size","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68447","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-12"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:C/C:N/I:N/A:H","cwe":["CWE-617"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-72284","cve":"CVE-2026-72284","aliases":[],"title":"Linux kernel (arch/x86/kvm): A guest that disables paravirtual EOI while KVM still has a pending PV-EOI request, and","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (arch/x86/kvm)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"A guest that disables paravirtual EOI while KVM still has a pending PV-EOI request, and arranges for the read of the PV-EOI area to fail, hits a kernel BUG in the LAPIC sync path. Denial of service rather than escape, but it is one tenant reliably crashing host kernel context - a node-level outage risk for everyone sharing the machine.","attack_vector":"Entirely guest-driven and reproduced by a small in-guest test program: unmap or invalidate the PV-EOI page so the host's read fails, then disable PV EOI via the KVM paravirt MSR. Applies to any guest using KVM paravirt EOI, which is the default for Linux guests on KVM.","remediation":"Update to a kernel with the referenced stable commits. Interim: disable the KVM PV EOI feature bit in the guest CPU model on unpatched nodes, and keep panic_on_oops off there so the crash stays inside the offending VM's kernel context.","references":["https://git.kernel.org/stable/c/038b9ce6fafda1babd1e33d52cbc6039747a6d87","https://git.kernel.org/stable/c/8e9f7a95279bf608cf4c331ed89612e28c04564f","https://nvd.nist.gov/vuln/detail/CVE-2026-72284"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-476","CWE-415"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-74395","cve":"CVE-2026-74395","aliases":[],"title":"NVIDIA/Mellanox ConnectX driver (mlx5_ib DEVX subscribe-event unwind): DEVX is the raw device-command escape hatch that","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA/Mellanox ConnectX driver (mlx5_ib DEVX subscribe-event unwind)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"DEVX is the raw device-command escape hatch that lets a userspace RDMA library drive the adapter directly. The subscribe-event handler linked its event object into the shared subscription list before initialising the fields the error path uses, so a failing eventfd_ctx_fdget() - trivially forced by passing a bad file descriptor - dereferences an unset ev_file and calls the xarray deallocator with an unset key. A tenant chooses when to fail, which makes the unwind path reachable on demand on the interface that speaks directly to adapter firmware.","attack_vector":"Local, unprivileged. A tenant calls the DEVX subscribe-event ioctl with an invalid eventfd.","remediation":"Kernel update ordering the initialisation before list insertion and making the xarray deallocation exactly-once. If DEVX is not needed by tenant workloads, restricting it is the sharpest control - but note that some accelerated userspace libraries require it, so check before disabling.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=1025dc2f7ba29b04b8687790fa91f9cd1a53141e","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-74395.json"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:L/A:H","cwe":["CWE-862"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2026-75109","cve":"CVE-2026-75109","aliases":[],"title":"Determined AI (master API, generic task kill/pause/unpause handlers): The generic task kill, pause and unpause","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Determined AI (master API, generic task kill/pause/unpause handlers)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"The generic task kill, pause and unpause endpoints run no ownership check, so any authenticated user terminates or suspends another user's tasks. On a shared GPU cluster that means one tenant can wipe out days of another team's training run and free their GPUs, with no permission escalation required.","attack_vector":"Any authenticated Determined user with network access to the master API. Affects versions up to and including 0.38.1.","remediation":"Upgrade Determined past 0.38.1 and restart the master. Until then, check master audit logs for kill/pause calls whose caller does not match the task owner - the behaviour is indistinguishable from a legitimate request in the task's own history.","references":["https://github.com/determined-ai/determined/issues/10270","https://nvd.nist.gov/vuln/detail/CVE-2026-75109"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:L/A:H","cwe":["CWE-400","CWE-923"],"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2026-057-flux-cd-allow-webhooks-networkpo","cve":null,"aliases":["GHSA-mwcp-qpcg-fr7c"],"title":"Flux CD (allow-webhooks NetworkPolicy, notification-controller event server): CROSS-TENANT EVENT FORGERY: the","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Flux CD (allow-webhooks NetworkPolicy, notification-controller event server)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"CROSS-TENANT EVENT FORGERY: the NetworkPolicy Flux installs by default, allow-webhooks, admits ingress to notification-controller from every namespace and names no port, so the event server on 9090 — whose only security boundary is namespace isolation — is reachable by every pod in the cluster. Any compromised or malicious workload in any namespace then injects forged Flux events. Concretely that means denial of service against the event server and the downstream notification providers, forged alerts landing in other tenants' Slack, Teams and PagerDuty channels, attacker-authored comments posted to a tenant's pull requests where the PR URL is known, and forged commit statuses where a commit hash in the target repo is known. The alerting channel operators rely on to tell them something is wrong becomes something a neighbour tenant can write to — and a forged green commit status is a nudge toward shipping code that was never actually validated.","attack_vector":"Network / in-cluster, low privileges: any pod in any namespace of a cluster bootstrapped with the Flux CLI defaults can reach notification-controller on port 9090 and post events. No credentials for the notification providers are needed.","remediation":"Upgrade to Flux v2.9.4, which patches the allow-webhooks policy and includes notification-controller v1.9.3 with the DoS mitigation. If you cannot upgrade now, apply the port restriction as a bootstrap customization: patch the allow-webhooks NetworkPolicy in your gotk-components kustomization so ingress is limited to the ports and sources that actually need it. Review notification channels and PR comment history for forged entries during the exposure window.","references":["https://github.com/fluxcd/flux2/security/advisories/GHSA-mwcp-qpcg-fr7c","https://github.com/fluxcd/flux2/pull/6028"],"status":"curated"},{"id":"CVE-2019-11983","cve":"CVE-2019-11983","aliases":["HPESBHF03917"],"title":"HPE iLO 4 / iLO 5 (remote buffer overflow): Remotely triggerable buffer overflow in the iLO firmware on both the Gen9","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HPE iLO 4 / iLO 5 (remote buffer overflow)","year":"2019","cvss_score":7,"severity":"high","kev":false,"impact":"Remotely triggerable buffer overflow in the iLO firmware on both the Gen9 (iLO 4) and Gen10 (iLO 5) generations. Memory corruption on a service processor is the highest-value bug class in a fleet, because success means code on the BMC and therefore power control, Virtual Media, console, and an implant that persists across host reinstalls. Worth noting alongside the older iLO 4 authentication-bypass work that made this platform a known research target - the Gen9 tier tends to be the part of a fleet that stopped receiving attention.","attack_vector":"Reachable over the network to the iLO address on the out-of-band management VLAN.","remediation":"Flash iLO 4 to v2.61b or later and iLO 5 to v1.39 or later. Out-of-band, per-node, no host reboot and no job drain. On Gen9 hardware the practical obstacle is inventory and change control rather than the flash itself - these are usually the nodes with the least recent firmware campaign. Interim control: hard-ACL the iLO management network.","references":["https://support.hpe.com/hpsc/doc/public/display?docLocale=en_US&docId=emr_na-hpesbhf03917en_us","https://nvd.nist.gov/vuln/detail/CVE-2019-11983"],"status":"curated","published":"2019-06-05"},{"id":"CVE-2019-19921","cve":"CVE-2019-19921","aliases":[],"title":"runc: Volume-mount race gives incorrect access control and privilege escalation to host","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"runc","year":"2019","cvss_score":7,"severity":"high","kev":false,"impact":"Volume-mount race gives incorrect access control and privilege escalation to host","attack_vector":"Any tenant workload able to spawn two containers with crafted mounts","remediation":"Replace runc binary; drain node","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-19921"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2020-02-12"},{"id":"CVE-2021-1099","cve":"CVE-2021-1099","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): Stack buffer overflow in the vGPU Manager with enough control for a guest to place a","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":7,"severity":"high","kev":false,"impact":"Stack buffer overflow in the vGPU Manager with enough control for a guest to place a ROP chain on the host stack. This is the most explicitly weaponisable of the 2021 vGPU set - guest-to-host code execution on the hypervisor. vGPU 12.x before 12.3, 11.x before 11.5, 8.x before 8.8.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU on the affected host.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1099"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2021-07-21"},{"id":"CVE-2021-1120","cve":"CVE-2021-1120","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): A guest-supplied string may not be null-terminated, and the host plugin reads past","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":7,"severity":"high","kev":false,"impact":"A guest-supplied string may not be null-terminated, and the host plugin reads past it - disclosure, tampering or unauthorised code execution on the host, though NVIDIA notes the guest cannot choose the content it pushes.","attack_vector":"Any user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1120"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2021-10-29"},{"id":"CVE-2021-20188","cve":"CVE-2021-20188","aliases":[],"title":"Podman: File permissions not checked for non-root users in a privileged container","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Podman","year":"2021","cvss_score":7,"severity":"high","kev":false,"impact":"File permissions not checked for non-root users in a privileged container; cross-user file access","attack_vector":"Any tenant workload in a privileged container","remediation":"Upgrade Podman","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-20188"],"status":"curated","published":"2021-02-11"},{"id":"CVE-2021-22543","cve":"CVE-2021-22543","aliases":[],"title":"KVM: Improper handling of VM_IO/VM_PFNMAP vmas in KVM lets a guest bypass read-only checks","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"KVM","year":"2021","cvss_score":7,"severity":"high","kev":false,"impact":"Improper handling of VM_IO/VM_PFNMAP vmas in KVM lets a guest bypass read-only checks - host privilege escalation","attack_vector":"Tenant VM guest / local user with /dev/kvm","remediation":"Kernel patch. Livepatchable on some vendors; KVM module changes often are not - budget drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2021-22543"],"status":"curated","published":"2021-05-26"},{"id":"CVE-2021-22600","cve":"CVE-2021-22600","aliases":[],"title":"Linux kernel (af_packet): Double free in packet_set_ring(), local privilege escalation","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (af_packet)","year":"2021","cvss_score":7,"severity":"high","kev":true,"impact":"Double free in packet_set_ring(), local privilege escalation [KEV]","attack_vector":"Any tenant process in a container with CAP_NET_RAW","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2021-22600"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2022-01-26"},{"id":"CVE-2021-4204","cve":"CVE-2021-4204","aliases":[],"title":"Linux kernel (eBPF): eBPF improper input validation leading to local privilege escalation","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (eBPF)","year":"2021","cvss_score":7,"severity":"high","kev":false,"impact":"eBPF improper input validation leading to local privilege escalation","attack_vector":"Any tenant process in a container with BPF access","remediation":"Livepatchable; otherwise drain + reboot. Set `kernel.unprivileged_bpf_disabled=1`","references":["https://access.redhat.com/security/cve/CVE-2021-4204"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2022-08-24"},{"id":"CVE-2022-0492","cve":"CVE-2022-0492","aliases":[],"title":"Linux kernel (cgroups v1): cgroups v1 release_agent lets a container with CAP_SYS_ADMIN (or an unconfined userns) run","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (cgroups v1)","year":"2022","cvss_score":7,"severity":"high","kev":true,"impact":"cgroups v1 release_agent lets a container with CAP_SYS_ADMIN (or an unconfined userns) run arbitrary host-root commands [KEV]","attack_vector":"Any tenant process in a container","remediation":"Livepatchable; otherwise drain + reboot. Strong compensating controls: cgroup v2 only, seccomp/AppArmor default profiles, drop CAP_SYS_ADMIN","references":["https://access.redhat.com/security/cve/CVE-2022-0492"],"status":"curated","fleet":{"ubiquity":"Universal - the `release_agent` path exists on every pre-5.17 kernel; cgroups v1 was still the default on most GPU host images","remediation_pain":"`node-reboot` - kernel upgrade; the only mitigation is relying on AppArmor/SELinux/seccomp being correctly applied, which is exactly what privileged AI workloads often disable","pain_class":"node-reboot","why_fleet_wide":"A root-in-container process abuses user namespaces to mount cgroups v1, sets `release_agent` and runs arbitrary commands as host root; GPU containers routinely run privileged or with `--cap-add`, which removes the default hardening that would have blocked it"},"published":"2022-03-03"},{"id":"CVE-2022-2602","cve":"CVE-2022-2602","aliases":[],"title":"Linux kernel (io_uring): Use-after-free between io_uring and the unix GC - local root","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (io_uring)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"Use-after-free between io_uring and the unix GC - local root","attack_vector":"Any tenant process in a container with io_uring enabled","remediation":"Livepatchable; otherwise drain + reboot. Durable control: block io_uring via seccomp in the default container profile","references":["https://access.redhat.com/security/cve/CVE-2022-2602"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-01-08"},{"id":"CVE-2022-31614","cve":"CVE-2022-31614","aliases":[],"title":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko): The vGPU plugin double-frees host","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"The vGPU plugin double-frees host resources; chained with another bug it reaches code execution on the hypervisor host. The attacker sits inside a tenant guest VM and the blast radius is the hypervisor host, which is running every other tenant's vGPU on the same physical GPU. This is precisely the boundary a vGPU-based multi-tenant offering is sold on.","attack_vector":"A tenant inside their own guest VM, driving the paravirtualised vGPU control interface. Several of these need only an unprivileged process in the guest; the rest need guest root, which a tenant already has on a VM they rented. No host credentials are involved at any point.","remediation":"Patch the vGPU Manager on the hypervisor host per bulletin 5383. Cost: the highest of any class here. The host driver cannot be reloaded while vGPUs are attached, so every tenant VM on that hypervisor must be live-migrated or powered off - a full host drain. NVIDIA also enforces a supported host/guest driver skew, so budget a matching guest-driver campaign in the same window or tenants lose their vGPU on next boot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31614","https://github.com/NVIDIA/product-security/tree/main/2022/5383"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-415"],"fleet":{"pain_class":"node-drain"},"tags":["tenant-isolation"],"published":"2022-08-05"},{"id":"CVE-2022-32469","cve":"CVE-2022-32469","aliases":["INSYDE-SA-2023001"],"title":"Insyde InsydeH2O (PnpSmm shared SMM/non-SMM buffer, DMA TOCTOU): A buffer shared between SMM and non-SMM code","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (PnpSmm shared SMM/non-SMM buffer, DMA TOCTOU)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"A buffer shared between SMM and non-SMM code in the plug-and-play driver can be rewritten by DMA between validation and use, corrupting SMRAM and escalating privilege. PnpSmm owns SMBIOS/platform description, so alongside ring -2 the attacker can poison the hardware inventory your fleet tooling trusts.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Insyde lists kernel 5.0 through 5.5 affected; take the per-kernel fixed version from the advisory.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. This is Insyde's second pass at the same defect class in a different set of buffers - a fleet that took the 2022 BIOS release is NOT covered for this batch, and OEM release notes rarely make that distinction clear. Verify by kernel version, not by 'we patched the Insyde DMA bugs'.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-32469","https://www.insyde.com/security-pledge/SA-2023001"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-02-15"},{"id":"CVE-2022-32470","cve":"CVE-2022-32470","aliases":["INSYDE-SA-2023002"],"title":"Insyde InsydeH2O (FwBlockServiceSmm shared buffer, DMA TOCTOU): The firmware block service's shared buffer is racy, so","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (FwBlockServiceSmm shared buffer, DMA TOCTOU)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"The firmware block service's shared buffer is racy, so DMA lands an attacker on the SPI flash write path with SMRAM corruption alongside it. Insyde's own suggested fix - copy the firmware block services data into SMRAM before checking it - tells you the shape of the bug: validation happening on memory the attacker still owns.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Insyde lists kernel 5.0 through 5.5 affected; take the per-kernel fixed version from the advisory. Confirm SPI flash write protection is enforced in the interim. The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. This is Insyde's second pass at the same defect class in a different set of buffers - a fleet that took the 2022 BIOS release is NOT covered for this batch, and OEM release notes rarely make that distinction clear. Verify by kernel version, not by 'we patched the Insyde DMA bugs'.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-32470","https://www.insyde.com/security-pledge/SA-2023002"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-02-15"},{"id":"CVE-2022-32471","cve":"CVE-2022-32471","aliases":["INSYDE-SA-2023003"],"title":"Insyde InsydeH2O (IhisiSmm / IhisiDxe command buffer): One representative of a family of roughly a dozen Insyde","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (IhisiSmm / IhisiDxe command buffer)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"One representative of a family of roughly a dozen Insyde advisories covering the same defect across different drivers: the SMI handler validates its parameters in a buffer that sits outside SMRAM, then uses them - and a DMA-capable device can rewrite that buffer in between. The result is SMRAM corruption and privilege escalation to ring -2 driven from a peripheral rather than from the CPU. In a GPU chassis the DMA-capable devices are GPUs, NICs and NVMe drives, several of which run tenant-flashable firmware, so this is not a theoretical adversary.","attack_vector":"An attacker with control of a DMA-capable device on the node - a compromised NIC or GPU firmware, a malicious PCIe device, or a tenant who can drive DMA from a passed-through device - racing the SMI handler. Does not require host root, which is what makes the family interesting.","remediation":"OEM BIOS update on the fixed Insyde kernel. Firmware flash, reboot per node. The real compensating control here is the IOMMU: enable VT-d/AMD-Vi and DMA protection (including pre-boot DMA protection where the platform supports it) so untrusted devices cannot reach arbitrary host memory. That is a config change and should be standard on any multi-tenant GPU node regardless of this CVE. Note the sibling advisories SA-2023001 through SA-2023015 cover the same bug in PnpSmm, FwBlockServiceSmm, HddPassword, AhciBusDxe, IdeBusDxe, NvmExpressDxe, SdHostDriver, SdMmcDevice, StorageSecurityCommandDxe, VariableRuntimeDxe and FvbServicesRuntimeDxe - patching one does not patch the rest.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-32471","https://www.insyde.com/security-pledge/SA-2023003"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-02-15"},{"id":"CVE-2022-32473","cve":"CVE-2022-32473","aliases":["INSYDE-SA-2023005"],"title":"Insyde InsydeH2O (HddPassword shared buffer, DMA TOCTOU): Racy shared buffer in the ATA security driver","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (HddPassword shared buffer, DMA TOCTOU)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"Racy shared buffer in the ATA security driver. What this one reaches is drive locking and unlock credentials, so a successful race yields ring -2 plus a foothold in the code holding drive secrets - which matters for any operator using ATA drive locks as part of the between-tenant wipe.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Insyde lists kernel 5.0 through 5.5 affected; take the per-kernel fixed version from the advisory. Do not rely on ATA HDD passwords as the at-rest control on unpatched nodes. The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. This is Insyde's second pass at the same defect class in a different set of buffers - a fleet that took the 2022 BIOS release is NOT covered for this batch, and OEM release notes rarely make that distinction clear. Verify by kernel version, not by 'we patched the Insyde DMA bugs'.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-32473","https://www.insyde.com/security-pledge/SA-2023005"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-02-15"},{"id":"CVE-2022-32474","cve":"CVE-2022-32474","aliases":["INSYDE-SA-2023006"],"title":"Insyde InsydeH2O (StorageSecurityCommandDxe shared buffer, DMA TOCTOU): The TCG/Opal security-command driver shares","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (StorageSecurityCommandDxe shared buffer, DMA TOCTOU)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"The TCG/Opal security-command driver shares a buffer between SMM and non-SMM code without protecting it from DMA. The reachable target is self-encrypting-drive authentication - the mechanism a GPU cloud points at when asked how tenant data is protected at rest between leases.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Insyde lists kernel 5.0 through 5.5 affected; take the per-kernel fixed version from the advisory.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. This is Insyde's second pass at the same defect class in a different set of buffers - a fleet that took the 2022 BIOS release is NOT covered for this batch, and OEM release notes rarely make that distinction clear. Verify by kernel version, not by 'we patched the Insyde DMA bugs'.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-32474","https://www.insyde.com/security-pledge/SA-2023006"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-02-15"},{"id":"CVE-2022-32475","cve":"CVE-2022-32475","aliases":["INSYDE-SA-2023007"],"title":"Insyde InsydeH2O (VariableRuntimeDxe shared buffer, DMA TOCTOU): Racy shared buffer in the UEFI variable driver","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (VariableRuntimeDxe shared buffer, DMA TOCTOU)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"Racy shared buffer in the UEFI variable driver - the store holding the Secure Boot key databases (PK/KEK/db/dbx). Corrupting SMRAM through this driver attacks the arbiter of what firmware and bootloaders are permitted to run, so the loss is the boot-integrity guarantee itself rather than any one workload. Insyde notes the fix also hardened chipset and OEM chipset code, which means OEM-specific builds needed their own rebase on top of the kernel fix.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Insyde lists kernel 5.0 through 5.5 affected; take the per-kernel fixed version from the advisory.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. This is Insyde's second pass at the same defect class in a different set of buffers - a fleet that took the 2022 BIOS release is NOT covered for this batch, and OEM release notes rarely make that distinction clear. Verify by kernel version, not by 'we patched the Insyde DMA bugs'.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-32475","https://www.insyde.com/security-pledge/SA-2023007"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-02-15"},{"id":"CVE-2022-32476","cve":"CVE-2022-32476","aliases":["INSYDE-SA-2023008"],"title":"Insyde InsydeH2O (AhciBusDxe shared buffer, DMA TOCTOU): DMA race on the SATA/AHCI driver's shared buffer produces","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (AhciBusDxe shared buffer, DMA TOCTOU)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"DMA race on the SATA/AHCI driver's shared buffer produces SMRAM corruption and privilege escalation. Reaches the SATA storage path, typically the node's boot device - firmware persistence plus a position on the disk the node starts from.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Insyde lists kernel 5.0 through 5.5 affected; take the per-kernel fixed version from the advisory.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. This is Insyde's second pass at the same defect class in a different set of buffers - a fleet that took the 2022 BIOS release is NOT covered for this batch, and OEM release notes rarely make that distinction clear. Verify by kernel version, not by 'we patched the Insyde DMA bugs'.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-32476","https://www.insyde.com/security-pledge/SA-2023008"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-02-15"},{"id":"CVE-2022-32477","cve":"CVE-2022-32477","aliases":["INSYDE-SA-2023009"],"title":"Insyde InsydeH2O (FvbServicesRuntimeDxe shared buffer, DMA TOCTOU): Firmware Volume Block services again, this time","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (FvbServicesRuntimeDxe shared buffer, DMA TOCTOU)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"Firmware Volume Block services again, this time via the shared SMM/non-SMM buffer. Same consequence as its 2022 twin: an attacker influencing SPI flash writes, which is the difference between a node you can clean by reimaging and a node you have to physically reflash before it goes back in the pool.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Insyde lists kernel 5.0 through 5.5 affected; take the per-kernel fixed version from the advisory. Verify SPI write protection is set. The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. This is Insyde's second pass at the same defect class in a different set of buffers - a fleet that took the 2022 BIOS release is NOT covered for this batch, and OEM release notes rarely make that distinction clear. Verify by kernel version, not by 'we patched the Insyde DMA bugs'.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-32477","https://www.insyde.com/security-pledge/SA-2023009"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-02-15"},{"id":"CVE-2022-32478","cve":"CVE-2022-32478","aliases":["INSYDE-SA-2023010"],"title":"Insyde InsydeH2O (IdeBusDxe shared buffer, DMA TOCTOU): Racy shared buffer in the legacy IDE/ATA driver leading","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (IdeBusDxe shared buffer, DMA TOCTOU)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"Racy shared buffer in the legacy IDE/ATA driver leading to SMRAM corruption. As with its 2022 counterpart, this driver is generally only active where CSM/legacy storage compatibility is enabled, so a UEFI-only server profile may not expose it at all.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Insyde lists kernel 5.0 through 5.5 affected; take the per-kernel fixed version from the advisory. Disable CSM / legacy storage on UEFI-only nodes to remove the driver rather than patch it. The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. This is Insyde's second pass at the same defect class in a different set of buffers - a fleet that took the 2022 BIOS release is NOT covered for this batch, and OEM release notes rarely make that distinction clear. Verify by kernel version, not by 'we patched the Insyde DMA bugs'.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-32478","https://www.insyde.com/security-pledge/SA-2023010"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-02-15"},{"id":"CVE-2022-32953","cve":"CVE-2022-32953","aliases":["INSYDE-SA-2023013"],"title":"Insyde InsydeH2O (SdHostDriver shared buffer, DMA TOCTOU): DMA race on the SD host controller's shared buffer","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (SdHostDriver shared buffer, DMA TOCTOU)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"DMA race on the SD host controller's shared buffer. Insyde's remediation note is more specific for this sub-group - copy the link data into SMRAM before checking it AND verify every pointer falls inside the buffer - which is worth knowing because it indicates the earlier fixes in this driver were incomplete rather than wrong.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Insyde lists kernel 5.0 through 5.5 affected; take the per-kernel fixed version from the advisory.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. This is Insyde's second pass at the same defect class in a different set of buffers - a fleet that took the 2022 BIOS release is NOT covered for this batch, and OEM release notes rarely make that distinction clear. Verify by kernel version, not by 'we patched the Insyde DMA bugs'.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-32953","https://www.insyde.com/security-pledge/SA-2023013"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-02-15"},{"id":"CVE-2022-32954","cve":"CVE-2022-32954","aliases":["INSYDE-SA-2023014"],"title":"Insyde InsydeH2O (SdMmcDevice shared buffer, DMA TOCTOU): Shared-buffer DMA race in the SD/MMC device layer, kernel 5.1","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (SdMmcDevice shared buffer, DMA TOCTOU)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"Shared-buffer DMA race in the SD/MMC device layer, kernel 5.1 through 5.5. Pairs with the SdHostDriver entry - Insyde consistently files the controller and device layers separately, so both need to be present in whatever BIOS release you accept as the fix.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Insyde lists kernel 5.0 through 5.5 affected; take the per-kernel fixed version from the advisory.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. This is Insyde's second pass at the same defect class in a different set of buffers - a fleet that took the 2022 BIOS release is NOT covered for this batch, and OEM release notes rarely make that distinction clear. Verify by kernel version, not by 'we patched the Insyde DMA bugs'.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-32954","https://www.insyde.com/security-pledge/SA-2023014"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-02-15"},{"id":"CVE-2022-32955","cve":"CVE-2022-32955","aliases":["INSYDE-SA-2023015"],"title":"Insyde InsydeH2O (NvmExpressDxe shared buffer, DMA TOCTOU): The NVMe driver's SMM/non-SMM shared buffer is racy, giving","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (NvmExpressDxe shared buffer, DMA TOCTOU)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"The NVMe driver's SMM/non-SMM shared buffer is racy, giving SMRAM corruption and ring -2 escalation. On a GPU node this is the driver sitting on the datasets, checkpoints and weights, and it is the second separately-filed NVMe DMA defect after SA-2022055 - strong evidence that this driver deserves standing attention in a firmware patch policy rather than case-by-case triage.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Insyde lists kernel 5.0 through 5.5 affected; take the per-kernel fixed version from the advisory.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. This is Insyde's second pass at the same defect class in a different set of buffers - a fleet that took the 2022 BIOS release is NOT covered for this batch, and OEM release notes rarely make that distinction clear. Verify by kernel version, not by 'we patched the Insyde DMA bugs'.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-32955","https://www.insyde.com/security-pledge/SA-2023015"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-02-15"},{"id":"CVE-2022-33905","cve":"CVE-2022-33905","aliases":["INSYDE-SA-2022047"],"title":"Insyde InsydeH2O (AhciBusDxe SMI input buffer, DMA TOCTOU): DMA race on the SATA/AHCI controller driver's SMI input","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (AhciBusDxe SMI input buffer, DMA TOCTOU)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"DMA race on the SATA/AHCI controller driver's SMI input buffer yields SMRAM corruption and escalation to ring -2. What this driver reaches is the SATA storage path - on a GPU node that is typically the boot drive or a bulk data volume, so an attacker lands both firmware persistence and a position astride the disk the node boots from.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in kernel 5.2 / 05.27.23, 5.3 / 05.36.23, 5.4 / 05.44.23, 5.5 / 05.52.23.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-33905","https://www.insyde.com/security-pledge/SA-2022047"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-11-15"},{"id":"CVE-2022-33908","cve":"CVE-2022-33908","aliases":["INSYDE-SA-2022050"],"title":"Insyde InsydeH2O (SdHostDriver SMI input buffer, DMA TOCTOU): DMA race on the SD host controller driver gives SMRAM","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (SdHostDriver SMI input buffer, DMA TOCTOU)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"DMA race on the SD host controller driver gives SMRAM corruption and ring -2 escalation. On server hardware the SD/eMMC controller is often wired to platform or BMC-adjacent boot media rather than to anything a tenant uses, which makes this an easy driver to forget about - and an attractive one for an attacker, because nobody is watching it.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in kernel 5.2 / 05.27.25, 5.3 / 05.36.25, 5.4 / 05.44.25, 5.5 / 05.52.25. Where the platform exposes it, disabling the unused SD/eMMC controller in BIOS removes the surface. The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-33908","https://www.insyde.com/security-pledge/SA-2022050"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-11-15"},{"id":"CVE-2022-33909","cve":"CVE-2022-33909","aliases":["INSYDE-SA-2022051"],"title":"Insyde InsydeH2O (HddPassword SMI input buffer, DMA TOCTOU): The HddPassword driver handles ATA security","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (HddPassword SMI input buffer, DMA TOCTOU)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"The HddPassword driver handles ATA security - drive locking and unlock credentials. A DMA race here corrupts SMRAM and puts the attacker inside the code that holds drive-unlock secrets in memory, so the reachable prize is not just ring -2 but the credentials protecting the drive. Relevant to any fleet that leans on ATA drive locking as part of its between-tenant wipe story.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in kernel 5.2 / 05.27.23, 5.3 / 05.36.23, 5.4 / 05.44.23, 5.5 / 05.52.23. Do not treat ATA HDD passwords as the confidentiality control on unpatched nodes - prefer SED keys or software FDE with keys held off the node. The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-33909","https://www.insyde.com/security-pledge/SA-2022051"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-11-15"},{"id":"CVE-2022-33983","cve":"CVE-2022-33983","aliases":["INSYDE-SA-2022053"],"title":"Insyde InsydeH2O (NvmExpressLegacy SMI input buffer, DMA TOCTOU): DMA race on the legacy NVMe SMI handler","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (NvmExpressLegacy SMI input buffer, DMA TOCTOU)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"DMA race on the legacy NVMe SMI handler. NVMe is the data path on a GPU node - datasets, checkpoints, model weights - so a driver that reaches NVMe reaches tenant data as well as SMRAM. This is the legacy-path twin of the NvmExpressDxe issue filed as SA-2022055; both ship, and patching one does not patch the other.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in kernel 5.2 / 05.27.25, 5.3 / 05.36.25, 5.4 / 05.44.25, 5.5 / 05.52.25.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-33983","https://www.insyde.com/security-pledge/SA-2022053"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-11-15"},{"id":"CVE-2022-33984","cve":"CVE-2022-33984","aliases":["INSYDE-SA-2022054"],"title":"Insyde InsydeH2O (SdMmcDevice SMI input buffer, DMA TOCTOU): SMRAM corruption through a DMA race on the SD/MMC device","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (SdMmcDevice SMI input buffer, DMA TOCTOU)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"SMRAM corruption through a DMA race on the SD/MMC device driver. Pairs with the SdHostDriver issue (SA-2022050) - Insyde filed the controller and the device layer separately, so a fleet that patched one BIOS release for 'the SD bug' may still be carrying the other.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in kernel 5.2 / 05.27.25, 5.3 / 05.36.25, 5.4 / 05.44.25, 5.5 / 05.52.25.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-33984","https://www.insyde.com/security-pledge/SA-2022054"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-11-15"},{"id":"CVE-2022-33985","cve":"CVE-2022-33985","aliases":["INSYDE-SA-2022055"],"title":"Insyde InsydeH2O (NvmExpressDxe SMI input buffer, DMA TOCTOU): DMA race on the primary NVMe driver's SMI input buffer","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (NvmExpressDxe SMI input buffer, DMA TOCTOU)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"DMA race on the primary NVMe driver's SMI input buffer gives SMRAM corruption and ring -2 escalation. This is the one that matters most on a modern GPU server: NVMe is where the training data, checkpoints and weights live, and the same driver that touches them is the one exposing a racy SMI handler.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in kernel 5.2 / 05.27.25, 5.3 / 05.36.25, 5.4 / 05.44.25, 5.5 / 05.52.25.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-33985","https://www.insyde.com/security-pledge/SA-2022055"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-11-15"},{"id":"CVE-2022-50079","cve":"CVE-2022-50079","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Check correct bounds for stream encoder instances for DCN303","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-50079","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-06-18"},{"cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-50129","cve":"CVE-2022-50129","aliases":[],"title":"Linux kernel SRP target (ib_srpt, LIO port lifetime vs RDMA port lifetime): The SRP target's port structures were owned","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel SRP target (ib_srpt, LIO port lifetime vs RDMA port lifetime)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"The SRP target's port structures were owned by the RDMA core while the LIO target port data inside them was owned by the SCSI target subsystem, and the two lifetimes were not decoupled - KASAN caught a use-after-free in srpt_enable_tpg. This is on the target side of SCSI-over-RDMA, the process exporting block devices to the cluster, so a use-after-free during target reconfiguration lands in the daemon that mediates every initiator's access to those devices.","attack_vector":"Local on the storage target, racing RDMA port teardown against LIO target-portal-group configuration. Requires the ability to drive target configuration or to time an RDMA port event against it.","remediation":"Kernel update decoupling srpt_port and srpt_port_id lifetimes. Operationally: do not reconfigure LIO target portal groups while RDMA ports are being brought up or down - sequence storage-target maintenance rather than overlapping it.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=388326bb1c32fcd09371c1d494af71471ef3a04b","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2022/CVE-2022-50129.json"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2023-0386","cve":"CVE-2023-0386","aliases":[],"title":"Linux kernel (OverlayFS/FUSE): OverlayFS copies setuid files from a nosuid FUSE mount","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (OverlayFS/FUSE)","year":"2023","cvss_score":7,"severity":"high","kev":true,"impact":"OverlayFS copies setuid files from a nosuid FUSE mount - unprivileged local user to root, a live container-escape chain [KEV]","attack_vector":"Any tenant process in a container with a user namespace","remediation":"Livepatchable; otherwise drain + reboot. Compensating control: disallow unprivileged FUSE mounts","references":["https://access.redhat.com/security/cve/CVE-2023-0386"],"status":"curated","fleet":{"ubiquity":"Universal - kernels 5.11-6.1.8, and OverlayFS *is* the container storage driver on every containerized GPU host","remediation_pain":"`node-reboot` - kernel upgrade; no meaningful runtime mitigation since disabling OverlayFS breaks container storage","pain_class":"node-reboot","why_fleet_wide":"Copying a setuid binary across a `nosuid` OverlayFS mount preserves capabilities, giving any regular user root; inside a container it gives container-root, which then chains into the cgroup/procfs escapes above on a shared GPU host"},"published":"2023-03-22"},{"id":"CVE-2023-27561","cve":"CVE-2023-27561","aliases":[],"title":"runc: Regression of CVE-2019-19921: incorrect access control leading to privilege escalation via volume mounts","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"runc","year":"2023","cvss_score":7,"severity":"high","kev":false,"impact":"Regression of CVE-2019-19921: incorrect access control leading to privilege escalation via volume mounts","attack_vector":"Any tenant workload able to spawn two containers with custom mounts","remediation":"Replace runc binary; drain node","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-27561"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2023-03-03"},{"id":"CVE-2023-41914","cve":"CVE-2023-41914","aliases":[],"title":"Slurm: Filesystem race conditions allow gaining ownership of, overwriting, or deleting files","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Slurm","year":"2023","cvss_score":7,"severity":"high","kev":false,"impact":"Filesystem race conditions allow gaining ownership of, overwriting, or deleting files","attack_vector":"Any user who can submit a job","remediation":"Upgrade Slurm","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-41914"],"status":"curated","fleet":{"ubiquity":"Very common - Slurm 22.05.x / 23.02.x branches","remediation_pain":"`daemon-restart` - upgrade slurmd/slurmctld on every node; running jobs survive a slurmd restart in many configs, so this is cheaper than a runc bug but still fleet-wide","pain_class":"daemon-restart","why_fleet_wide":"TOCTOU race in Slurm file handling lets a low-privilege job take ownership of, overwrite or delete arbitrary files, escalating on any compute node the scheduler touches"},"published":"2023-11-03"},{"id":"CVE-2023-46813","cve":"CVE-2023-46813","aliases":[],"title":"Linux kernel SEV-ES #VC handler - MMIO access checking: Incorrect access checking in the SEV-ES #VC handler and","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel SEV-ES #VC handler - MMIO access checking","year":"2023","cvss_score":7,"severity":"high","kev":false,"impact":"Incorrect access checking in the SEV-ES #VC handler and instruction emulation lets a local user with userspace access to MMIO registers escalate. Inside a confidential VM, a merely local user reaches privileged guest state through the exception handler that SEV-ES uses to virtualise MMIO - so the confidential VM's own internal privilege boundary breaks, not just the host/guest one.","attack_vector":"Local, from userspace inside an SEV-ES guest that has userspace-accessible MMIO. Affects Linux before 6.5.9.","remediation":"Fixed in the Linux kernel. Take the distro kernel update (RHEL/Rocky, Ubuntu, SLES) and reboot the host - no firmware, VBIOS or AGESA step. On a GPU fleet this is a cordon, drain and rolling reboot; plan it as normal kernel maintenance. The fix belongs in the **guest** kernel, so update your confidential-VM images (or publish a minimum guest kernel to tenants) rather than assuming host patching covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-46813"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2023-10-27"},{"id":"CVE-2023-51767","cve":"CVE-2023-51767","aliases":["OpenSSH single-bit auth bypass","Rowhammer-assisted authentication bypass"],"title":"OpenSSH through 10.0 - mm_answer_authpassword uses an integer 'authenticated' flag that does not resist a single bit","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"OpenSSH through 10.0 - mm_answer_authpassword uses an integer 'authenticated' flag that does not resist a single bit…","year":"2023","cvss_score":7,"severity":"high","kev":false,"impact":"Shows the second half of the Rowhammer chain that operators usually skip: the flip has to land somewhere useful, and a lot of privileged software stores 'is this user authenticated' as a plain integer that a one-bit change turns from false to true. A co-resident attacker who can hammer the sshd privilege-separation monitor's memory authenticates without credentials. Sudo had the identical pattern. For a bare-metal or shared-host fleet, this converts 'Rowhammer is a research curiosity' into 'a co-tenant logs in as root on the host'.","attack_vector":"Requires attacker-victim co-location - unprivileged local code on the same machine as the sshd being attacked, with hammerable DRAM. It is a threat model that only exists on shared hosts, which is exactly the neocloud and multi-tenant cluster model.","remediation":"Patch OpenSSH and sudo to versions with the hardened comparisons (sudo 1.9.15 and later; OpenSSH per your distribution's backport) - that is a normal package update with an sshd restart, no reboot. But understand what it buys you: it removes two known landing spots, not the underlying ability to flip bits. Treat this as a prompt to audit other privileged daemons in your control plane for single-integer authorisation flags, and to fix the DRAM-sharing policy underneath.","references":["https://arxiv.org/abs/2309.02545","https://nvd.nist.gov/vuln/detail/CVE-2023-51767","https://nvd.nist.gov/vuln/detail/CVE-2023-42465"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2023-12-24"},{"cwe":["CWE-415","CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-52851","cve":"CVE-2023-52851","aliases":[],"title":"NVIDIA/Mellanox ConnectX driver (mlx5_ib UMR resource init/cleanup): A failed workqueue allocation during mkey-cache","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA/Mellanox ConnectX driver (mlx5_ib UMR resource init/cleanup)","year":"2023","cvss_score":7,"severity":"high","kev":false,"impact":"A failed workqueue allocation during mkey-cache init caused the UMR queue pair to be destroyed twice, giving a double free and a syzkaller-confirmed use-after-free in ib_destroy_qp_user. The UMR queue pair is the driver's own privileged channel for rewriting memory-key translation tables, so a freed-and-reused UMR QP is a control channel over every tenant's memory keys on that adapter.","attack_vector":"Local. Triggered on the mlx5_ib device initialisation path under memory pressure or interrupted allocation - reachable in practice by a tenant that can drive the node into allocation failure while the RDMA stack is initialising.","remediation":"Kernel update removing the duplicate cleanup call in mlx5_ib_stage_post_ib_reg_umr_init(). No workaround.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=2ef422f063b74adcc4a4a9004b0a87bb55e0a836","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2023/CVE-2023-52851.json"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-53123","cve":"CVE-2023-53123","aliases":[],"title":"Linux kernel (arch/s390/pci): When an SR-IOV VF is hot-unplugged its MMIO resources are freed, but the parent bus keeps","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (arch/s390/pci)","year":"2023","cvss_score":7,"severity":"high","kev":false,"impact":"When an SR-IOV VF is hot-unplugged its MMIO resources are freed, but the parent bus keeps pointers to them in its resource list. When the VF is plugged back in - the normal churn of reclaiming a VF from one tenant and issuing it to the next - the stale resources are claimed again, giving a use-after-free on the structures that describe which MMIO window belongs to which function. Resource metadata is the wrong thing to have dangling on a machine that partitions devices between tenants.","attack_vector":"Host-side, and s390 only - the bug lives in the s390 PCI implementation, where individual functions of a multi-function/SR-IOV device can be hotplugged independently (the core change to drivers/pci/bus.c is the supporting API, not the flaw). x86_64 and arm64 GPU nodes are not affected. Triggered by the VF remove/re-add cycle, i.e. by the operator's own VF lifecycle automation rather than directly by a tenant; a tenant that can make its VF drop and reappear influences the timing.","remediation":"Boot a patched kernel on any s390 node (the fix adds pci_bus_remove_resource() and drops per-function resources on unplug). Interim on s390: avoid the VF remove/re-add cycle - reboot the LPAR rather than recycling individual VFs between workloads. No action needed on x86_64/arm64 GPU nodes.","references":["https://git.kernel.org/stable/c/437bb839e36cc9f35adc6d2a2bf113b7a0fc9985","https://git.kernel.org/stable/c/a2410d0c3d2d714ed968a135dfcbed6aa3ff7027","https://nvd.nist.gov/vuln/detail/CVE-2023-53123"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2023-6931","cve":"CVE-2023-6931","aliases":[],"title":"Linux kernel (perf): Out-of-bounds write in perf_read_group() via read_size overflow - local root","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (perf)","year":"2023","cvss_score":7,"severity":"high","kev":false,"impact":"Out-of-bounds write in perf_read_group() via read_size overflow - local root","attack_vector":"Any tenant process in a container with perf access","remediation":"Livepatchable; otherwise drain + reboot. Set `kernel.perf_event_paranoid=3` - but note profiling access is a feature many AI tenants expect","references":["https://access.redhat.com/security/cve/CVE-2023-6931"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2023-12-19"},{"id":"CVE-2023-6932","cve":"CVE-2023-6932","aliases":[],"title":"Linux kernel (IGMP): Use-after-free in IPv4 IGMP - local privilege escalation","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (IGMP)","year":"2023","cvss_score":7,"severity":"high","kev":false,"impact":"Use-after-free in IPv4 IGMP - local privilege escalation","attack_vector":"Any tenant process in a container","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2023-6932"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2023-12-19"},{"id":"CVE-2024-0646","cve":"CVE-2024-0646","aliases":[],"title":"Linux kernel (kTLS): splice() into a kTLS socket overwrites read-only kernel pages - local privilege escalation","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (kTLS)","year":"2024","cvss_score":7,"severity":"high","kev":false,"impact":"splice() into a kTLS socket overwrites read-only kernel pages - local privilege escalation","attack_vector":"Any tenant process in a container","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2024-0646"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-01-17"},{"id":"CVE-2024-22428","cve":"CVE-2024-22428","aliases":[],"title":"Dell iDRAC Service Module (incorrect default permissions): Weak default folder permissions let an unprivileged local","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC Service Module (incorrect default permissions)","year":"2024","cvss_score":7,"severity":"high","kev":false,"impact":"Weak default folder permissions let an unprivileged local user escalate and execute code on the host.","attack_vector":"Any local unprivileged user on a server running iSM 5.2.0.0 or earlier.","remediation":"Upgrade iSM past 5.2.0.0. Host package update; no reboot required.","references":["https://www.dell.com/support/kbdoc/en-us/000221129/dsa-2024-018-security-update-for-dell-idrac-service-module-for-weak-folder-permission-vulnerabilities"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2024-26660","cve":"CVE-2024-26660","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Implement bounds check for stream encoder creation in DCN301","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26660","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-04-02"},{"id":"CVE-2024-27045","cve":"CVE-2024-27045","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Fix a potential buffer overflow in 'dp_dsc_clock_en_read()'","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-27045","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-01"},{"id":"CVE-2024-27134","cve":"CVE-2024-27134","aliases":[],"title":"MLflow (`spark_udf` dir perms): Excessive directory permissions","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"MLflow (`spark_udf` dir perms)","year":"2024","cvss_score":7,"severity":"high","kev":false,"impact":"Excessive directory permissions → local privilege escalation","attack_vector":"Co-tenant local user on the same node","remediation":"Upgrade; matters on shared bare-metal nodes","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-27134"],"status":"curated","published":"2024-11-25"},{"id":"CVE-2024-36334","cve":"CVE-2024-36334","aliases":[],"title":"AMD Radeon RGB tool - signature verification on files in the installation directory: The Radeon RGB tool does","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD Radeon RGB tool - signature verification on files in the installation directory","year":"2024","cvss_score":7,"severity":"high","kev":false,"impact":"The Radeon RGB tool does not verify signatures on files placed in its installation directory, so a planted file runs with elevated privileges. A cosmetic utility that escalates to code execution - the reason it appears here is that vendor GPU tooling gets installed wholesale on GPU hosts without anyone asking what the LED control daemon is doing running as root.","attack_vector":"Local, requires write access to the tool's installation directory.","remediation":"Update or, better, uninstall - RGB lighting control has no business on a datacenter GPU node. Removing unnecessary vendor tooling from the golden image is the durable fix and costs nothing at runtime. No reboot needed to uninstall.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36334","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2026-05-15"},{"id":"CVE-2024-39766","cve":"CVE-2024-39766","aliases":[],"title":"Intel Neural Compressor (SQL injection, second instance): A second SQL-injection path in Neural Compressor reachable","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel Neural Compressor (SQL injection, second instance)","year":"2024","cvss_score":7,"severity":"high","kev":false,"impact":"A second SQL-injection path in Neural Compressor reachable by an authenticated user. Same consequence as the first: control of the service's backing store.","attack_vector":"Any authenticated user of the service.","remediation":"Upgrade to v3.0 or later. Userspace, restart only.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-39766","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01219.html"],"status":"curated","published":"2024-11-13"},{"id":"CVE-2024-41022","cve":"CVE-2024-41022","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): A correctness defect in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2024","cvss_score":7,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu: Fix signedness bug in sdma_v4_0_process_trap_irq()","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-41022","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-07-29"},{"id":"CVE-2024-42123","cve":"CVE-2024-42123","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): A double free in the amdgpu firmware, ACPI","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2024","cvss_score":7,"severity":"high","kev":false,"impact":"A double free in the amdgpu firmware, ACPI and IP-block initialisation. The same allocation is released twice, corrupting the slab allocator's freelist. This is a classic heap-corruption primitive: with slab grooming it becomes arbitrary kernel memory write and therefore host compromise from an unprivileged GPU workload. The cheap outcome is a node panic. Upstream fix: drm/amdgpu: fix double free err_addr pointer warnings","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-42123","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-07-30"},{"id":"CVE-2024-42228","cve":"CVE-2024-42228","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): Memory is handed to a consumer without being initialised or","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2024","cvss_score":7,"severity":"high","kev":false,"impact":"Memory is handed to a consumer without being initialised or cleared in the amdgpu kernel driver core. Whatever the previous owner left behind is readable - and on a GPU node the previous owner is very often a different tenant's job. This is the classic residual-data leak between workloads sharing a card: model weights, activations, keys or tokens from the prior tenant can surface in a fresh allocation. Upstream fix: drm/amdgpu: Using uninitialized value *size when calling amdgpu_vce_cs_reloc","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-42228","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-07-30"},{"id":"CVE-2024-47975","cve":"CVE-2024-47975","aliases":["Solidigm SA-000563"],"title":"Solidigm DC SSDs with TCG Opal (DC P4510/P4511/P4610 Opal, D5-P4320/P4326 Opal, D5-P5316 Opal, D7-P5510/P5520/P5620","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Solidigm DC SSDs with TCG Opal (DC P4510/P4511/P4610 Opal, D5-P4320/P4326 Opal, D5-P5316 Opal, D7-P5510/P5520/P5620…","year":"2024","cvss_score":7,"severity":"high","kev":false,"impact":"Improper access-control validation in the Opal-enabled firmware lets an attacker with physical access gain unauthorized access to the drive, or a local attacker knock it offline. The highest-scored entry in Solidigm's 2024 advisory set, and it hits exactly the SKUs an operator chose specifically BECAUSE they are Opal drives - the ones bought so that locking ranges would separate tenants and make decommissioning safe. BREAKS TENANT HANDOFF: the locking that your reclaim process depends on can be walked around, so a drive you believed was cryptographically locked between customers is readable. Also carries an availability tail - a local attacker can take the drive down, and on a shared bare-metal node that is a tenant-triggered outage.","attack_vector":"An attacker with physical access to the drive - the RMA return path, a decommissioned node in the resale channel, or a colo/rack tech - for the unauthorized-access half; a tenant with local access on the host for the denial-of-service half.","remediation":"Firmware flash per SKU with drive offline and node drained, using Solidigm Storage Tool: VEV10294/VDV10194/VEV10394 for the P4510/P4511/P4610 Opal variants, 3DV10132 for D5-P4320 Opal, 8DV10564 for D5-P4326 Opal, ACV10310 for D5-P5316 Opal, JCV10300 for D7-P5510 Opal, 9CV10410 for D7-P5520/P5620 Opal. Stop treating an Opal locking range as your tenant boundary on its own - it is a vendor-attested control you cannot audit. Layer LUKS/dm-crypt above it so a locking-range bypass yields ciphertext. Chain-of-custody matters as much as the flash here: because the attack is physical, tighten RMA and decommission handling (destroy rather than return where contract allows, or degauss/shred media that held tenant data) - a firmware update on drives still in the rack does nothing for the ones already on a pallet heading back to the vendor.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-47975","https://www.solidigm.com/support-page/support-security.html","https://www.solidigm.com/content/dam/solidigm/en/site/support/support-community/cve-(security)/documents/public-security-advisory-v2.pdf"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-10-07"},{"id":"CVE-2024-6409","cve":"CVE-2024-6409","aliases":[],"title":"OpenSSH (sshd, RHEL9): Signal-handling race in the privsep child - possible RCE, RHEL 9 specific","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"OpenSSH (sshd, RHEL9)","year":"2024","cvss_score":7,"severity":"high","kev":false,"impact":"Signal-handling race in the privsep child - possible RCE, RHEL 9 specific","attack_vector":"Unauthenticated network","remediation":"Package update + sshd restart","references":["https://access.redhat.com/security/cve/CVE-2024-6409"],"status":"curated","published":"2024-07-08"},{"cwe":["CWE-362","CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-21732","cve":"CVE-2025-21732","aliases":[],"title":"NVIDIA/Mellanox ConnectX driver (mlx5_ib ODP invalidation vs MR deregistration race): During deregistration the lkey is","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA/Mellanox ConnectX driver (mlx5_ib ODP invalidation vs MR deregistration race)","year":"2025","cvss_score":7,"severity":"high","kev":false,"impact":"During deregistration the lkey is revoked in hardware first, but mlx5_ib_invalidate_range() can concurrently post a UMR update for that same freed lkey. The window is a race between a tenant tearing down a memory region and the MMU notifier invalidating a range of it - two paths the tenant drives - operating on a key the hardware has already released. Keys being operated on after revocation is exactly the state in which a stale key can be observed doing work it should no longer be able to do.","attack_vector":"Local, unprivileged. Concurrent munmap/deregistration of an ODP memory region by a tenant process.","remediation":"Kernel update serialising revocation against invalidation. No configuration mitigation.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=5297f5ddffef47b94172ab0d3d62270002a3dcc1","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2025/CVE-2025-21732.json"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-23279","cve":"CVE-2025-23279","aliases":[],"title":"GPU Display Driver: Local privesc (race condition)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2025","cvss_score":7,"severity":"high","kev":false,"impact":"Local privesc (race condition)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23279","https://github.com/NVIDIA/product-security/tree/main/2025/5670"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-367"],"fleet":{"pain_class":"node-reboot"},"published":"2025-08-02"},{"id":"CVE-2025-23280","cve":"CVE-2025-23280","aliases":[],"title":"GPU Display Driver: Local privesc (use-after-free)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2025","cvss_score":7,"severity":"high","kev":false,"impact":"Local privesc (use-after-free)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23280","https://github.com/NVIDIA/product-security/tree/main/2025/5703"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"published":"2025-10-10"},{"id":"CVE-2025-23281","cve":"CVE-2025-23281","aliases":[],"title":"GPU Display Driver: Local privesc (use-after-free)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2025","cvss_score":7,"severity":"high","kev":false,"impact":"Local privesc (use-after-free)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23281","https://github.com/NVIDIA/product-security/tree/main/2025/5670"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"published":"2025-08-02"},{"id":"CVE-2025-23282","cve":"CVE-2025-23282","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): A race condition in the Linux display driver is winnable","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2025","cvss_score":7,"severity":"high","kev":false,"impact":"A race condition in the Linux display driver is winnable by a local attacker and escalates to code execution in kernel context - the strongest container-escape primitive in this batch of driver bugs. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5703. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23282","https://github.com/NVIDIA/product-security/tree/main/2025/5703"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-415"],"fleet":{"pain_class":"node-drain"},"tags":["tenant-isolation"],"published":"2025-10-10"},{"id":"CVE-2025-32023","cve":"CVE-2025-32023","aliases":[],"title":"Redis: Authenticated user triggers a stack/heap out-of-bounds write in hyperloglog ops","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Redis","year":"2025","cvss_score":7,"severity":"high","kev":false,"impact":"Authenticated user triggers a stack/heap out-of-bounds write in hyperloglog ops -> potential RCE","attack_vector":"Local","remediation":"Control-plane: upgrade to 8.0.3/7.4.5/7.2.10/6.2.19; restrict the command surface","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-32023"],"status":"curated","published":"2025-07-07"},{"id":"CVE-2025-32462","cve":"CVE-2025-32462","aliases":[],"title":"sudo: Local privilege escalation via the `--host` option against host-specific sudoers rules","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"sudo","year":"2025","cvss_score":7,"severity":"high","kev":false,"impact":"Local privilege escalation via the `--host` option against host-specific sudoers rules","attack_vector":"Local user with any sudoers entry","remediation":"Package update; no reboot","references":["https://access.redhat.com/security/cve/CVE-2025-32462"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2025-06-30"},{"id":"CVE-2025-3770","cve":"CVE-2025-3770","aliases":["GHSA-vx5v-4gg6-6qxr"],"title":"EDK II (SMM environment, Machine Check Exception handling): Machine Check Exceptions are enabled before SMM installs","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"EDK II (SMM environment, Machine Check Exception handling)","year":"2025","cvss_score":7,"severity":"high","kev":false,"impact":"Machine Check Exceptions are enabled before SMM installs a handler for them, so an MCE fired in that window is delivered to whatever the IDT happens to point at - which an attacker can arrange. That is arbitrary code execution at ring -2. SMM sits below the hypervisor and below the OS: an implant there survives OS reinstall and node reimaging, can forge or suppress the measurements that feed remote attestation, and is invisible to every agent a tenant or operator runs. On multi-tenant GPU hardware it is the difference between wiping a node between customers and believing you wiped it.","attack_vector":"Local attacker with sufficient privilege on the host OS to trigger a machine check at the right moment - realistically root/admin on the node, or a tenant with kernel-level access on bare metal.","remediation":"Firmware flash from the server OEM. This is a core edk2 fix published August 2025, so the IBV rebase into AMI/Insyde/Phoenix trees and then into OEM BIOS payloads is the long pole - budget one to two OEM BIOS release cycles and track it per platform generation. One reboot per node, drain first. No configuration workaround: SMM is always present and cannot be disabled. Compensating control is to not hand kernel-level access on shared bare metal to untrusted tenants without a full firmware re-flash between leases.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-3770","https://github.com/tianocore/edk2/security/advisories/GHSA-vx5v-4gg6-6qxr"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2025-08-07"},{"id":"CVE-2025-4802","cve":"CVE-2025-4802","aliases":[],"title":"glibc: Static setuid binaries incorrectly search LD_LIBRARY_PATH during dlopen - local privilege escalation","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"glibc","year":"2025","cvss_score":7,"severity":"high","kev":false,"impact":"Static setuid binaries incorrectly search LD_LIBRARY_PATH during dlopen - local privilege escalation","attack_vector":"Local user","remediation":"Package update; restart or reboot for full coverage of long-running processes","references":["https://access.redhat.com/security/cve/CVE-2025-4802"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-05-16"},{"id":"CVE-2025-54518","cve":"CVE-2025-54518","aliases":["XSA-490"],"title":"Xen (x86): CPU opcode cache corruption - host instability triggerable from a guest","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen (x86)","year":"2025","cvss_score":7,"severity":"high","kev":false,"impact":"CPU opcode cache corruption - host instability triggerable from a guest","attack_vector":"Tenant VM guest","remediation":"Hypervisor patch + host reboot","references":["https://xenbits.xen.org/xsa/advisory-490.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-05-15"},{"id":"CVE-2025-6019","cve":"CVE-2025-6019","aliases":[],"title":"libblockdev / udisks: allow_active to root via libblockdev through udisks","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"libblockdev / udisks","year":"2025","cvss_score":7,"severity":"high","kev":false,"impact":"allow_active to root via libblockdev through udisks - second half of the chain, works on nearly every mainstream distro","attack_vector":"Local user","remediation":"Package update + restart udisksd; no reboot","references":["https://access.redhat.com/security/cve/CVE-2025-6019"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2025-06-19"},{"cwe":["CWE-1188","CWE-665"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-68209","cve":"CVE-2025-68209","aliases":[],"title":"NVIDIA/Mellanox ConnectX driver (mlx5 core completion-queue creation defaults): Every CQ created without an explicit","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA/Mellanox ConnectX driver (mlx5 core completion-queue creation defaults)","year":"2025","cvss_score":7,"severity":"high","kev":false,"impact":"Every CQ created without an explicit completion handler was defaulted to mlx5_add_cq_to_tasklet, a function only user CQs created through mlx5_ib are meant to use, and the default creation path left a valid value in the CQ's arm_db so firmware could raise interrupts on completion queues that were only ever meant to be polled. The result is a kernel-internal, polling-only completion queue being driven by a hardware interrupt into a completion handler intended for a different class of object - a firmware-triggered control-flow crossing between the user-facing and kernel-internal halves of the same adapter.","attack_vector":"Local/adapter-level. Requires the corner-case interrupt condition on a polling-only kernel CQ; the tenant influences adapter interrupt load, not the handler choice directly.","remediation":"Kernel update that stops defaulting the completion function and clears arm_db for polling-only CQs. No workaround.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=08469f5393a1a39f26a6e2eb2e8c33187665c1f4","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2025/CVE-2025-68209.json"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-20885","cve":"CVE-2026-20885","aliases":["TDX module improper authentication"],"title":"Intel TDX module, Ring 0 / Trust Domain context, multiple Intel platforms - INTEL-SA-01436: Improper authentication","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel TDX module, Ring 0 / Trust Domain context, multiple Intel platforms - INTEL-SA-01436","year":"2026","cvss_score":7,"severity":"high","kev":false,"impact":"Improper authentication in the TDX module allows both information disclosure and privilege escalation from a system-software adversary. Authentication failures in the TDX module are the worst shape of TDX bug because the module's job is to decide who is allowed to ask it for what - a flaw there means the untrusted hypervisor can present itself as authorized for operations reserved to the trust domain. Result is tenant data disclosure plus escalation into the protected domain. This is one of the most recent items in the set, from Intel's August 2026 advisory batch, so many fleets have not yet deployed the fix.","attack_vector":"A system-software adversary with privileged host access - the hypervisor or host root - against the TDX module. Intel rates attack complexity as high, but the required position is one every infrastructure operator already occupies.","remediation":"Update the TDX module to the fixed version per INTEL-SA-01436, plus the associated platform firmware. Host reboot and full drain of trust domains. Because this landed in the August 2026 batch alongside several other TDX and transient-execution CVEs, deploy it as one consolidated firmware and TDX-module campaign rather than a series of single-CVE windows - each window costs a full fleet drain. Then raise the accepted TDX module SVN in your attestation policy; until you do, patched and unpatched hosts are indistinguishable to relying parties.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-20885","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01436.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-11"},{"cwe":["CWE-362","CWE-754"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-23005","cve":"CVE-2026-23005","aliases":[],"title":"Linux kernel (arch/x86/kernel/fpu): A guest that disables an XSAVE feature through XFD while the saved XSTATE_BV still","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (arch/x86/kernel/fpu)","year":"2026","cvss_score":7,"severity":"high","kev":false,"impact":"A guest that disables an XSAVE feature through XFD while the saved XSTATE_BV still advertises it makes the host execute XRSTOR with a state combination that raises #NM in kernel context and panics the machine. One tenant's WRMSR takes the entire node down, which in a shared GPU cluster is every co-resident tenant's outage.","attack_vector":"Guest-side: the guest writes MSR_IA32_XFD to disable a feature (AMX and friends) and wins a race against a host interrupt that triggers kernel_fpu_begin() before KVM updates the guest XFD. Requires an XFD-capable CPU (Sapphire Rapids-class AMX hardware) and any tenant VM on it. A second path is a VMM stuffing XSTATE_BV via KVM_SET_XSAVE, which needs /dev/kvm.","remediation":"Update to a kernel with the referenced stable commits. Interim: do not expose AMX / XFD-gated features in the guest CPU model on unpatched nodes, which removes the guest's ability to set XFD at all.","references":["https://git.kernel.org/stable/c/b5995c01ba53d84182ecb9492fc4d91cfe8a362d","https://git.kernel.org/stable/c/1e2848bda819af569dfe7ab186223855e092a2cb","https://nvd.nist.gov/vuln/detail/CVE-2026-23005"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-24200","cve":"CVE-2026-24200","aliases":[],"title":"vGPU Manager: Guest-to-host escape (use-after-free in GPU context mgmt)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU Manager","year":"2026","cvss_score":7,"severity":"high","kev":false,"impact":"Guest-to-host escape (use-after-free in GPU context mgmt)","attack_vector":"Tenant VM guest","remediation":"vGPU Manager upgrade; evacuate all guest VMs, reboot host","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24200","https://github.com/NVIDIA/product-security/tree/main/2026/5821"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"published":"2026-05-26"},{"cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-52939","cve":"CVE-2026-52939","aliases":[],"title":"Linux kernel (net/rds): RDS always programs the masked variants of the RDMA atomic opcodes, but the send-completion","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/rds)","year":"2026","cvss_score":7,"severity":"high","kev":false,"impact":"RDS always programs the masked variants of the RDMA atomic opcodes, but the send-completion path only recognises the unmasked ones, so every atomic operation completes with a NULL message pointer that is then dereferenced. A tenant sending one atomic control message over an RDS/IB connection panics the node from softirq - the commit states plainly that an unprivileged AF_RDS sendmsg() triggers it with no extra setup on mlx4/mlx5.","attack_vector":"Local and unprivileged, single syscall: socket(AF_RDS, SOCK_SEQPACKET, 0) - which autoloads the rds and rds_rdma modules through the net-pf-21 alias with no capability check - then sendmsg() with an RDS atomic cmsg over an active RDS/IB connection. Requires an RDMA device the tenant's traffic can use, which is the normal case on a GPU node, and native masked-atomic support (mlx4/mlx5). The fault is in the completion tasklet, so it is a fatal exception in interrupt context: the whole node goes down.","remediation":"Boot a kernel carrying the fix commits (handles the masked atomic opcodes in the completion unmap path). Interim: blacklist the rds, rds_rdma and rds_tcp modules (`install rds /bin/false`) or deny socket family 21 in tenant seccomp profiles - RDS is almost never intentionally used on a GPU cluster and is a good candidate for blanket removal.","references":["https://git.kernel.org/stable/c/a0148342badd8c9b2e46551766a27cb76c82e715","https://git.kernel.org/stable/c/6e4615164d185a26badb2f376a2449f4d174a5f0","https://nvd.nist.gov/vuln/detail/CVE-2026-52939"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-53329","cve":"CVE-2026-53329","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":7,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Use krealloc_array() in dal_vector_reserve()","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53329","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-07-01"},{"id":"CVE-2026-64219","cve":"CVE-2026-64219","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":7,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Validate payload length and link_index in dc_process_dmub_aux_transfer_async","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-64219","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-07-24"},{"cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:C/C:L/I:L/A:H","cwe":["CWE-125","CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-64460","cve":"CVE-2026-64460","aliases":[],"title":"Linux kernel (drivers/pci): An SR-IOV device that stops answering config reads makes the VF Resizable BAR restore path","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/pci)","year":"2026","cvss_score":7,"severity":"high","kev":false,"impact":"An SR-IOV device that stops answering config reads makes the VF Resizable BAR restore path read index 7 out of a 6-entry BAR-size array, so the host kernel reads past the end of the VF's resource table and then restores BAR windows from whatever it found. The reported case is exactly the fleet shape that matters - an NVIDIA GPU that stopped responding on a power-state exit - and the vendor scores it scope-changed, meaning the bad restore reaches beyond the failing device.","attack_vector":"Host-side, on a node with SR-IOV enabled and VFs created on a device that supports VF Resizable BAR (NVIDIA GPUs among them). It fires when pci_restore_state() runs while the device is unreachable and config reads return all-ones - a wedged or power-state-stuck GPU, a link that dropped, or a device a tenant has driven into a bad state through its assigned VF. A tenant does not call it directly; a tenant that can hang its assigned device can make the host walk into it.","remediation":"Boot a kernel with the sriov_restore_vf_rebar_state() error-response guard. Interim: on nodes where you do not need them, disable SR-IOV VFs (sriov_numvfs = 0) so the VF ReBAR restore path is never entered, and drain nodes whose GPUs are logging config-read failures or GC6/power-state exit errors rather than letting the recovery path run.","references":["https://git.kernel.org/stable/c/b77524621250407386f44c6eea7e5e4619ada1ce","https://git.kernel.org/stable/c/55fd485e66d0ad5c762c23dba1461fe9c741cd96","https://nvd.nist.gov/vuln/detail/CVE-2026-64460"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:4.0/AV:N/AC:L/AT:P/PR:L/UI:N/VC:N/VI:N/VA:H/SC:N/SI:N/SA:H","fleet":{"pain_class":"unpatchable / mitigate-only"},"id":"NCVD-2025-021-redis-multi-bulk-command-protoco","cve":null,"aliases":["GHSA-2r7g-8hpc-rpq9"],"title":"Redis (multi-bulk command protocol handling): PERMANENT, VENDOR-ACKNOWLEDGED DENIAL OF SERVICE WITH NO FIX PLANNED. An","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Redis (multi-bulk command protocol handling)","year":"2025","cvss_score":7,"severity":"high","kev":false,"impact":"PERMANENT, VENDOR-ACKNOWLEDGED DENIAL OF SERVICE WITH NO FIX PLANNED. An authenticated client can abuse the multi-bulk command network protocol to degrade or halt the database it has access to. Redis's position is that this does not violate their security model because it requires authentication, so no CVE was assigned, and that a code fix would harm legitimate functionality and performance — so they published a Security Advisory instead and are shipping no patch. For a cluster operator that changes the nature of the item: it is not a patch-and-forget entry but a standing architectural constraint on every Redis instance shared between tenants. In AI infrastructure Redis is rarely just a cache — it backs Celery and RQ job queues, feature stores, KV caches and session state for inference gateways — so a single authenticated tenant can stall the queue or state layer that other tenants' jobs depend on, and there will never be a version number that resolves it.","attack_vector":"Network, authenticated: any client that successfully authenticates to the Redis instance. Reported against Redis 8.0 and earlier. No privilege escalation and no code execution.","remediation":"There is no patch and none is planned — treat this as an architectural control. Do not share a Redis instance across trust boundaries; give each tenant its own instance or at minimum its own ACL user with a tightly scoped command set. Enforce strong access controls and identity-provider integration per the vendor's security guidance, apply connection and resource limits, and monitor for availability degradation. Assume any party you grant authenticated Redis access can deny service to everyone else on that instance.","references":["https://github.com/redis/redis/security/advisories/GHSA-2r7g-8hpc-rpq9"],"status":"curated"},{"id":"CVE-2020-25647","cve":"CVE-2020-25647","aliases":[],"title":"GRUB2 (USB device initialization): Out-of-bounds write in grub_usb_device_initialize from a malicious USB descriptor","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (USB device initialization)","year":"2020","cvss_score":6.9,"severity":"medium","kev":false,"impact":"Out-of-bounds write in grub_usb_device_initialize from a malicious USB descriptor. In a datacenter this is not a 'someone walks up with a USB stick' story - the BMC presents virtual media as a USB device, so anyone with BMC credentials can trigger it entirely remotely.","attack_vector":"Physical USB, or - the one that matters - BMC virtual media, which turns this into a remote attack for anyone on the management VLAN with iDRAC/iLO/XCC credentials.","remediation":"grub2 package update + reboot. Meaningful compensating control: disable virtual media on the BMC where you do not use it for provisioning, and keep the management network off any tenant-reachable path.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-25647","https://access.redhat.com/security/cve/CVE-2020-25647"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-03-03"},{"id":"CVE-2023-41333","cve":"CVE-2023-41333","aliases":[],"title":"Cilium: A user who can create CiliumNetworkPolicy in one namespace affects traffic cluster-wide","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2023","cvss_score":6.9,"severity":"medium","kev":false,"impact":"A user who can create CiliumNetworkPolicy in one namespace affects traffic cluster-wide; cross-tenant policy tampering","attack_vector":"Cluster user with namespace access","remediation":"Rolling Cilium upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-41333"],"status":"curated","published":"2023-09-27"},{"id":"CVE-2024-24557","cve":"CVE-2024-24557","aliases":[],"title":"Docker / moby: Classic builder cache poisoning for images built FROM scratch","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2024","cvss_score":6.9,"severity":"medium","kev":false,"impact":"Classic builder cache poisoning for images built FROM scratch","attack_vector":"Anyone sharing a build cache, e.g. a multi-tenant CI builder","remediation":"Upgrade moby; give each tenant an isolated build cache","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-24557"],"status":"curated","published":"2024-02-01"},{"id":"CVE-2025-0045","cve":"CVE-2025-0045","aliases":[],"title":"AMD Secure Processor PCI driver - input validation: Improper input validation in the ASP PCI driver lets a local","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor PCI driver - input validation","year":"2025","cvss_score":6.9,"severity":"medium","kev":false,"impact":"Improper input validation in the ASP PCI driver lets a local attacker trigger a buffer overflow and crash the node. Practically this is availability: a tenant-adjacent process that can reach the driver can take the host down, and on a GPU node that means every co-resident training job dies with it.","attack_vector":"Local, through the ASP PCI driver interface. Not tenant-container reachable on a normally configured node - the ccp/PSP device is not handed to workload containers.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. There is also a kernel-side component (the ccp driver); take the distro kernel update as well as the BIOS, since the OEM firmware will lag.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0045","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2026-05-15"},{"id":"CVE-2025-29939","cve":"CVE-2025-29939","aliases":[],"title":"AMD SEV firmware - RMP write during SNP initialization: A privileged attacker can write to the reverse map page during","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV firmware - RMP write during SNP initialization","year":"2025","cvss_score":6.9,"severity":"medium","kev":false,"impact":"A privileged attacker can write to the reverse map page during secure nested paging initialization, corrupting the ownership map before any guest launches. Because the RMP is initialised once and then trusted, poisoning it at init time means every guest that subsequently launches on that host inherits a compromised isolation boundary.","attack_vector":"Local, privileged, during SNP init - so an attacker who controls the host's boot or SNP initialisation path.","remediation":"Fixed in AMD reference firmware (AGESA / PSP / SEV firmware) and delivered only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo and the ODMs each rebuild and requalify AMD's AGESA drop before shipping. **Expect one to six months of OEM lag**, and on end-of-support platforms expect nothing. Applying it is a drain plus full power cycle, not a driver reload. Verify by reading back the PSP/SMU firmware version afterwards rather than trusting the BIOS version string. This sits inside the SEV-SNP trust boundary, so the update moves the platform's reported TCB version: refresh VCEK certificates from AMD's KDS and update any attestation policy your tenants pin, or confidential guest launches will start failing right after the BIOS lands. Pair with host measured boot: if you cannot attest the boot sequence, you cannot rule out that SNP was initialised under an attacker's influence.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-29939","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2026-02-10"},{"id":"CVE-2025-48516","cve":"CVE-2025-48516","aliases":[],"title":"AMD AGESA bootloader - DDR5 PMIC default configuration: The AGESA bootloader leaves DDR5 memory modules in an insecure","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD AGESA bootloader - DDR5 PMIC default configuration","year":"2025","cvss_score":6.9,"severity":"medium","kev":false,"impact":"The AGESA bootloader leaves DDR5 memory modules in an insecure default state with the on-DIMM power management IC interface unprotected. A local user can then reprogram the PMIC and destroy the module - a **permanent**, physical denial of service. This is one of the rare software bugs whose remediation is an RMA: an attacker who runs this across a fleet does not take your nodes offline for a reboot, they take them offline for a parts order.","attack_vector":"Local user privilege on the host. No physical access needed - the PMIC is reachable over the platform's memory-module management interface.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. There is no software undo once a DIMM is bricked. Until the OEM BIOS lands, the mitigation is access control: this needs local execution on the host, so it is a strong argument for not giving semi-trusted workloads a shell on bare metal. Budget for spare DIMMs on any fleet where you cannot patch quickly.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-48516","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2026-05-15"},{"id":"CVE-2025-48518","cve":"CVE-2025-48518","aliases":[],"title":"AMD Graphics Driver - out-of-bounds write: Improper input validation lets a local attacker write out of bounds through","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD Graphics Driver - out-of-bounds write","year":"2025","cvss_score":6.9,"severity":"medium","kev":false,"impact":"Improper input validation lets a local attacker write out of bounds through the AMD graphics driver, costing integrity or availability. Out-of-bounds kernel writes reachable from a GPU device handle are the standard escape route out of a GPU container.","attack_vector":"Local, via the graphics driver interface - reachable from a tenant container with a GPU device node.","remediation":"Update the AMD graphics driver package and reload the driver or reboot the node.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-48518","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-02-11"},{"id":"CVE-2025-48521","cve":"CVE-2025-48521","aliases":[],"title":"AMD Secure Processor PCI driver - use-after-free: A use-after-free reachable through the ASP PCI driver","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor PCI driver - use-after-free","year":"2025","cvss_score":6.9,"severity":"medium","kev":false,"impact":"A use-after-free reachable through the ASP PCI driver. Beyond the crash, freed-then-reused kernel memory is a corruption primitive, so the honest read is loss of platform integrity rather than simple denial of service.","attack_vector":"Local, via the ASP PCI driver interface; requires access to the crypto/PSP device, i.e. host-level not tenant-level.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. Pair the BIOS update with the corresponding kernel ccp driver fix.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-48521","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2026-05-15"},{"id":"CVE-2025-58354","cve":"CVE-2025-58354","aliases":[],"title":"Kata Containers: A malicious host can circumvent guest protections","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kata Containers","year":"2025","cvss_score":6.9,"severity":"medium","kev":false,"impact":"A malicious host can circumvent guest protections","attack_vector":"A compromised or hostile host operator","remediation":"Upgrade Kata; relevant when Kata is used to isolate tenants from each other, not from you","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-58354"],"status":"curated","published":"2025-09-23"},{"id":"CVE-2025-64329","cve":"CVE-2025-64329","aliases":[],"title":"containerd: CRI Attach implementation bug lets a user attach to a container they should not reach","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2025","cvss_score":6.9,"severity":"medium","kev":false,"impact":"CRI Attach implementation bug lets a user attach to a container they should not reach","attack_vector":"Cluster user with namespace access","remediation":"Rolling containerd upgrade with node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-64329"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2025-11-07"},{"id":"CVE-2025-64436","cve":"CVE-2025-64436","aliases":[],"title":"KubeVirt: virt-handler service-account permissions (update VMI, patch nodes) can be abused to force VMI migration","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"KubeVirt","year":"2025","cvss_score":6.9,"severity":"medium","kev":false,"impact":"virt-handler service-account permissions (update VMI, patch nodes) can be abused to force VMI migration","attack_vector":"An attacker with virt-handler credentials","remediation":"Upgrade KubeVirt; scope down the service account","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-64436"],"status":"curated","published":"2025-11-07"},{"id":"CVE-2026-15789","cve":"CVE-2026-15789","aliases":[],"title":"BuildKit: Crafted upload request lets files escape the BuildKit state directory onto the host","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"BuildKit","year":"2026","cvss_score":6.9,"severity":"medium","kev":false,"impact":"Crafted upload request lets files escape the BuildKit state directory onto the host","attack_vector":"Anyone with build control API access","remediation":"Upgrade BuildKit","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-15789"],"status":"curated","published":"2026-07-21"},{"id":"CVE-2026-31838","cve":"CVE-2026-31838","aliases":[],"title":"Istio: Envoy RBAC header matching flaw bypasses header-based authorization policy","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Istio","year":"2026","cvss_score":6.9,"severity":"medium","kev":false,"impact":"Envoy RBAC header matching flaw bypasses header-based authorization policy","attack_vector":"Unauthenticated network","remediation":"Rolling istiod and proxy upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-31838"],"status":"curated","published":"2026-03-10"},{"id":"CVE-2026-8810","cve":"CVE-2026-8810","aliases":["INSYDE-SA-2026005"],"title":"Insyde InsydeH2O on ARM platforms (HDD password storage in UEFI variables): HDD passwords are recoverable from UEFI","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O on ARM platforms (HDD password storage in UEFI variables)","year":"2026","cvss_score":6.9,"severity":"medium","kev":false,"impact":"HDD passwords are recoverable from UEFI variables on affected ARM platforms - insufficiently protected credentials, in CWE terms. The operator concern is drive-level secrets surviving in firmware storage where the next person to hold the box can read them, which matters for any fleet where drive locking is part of the between-tenant wipe story, and for ARM-based nodes appearing in AI inference and edge fleets.","attack_vector":"Requires physical access plus local privilege and user interaction per Insyde's own scoring - so this is a returned-hardware, decommissioning, or colo-access risk rather than a remote one.","remediation":"OEM firmware update on Insyde kernel 5.6 / 05.63.21 or 5.7 / 05.72.21. Firmware flash, reboot per node. Advisory dated 2026-08-18, so OEM images are not yet widely available. Operational mitigation that does not wait for the flash: do not rely on ATA HDD passwords as the confidentiality control on affected ARM platforms - use self-encrypting-drive keys or software full-disk encryption with keys held off the node, and cryptographically erase rather than password-lock drives at decommission.","references":["https://www.insyde.com/security-pledge/SA-2026005/"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2026-08-19"},{"id":"CVE-2017-8371","cve":"CVE-2017-8371","aliases":[],"title":"Schneider Electric StruxureWare Data Center Expert before 7.4.0: Passwords held in cleartext in RAM on the DCIM","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Schneider Electric StruxureWare Data Center Expert before 7.4.0","year":"2017","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Passwords held in cleartext in RAM on the DCIM appliance, recoverable remotely. Included here because it is the earliest entry in a seven-year pattern: DCE has repeatedly failed to protect the device credentials it must hold, and any operator running an old DCE build should assume the facility credential set is compromised rather than assume otherwise.","attack_vector":"Remote, per the advisory; unspecified vectors, but the practical read is that a foothold on or near the appliance yields the credentials.","remediation":"Upgrade to 7.4.0 or later - though anyone still on a pre-7.4 build has far larger problems from the 2021-2024 RCEs above. Rotate all device credentials.","references":["https://nvd.nist.gov/vuln/detail/CVE-2017-8371"],"status":"curated","published":"2017-04-30"},{"id":"CVE-2018-15776","cve":"CVE-2018-15776","aliases":[],"title":"Dell iDRAC (u-boot): Improper error handling grants access to the u-boot shell — pre-BMC-OS control, i.e","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC (u-boot)","year":"2018","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Improper error handling grants access to the u-boot shell — pre-BMC-OS control, i.e. below even the BMC firmware image","attack_vector":"Local / serial-adjacent","remediation":"Firmware update; a u-boot foothold survives BMC firmware reflash, so affected nodes need physical verification rather than a remote fix","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-15776"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2018-12-13"},{"id":"CVE-2018-18095","cve":"CVE-2018-18095","aliases":["INTEL-SA-00267","LEN-28116"],"title":"Intel SSD DC S4500 and SSD DC S4600 series firmware before SCV10150 - improper authentication: Improper authentication","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SSD DC S4500 and SSD DC S4600 series firmware before SCV10150 - improper authentication","year":"2018","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Improper authentication in the drive firmware lets an UNPRIVILEGED user with physical access escalate privilege on the drive, with full confidentiality, integrity and availability impact. These are mainstream datacenter SATA SSDs that shipped in enormous volume as boot and scratch media in ProLiant/PowerEdge/ThinkSystem-class nodes, which is exactly the tier of hardware that gets recycled into GPU builds and rented out bare-metal. BREAKS TENANT HANDOFF: authentication on the drive is the mechanism that is supposed to keep a departing tenant from reaching a reprovisioned drive's contents, and it is bypassable by someone with no privilege at all. The integrity impact matters as much as the read - an attacker who owns the drive controller can persist there through your reimage.","attack_vector":"An unprivileged attacker with physical access to the drive. No credentials of any kind required. Realistically: anyone in the hardware's physical path - decommission, RMA, colo, or a rack tech - and any second-hand S4500/S4600 you bought to fill out a build.","remediation":"Flash to firmware SCV10150 or later; drive offline, node drained, vendor tooling (Intel MAS, or the OEM's own bundle - Lenovo shipped this as LEN-28116 and F5 issued its own bulletin, so if these drives came inside an OEM chassis check the OEM's firmware bundle rather than Intel's, because the OEM-qualified image is often the only one their controller will accept). This is a 2018 fix, so the operational question in 2026 is not whether a patch exists but whether anyone ever applied it to the second-hand S4500/S4600 inventory now sitting in your fleet - assume not, and audit. Any drive you cannot confirm was flashed and cannot account for the custody of should be treated as potentially firmware-implanted and destroyed rather than redeployed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-18095","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00267.html","https://support.lenovo.com/us/en/product_security/LEN-28116"],"status":"curated","published":"2019-07-11"},{"id":"CVE-2019-18424","cve":"CVE-2019-18424","aliases":["XSA-302","Xen passthrough DMA host privilege escalation"],"title":"Xen through 4.12.x - passed-through PCI devices left able to DMA into host memory after being handed to an untrusted","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen through 4.12.x - passed-through PCI devices left able to DMA into host memory after being handed to an untrusted…","year":"2019","cvss_score":6.8,"severity":"medium","kev":false,"impact":"A guest that has been assigned a physical device gains host privileges through DMA, because the device retains reach into memory it should have lost. For a GPU-passthrough cloud this is the direct form of the failure everyone worries about: the tenant you gave a GPU to uses that GPU to read and write the hypervisor. It also breaks tenant handoff, since the state that makes the device dangerous is set up around assignment and de-assignment - the exact transition that happens between customers.","attack_vector":"A tenant in a guest domain with a physical device assigned to it. That is the normal configuration of a GPU-passthrough product, not an unusual one.","remediation":"Patch Xen (XSA-302) and reboot the hypervisor - a rolling drain across the fleet. Structurally, use the assignable-add workflow so devices are explicitly quarantined before and after assignment rather than being handed straight from host to guest. Beyond this specific CVE, the general lesson holds for every passthrough platform including KVM/VFIO: reset the device, flush its DMA mappings and re-verify its firmware between tenants, and treat 'the device was assigned to someone else five minutes ago' as untrusted state.","references":["https://xenbits.xen.org/xsa/advisory-302.html","https://nvd.nist.gov/vuln/detail/CVE-2019-18424"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2019-10-31"},{"id":"CVE-2020-12355","cve":"CVE-2020-12355","aliases":["INTEL-SA-00391"],"title":"RPMB protocol message authentication subsystem in Intel TXE before 4.0.30 (replay-protected memory block)","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"RPMB protocol message authentication subsystem in Intel TXE before 4.0.30 (replay-protected memory block)","year":"2020","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Capture-replay authentication bypass in RPMB - the mechanism that is supposed to make firmware anti-rollback and monotonic-counter state tamper-evident. Break RPMB and you break rollback protection: an attacker can replay old, signed state to roll firmware back to a known-vulnerable version, or to reset counters that firmware relies on to detect tampering. The operator consequence is that 'we patched that' stops being verifiable from the platform itself, and a node that you believe is on current firmware can be silently downgraded and left that way through a tenant handoff.","attack_vector":"Physical access to the platform, capturing and replaying RPMB traffic. Relevant for hardware that passes through untrusted hands - shared cages, remote-hands, RMA and resale channels, and any secondhand GPU capacity you have taken on.","remediation":"TXE/CSME firmware update from the OEM bundle, host reboot and drain. More importantly, stop treating the platform's own report of its firmware version as authoritative: read firmware versions out of band via the BMC and compare against an externally held inventory, and re-flash the full firmware stack on any node that has been out of your physical custody before it re-enters a tenant pool.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-12355","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00391.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2020-13799","cve":"CVE-2020-13799","aliases":["VU#231329","WDC-20008","RPMB replay"],"title":"Replay Protected Memory Block (RPMB) protocol as specified for eMMC, UFS and ALL versions of NVMe","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Replay Protected Memory Block (RPMB) protocol as specified for eMMC, UFS and ALL versions of NVMe - multi-vendor…","year":"2020","cvss_score":6.8,"severity":"medium","kev":false,"impact":"RPMB is the small authenticated region storage devices provide so a host can keep trusted firmware and anti-rollback state where software cannot forge it. The protocol as SPECIFIED - not one vendor's bug, but the standard itself, across eMMC, UFS and every version of NVMe - permits replay attacks that let an attacker roll the protected region back to an earlier state. That undermines the anti-rollback guarantee that firmware-integrity schemes are built on, so an attacker can reinstate previously-revoked firmware or state. This is a foundational-trust issue rather than a data-read issue: the mechanism your platform uses to prove firmware has not been downgraded can itself be replayed. Because it is written into the standard, it is present across vendors and generations simultaneously - the widest blast radius of anything in this category.","attack_vector":"An attacker with physical access to the device, or with the ability to interpose on the host-to-device command path, capturing and replaying RPMB message sequences. Multi-vendor by construction, since the flaw is in the specification that every implementer followed.","remediation":"No single patch exists - remediation is per-implementation and depends on the host software and the device firmware cooperating, which is why CERT/CC coordinated it across vendors rather than issuing one fix. Check each storage vendor's advisory for your specific SKUs (Western Digital published WDC-20008; CERT/CC VU#231329 tracks the multi-vendor response) and apply device firmware plus any host-side platform firmware updates they name. Practically, most operators will not be able to close this on existing fleet hardware, so treat RPMB-backed anti-rollback as a control you cannot fully rely on: do not let it be the only thing preventing a firmware downgrade. Keep an independent record of expected firmware versions per drive serial, and alert on any drive whose reported firmware version goes BACKWARDS between inventory scans - a downgrade you detect is far more useful than an anti-rollback guarantee you cannot verify.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-13799","https://www.kb.cert.org/vuls/id/231329"],"status":"curated","published":"2020-11-18"},{"id":"CVE-2020-16844","cve":"CVE-2020-16844","aliases":[],"title":"Istio: DENY AuthorizationPolicy with wildcard-suffix principals silently fails to deny","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Istio","year":"2020","cvss_score":6.8,"severity":"medium","kev":false,"impact":"DENY AuthorizationPolicy with wildcard-suffix principals silently fails to deny","attack_vector":"Any pod on the mesh","remediation":"Rolling istiod upgrade; re-verify deny policies","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-16844"],"status":"curated","published":"2020-10-01"},{"id":"CVE-2020-8705","cve":"CVE-2020-8705","aliases":[],"title":"Intel Boot Guard in Intel CSME / TXE / SPS: Insecure default initialisation in Boot Guard means the S3 resume path does","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Boot Guard in Intel CSME / TXE / SPS","year":"2020","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Insecure default initialisation in Boot Guard means the S3 resume path does not re-verify the boot chain, so an attacker who can modify firmware while the machine is suspended defeats verified boot. Boot Guard is the hardware root of trust the rest of your platform attestation chains to - if it does not hold on resume, it does not hold.","attack_vector":"An attacker able to modify platform firmware, typically with physical access or an existing firmware-write primitive, against a machine that suspends.","remediation":"Fixed in Intel CSME/SPS firmware, which reaches you as an OEM BIOS or firmware package - not as a microcode or OS update. That means: wait for your server vendor to ship it, drain the node, flash, and reboot. OEM availability is the long pole and routinely lags the Intel advisory by one or more quarters on server platforms. Track it per platform SKU, because vendors ship these unevenly across their own product lines. Servers that never suspend are largely out of scope, which is most of a datacenter fleet - but verify rather than assume, because management controllers do use low-power states.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-8705","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00391"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2020-11-12"},{"id":"CVE-2021-21284","cve":"CVE-2021-21284","aliases":[],"title":"Docker / moby: With --userns-remap, remapped root can escalate to real host root","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2021","cvss_score":6.8,"severity":"medium","kev":false,"impact":"With --userns-remap, remapped root can escalate to real host root","attack_vector":"Any tenant workload on a userns-remapped daemon","remediation":"Upgrade Docker Engine; daemon restart","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-21284"],"status":"curated","published":"2021-02-02"},{"id":"CVE-2021-28694","cve":"CVE-2021-28694","aliases":[],"title":"Xen on AMD-Vi (AMD IOMMU) - ACPI IVMD unity-map page permissions: Xen honours ACPI-described IOMMU unity mappings but","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen on AMD-Vi (AMD IOMMU) - ACPI IVMD unity-map page permissions","year":"2021","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Xen honours ACPI-described IOMMU unity mappings but applied page permissions inconsistently, so a passed-through device can end up with wider DMA access than the hypervisor intended. On a GPU cloud that does device passthrough, this is the boundary that stops one tenant's assigned accelerator from DMA-ing into another guest's memory - and it was not holding.","attack_vector":"Requires a guest with a passed-through PCI device, which on a GPU cloud is the standard configuration. The malicious guest drives DMA from its own assigned device.","remediation":"Fixed in Xen via XSA-378. Update the hypervisor and reboot the host; no firmware step. If you run GPU passthrough on Xen, this is a top-priority class - the whole security model of passthrough rests on the IOMMU restricting device DMA correctly.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-28694","https://xenbits.xen.org/xsa/advisory-378.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2021-08-27"},{"id":"CVE-2021-28695","cve":"CVE-2021-28695","aliases":[],"title":"Xen on AMD-Vi - IOMMU page mapping permissions: Second of the XSA-378 IOMMU page-mapping issues on AMD-Vi. Incorrect","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen on AMD-Vi - IOMMU page mapping permissions","year":"2021","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Second of the XSA-378 IOMMU page-mapping issues on AMD-Vi. Incorrect mapping permissions give a passed-through device DMA reach beyond its guest, which is guest-to-host and guest-to-guest memory access via a device rather than via the CPU.","attack_vector":"Guest with a passed-through PCI device - the normal GPU-passthrough configuration.","remediation":"Fixed in Xen (XSA-378). Hypervisor update plus host reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-28695","https://xenbits.xen.org/xsa/advisory-378.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2021-08-27"},{"id":"CVE-2021-28696","cve":"CVE-2021-28696","aliases":[],"title":"Xen on AMD-Vi - IOMMU page mapping permissions: Third of the XSA-378 AMD-Vi mapping issues. Same practical consequence","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen on AMD-Vi - IOMMU page mapping permissions","year":"2021","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Third of the XSA-378 AMD-Vi mapping issues. Same practical consequence: a device assigned to one guest can reach memory it should not, defeating passthrough isolation.","attack_vector":"Guest with an assigned PCI device.","remediation":"Fixed in Xen (XSA-378). Update and reboot; patch all three of the XSA-378 CVEs together.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-28696","https://xenbits.xen.org/xsa/advisory-378.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2021-08-27"},{"id":"CVE-2021-32690","cve":"CVE-2021-32690","aliases":[],"title":"Helm: Helm repository credentials leaked to a redirected third-party host","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Helm","year":"2021","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Helm repository credentials leaked to a redirected third-party host","attack_vector":"Malicious or compromised chart repository","remediation":"Upgrade Helm; rotate repo credentials","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-32690"],"status":"curated","published":"2021-06-16"},{"id":"CVE-2021-33077","cve":"CVE-2021-33077","aliases":["CVE-2021-33080","INTEL-SA-00563"],"title":"Intel / Solidigm SSD, SSD DC and Optane SSD firmware","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel / Solidigm SSD, SSD DC and Optane SSD firmware - control-flow flaw and uncleared debug data reachable over…","year":"2021","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Two firmware defects in the same advisory, both scored 6.8 and both giving full confidentiality-plus-integrity impact to someone holding the drive. The control-flow flaw escalates privilege inside the drive controller; the second leaves debug information in the firmware that was never cleared for production, exposing internal state and enabling further escalation. Together they mean the drive controller itself is takeable - and a controller you control is a controller that can lie about sanitize, lie about encryption, and read every block regardless of Opal locking. BREAKS TENANT HANDOFF at the root: every erase-and-report guarantee in your reclaim pipeline is only as trustworthy as the firmware making the report, and here that firmware is compromisable. Also a leftover-debug-in-shipping-firmware finding, which tells you what the vendor's production hardening was actually worth on these SKUs.","attack_vector":"An unauthenticated attacker with physical access to the drive - no password, no host credential, no prior privilege. In an operator's world that means the RMA return path, a decommissioned node, a colo cage with shared access, or a supply-chain touchpoint before the drive was ever racked.","remediation":"Firmware flash per SKU using Intel MAS / Solidigm Storage Tool, drive offline and node drained. Match your inventory against INTEL-SA-00563 - it covers a wide spread of Intel SSD, SSD DC, Optane SSD and Optane SSD DC families with different fixed-firmware versions each, and some end-of-support SKUs get no fix at all. Beyond patching, this is the entry that should drive a chain-of-custody policy rather than a technical one: because exploitation needs only physical possession, the highest-leverage change is to stop letting drives that held tenant data leave your control intact. Destroy media on decommission instead of reselling, negotiate destroy-in-place terms with the vendor for RMA, and treat any drive with an unexplained gap in custody as compromised at the firmware level rather than re-racking it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-33077","https://nvd.nist.gov/vuln/detail/CVE-2021-33080","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00563.html","https://www.solidigm.com/support-page/support-security.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"published":"2022-05-12"},{"cwe":["CWE-476","CWE-843"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2021-47230","cve":"CVE-2021-47230","aliases":[],"title":"Linux kernel (arch/x86/kvm): A failed RSM leaves the vCPU's SMM flag and the MMU role out of sync, so KVM resolves a","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (arch/x86/kvm)","year":"2021","cvss_score":6.8,"severity":"medium","kev":false,"impact":"A failed RSM leaves the vCPU's SMM flag and the MMU role out of sync, so KVM resolves a GFN against the non-SMM memslot set and then looks the rmap up in the SMM set. The mismatched lookup faults in host kernel context (KASAN reports a null-pointer dereference / general protection fault); more broadly it is memslot-set confusion inside the shadow MMU, driven by the guest.","attack_vector":"Guest-side: execute RSM in a way that fails emulation, then take a page fault. Found by syzkaller from an unprivileged process running a guest. Applies to guests with SMM enabled (the default for OVMF/secure-boot guest firmware) on the shadow MMU path.","remediation":"Update to a kernel with the referenced stable commits (no fixed release string in the record - match by commit). Interim: disable SMM for tenant VMs on unpatched nodes where the guest firmware does not require it.","references":["https://git.kernel.org/stable/c/cbb425f62df9df7abee4b3f068f7ed6ffc3561e2","https://git.kernel.org/stable/c/669a8866e468fd020d34eb00e08cb41d3774b71b","https://nvd.nist.gov/vuln/detail/CVE-2021-47230"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2022-0004","cve":"CVE-2022-0004","aliases":[],"title":"Intel Boot Guard and Intel TXT (hardware debug / INIT): Hardware debug modes and processor INIT handling can override","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Boot Guard and Intel TXT (hardware debug / INIT)","year":"2022","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Hardware debug modes and processor INIT handling can override the locks that Boot Guard and TXT rely on, letting an unauthenticated attacker escalate privilege and defeat measured/verified boot. This reaches both of Intel's boot-integrity technologies at once - the two things a remote attestation claim about a bare-metal node ultimately rests on.","attack_vector":"An attacker with the ability to trigger the relevant debug or INIT conditions - realistically physical or deep platform access.","remediation":"Platform BIOS/firmware update from the OEM. Drain and reboot per node, and wait on OEM packaging. Until then, treat Boot Guard and TXT measurements from affected platforms as advisory rather than authoritative in any attestation policy you enforce.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-0004","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00613.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-05-12"},{"id":"CVE-2022-21657","cve":"CVE-2022-21657","aliases":[],"title":"Envoy: Envoy accepts any peer certificate rather than restricting to configured CAs","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy","year":"2022","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Envoy accepts any peer certificate rather than restricting to configured CAs; mTLS trust bypass","attack_vector":"Unauthenticated network with any valid-looking cert","remediation":"Upgrade Envoy; sidecar restart","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-21657"],"status":"curated","published":"2022-02-22"},{"id":"CVE-2022-28185","cve":"CVE-2022-28185","aliases":[],"title":"NVIDIA GPU Display Driver - kernel mode layer (Windows nvlddmkm.sys and Linux nvidia.ko): An out-of-bounds write","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - kernel mode layer (Windows nvlddmkm.sys and Linux nvidia.ko)","year":"2022","cvss_score":6.8,"severity":"medium","kev":false,"impact":"An out-of-bounds write in the driver's ECC layer, reachable by an unprivileged user, corrupts state and crashes the node. Relevant because ECC handling is the layer operators lean on for memory integrity on Tesla parts. Both the Windows and Linux datacenter drivers are affected, so a mixed fleet needs two separate rollouts.","attack_vector":"Local and unprivileged on either OS. On Linux it is reachable from any GPU container via /dev/nvidia*; on Windows from any session holding a GPU handle.","remediation":"Upgrade both the Linux and the Windows datacenter driver branches listed in bulletin 5353. Cost: Linux needs a drain and nvidia.ko reload per node; Windows needs a reboot per node. Two change windows unless your fleet is homogeneous.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28185","https://github.com/NVIDIA/product-security/tree/main/2022/5353"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:U/C:N/I:L/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-reboot"},"published":"2022-05-17"},{"id":"CVE-2022-31611","cve":"CVE-2022-31611","aliases":[],"title":"GeForce Experience installer: Local privesc during install","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GeForce Experience installer","year":"2022","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Local privesc during install","attack_vector":"Local user","remediation":"Consumer-only; no DC action","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31611","https://github.com/NVIDIA/product-security/tree/main/2023/5384"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:L/I:H/A:H","cwe":["CWE-427"],"published":"2023-02-07"},{"id":"CVE-2022-34674","cve":"CVE-2022-34674","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): A kernel helper maps more physical pages than were","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":6.8,"severity":"medium","kev":false,"impact":"A kernel helper maps more physical pages than were actually requested, so a caller can read physical memory that was never meant to be exposed to it - a direct residual-memory read from an unprivileged GPU workload. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5415. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34674","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:U/C:H/I:L/A:N","cwe":["CWE-200"],"fleet":{"pain_class":"node-drain"},"tags":["tenant-isolation"],"published":"2022-12-30"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:H/UI:N/S:C/C:H/I:N/A:N","cwe":["CWE-22"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2022-40607","cve":"CVE-2022-40607","aliases":[],"title":"IBM Spectrum Scale Container Native Storage Access (CSI volume handling): Anyone who can create a pod plus a PV/PVC","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"IBM Spectrum Scale Container Native Storage Access (CSI volume handling)","year":"2022","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Anyone who can create a pod plus a PV/PVC reads files outside their own volume, including files on the host filesystem. On a shared Kubernetes GPU cluster that is a straight escape from a namespace's storage to everyone else's data and to node secrets.","attack_vector":"Kubernetes API access sufficient to create pods and persistent volume claims in any namespace served by Storage Scale CNSA 5.1. That is the normal permission set handed to a tenant team.","remediation":"Upgrade Storage Scale CNSA to the fixed level in IBM's bulletin and roll the driver pods. Then check whether tenants actually need PV creation rights, or whether pre-provisioned volumes bound by an admin would do.","references":["https://www.ibm.com/support/pages/node/6848231","https://nvd.nist.gov/vuln/detail/CVE-2022-40607"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2023-0778","cve":"CVE-2023-0778","aliases":[],"title":"Podman: TOCTOU during volume export lets a symlink swap expose arbitrary host files","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Podman","year":"2023","cvss_score":6.8,"severity":"medium","kev":false,"impact":"TOCTOU during volume export lets a symlink swap expose arbitrary host files","attack_vector":"Any tenant workload with volume access","remediation":"Upgrade Podman","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0778"],"status":"curated","published":"2023-03-27"},{"id":"CVE-2023-20589","cve":"CVE-2023-20589","aliases":["faulTPM"],"title":"AMD Secure Processor secure boot - voltage fault injection (AMD-SB-4005): Voltage fault injection against the ASP","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor secure boot - voltage fault injection (AMD-SB-4005)","year":"2023","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Voltage fault injection against the ASP defeats its secure boot, yielding arbitrary code execution on the secure processor and extraction of fTPM-sealed secrets - the published work pulls BitLocker keys out. The affected list is Zen 1/2/3 **client** parts and EPYC is not on it, but the technique is the reason to read this entry: a datacenter operator with hardware in colocation, in transit, or coming back from RMA has to assume physical access to some nodes by someone.","attack_vector":"Physical access plus specialised fault-injection hardware. Not remote, not local-software.","remediation":"Fixed in AGESA/PI firmware for the affected client parts - OEM BIOS package, drain and power cycle. For a server fleet the practical answer is not patching but physical control: tamper-evident handling, chain of custody for RMAs and redeployments, and not trusting fTPM-sealed secrets on any node that has left your custody. AMD's position on the server analogue (AMD-SB-3028, voltage fault injection against SEV VMs on EPYC 7272) is **WONTFIX** - physical attacks are declared outside the SEV-SNP threat model, so there is no patch coming for the server case at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-20589","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-4005.html","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3028.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"tags":["tenant-isolation"],"published":"2023-08-08"},{"id":"CVE-2023-28005","cve":"CVE-2023-28005","aliases":[],"title":"Trend Micro Endpoint Encryption Full Disk Encryption (UEFI pre-boot): A signed pre-boot component that allows Secure","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Trend Micro Endpoint Encryption Full Disk Encryption (UEFI pre-boot)","year":"2023","cvss_score":6.8,"severity":"medium","kev":false,"impact":"A signed pre-boot component that allows Secure Boot bypass on affected machines. Relevant to any fleet where an endpoint-security vendor's UEFI component is part of the boot chain - the security product itself becomes the way in.","attack_vector":"Local access to a machine running the affected pre-boot component.","remediation":"Vendor product update plus, where the binary is revoked, a dbx update. Worth a general lesson for fleet operators: every third-party signed EFI component you allow into the boot chain is a permanent addition to your Secure Boot attack surface, and you cannot remove it later without a revocation rollout.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-28005"],"status":"curated","published":"2023-03-22"},{"id":"CVE-2023-28841","cve":"CVE-2023-28841","aliases":[],"title":"Docker / moby: Encrypted overlay network traffic can be unencrypted due to missing rules","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2023","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Encrypted overlay network traffic can be unencrypted due to missing rules","attack_vector":"Unauthenticated network adjacent to the underlay","remediation":"Upgrade moby","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-28841"],"status":"curated","published":"2023-04-04"},{"id":"CVE-2023-28842","cve":"CVE-2023-28842","aliases":[],"title":"Docker / moby: Unauthenticated injection of traffic into an encrypted overlay network","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2023","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Unauthenticated injection of traffic into an encrypted overlay network","attack_vector":"Unauthenticated network adjacent to the underlay","remediation":"Upgrade moby","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-28842"],"status":"curated","published":"2023-04-04"},{"id":"CVE-2023-31010","cve":"CVE-2023-31010","aliases":[],"title":"DGX H100 BMC (IPMI): DoS of the BMC","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX H100 BMC (IPMI)","year":"2023","cvss_score":6.8,"severity":"medium","kev":false,"impact":"DoS of the BMC","attack_vector":"Network-adjacent IPMI client","remediation":"Flash BMC 23.08.18 out-of-band","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5473/5473.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-20"],"published":"2023-09-20"},{"id":"CVE-2023-31033","cve":"CVE-2023-31033","aliases":[],"title":"DGX A100 BMC: Missing authentication on BMC service","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX A100 BMC","year":"2023","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Missing authentication on BMC service","attack_vector":"Network-adjacent unauthenticated","remediation":"Flash BMC 00.22.05+ out-of-band","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31033","https://github.com/NVIDIA/product-security/tree/main/2024/5510"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-306"],"published":"2024-01-12"},{"id":"CVE-2024-0111","cve":"CVE-2024-0111","aliases":[],"title":"CUDA Toolkit: Memory-safety issue","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2024","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Memory-safety issue -> code exec","attack_vector":"Malicious model/binary artifact","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0111","https://github.com/NVIDIA/product-security/tree/main/2024/5564"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:L/A:L","cwe":["CWE-1284"],"published":"2024-08-31"},{"id":"CVE-2024-0140","cve":"CVE-2024-0140","aliases":[],"title":"RAPIDS cuDF / cuML: RCE via unsafe deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"RAPIDS cuDF / cuML","year":"2024","cvss_score":6.8,"severity":"medium","kev":false,"impact":"RCE via unsafe deserialization","attack_vector":"Malicious model/dataset artifact loaded by a tenant job","remediation":"Bump RAPIDS packages; rebuild data-science base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0140","https://github.com/NVIDIA/product-security/tree/main/2025/5597"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:L/I:H/A:H","cwe":["CWE-502"],"published":"2025-01-28"},{"id":"CVE-2024-0141","cve":"CVE-2024-0141","aliases":[],"title":"Hopper HGX 8-GPU: DoS (inadequate loop termination in firmware)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Hopper HGX 8-GPU","year":"2024","cvss_score":6.8,"severity":"medium","kev":false,"impact":"DoS (inadequate loop termination in firmware)","attack_vector":"Local privileged host access","remediation":"Flash HGX baseboard firmware; node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0141","https://github.com/NVIDIA/product-security/tree/main/2025/5561"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:H/UI:N/S:C/C:N/I:N/A:H","cwe":["CWE-782"],"fleet":{"pain_class":"node-drain"},"published":"2025-03-05"},{"id":"CVE-2024-0142","cve":"CVE-2024-0142","aliases":[],"title":"nvJPEG2000: Code exec via OOB write on malicious image","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"nvJPEG2000","year":"2024","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Code exec via OOB write on malicious image","attack_vector":"Malicious dataset / user-uploaded image","remediation":"Bump nvJPEG2000 in preprocessing images; rebuild","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0142","https://github.com/NVIDIA/product-security/tree/main/2025/5596"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:R/S:U/C:N/I:H/A:H","cwe":["CWE-787"],"published":"2025-02-12"},{"id":"CVE-2024-0143","cve":"CVE-2024-0143","aliases":[],"title":"nvJPEG2000: Code exec via OOB write","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"nvJPEG2000","year":"2024","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Code exec via OOB write","attack_vector":"Malicious dataset","remediation":"Bump nvJPEG2000; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0143","https://github.com/NVIDIA/product-security/tree/main/2025/5596"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:R/S:U/C:N/I:H/A:H","cwe":["CWE-787"],"published":"2025-02-12"},{"id":"CVE-2024-0144","cve":"CVE-2024-0144","aliases":[],"title":"nvJPEG2000: Code exec via buffer overflow","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"nvJPEG2000","year":"2024","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Code exec via buffer overflow","attack_vector":"Malicious dataset","remediation":"Bump nvJPEG2000; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0144","https://github.com/NVIDIA/product-security/tree/main/2025/5596"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:R/S:U/C:N/I:H/A:H","cwe":["CWE-120"],"published":"2025-02-12"},{"id":"CVE-2024-0145","cve":"CVE-2024-0145","aliases":[],"title":"nvJPEG2000: Code exec via heap overflow","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"nvJPEG2000","year":"2024","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Code exec via heap overflow","attack_vector":"Malicious dataset","remediation":"Bump nvJPEG2000; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0145","https://github.com/NVIDIA/product-security/tree/main/2025/5596"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:R/S:U/C:N/I:H/A:H","cwe":["CWE-122"],"published":"2025-02-12"},{"id":"CVE-2024-2315","cve":"CVE-2024-2315","aliases":["AMI-SA-2024004"],"title":"AMI AptioV UEFI BIOS (SPI flash access control): Improper access control in the BIOS that lets a local attacker make","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI AptioV UEFI BIOS (SPI flash access control)","year":"2024","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Improper access control in the BIOS that lets a local attacker make unexpected SPI flash modifications and launch a BIOS bootkit. This is the direct route to firmware persistence: write the flash, own every subsequent boot, and become invisible to the OS and to every agent running in it. AMI also calls out an availability impact - a bad write bricks the board, which on a GPU node means an RMA and weeks of lost capacity rather than a reboot.","attack_vector":"Local access with low privileges, no user interaction. Code on the host OS is enough; it does not require root by AMI's scoring. That makes it one of the cheaper firmware-persistence paths in this cluster for an attacker who has landed anywhere on the node.","remediation":"BIOS update to BKC_5.37 or later - firmware flash plus a host reboot, per node, gated on your server vendor's rebase. Alongside the update, verify that the platform's flash write protections are actually enabled in your BIOS configuration (flash descriptor lock, BIOS write-protect, boot guard where the platform supports it) - operators routinely find these left open by an OEM's default profile, and that config check costs one reboot rather than a flash campaign.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/2024/AMI-SA-2024004.pdf","https://nvd.nist.gov/vuln/detail/CVE-2024-2315"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-11-12"},{"id":"CVE-2024-37085","cve":"CVE-2024-37085","aliases":[],"title":"VMware ESXi: AD-integrated ESXi grants full host admin to any member of a re-created \"ESX Admins\" group","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"VMware ESXi","year":"2024","cvss_score":6.8,"severity":"medium","kev":true,"impact":"AD-integrated ESXi grants full host admin to any member of a re-created \"ESX Admins\" group - used by Akira and Black Basta to mass-encrypt VMs [KEV]","attack_vector":"Attacker with Active Directory write access (post-initial-access, not tenant-facing)","remediation":"Patch ESXi and stop using AD for ESXi user management. Configuration change, not just a binary update - the real fix is removing the AD trust from the hypervisor plane","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-37085"],"status":"curated","published":"2024-06-25"},{"cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-42302","cve":"CVE-2024-42302","aliases":[],"title":"Linux kernel (drivers/pci): A Downstream Port Containment event and a device removal happening at the same time leave","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/pci)","year":"2024","cvss_score":6.8,"severity":"medium","kev":false,"impact":"A Downstream Port Containment event and a device removal happening at the same time leave the DPC handler polling the config space of a struct pci_dev that has already been freed. That is a use-after-free in an interrupt thread on the host - it panics the node, and a freed-and-reallocated pci_dev is the kind of primitive that turns a device-triggered error into host memory corruption rather than just a crash.","attack_vector":"The DPC half is device-driven, not admin-driven: an endpoint that emits an uncorrectable error - a wedged or deliberately misbehaving NVMe drive, a passthrough GPU or NIC a tenant is hammering with malformed transactions - makes the Downstream Port fire DPC, and dpc_handler() then walks the child device without holding a reference. If a hot-removal (pciehp, surprise removal of an NVMe, or an admin-triggered remove) races that window, the handler dereferences freed memory. Reachable on any node with DPC-capable root/downstream ports and hot-pluggable devices; needs no tenant credentials on the host, only a device that can be pushed into generating errors.","remediation":"Update to 5.10.224 / 5.15.165 / 6.1.103 / 6.3 or later. Interim: avoid concurrent hot-remove operations on ports that have DPC enabled, and drain a node before servicing or removing PCIe devices rather than surprise-pulling them under load.","references":["https://git.kernel.org/stable/c/c52f9e1a9eb40f13993142c331a6cfd334d4b91d","https://git.kernel.org/stable/c/2c111413f38ca5cf87557cab89f6d82b0e3433e7","https://nvd.nist.gov/vuln/detail/CVE-2024-42302"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-42488","cve":"CVE-2024-42488","aliases":[],"title":"Cilium: Agent race condition drops pod labels, so the wrong (often more permissive) policy applies","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2024","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Agent race condition drops pod labels, so the wrong (often more permissive) policy applies","attack_vector":"Any tenant workload, timing-dependent","remediation":"Rolling Cilium upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-42488"],"status":"curated","published":"2024-08-15"},{"id":"CVE-2024-45101","cve":"CVE-2024-45101","aliases":["LEN-154748"],"title":"Lenovo XClarity Administrator (LXCA) - single sign-on to XCC: Where LXCA acts as the single sign-on provider for XCC","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lenovo XClarity Administrator (LXCA) - single sign-on to XCC","year":"2024","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Where LXCA acts as the single sign-on provider for XCC, an attacker who gets an authenticated LXCA user to click a crafted URL can intercept that user's XCC session. The stolen session is a live, authenticated connection to a node's service processor, carrying whatever privileges the victim held - typically enough for power control, Virtual Media and console. The interesting property is that SSO, which operators deploy to reduce credential sprawl across BMCs, becomes the mechanism that spreads one phished click into out-of-band access. Affects LXCA before 4.1.","attack_vector":"Requires an authenticated LXCA user to click an attacker-supplied link, and requires that SSO between LXCA and XCC is enabled. The attacker needs no reachability to the management VLAN themselves - the operator's browser session is the bridge.","remediation":"Upgrade LXCA to 4.1 or later - a single appliance upgrade, no per-node work, no reboots and no job drain. Config-only mitigation if you cannot upgrade immediately: disable LXCA-to-XCC single sign-on and fall back to direct XCC authentication, accepting the credential-management cost. Independently, administer LXCA and XCC from a dedicated browser profile or privileged access workstation so a crafted link cannot reach a live management session.","references":["https://support.lenovo.com/us/en/product_security/LEN-154748","https://nvd.nist.gov/vuln/detail/CVE-2024-45101"],"status":"curated","published":"2024-09-13"},{"cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-53194","cve":"CVE-2024-53194","aliases":[],"title":"Linux kernel (drivers/pci): A pci_slot holds an uncounted pointer to the pci_bus below it, and on hot removal the bus","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/pci)","year":"2024","cvss_score":6.8,"severity":"medium","kev":false,"impact":"A pci_slot holds an uncounted pointer to the pci_bus below it, and on hot removal the bus can be destroyed before the slot is. The slot release path then dereferences freed memory, crashing the host. The window exists because nothing prevents a driver binding between the stop and remove halves of PCI teardown, so a device that drops its link at the wrong moment gets a use-after-free rather than a clean removal.","attack_vector":"Driven by the device, not by tenant software: any event that clears Presence Detect / Data Link Layer Link Active on a hotplug-capable downstream port starts the removal - a link drop, a card reset, a sled being pulled, or in the reported case a host-router reset during driver probe. Needs pciehp managing the slot and a hotplug hierarchy below it, which is the normal shape for NVMe bays, PCIe switch fabrics and any externally cabled expansion. A tenant that can wedge or reset its assigned card hard enough to bounce the link influences when it fires; it cannot reach the code directly.","remediation":"Boot a kernel where pci_slot takes a counted reference on its pci_bus. Interim: avoid driver bind/unbind churn on hotplug hierarchies while removals are in flight, and treat repeated link-down events on a node as a drain signal rather than something to let retry in place.","references":["https://git.kernel.org/stable/c/50473dd3b2a08601a078f852ea05572de9b1f86c","https://git.kernel.org/stable/c/d0ddd2c92b75a19a37c887154223372b600fed37","https://nvd.nist.gov/vuln/detail/CVE-2024-53194"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-7726","cve":"CVE-2024-7726","aliases":["GHSA-3hh8-94j4-62rh","Kioxia JTAG"],"title":"Kioxia CM6 (GPK5 and earlier), PM6 (BD0D and earlier), PM7 (C40A and earlier) enterprise NVMe/SAS SSDs","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Kioxia CM6 (GPK5 and earlier), PM6 (BD0D and earlier), PM7 (C40A and earlier) enterprise NVMe/SAS SSDs…","year":"2024","cvss_score":6.8,"severity":"medium","kev":false,"impact":"An open, unauthenticated JTAG debug port on the drive's PCB exposes both main SoC CPU cores - and the enclosure cutout is wide enough that you do not even need to open the drive to reach it. With a cheap ARM JTAG probe an attacker executes arbitrary code on the controller, reads firmware and memory, and bypasses the RSA firmware signature check at boot. These are Kioxia's flagship DATACENTER drives; CM6/PM6/PM7 are standard fitment in the exact GPU and AI-server platforms this database is about. BREAKS TENANT HANDOFF and does it in the worst direction: an attacker who owns the controller can read everything regardless of Opal state, can make sanitize report success while preserving data, and can attempt to leave an implant that survives every reimage you perform. Note the researchers' own caveat - fully PERSISTENT firmware modification additionally requires a shared secret used to compute the firmware MAC, so persistence is not demonstrated as trivially achievable, but transient full control of the controller is.","attack_vector":"Anyone with brief physical access to the drive and a low-cost JTAG probe. The enclosure does not have to be opened, so this is minutes of unsupervised contact, not a lab teardown - a rack tech, a colo neighbour with cage access, a courier in the RMA path, a decommission handler, or anyone in the supply chain before the drive reached you.","remediation":"UNPATCHABLE on the affected SKUs - the Google advisory records no fixed version, and an exposed JTAG pad is a board-design property that no firmware update removes. This is therefore a physical-security and procurement control, not a patch: enforce tamper-evident seals and chain of custody on CM6/PM6/PM7 media, never return or resell a drive that held tenant data (destroy it), and treat any of these drives with an unexplained custody gap as compromised at controller level rather than reimaging and re-renting it. Because you cannot trust the controller's own attestations on these drives, run LUKS/dm-crypt with an operator-held key so the drive only ever sees ciphertext and reclaim means destroying your key. Raise it with Kioxia as a procurement question for future SKUs - the fix has to come in hardware.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-7726","https://github.com/google/security-research/security/advisories/GHSA-3hh8-94j4-62rh"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"published":"2024-12-20"},{"id":"CVE-2025-0012","cve":"CVE-2025-0012","aliases":[],"title":"AMD - overlap between segmented reverse map table (RMP) and SMM memory: Improper handling of overlap between the","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD - overlap between segmented reverse map table (RMP) and SMM memory","year":"2025","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Improper handling of overlap between the segmented RMP and System Management Mode memory lets a privileged attacker corrupt or partially infer SMM memory. SMM is the most privileged execution context on x86 - above the hypervisor - so reaching it from the RMP path is both a route to total platform control and, in the inference direction, a leak out of the one context nothing else can inspect.","attack_vector":"Local, privileged attacker on a platform using segmented RMP (large-memory SEV-SNP configurations).","remediation":"Fixed in AMD firmware/AGESA, delivered as an OEM SBIOS package with **one to six months of OEM lag** and a drained-node power cycle. Because it touches the SEV-SNP trust boundary, refresh VCEK certificates and update tenant attestation policy after the TCB version moves. Segmented RMP is used on very large memory configurations - exactly the shape of an AI training host - so do not assume this is an edge case on a GPU fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0012","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2026-02-10"},{"id":"CVE-2025-14304","cve":"CVE-2025-14304","aliases":[],"title":"Motherboards from ASRock and its subsidiaries ASRockRack and ASRockInd built on Intel 500-series chipsets","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Motherboards from ASRock and its subsidiaries ASRockRack and ASRockInd built on Intel 500-series chipsets","year":"2025","cvss_score":6.8,"severity":"medium","kev":false,"impact":"A DMA-capable PCIe device reads and writes system memory without restriction. On a GPU platform this is a direct hit on the assumption the whole tenant-isolation model rests on: IOMMU enforcement is what stops a device - or a tenant-controlled device - from reading arbitrary host memory, and it is what makes GPU passthrough to a VM safe. With it off, an attacker with a malicious or reprogrammed peripheral, or with control of a passed-through device, reads host kernel memory, extracts keys, and writes to memory to escalate. For any operator doing GPU passthrough or accepting tenant-supplied hardware, this invalidates the isolation guarantee. The IOMMU is not properly enabled, so the protection that is supposed to confine what a PCIe device can reach in system memory is simply not active.","attack_vector":"Physical access sufficient to attach a DMA-capable PCIe device - which includes Thunderbolt/USB4 ports, open PCIe slots, and any peripheral in a colocation or edge environment where the chassis is not under your exclusive control. Also relevant wherever a device is passed through to an untrusted guest, since the confinement that passthrough relies on is absent.","remediation":"BIOS/UEFI update from ASRock, then verify - do not assume. After flashing, confirm the IOMMU is actually active at runtime: check for DMAR/IVRS tables and that the kernel reports IOMMU groups, rather than trusting a BIOS setting. ASRock and ASRockRack publish a security page, but it sits behind an Imperva challenge and returns unreadable content to automated clients, so advisory tracking has to be manual or via NVD and TWCERT. Config-only hardening in the meantime: enable IOMMU explicitly in BIOS and on the kernel command line, and disable unused Thunderbolt/PCIe hot-plug paths.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-14304","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2025/14xxx/CVE-2025-14304.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-23216","cve":"CVE-2025-23216","aliases":[],"title":"Argo CD: Secret values exposed in error messages and the diff view when an invalid Secret is synced","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Argo CD","year":"2025","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Secret values exposed in error messages and the diff view when an invalid Secret is synced","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade + rotate any secret rendered in the UI","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23216"],"status":"curated","published":"2025-01-30"},{"id":"CVE-2025-26465","cve":"CVE-2025-26465","aliases":[],"title":"OpenSSH (client): Machine-in-the-middle against the client when VerifyHostKeyDNS is enabled","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"OpenSSH (client)","year":"2025","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Machine-in-the-middle against the client when VerifyHostKeyDNS is enabled","attack_vector":"Unauthenticated network (MITM)","remediation":"Package update; no reboot. Audit client configs for VerifyHostKeyDNS","references":["https://access.redhat.com/security/cve/CVE-2025-26465"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2025-02-18"},{"id":"CVE-2025-2713","cve":"CVE-2025-2713","aliases":[],"title":"gVisor: runsc mishandles file access permissions, letting unprivileged users read restricted files","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"gVisor","year":"2025","cvss_score":6.8,"severity":"medium","kev":false,"impact":"runsc mishandles file access permissions, letting unprivileged users read restricted files; local privilege escalation","attack_vector":"Any tenant workload inside the sandbox","remediation":"Upgrade runsc; restart sandboxed pods","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-2713"],"status":"curated","published":"2025-03-28"},{"id":"CVE-2025-33215","cve":"CVE-2025-33215","aliases":[],"title":"SNAP-4 container (BlueField storage): DoS of the storage dataplane","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"SNAP-4 container (BlueField storage)","year":"2025","cvss_score":6.8,"severity":"medium","kev":false,"impact":"DoS of the storage dataplane","attack_vector":"Authenticated network attacker","remediation":"Bump SNAP container image; restart DPU storage service","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33215","https://github.com/NVIDIA/product-security/tree/main/2026/5744"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:L/UI:N/S:C/C:N/I:N/A:H","cwe":["CWE-823"],"published":"2026-03-24"},{"id":"CVE-2025-33216","cve":"CVE-2025-33216","aliases":[],"title":"SNAP-4 container: DoS (integer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"SNAP-4 container","year":"2025","cvss_score":6.8,"severity":"medium","kev":false,"impact":"DoS (integer overflow)","attack_vector":"Network-accessible client","remediation":"Bump SNAP container image","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33216","https://github.com/NVIDIA/product-security/tree/main/2026/5744"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:L/UI:N/S:C/C:N/I:N/A:H","cwe":["CWE-131"],"published":"2026-03-24"},{"id":"CVE-2025-35979","cve":"CVE-2025-35979","aliases":["guest-mode predictor state leak","VMX non-root transient execution disclosure"],"title":"Intel processors, exploitable from within VMX non-root (guest) operation - INTEL-SA-01420: Shared microarchitectural","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors, exploitable from within VMX non-root (guest) operation - INTEL-SA-01420","year":"2025","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Shared microarchitectural predictor state influences transient execution inside guest (VMX non-root) operation, letting unprivileged software in a VM observe data it should not. This is the shape of bug that matters most to anyone renting VMs on shared hosts: the leak is reachable from inside a guest by ordinary unprivileged code, targeting state shared with whatever else the host is running. For a neocloud running multiple tenant VMs per physical machine, it is a tenant-boundary issue by construction; for a bare-metal-per-tenant product it is contained to that tenant.","attack_vector":"Unprivileged software inside a guest VM. Intel rates the attack complexity as high and notes attack requirements must be present, so this is a capable-adversary scenario rather than a commodity exploit - but the position required is just 'a customer with a VM'.","remediation":"Microcode/BIOS update via OEM firmware - firmware flash, host reboot, job drain - plus hypervisor updates where the VMM must invoke the new controls. This is part of Intel's 2026 quarterly advisory batch, so bundle it with the other CVEs in that IPU rather than scheduling separately; the marginal cost of adding it to an existing firmware window is zero and the cost of its own window is a full fleet drain. No SMT decision attached. Verify by microcode revision and by the hypervisor's own mitigation reporting.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-35979","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01420.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2026-05-12"},{"id":"CVE-2025-9708","cve":"CVE-2025-9708","aliases":[],"title":"Kubernetes C# client: Improper certificate validation in custom-CA mode enables MITM on the Kubernetes API connection","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes C# client","year":"2025","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Improper certificate validation in custom-CA mode enables MITM on the Kubernetes API connection","attack_vector":"Unauthenticated network in a MITM position","remediation":"Upgrade any internal tooling built on the C# client","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated","published":"2025-09-16"},{"id":"CVE-2026-24234","cve":"CVE-2026-24234","aliases":[],"title":"TensorRT-LLM: SSRF via plugin loading","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT-LLM","year":"2026","cvss_score":6.8,"severity":"medium","kev":false,"impact":"SSRF via plugin loading","attack_vector":"Tenant-supplied plugin reference","remediation":"Bump TensorRT-LLM; add egress policy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24234","https://github.com/NVIDIA/product-security/tree/main/2026/5840"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:C/C:L/I:L/A:L","cwe":["CWE-918"],"published":"2026-07-14"},{"id":"CVE-2026-27765","cve":"CVE-2026-27765","aliases":[],"title":"Intel Gaudi / vLLM hardware plugin: Malformed input to the Gaudi vLLM plugin crashes or wedges the serving process","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Intel Gaudi / vLLM hardware plugin","year":"2026","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Malformed input to the Gaudi vLLM plugin crashes or wedges the serving process. On an inference fleet this is a request-triggered outage of the model server rather than a data-confidentiality problem, but a single tenant can repeatedly take down a shared endpoint that is pinned to expensive accelerators.","attack_vector":"Anyone who can reach the vLLM endpoint with an authenticated request - so any tenant of a shared inference service, or any workload inside the cluster if the endpoint is not network-segmented.","remediation":"Upgrade the vLLM hardware plugin for Gaudi to 0.16.0 or later. Pure Python/userspace package update - restart the serving process, no node reboot, no firmware.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-27765","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01487.html"],"status":"curated","published":"2026-08-11"},{"cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:L/UI:N/S:C/C:N/I:N/A:H","cwe":["CWE-400","CWE-770"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2026-44247","cve":"CVE-2026-44247","aliases":["GHSA-8wxp-xxp2-rcgx"],"title":"Volcano (admission webhook server, unbounded HTTP request body): The Volcano webhook server accepts request bodies of","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Volcano (admission webhook server, unbounded HTTP request body)","year":"2026","cvss_score":6.8,"severity":"medium","kev":false,"impact":"The Volcano webhook server accepts request bodies of any size, so any pod in the cluster can OOM-kill it. When the admission webhook is down, Volcano job and pod admission fails, which stalls GPU scheduling for every tenant that goes through Volcano.","attack_vector":"Any in-cluster pod that can reach the Volcano webhook service endpoint. Requires only network reach, not Volcano permissions.","remediation":"Upgrade Volcano to 1.12.4, 1.13.3 or 1.14.2 and restart the webhook deployment. In the meantime restrict the webhook Service with a NetworkPolicy so only the API server can reach it, and set a memory limit on the webhook pod.","references":["https://github.com/volcano-sh/volcano/security/advisories/GHSA-8wxp-xxp2-rcgx","https://nvd.nist.gov/vuln/detail/CVE-2026-44247"],"status":"curated","tags":["tenant-isolation"]},{"id":"NCVD-2026-004-openbmc-bmcweb-mtls-client-certi","cve":null,"aliases":["AUTH-F6"],"title":"OpenBMC bmcweb mTLS client-certificate UPN validation: Where mTLS is configured, bmcweb matches the certificate's UPN","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"OpenBMC bmcweb mTLS client-certificate UPN validation","year":"2026","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Where mTLS is configured, bmcweb matches the certificate's UPN by walking dot-separated labels with no bound. A certificate issued for user@com authenticates against any host under that TLD, and a parent-domain certificate authenticates against every child deployment. For an operator running mTLS across a fleet - which is the hardened configuration, chosen by the most security-conscious teams - this means one certificate from anywhere in the hierarchy authenticates to every BMC in it, silently, with no failed-auth log to notice. It is the rare bug that punishes you specifically for having done the harder thing.","attack_vector":"Requires mTLS to be enabled on the BMC (not the default) and possession of any certificate the BMC's trust store chains to, including one issued for a parent domain or a different deployment. Network access to the BMC's HTTPS port.","remediation":"Unpatched at disclosure. An April 2026 commit fixed only case-insensitivity in the comparison and left the suffix-walking logic intact, so do not assume a recent bmcweb clears it. If you run mTLS on BMCs, audit which CAs are in each BMC's trust store and narrow them to a per-fleet issuing CA that signs nothing else - that is a config-only change and is the effective mitigation today. Do not put a broadly-scoped corporate CA in a BMC trust store. A real fix will require a BMC firmware flash once upstream lands one.","references":["https://seclists.org/fulldisclosure/2026/May/24","https://binreaper.pages.dev/posts/2026-05-27-bmcweb-disclosure/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"cvss_vector":"CVSS:3.1/AV:A/AC:H/PR:N/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-522"],"fleet":{"pain_class":"hot-patch"},"id":"NCVD-2026-043-rclone-s3-backend-redirect-sanit","cve":null,"aliases":["GHSA-8mxv-9xhp-86h4"],"title":"rclone (S3 backend, redirect sanitization): When rclone's S3 backend follows a redirect it strips some sensitive","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"rclone (S3 backend, redirect sanitization)","year":"2026","cvss_score":6.8,"severity":"medium","kev":false,"impact":"When rclone's S3 backend follows a redirect it strips some sensitive headers but not IBM IAM bearer tokens or SSE-C customer-supplied encryption keys. An attacker who can steer a redirect harvests the bearer token, and the SSE-C key is the thing standing between them and the plaintext of the encrypted objects - so this leaks both the credential and the key material for a dataset store in one shot.","attack_vector":"An attacker positioned to influence the HTTP redirect chain between rclone and the object store - a hostile or compromised endpoint, or a network position on the cluster's egress path.","remediation":"Upgrade rclone to the fixed release. Rotate any IBM IAM credentials and SSE-C keys that were used with a redirecting endpoint. Pin the S3 endpoint and disable redirect following where your object store does not need it. No CVE ID has been assigned; track it by the GHSA.","references":["https://github.com/rclone/rclone/security/advisories/GHSA-8mxv-9xhp-86h4"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:C/C:N/I:L/A:H","cwe":["CWE-264"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2013-0153","cve":"CVE-2013-0153","aliases":["XSA-36"],"title":"Xen AMD-Vi (AMD IOMMU) interrupt remapping table handling: On AMD-Vi platforms Xen used a single interrupt remapping","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen AMD-Vi (AMD IOMMU) interrupt remapping table handling","year":"2013","cvss_score":6.7,"severity":"medium","kev":false,"impact":"On AMD-Vi platforms Xen used a single interrupt remapping table shared between the host and every guest, and did not clear stale entries when a device was reassigned. A tenant with a passed-through device therefore programs interrupt vectors that land in other tenants' domains or in the hypervisor. This is a cross-tenant channel that exists at the level below the hypervisor's own bookkeeping - a shared hardware table doing exactly what a shared hardware table does. Impact per the advisory is denial of service against other guests, but the stale-entry half means a device reassigned from tenant A to tenant B inherits A's remapping state, which is a reuse-hygiene failure of the same family as unzeroed framebuffers.","attack_vector":"A guest with a passed-through device on an AMD-Vi host. Affects the other guests on that host, not just the attacker's own.","remediation":"Apply the XSA-36 patches for the Xen 4.1/4.2 branches and reboot; per-device interrupt remapping tables are a hypervisor structural change, so there is no configuration workaround. If you are running AMD platforms with GPU passthrough on an unpatched Xen, the honest interim posture is to treat every host as single-tenant until it is patched - co-tenancy on a shared IRT is not a boundary you can compensate for at the network or scheduler layer.","references":["https://xenbits.xen.org/xsa/advisory-36.html","https://nvd.nist.gov/vuln/detail/CVE-2013-0153"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2017-12334","cve":"CVE-2017-12334","aliases":[],"title":"Cisco NX-OS CLI: CLI command injection giving root-level execution on the switch OS for an authenticated admin","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco NX-OS CLI","year":"2017","cvss_score":6.7,"severity":"medium","kev":false,"impact":"CLI command injection giving root-level execution on the switch OS for an authenticated admin — the same pattern later re-appeared as the exploited CVE-2024-20399","attack_vector":"Network, authenticated admin","remediation":"NX-OS upgrade with fabric failover; recurrence of the class means CLI-injection hardening on the switch OS is a standing item, not a one-off patch","references":["https://nvd.nist.gov/vuln/detail/CVE-2017-12334"],"status":"curated","published":"2017-11-30"},{"id":"CVE-2018-12204","cve":"CVE-2018-12204","aliases":[],"title":"Intel Server Board / Server System / Compute Module platform firmware: Improper memory initialisation in platform","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Server Board / Server System / Compute Module platform firmware","year":"2018","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Improper memory initialisation in platform sample/silicon reference firmware on Intel server boards, allowing privilege escalation. Reference firmware defects propagate into whatever the OEM built on top of it, so the affected population is wider than Intel-branded boards.","attack_vector":"Privileged local access on the host.","remediation":"Fixed in platform BIOS/UEFI firmware. That means an OEM release, a per-node drain, a flash and a cold reboot - and OEM availability commonly lags the Intel advisory by quarters on server boards. There is no microcode or OS-level shortcut for this class; budget it as a fleet-wide maintenance campaign, not a patch.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-12204","https://www.intel.com/content/www/us/en/security-center/advisory/INTEL-SA-00191.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2019-03-14"},{"id":"CVE-2018-13787","cve":"CVE-2018-13787","aliases":[],"title":"SPI flash descriptor region configuration on a wide range of Supermicro boards: Any software running with sufficient","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"SPI flash descriptor region configuration on a wide range of Supermicro boards","year":"2018","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Any software running with sufficient privilege on the host operating system can rewrite the platform firmware. That is a UEFI/BIOS implant written from userland-reachable code - it survives disk wipes, OS reinstalls, and node reprovisioning, and it is invisible to every host-based security tool the operator runs. On bare-metal GPU rental this is the canonical tenant-persistence attack: a tenant with root on their leased node writes firmware, releases the node, and retains a foothold on hardware that is subsequently handed to someone else. X11S, X10, X9, X8SI, K1SP, C9X299, C7, B1, A2 and A1 families. The descriptor is what tells the chipset which flash regions the host CPU may write; Supermicro shipped it misconfigured so the OS could write firmware.","attack_vector":"Host-side privileged code execution - root on Linux or an equivalent. No network position on the management VLAN is required at all, which makes it the mirror image of the BMC bugs in this list: the threat comes from inside the node, from whoever you rented it to.","remediation":"BIOS/firmware flash with a Supermicro image that ships a locked descriptor region, per board family. Verification matters more than the flash here: after updating, confirm the descriptor is actually locked - Intel's chipsec `common.bios_wp` and descriptor checks will tell you - because the fix is a configuration inside the firmware image and it is easy to assume it landed when it did not. For a bare-metal rental fleet, add a descriptor/flash-lock check to the node turnup and node return pipeline rather than treating it as a one-time patch.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-13787","https://eclypsium.com/blog/firmware-vulnerabilities-in-supermicro-systems/","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2018/13xxx/CVE-2018-13787.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2019-0119","cve":"CVE-2019-0119","aliases":[],"title":"Intel Xeon D / Xeon Scalable system firmware, Server Board and Server System: A buffer overflow in system firmware","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Xeon D / Xeon Scalable system firmware, Server Board and Server System","year":"2019","cvss_score":6.7,"severity":"medium","kev":false,"impact":"A buffer overflow in system firmware across Xeon D and Xeon Scalable server boards, letting a privileged user escalate into firmware. Affects the exact generations that built out the first wave of GPU datacenter capacity, and many of those nodes are still running.","attack_vector":"Privileged local access on the host.","remediation":"Fixed in platform BIOS/UEFI firmware. That means an OEM release, a per-node drain, a flash and a cold reboot - and OEM availability commonly lags the Intel advisory by quarters on server boards. There is no microcode or OS-level shortcut for this class; budget it as a fleet-wide maintenance campaign, not a patch.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-0119","https://www.intel.com/content/www/us/en/security-center/advisory/INTEL-SA-00223.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2019-05-17"},{"id":"CVE-2019-11157","cve":"CVE-2019-11157","aliases":["Plundervolt","V0ltPwn"],"title":"Intel SGX / dynamic voltage and frequency scaling interface: Undervolting the CPU through the privileged","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SGX / dynamic voltage and frequency scaling interface","year":"2019","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Undervolting the CPU through the privileged voltage-scaling MSR induces faults inside SGX enclaves, which is enough to extract AES keys and to corrupt RSA-CRT signatures into key-recovering faulty signatures. The attacker is the host operator - which is precisely the party SGX is supposed to be defended against - so this collapses confidential compute on the affected generations.","attack_vector":"Privileged local access on the host (ring 0). That is the SGX threat model: the platform owner attacking a tenant's enclave, or a compromised hypervisor attacking a confidential workload.","remediation":"BIOS/firmware update that locks the undervolting MSR, and a TCB recovery: enclaves must re-attest against the new SVN and any secrets sealed under the old TCB should be considered compromised. Because the lock is applied by platform firmware, this needs an OEM BIOS release and a full drain-and-reboot per node - OEM availability historically lagged Intel's advisory by months on server boards.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-11157","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00289.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2019-12-16"},{"id":"CVE-2019-5676","cve":"CVE-2019-5676","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): The driver loads Windows system DLLs without validating path","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2019","cvss_score":6.7,"severity":"medium","kev":false,"impact":"The driver loads Windows system DLLs without validating path or signature. A local user who can write to a searched directory gets code execution in a privileged NVIDIA process.","attack_vector":"Any local user with write access to a directory in the driver process's DLL search path.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://support.lenovo.com/us/en/product_security/LEN-27815","https://nvd.nist.gov/vuln/detail/CVE-2019-5676"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2019-05-10"},{"id":"CVE-2019-5688","cve":"CVE-2019-5688","aliases":[],"title":"NVIDIA NVFlash / NVUFlash / GPUModeSwitch kernel driver (nvflash.sys, nvflsh32/64.sys): NVIDIA's signed flashing driver","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NVFlash / NVUFlash / GPUModeSwitch kernel driver (nvflash.sys, nvflsh32/64.sys)","year":"2019","cvss_score":6.7,"severity":"medium","kev":false,"impact":"NVIDIA's signed flashing driver hands an administrator raw access to device memory and registers of arbitrary PCIe devices - including hardware NVIDIA does not own. That is a general-purpose physical-memory and device-register primitive wrapped in a legitimately signed driver, so it defeats driver signature enforcement and lets an admin (or malware that reached admin) reach across to the BMC, NICs, NVMe and other tenants' passthrough devices on the same host. Even on hosts that never flashed a GPU, its presence is a lasting bring-your-own-vulnerable-driver weapon.","attack_vector":"An authenticated administrator on the host, or any code that reached admin - which is exactly the boundary a signed kernel driver is supposed to hold.","remediation":"Update NVFlash/NVUFlash to 5.588.0 or later and GPUModeSwitch to the 2019-11 build, and remove the old nvflash.sys / nvflsh32.sys / nvflsh64.sys drivers from every host they were ever installed on. These are flashing tools, not runtime components: the durable fix is to keep the signed vulnerable driver off production hosts entirely and add its hashes to your driver blocklist, since it is a bring-your-own-vulnerable-driver primitive independent of NVIDIA. Removing the driver needs a reboot; no GPU firmware flash.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-5688"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2019-11-18"},{"id":"CVE-2020-11488","cve":"CVE-2020-11488","aliases":[],"title":"NVIDIA DGX BMC (AMI firmware): The BMC does not validate the RSA-1024 public key used to verify firmware signatures","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVIDIA DGX BMC (AMI firmware)","year":"2020","cvss_score":6.7,"severity":"medium","kev":false,"impact":"The BMC does not validate the RSA-1024 public key used to verify firmware signatures. That breaks the root of trust for BMC firmware updates: an attacker who can push an update can install their own firmware image and persist below the host OS indefinitely. DGX-1 before 3.38.30, DGX-2 before 1.06.06.","attack_vector":"An attacker with access to the BMC's firmware update path - network reach to the management interface, or an administrator session obtained through the other bugs in this bulletin.","remediation":"Flash the DGX BMC firmware from NVIDIA's DGX firmware update container (DGX-1 to 3.38.30 or later, DGX-2 to 1.06.06 or later; DGX A100 per the bulletin's table). A BMC flash does not require the host OS to reboot but drops out-of-band management for several minutes and NVIDIA recommends a host power cycle afterwards, so treat it as a per-node maintenance window. Rotate every BMC and IPMI credential after the flash - flashing does not invalidate secrets an attacker already pulled. Keep BMCs on an isolated management VLAN with no route from tenant or job networks.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-11488"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2020-10-29"},{"id":"CVE-2020-27779","cve":"CVE-2020-27779","aliases":[],"title":"GRUB2 (cutmem command): The cutmem command was not gated by Secure Boot lockdown, so a privileged user could carve","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (cutmem command)","year":"2020","cvss_score":6.7,"severity":"medium","kev":false,"impact":"The cutmem command was not gated by Secure Boot lockdown, so a privileged user could carve memory regions out of the map GRUB hands the kernel. Used to remove the regions that hold verification state, which downgrades a verified boot to an unverified one without tripping anything.","attack_vector":"Local privileged user at the GRUB shell.","remediation":"grub2 package update + reboot. This is the lockdown-coverage class of bug: the fix is that the command is now refused when Secure Boot is on, so there is no config workaround short of a GRUB password.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-27779","https://access.redhat.com/security/cve/CVE-2020-27779"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-03-03"},{"id":"CVE-2020-8692","cve":"CVE-2020-8692","aliases":[],"title":"Intel Ethernet 700 Series Controller firmware (access control): Insufficient access control inside 700-series NIC","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Ethernet 700 Series Controller firmware (access control)","year":"2020","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Insufficient access control inside 700-series NIC firmware lets a privileged host user escalate further or deny service. On bare-metal GPU rental the 'privileged host user' is the tenant, and the thing they are escalating into is firmware that the next tenant will inherit. This is the concrete mechanism behind the tenant-handoff problem: patching the host OS between customers does nothing about the NIC.","attack_vector":"A privileged local user on the host — in a bare-metal rental model, the customer with root.","remediation":"Flash 700-series firmware to 7.3 or later; cold power cycle. Operationally the stronger control is to reflash NIC firmware from a known-good image at every tenant handoff and verify the resulting version, rather than trusting whatever the previous tenant left behind. Related issues fixed in the same family: CVE-2020-8691, CVE-2020-8693, CVE-2019-0139, CVE-2019-0144.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-8692","https://nvd.nist.gov/vuln/detail/CVE-2020-8693","https://nvd.nist.gov/vuln/detail/CVE-2019-0139"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2020-11-12"},{"id":"CVE-2021-0157","cve":"CVE-2021-0157","aliases":[],"title":"Intel BIOS firmware: Insufficient control-flow management in Intel BIOS firmware lets a privileged user escalate","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel BIOS firmware","year":"2021","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Insufficient control-flow management in Intel BIOS firmware lets a privileged user escalate. Broad advisory covering many platform SKUs - check your specific board rather than assuming you are out of scope.","attack_vector":"Privileged local access on the host.","remediation":"Fixed in platform BIOS/UEFI firmware. That means an OEM release, a per-node drain, a flash and a cold reboot - and OEM availability commonly lags the Intel advisory by quarters on server boards. There is no microcode or OS-level shortcut for this class; budget it as a fleet-wide maintenance campaign, not a patch.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-0157","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00562.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-11-17"},{"id":"CVE-2021-0186","cve":"CVE-2021-0186","aliases":["SmashEx"],"title":"Intel SGX SDK (asynchronous exit / exception handling): SmashEx: an asynchronous exception delivered at the right","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SGX SDK (asynchronous exit / exception handling)","year":"2021","cvss_score":6.7,"severity":"medium","kev":false,"impact":"SmashEx: an asynchronous exception delivered at the right moment during an enclave entry/exit leaves the enclave's internal state inconsistent, and the SDK's own exception handling can then be steered into an in-enclave control-flow hijack. Because the host controls interrupt delivery, the attacker in this model is the platform - so it is a direct break of the confidential-compute promise, and it recovers enclave secrets in the published attack.","attack_vector":"Privileged host code that can inject exceptions/interrupts into a running enclave - i.e. the hypervisor or host OS on a confidential-compute node.","remediation":"Rebuild enclaves against a fixed SGX SDK and re-attest. This is a software fix in the SDK's AEX handling, so it needs a new enclave binary from the enclave author; no microcode, BIOS or reboot on the operator side, but also nothing the operator can do alone.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-0186","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00548.html"],"status":"curated","tags":["tenant-isolation"],"published":"2021-11-17"},{"id":"CVE-2021-20225","cve":"CVE-2021-20225","aliases":[],"title":"GRUB2 (short-form option parser): Heap out-of-bounds write in the short-form option parser","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (short-form option parser)","year":"2021","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Heap out-of-bounds write in the short-form option parser. Another arbitrary-write primitive inside the signed bootloader, usable to load unsigned code with Secure Boot enabled.","attack_vector":"Local, via GRUB command line or grub.cfg.","remediation":"grub2 package update + reboot per node, followed by the dbx pass if you are actually revoking old binaries.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-20225","https://access.redhat.com/security/cve/CVE-2021-20225"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-03-03"},{"id":"CVE-2021-20233","cve":"CVE-2021-20233","aliases":[],"title":"GRUB2 (option quoting): Miscalculated buffer size when quoting options produces a heap out-of-bounds write","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (option quoting)","year":"2021","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Miscalculated buffer size when quoting options produces a heap out-of-bounds write. Same outcome: pre-boot code execution and a Secure Boot bypass on a node the attacker already touched once.","attack_vector":"Local, via GRUB command line or a modified grub.cfg.","remediation":"grub2 package update + reboot per node.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-20233","https://access.redhat.com/security/cve/CVE-2021-20233"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-03-03"},{"id":"CVE-2021-29213","cve":"CVE-2021-29213","aliases":["HPESBHF04197"],"title":"HPE ProLiant Gen10 System ROM (security restriction bypass): A local bypass of security restrictions in the System ROM","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HPE ProLiant Gen10 System ROM (security restriction bypass)","year":"2021","cvss_score":6.7,"severity":"medium","kev":false,"impact":"A local bypass of security restrictions in the System ROM (BIOS) of ProLiant DL20 Gen10, ML30 Gen10 and MicroServer Gen10 Plus. Firmware-level restriction bypasses matter more than their score suggests, because the restrictions being bypassed are the ones enforcing secure configuration at boot - and anything an attacker changes there persists across OS reinstalls. These are edge and utility SKUs rather than GPU nodes, but they are the boxes that end up as bastion hosts, PXE servers and DHCP/DNS for a GPU cluster, which makes them a useful staging point.","attack_vector":"Local to the host with elevated privileges - an administrator-level account on the operating system, or physical access during maintenance. Not reachable from the management VLAN.","remediation":"Update the System ROM to v2.52 or later. Unlike an iLO flash, a System ROM update only takes effect on the next host reboot, so it needs a maintenance window per node - though for this SKU set that is far cheaper than draining a GPU node. Stage the ROM through iLO or Service Pack for ProLiant and let it apply at the next scheduled restart. No config-only mitigation.","references":["https://support.hpe.com/hpsc/doc/public/display?docLocale=en_US&docId=emr_na-hpesbhf04197en_us","https://nvd.nist.gov/vuln/detail/CVE-2021-29213"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-11-01"},{"id":"CVE-2022-23222","cve":"CVE-2022-23222","aliases":[],"title":"Linux kernel (eBPF verifier): kernel/bpf/verifier.c mishandles pointer types - unprivileged BPF to local root","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (eBPF verifier)","year":"2022","cvss_score":6.7,"severity":"medium","kev":false,"impact":"kernel/bpf/verifier.c mishandles pointer types - unprivileged BPF to local root","attack_vector":"Any tenant process in a container where unprivileged BPF is enabled","remediation":"Livepatchable; otherwise drain + reboot. `kernel.unprivileged_bpf_disabled=1`","references":["https://access.redhat.com/security/cve/CVE-2022-23222"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2022-01-14"},{"id":"CVE-2022-2586","cve":"CVE-2022-2586","aliases":[],"title":"Linux kernel (nf_tables): Cross-table use-after-free in nf_tables leading to local privilege escalation","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (nf_tables)","year":"2022","cvss_score":6.7,"severity":"medium","kev":true,"impact":"Cross-table use-after-free in nf_tables leading to local privilege escalation [KEV]","attack_vector":"Any tenant process in a container with CAP_NET_ADMIN in a userns","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2022-2586"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-01-08"},{"id":"CVE-2022-25905","cve":"CVE-2022-25905","aliases":[],"title":"Intel oneAPI Data Analytics Library (oneDAL): An uncontrolled library search path: the component loads a shared library","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel oneAPI Data Analytics Library (oneDAL)","year":"2022","cvss_score":6.7,"severity":"medium","kev":false,"impact":"An uncontrolled library search path: the component loads a shared library by name from a directory a non-root user can write. Anyone who can drop a file in that directory gets code execution in the context of whoever next runs the tool - which on an AI node is usually a privileged installer, a service account, or root.","attack_vector":"A local authenticated user on a node that has the toolkit installed. On shared build/dev nodes and on container images built from the Intel toolkits, that is a broad set of people.","remediation":"Upgrade the affected component and, just as importantly, audit directory permissions on already-provisioned nodes and container images - upgrading the package does not remove a writable directory an earlier install created. Userspace only: no reboot, no BIOS, no microcode. Rebuild base images rather than patching running nodes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-25905","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00674.html"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2023-02-16"},{"id":"CVE-2022-26052","cve":"CVE-2022-26052","aliases":[],"title":"Intel MPI Library (oneAPI HPC Toolkit): An uncontrolled library search path: the component loads a shared library","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel MPI Library (oneAPI HPC Toolkit)","year":"2022","cvss_score":6.7,"severity":"medium","kev":false,"impact":"An uncontrolled library search path: the component loads a shared library by name from a directory a non-root user can write. Anyone who can drop a file in that directory gets code execution in the context of whoever next runs the tool - which on an AI node is usually a privileged installer, a service account, or root.","attack_vector":"A local authenticated user on a node that has the toolkit installed. On shared build/dev nodes and on container images built from the Intel toolkits, that is a broad set of people.","remediation":"Upgrade the affected component and, just as importantly, audit directory permissions on already-provisioned nodes and container images - upgrading the package does not remove a writable directory an earlier install created. Userspace only: no reboot, no BIOS, no microcode. Rebuild base images rather than patching running nodes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-26052","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00674.html"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2023-02-16"},{"id":"CVE-2022-26076","cve":"CVE-2022-26076","aliases":[],"title":"Intel oneAPI Deep Neural Network Library (oneDNN): An uncontrolled library search path: the component loads a shared","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel oneAPI Deep Neural Network Library (oneDNN)","year":"2022","cvss_score":6.7,"severity":"medium","kev":false,"impact":"An uncontrolled library search path: the component loads a shared library by name from a directory a non-root user can write. Anyone who can drop a file in that directory gets code execution in the context of whoever next runs the tool - which on an AI node is usually a privileged installer, a service account, or root.","attack_vector":"A local authenticated user on a node that has the toolkit installed. On shared build/dev nodes and on container images built from the Intel toolkits, that is a broad set of people.","remediation":"Upgrade the affected component and, just as importantly, audit directory permissions on already-provisioned nodes and container images - upgrading the package does not remove a writable directory an earlier install created. Userspace only: no reboot, no BIOS, no microcode. Rebuild base images rather than patching running nodes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-26076","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00674.html"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2023-02-16"},{"id":"CVE-2022-26345","cve":"CVE-2022-26345","aliases":[],"title":"Intel oneAPI OpenMP runtime: An uncontrolled library search path: the component loads a shared library by name","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel oneAPI OpenMP runtime","year":"2022","cvss_score":6.7,"severity":"medium","kev":false,"impact":"An uncontrolled library search path: the component loads a shared library by name from a directory a non-root user can write. Anyone who can drop a file in that directory gets code execution in the context of whoever next runs the tool - which on an AI node is usually a privileged installer, a service account, or root.","attack_vector":"A local authenticated user on a node that has the toolkit installed. On shared build/dev nodes and on container images built from the Intel toolkits, that is a broad set of people.","remediation":"Upgrade the affected component and, just as importantly, audit directory permissions on already-provisioned nodes and container images - upgrading the package does not remove a writable directory an earlier install created. Userspace only: no reboot, no BIOS, no microcode. Rebuild base images rather than patching running nodes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-26345","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00674.html"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2023-02-16"},{"id":"CVE-2022-26363","cve":"CVE-2022-26363","aliases":["XSA-402"],"title":"Xen (x86 PV): Insufficient care with non-coherent mappings - PV guest to host compromise","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen (x86 PV)","year":"2022","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Insufficient care with non-coherent mappings - PV guest to host compromise","attack_vector":"Tenant VM guest (PV)","remediation":"Hypervisor patch; livepatchable via Xen livepatch, otherwise host reboot + guest evacuation","references":["https://xenbits.xen.org/xsa/advisory-402.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2022-06-09"},{"id":"CVE-2022-26421","cve":"CVE-2022-26421","aliases":[],"title":"Intel oneAPI DPC++/C++ compiler runtime: An uncontrolled library search path: the component loads a shared library","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel oneAPI DPC++/C++ compiler runtime","year":"2022","cvss_score":6.7,"severity":"medium","kev":false,"impact":"An uncontrolled library search path: the component loads a shared library by name from a directory a non-root user can write. Anyone who can drop a file in that directory gets code execution in the context of whoever next runs the tool - which on an AI node is usually a privileged installer, a service account, or root.","attack_vector":"A local authenticated user on a node that has the toolkit installed. On shared build/dev nodes and on container images built from the Intel toolkits, that is a broad set of people.","remediation":"Upgrade the affected component and, just as importantly, audit directory permissions on already-provisioned nodes and container images - upgrading the package does not remove a writable directory an earlier install created. Userspace only: no reboot, no BIOS, no microcode. Rebuild base images rather than patching running nodes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-26421","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00674.html"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2023-02-16"},{"id":"CVE-2022-26425","cve":"CVE-2022-26425","aliases":[],"title":"Intel oneAPI Collective Communications Library (oneCCL): An uncontrolled library search path: the component loads","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel oneAPI Collective Communications Library (oneCCL)","year":"2022","cvss_score":6.7,"severity":"medium","kev":false,"impact":"An uncontrolled library search path: the component loads a shared library by name from a directory a non-root user can write. Anyone who can drop a file in that directory gets code execution in the context of whoever next runs the tool - which on an AI node is usually a privileged installer, a service account, or root.","attack_vector":"A local authenticated user on a node that has the toolkit installed. On shared build/dev nodes and on container images built from the Intel toolkits, that is a broad set of people.","remediation":"Upgrade the affected component and, just as importantly, audit directory permissions on already-provisioned nodes and container images - upgrading the package does not remove a writable directory an earlier install created. Userspace only: no reboot, no BIOS, no microcode. Rebuild base images rather than patching running nodes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-26425","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00674.html"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2023-02-16"},{"id":"CVE-2022-28736","cve":"CVE-2022-28736","aliases":[],"title":"GRUB2 (chainloader): Use-after-free in grub_cmd_chainloader when a chainloaded image fails to start","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (chainloader)","year":"2022","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Use-after-free in grub_cmd_chainloader when a chainloaded image fails to start. Gives pre-boot code execution and, in combination with the other 2022 bugs, a full Secure Boot bypass chain.","attack_vector":"Local, via GRUB command line or grub.cfg on a node the attacker has touched.","remediation":"grub2 package update + reboot per node.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28736","https://access.redhat.com/security/cve/CVE-2022-28736"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2023-07-20"},{"id":"CVE-2022-31601","cve":"CVE-2022-31601","aliases":[],"title":"NVIDIA DGX A100 - SBIOS / SMM firmware: An out-of-bounds write in the SmbiosPei module gives a highly privileged local","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX A100 - SBIOS / SMM firmware","year":"2022","cvss_score":6.7,"severity":"medium","kev":false,"impact":"An out-of-bounds write in the SmbiosPei module gives a highly privileged local attacker firmware-phase code execution. This is firmware-level persistence: it survives OS reinstall, image re-flash and tenant handoff, and it is invisible to anything running above it. On a bare-metal GPU rental business it is the difference between wiping a node between tenants and not actually being able to.","attack_vector":"Local and already privileged - host root, or code that has reached the platform firmware/SMM path. It is not a first foothold; it is what turns a one-time root compromise into something you cannot remediate by reimaging.","remediation":"Flash the fixed SBIOS from bulletin 5367. Cost: not live-patchable. Full node drain, host power cycle, and on DGX the SBIOS ships inside a firmware bundle alongside BMC and CPLD components, so budget 30-60 minutes of node downtime plus a post-flash health check. Firmware rollback protection means you cannot cleanly revert - stage on one node before the fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31601","https://github.com/NVIDIA/product-security/tree/main/2022/5367"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"firmware-flash"},"published":"2022-07-04"},{"id":"CVE-2022-34301","cve":"CVE-2022-34301","aliases":["Three more bootloaders"],"title":"CryptoPro Secure Disk (signed UEFI bootloader): A Microsoft-signed bootloader that can be made to execute arbitrary","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"CryptoPro Secure Disk (signed UEFI bootloader)","year":"2022","cvss_score":6.7,"severity":"medium","kev":false,"impact":"A Microsoft-signed bootloader that can be made to execute arbitrary pre-boot code. Same portable-bypass shape as the Howyar case: the attacker brings the signed binary with them, so a fleet that never deployed this product is still exploitable simply because the firmware trusts the signature.","attack_vector":"Write access to the EFI System Partition on the target node.","remediation":"dbx revocation update pushed to every node via firmware update or OS vendor channel - not a package upgrade. Verify the dbx entry is present afterwards; revocation is the only fix because the vulnerable binary is signed and portable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34301","https://kb.cert.org/vuls/id/309662"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-08-26"},{"id":"CVE-2022-34302","cve":"CVE-2022-34302","aliases":["Three more bootloaders"],"title":"New Horizon Datasys (signed UEFI bootloader): Signed bootloader with a built-in mechanism to bypass Secure Boot","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"New Horizon Datasys (signed UEFI bootloader)","year":"2022","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Signed bootloader with a built-in mechanism to bypass Secure Boot enforcement, described by researchers as more subtle than its siblings because it disables verification without obviously tampering with the chain. Yields a bootkit that measured boot will not flag.","attack_vector":"EFI System Partition write access on the target node.","remediation":"dbx revocation, delivered via firmware or OS vendor update. No package fix exists because the vulnerable artifact is a signed third-party binary.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34302","https://kb.cert.org/vuls/id/309662"],"status":"curated","published":"2022-08-26"},{"id":"CVE-2022-34303","cve":"CVE-2022-34303","aliases":["Three more bootloaders"],"title":"Eurosoft (UK) Ltd (signed UEFI bootloader): Signed UEFI bootloader containing a shell that executes arbitrary code","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Eurosoft (UK) Ltd (signed UEFI bootloader)","year":"2022","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Signed UEFI bootloader containing a shell that executes arbitrary code, bypassing Secure Boot on any machine trusting the Microsoft third-party CA.","attack_vector":"EFI System Partition write access.","remediation":"dbx revocation update. Track revocation-list version per node as a fleet health metric - this class recurs roughly annually and package inventories will never show it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34303","https://kb.cert.org/vuls/id/309662"],"status":"curated","published":"2022-08-26"},{"id":"CVE-2022-42281","cve":"CVE-2022-42281","aliases":[],"title":"NVIDIA DGX A100 - SBIOS / SMM firmware: An out-of-bounds write in the FsRecovery module reaches firmware code execution","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX A100 - SBIOS / SMM firmware","year":"2022","cvss_score":6.7,"severity":"medium","kev":false,"impact":"An out-of-bounds write in the FsRecovery module reaches firmware code execution from a highly privileged local account. This is firmware-level persistence: it survives OS reinstall, image re-flash and tenant handoff, and it is invisible to anything running above it. On a bare-metal GPU rental business it is the difference between wiping a node between tenants and not actually being able to.","attack_vector":"Local and already privileged - host root, or code that has reached the platform firmware/SMM path. It is not a first foothold; it is what turns a one-time root compromise into something you cannot remediate by reimaging.","remediation":"Flash the fixed SBIOS from bulletin 5435. Cost: not live-patchable. Full node drain, host power cycle, and on DGX the SBIOS ships inside a firmware bundle alongside BMC and CPLD components, so budget 30-60 minutes of node downtime plus a post-flash health check. Firmware rollback protection means you cannot cleanly revert - stage on one node before the fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42281","https://github.com/NVIDIA/product-security/tree/main/2022/5435"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"firmware-flash"},"published":"2023-01-13"},{"id":"CVE-2023-0185","cve":"CVE-2023-0185","aliases":[],"title":"GPU Display Driver: DoS / data tampering (integer underflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2023","cvss_score":6.7,"severity":"medium","kev":false,"impact":"DoS / data tampering (integer underflow)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; rolling node reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0185","https://github.com/NVIDIA/product-security/tree/main/2023/5452"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:C/C:L/I:L/A:H","cwe":["CWE-196"],"fleet":{"pain_class":"node-reboot"},"published":"2023-04-01"},{"id":"CVE-2023-0201","cve":"CVE-2023-0201","aliases":[],"title":"NVIDIA DGX-2 - SBIOS / SMM firmware: An out-of-bounds write in the Bds phase gives a privileged local user firmware","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX-2 - SBIOS / SMM firmware","year":"2023","cvss_score":6.7,"severity":"medium","kev":false,"impact":"An out-of-bounds write in the Bds phase gives a privileged local user firmware code execution. This is firmware-level persistence: it survives OS reinstall, image re-flash and tenant handoff, and it is invisible to anything running above it. On a bare-metal GPU rental business it is the difference between wiping a node between tenants and not actually being able to.","attack_vector":"Local and already privileged - host root, or code that has reached the platform firmware/SMM path. It is not a first foothold; it is what turns a one-time root compromise into something you cannot remediate by reimaging.","remediation":"Flash the fixed SBIOS from bulletin 5449. Cost: not live-patchable. Full node drain, host power cycle, and on DGX the SBIOS ships inside a firmware bundle alongside BMC and CPLD components, so budget 30-60 minutes of node downtime plus a post-flash health check. Firmware rollback protection means you cannot cleanly revert - stage on one node before the fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0201","https://github.com/NVIDIA/product-security/tree/main/2023/5449"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-118"],"fleet":{"pain_class":"firmware-flash"},"published":"2023-04-22"},{"id":"CVE-2023-20567","cve":"CVE-2023-20567","aliases":[],"title":"AMD Radeon RX Vega M graphics driver installer - signature verification: The driver package launches","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD Radeon RX Vega M graphics driver installer - signature verification","year":"2023","cvss_score":6.7,"severity":"medium","kev":false,"impact":"The driver package launches AMDSoftwareInstaller.exe without validating its signature, so an attacker with admin privileges can substitute the binary and get their code run by a trusted installer flow. It is a signed-update-chain failure rather than a memory-safety bug: the mechanism you use to keep drivers current is the mechanism that runs the attacker's payload.","attack_vector":"Local, requires admin privilege to place the substituted binary. Windows driver packaging.","remediation":"Update the AMD driver package. Relevant only where you deploy AMD's Windows driver installer; Linux ROCm deployments are unaffected.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-20567","https://www.amd.com/en/resources/product-security.html"],"status":"curated","published":"2023-11-14"},{"id":"CVE-2023-24932","cve":"CVE-2023-24932","aliases":["BlackLotus"],"title":"Windows Boot Manager (Secure Boot bypass): The bypass the BlackLotus UEFI bootkit used in the wild","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Windows Boot Manager (Secure Boot bypass)","year":"2023","cvss_score":6.7,"severity":"medium","kev":false,"impact":"The bypass the BlackLotus UEFI bootkit used in the wild. An attacker with admin or physical access installs a bootkit that survives OS reinstall and disk replacement, disables Secure Boot enforcement from inside the boot chain, and hides from every in-OS security agent. On a mixed Windows/Linux estate this also poisons attestation for anything downstream.","attack_vector":"Local administrator or physical access - which on bare metal means any tenant that rented the node, and on a colo floor means anyone with remote-hands.","remediation":"The most operationally painful entry in this cluster. The security update alone does nothing: Microsoft shipped it behind a manual, multi-stage opt-in requiring boot manager updates, revocation of the old boot manager, and a UEFI CA/dbx update, staged over roughly two years precisely because enabling revocation early bricks machines that still boot old media. Budget for a phased rollout with per-node verification, keep known-good recovery media that is still trusted, and expect to touch firmware settings on some boards. Not a patch-and-forget.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-24932","https://msrc.microsoft.com/update-guide/vulnerability/CVE-2023-24932"],"status":"curated","published":"2023-05-09"},{"id":"CVE-2023-25508","cve":"CVE-2023-25508","aliases":[],"title":"NVIDIA DGX BMC (AMI-derived management controller): The DGX-1 BMC's IPMI handler allows an authorised attacker","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVIDIA DGX BMC (AMI-derived management controller)","year":"2023","cvss_score":6.7,"severity":"medium","kev":false,"impact":"The DGX-1 BMC's IPMI handler allows an authorised attacker to upload and download arbitrary files, which is enough to plant persistence on the controller or exfiltrate its configuration and credentials. A BMC compromise on a DGX gives an attacker power control, virtual media, serial console and a persistent foothold under the host OS on a node holding eight GPUs.","attack_vector":"Network access to the BMC management interface holding credentials at some authorised level. Whether that is 'remote' depends entirely on how genuinely isolated your OOB network is - in practice most fleets have a jump host, a DCIM integration or a monitoring collector that bridges it.","remediation":"Update the DGX BMC firmware bundle per bulletin 5458. Cost: BMC firmware usually updates without draining the GPUs, but the BMC resets and out-of-band access drops for a few minutes. Pair the patch with an actual audit of who can route to the BMC subnet - that control is worth more than the patch.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25508","https://github.com/NVIDIA/product-security/tree/main/2023/5458"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-22"],"fleet":{"pain_class":"firmware-flash"},"published":"2023-04-22"},{"id":"CVE-2023-3264","cve":"CVE-2023-3264","aliases":[],"title":"CyberPower PowerPanel Enterprise DCIM: Hard-coded credentials in the DCIM platform","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"CyberPower PowerPanel Enterprise DCIM","year":"2023","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Hard-coded credentials in the DCIM platform; chained with the other PowerPanel flaws for full DCIM takeover and, from there, control of every managed power device in the facility","attack_vector":"Network","remediation":"Upgrade PowerPanel Enterprise to 2.6.9; DCIM is often facility-operator-owned rather than tenant-owned, so a colocated neocloud may not control the remediation at all","references":["https://thehackernews.com/2023/08/multiple-flaws-in-cyberpower-and.html"],"status":"curated","published":"2023-08-14"},{"id":"CVE-2023-48733","cve":"CVE-2023-48733","aliases":[],"title":"EDK II / OVMF (UEFI Shell left enabled in downstream Ubuntu and LXD firmware builds): Not a memory-safety bug","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"EDK II / OVMF (UEFI Shell left enabled in downstream Ubuntu and LXD firmware builds)","year":"2023","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Not a memory-safety bug - a packaging default. The interactive UEFI Shell is compiled into the shipped firmware image, and from it an attacker who already has the OS can load unsigned code and walk straight past Secure Boot. For anyone running GPU workloads inside VMs on OVMF, this quietly voids the guest's Secure Boot guarantee and any attestation chain rooted in it, which is precisely the guarantee confidential-compute and tenant-isolation stories are sold on.","attack_vector":"An attacker who already has administrative control of the guest OS (or the VM's boot configuration) and can reach the UEFI Shell on the next boot. Local to the guest; no firmware flash needed by the attacker.","remediation":"Distribution package update, not an OEM BIOS flash - this is a build-flag fix in the edk2/OVMF firmware images shipped by the distro (Ubuntu, LXD). Update the ovmf/edk2 packages on your hypervisor hosts and restart the affected guests to pick up the new firmware blob; no host reboot required. Config workaround: remove the UEFI Shell from the boot order and enforce a firmware password / locked boot order in the guest's variable store, though a determined admin-level attacker inside the guest can often undo that.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-48733","https://ubuntu.com/security/CVE-2023-48733"],"status":"curated","published":"2024-02-14"},{"id":"CVE-2023-49144","cve":"CVE-2023-49144","aliases":["INTEL-SA-01078"],"title":"Intel Server OpenBMC firmware (before egs-1.15-0 / bhs-0.27): An out-of-bounds read reachable by a privileged BMC user","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Server OpenBMC firmware (before egs-1.15-0 / bhs-0.27)","year":"2023","cvss_score":6.7,"severity":"medium","kev":false,"impact":"An out-of-bounds read reachable by a privileged BMC user leaks memory contents across a scope boundary - CVSS v4 scores it 8.1, notably higher than the v3 6.7, because the disclosure crosses out of the vulnerable component. What comes back is BMC process memory, which on this stack means session tokens, credential material and configuration. Combined with the privilege-escalation entry from the same product line, an attacker with a modest BMC account has a path to reading things that let them keep the access permanently.","attack_vector":"Local access on the BMC with a privileged account. Requires an existing high-privilege BMC credential, so this is a post-compromise deepening tool rather than an entry point.","remediation":"Fixed in Intel Server OpenBMC egs-1.15-0 / bhs-0.27 and later - per-node out-of-band BMC firmware update via Intel platform packages, subject to OEM rebase lag on boards that derive from the same base. No config-only mitigation for the bug itself; limit the blast radius by minimizing the number of accounts holding BMC admin and by rotating BMC credentials after any suspected node compromise.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-49144","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01078.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-08-14"},{"cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-53446","cve":"CVE-2023-53446","aliases":[],"title":"Linux kernel (drivers/pci/pcie): The ASPM link state keeps a raw pointer to function 0 of a multi-function device.","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/pci/pcie)","year":"2023","cvss_score":6.7,"severity":"medium","kev":false,"impact":"The ASPM link state keeps a raw pointer to function 0 of a multi-function device. Remove that function and the pointer dangles; any later ASPM policy change dereferences freed memory, which KASAN catches as a slab use-after-free in the link-state walk. This matters on GPU nodes because the cards themselves are multi-function - a GPU with its companion audio function is the standard layout - so the removal that arms the bug is an ordinary device-management action.","attack_vector":"Needs write access to PCI sysfs on the host: remove a function via /sys/bus/pci/devices/<dev>/remove, then write the ASPM policy knob (or let anything else walk the link state). That is root, or a privileged container with /sys mounted writable - not a plain tenant container and not a guest. Worth tracking anyway because driver-rebind and device-reclaim automation performs exactly this remove step on multi-function cards, so the dangling pointer can be left armed by routine operations and tripped later by an unrelated power-management change.","remediation":"Boot a kernel that disables ASPM and frees the pcie_link_state when any child function is removed. Interim: do not expose writable PCI sysfs (remove, and the pcie_aspm policy module parameter) inside containers, and avoid per-function removal on multi-function cards - remove the whole device instead.","references":["https://git.kernel.org/stable/c/666e7f9d60cee23077ea3e6331f6f8a19f7ea03f","https://git.kernel.org/stable/c/7badf4d6f49a358a01ab072bbff88d3ee886c33b","https://nvd.nist.gov/vuln/detail/CVE-2023-53446"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-21766","cve":"CVE-2024-21766","aliases":[],"title":"Intel oneAPI Math Kernel Library (oneMKL): An uncontrolled library search path: the component loads a shared library","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel oneAPI Math Kernel Library (oneMKL)","year":"2024","cvss_score":6.7,"severity":"medium","kev":false,"impact":"An uncontrolled library search path: the component loads a shared library by name from a directory a non-root user can write. Anyone who can drop a file in that directory gets code execution in the context of whoever next runs the tool - which on an AI node is usually a privileged installer, a service account, or root.","attack_vector":"A local authenticated user on a node that has the toolkit installed. On shared build/dev nodes and on container images built from the Intel toolkits, that is a broad set of people.","remediation":"Upgrade the affected component and, just as importantly, audit directory permissions on already-provisioned nodes and container images - upgrading the package does not remove a writable directory an earlier install created. Userspace only: no reboot, no BIOS, no microcode. Rebuild base images rather than patching running nodes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21766","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01072.html"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2024-08-14"},{"id":"CVE-2024-21857","cve":"CVE-2024-21857","aliases":[],"title":"Intel oneAPI compiler: An uncontrolled library search path: the component loads a shared library by name","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel oneAPI compiler","year":"2024","cvss_score":6.7,"severity":"medium","kev":false,"impact":"An uncontrolled library search path: the component loads a shared library by name from a directory a non-root user can write. Anyone who can drop a file in that directory gets code execution in the context of whoever next runs the tool - which on an AI node is usually a privileged installer, a service account, or root.","attack_vector":"A local authenticated user on a node that has the toolkit installed. On shared build/dev nodes and on container images built from the Intel toolkits, that is a broad set of people.","remediation":"Upgrade the affected component and, just as importantly, audit directory permissions on already-provisioned nodes and container images - upgrading the package does not remove a writable directory an earlier install created. Userspace only: no reboot, no BIOS, no microcode. Rebuild base images rather than patching running nodes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21857","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01057.html"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2024-08-14"},{"id":"CVE-2024-31073","cve":"CVE-2024-31073","aliases":[],"title":"Intel oneAPI Level Zero software: An uncontrolled search path in Level Zero lets an authenticated local user get code","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel oneAPI Level Zero software","year":"2024","cvss_score":6.7,"severity":"medium","kev":false,"impact":"An uncontrolled search path in Level Zero lets an authenticated local user get code loaded at higher privilege. Level Zero is the low-level runtime under oneAPI and under Intel GPU compute generally, so it sits in the same position in an Intel accelerator stack that the CUDA driver API occupies in an NVIDIA one.","attack_vector":"Local, authenticated. A user able to place a file on a path the runtime searches - which on a shared build host is a low bar.","remediation":"Update the oneAPI Level Zero package. Cost: package update and job restart; no driver reload or node drain. Audit writable directories on shared build and inference hosts as the standing control.","references":["https://www.intel.com/content/www/us/en/security-center/default.html","https://nvd.nist.gov/vuln/detail/CVE-2024-31073"],"status":"curated","published":"2025-05-13"},{"id":"CVE-2024-42642","cve":"CVE-2024-42642","aliases":[],"title":"Micron Crucial MX500 series SSD, firmware M3CR046","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Micron Crucial MX500 series SSD, firmware M3CR046 - buffer overflow in the drive controller reachable from host ATA…","year":"2024","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Specially crafted ATA packets sent from the HOST to the drive controller overflow a buffer in the firmware, with high confidentiality, integrity and availability impact. This is the shape of bug that matters most for bare-metal multi-tenancy: the attack surface is the ordinary storage command path, reachable from the operating system, not a JTAG pad or a soldering iron. A tenant with root on the machine can reach the controller's memory. TOP TENANT-HANDOFF RISK PATTERN: a departing tenant who achieves code execution on the drive controller can attempt to leave something behind that outlives your reimage entirely, because your reimage rewrites NAND contents and never touches controller firmware. Even short of a persistent implant, controller-level code execution means the drive's own reports about sanitize, lock state and encryption become worthless.","attack_vector":"A tenant with root (high privilege) on the bare-metal host, issuing crafted ATA commands down the normal storage path. No physical access, no chassis entry, no special hardware - this is reachable from a shell on the rented machine.","remediation":"Flash past M3CR046; Micron states the issue was fully remediated in December 2024 and firmware is on Crucial's MX500 support page. Drive offline for the flash, plus Crucial's own tooling. The broader point for an operator: the MX500 is a consumer SSD and has no business being tenant-writable media in a bare-metal fleet - if you find these in GPU nodes (they turn up as cheap boot/scratch drives in budget builds), the correct remediation is to replace them with datacenter SKUs, not just to patch. Where you cannot replace, deny tenants the ability to issue raw ATA/NVMe pass-through: do not hand out unfiltered block devices, and if the tenant workload does not need direct device access, put a virtualization or filesystem layer between them and the controller. Note also that Crucial Storage Executive, the management tool you would use to flash this, has its own separate installer DLL-preloading flaw (CVE-2025-71178) - patch the tool before you trust it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-42642","https://www.crucial.com/support/ssd-support/mx500-support","https://github.com/VL4DR/CVE-2024-42642/tree/main"],"status":"curated","published":"2024-09-04"},{"id":"CVE-2024-45105","cve":"CVE-2024-45105","aliases":["LEN-165524"],"title":"Lenovo ThinkSystem UEFI/BIOS (SMM callout): A System Management Mode callout vulnerability in ThinkSystem UEFI","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lenovo ThinkSystem UEFI/BIOS (SMM callout)","year":"2024","cvss_score":6.7,"severity":"medium","kev":false,"impact":"A System Management Mode callout vulnerability in ThinkSystem UEFI - SMM code calls out to memory it does not control, letting a local attacker with elevated privileges execute code in SMM. SMM sits above the hypervisor and is invisible to it, so this is a persistent-implant primitive: what an attacker installs there is not removed by reimaging, disk replacement or hypervisor reinstall. The affected list runs to roughly 99 platforms and explicitly includes the GPU boxes - SR670 V2 and SR675 V3 - alongside SR630/SR650/SR645/SR665 V3 and the ThinkAgile appliances.","attack_vector":"Local to the host with elevated privileges - root or administrator on the operating system. Not reachable from the management VLAN; the path is a tenant or workload that already holds privileged host access.","remediation":"UEFI/BIOS update on each affected node, per the per-model version table in LEN-165524. This is the expensive kind: the payload can be staged through XCC, but it applies only on the next host reboot, so it needs a drain of running training jobs and a maintenance window per node. For SR670 V2 / SR675 V3 GPU nodes that is real lost capacity - plan it as a rolling campaign against spare capacity, not an emergency sweep. No config-only mitigation exists for an SMM defect.","references":["https://support.lenovo.com/us/en/product_security/LEN-165524","https://nvd.nist.gov/vuln/detail/CVE-2024-45105"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-09-13"},{"id":"CVE-2024-4550","cve":"CVE-2024-4550","aliases":[],"title":"Lenovo ThinkSystem / ThinkStation (firmware buffer overflow): A local attacker with elevated privileges executes","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lenovo ThinkSystem / ThinkStation (firmware buffer overflow)","year":"2024","cvss_score":6.7,"severity":"medium","kev":false,"impact":"A local attacker with elevated privileges executes arbitrary code via a firmware buffer overflow on ThinkSystem servers.","attack_vector":"Local elevated-privilege access to the server.","remediation":"Apply the Lenovo firmware update per LEN-165524. Firmware flash requiring a reboot - batch with other node firmware work.","references":["https://support.lenovo.com/us/en/product_security/LEN-165524"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-45774","cve":"CVE-2024-45774","aliases":[],"title":"GRUB2 (JPEG parser): Out-of-bounds write in GRUB's JPEG parser from a crafted image","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (JPEG parser)","year":"2024","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Out-of-bounds write in GRUB's JPEG parser from a crafted image; possible Secure Boot bypass — part of the 2024-2025 GRUB2 vulnerability wave (73 issues)","attack_vector":"Local, ESP write","remediation":"Coordinated GRUB2 + shim + dbx rollout across every distro image in the fleet. The distro-by-distro fan-out is the real cost: a neocloud offering multiple guest images has to rebuild all of them","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-45774"],"status":"curated","published":"2025-02-18"},{"id":"CVE-2024-45775","cve":"CVE-2024-45775","aliases":["GRUB2 2025 batch"],"title":"GRUB2 (commands/extcmd): A failed allocation goes unchecked, so GRUB proceeds on a NULL pointer and its state","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (commands/extcmd)","year":"2024","cvss_score":6.7,"severity":"medium","kev":false,"impact":"A failed allocation goes unchecked, so GRUB proceeds on a NULL pointer and its state becomes attacker-influenced. Practical outcome is the same as the rest of this batch: a foothold below the OS that reimaging does not remove.","attack_vector":"Local, through GRUB command processing on a node the attacker can already influence.","remediation":"grub2 package update + reboot. This whole 2025 batch lands as one distro update - do it as a single pass rather than 20 separate changes, and remember to rebuild the netboot image.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-45775","https://www.openwall.com/lists/oss-security/2025/02/18/3"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-02-18"},{"id":"CVE-2024-45778","cve":"CVE-2024-45778","aliases":["GRUB2 2025 batch"],"title":"GRUB2 (BFS filesystem parser): Integer overflow in the BeFS parser leads to heap corruption","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (BFS filesystem parser)","year":"2024","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Integer overflow in the BeFS parser leads to heap corruption. Nobody runs BFS in production, which is exactly the point: GRUB compiles in filesystem modules you will never mount, and each one is reachable from an attacker-supplied disk image.","attack_vector":"Attacker-supplied filesystem image on a disk or virtual media the node will parse.","remediation":"grub2 package update + reboot. The structural fix is to build GRUB with only the filesystem modules you actually boot from - it deletes a whole recurring CVE class from your fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-45778","https://www.openwall.com/lists/oss-security/2025/02/18/3"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-03-03"},{"id":"CVE-2024-45780","cve":"CVE-2024-45780","aliases":["GRUB2 2025 batch"],"title":"GRUB2 (tar filesystem parser): Integer overflow in the tarfs module writes out of bounds","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (tar filesystem parser)","year":"2024","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Integer overflow in the tarfs module writes out of bounds. Tar archives are a normal part of initrd and provisioning workflows, so this is more reachable in practice than the exotic filesystem parsers.","attack_vector":"Attacker-supplied tar image parsed by GRUB during boot.","remediation":"grub2 package update + reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-45780","https://www.openwall.com/lists/oss-security/2025/02/18/3"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-03-03"},{"id":"CVE-2024-45781","cve":"CVE-2024-45781","aliases":["GRUB2 2025 batch"],"title":"GRUB2 (UFS filesystem parser): Symlink name length is never validated, giving a heap out-of-bounds write in the UFS","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (UFS filesystem parser)","year":"2024","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Symlink name length is never validated, giving a heap out-of-bounds write in the UFS parser and a route to circumventing Secure Boot.","attack_vector":"Attacker-supplied UFS image on an attached or virtual disk.","remediation":"grub2 package update + reboot. Red Hat noted no viable mitigation short of the update, so this is a patch-or-accept decision.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-45781","https://access.redhat.com/security/cve/CVE-2024-45781"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-02-18"},{"id":"CVE-2024-47795","cve":"CVE-2024-47795","aliases":[],"title":"Intel oneAPI DPC++/C++ compiler: An uncontrolled library search path: the component loads a shared library by name","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel oneAPI DPC++/C++ compiler","year":"2024","cvss_score":6.7,"severity":"medium","kev":false,"impact":"An uncontrolled library search path: the component loads a shared library by name from a directory a non-root user can write. Anyone who can drop a file in that directory gets code execution in the context of whoever next runs the tool - which on an AI node is usually a privileged installer, a service account, or root.","attack_vector":"A local authenticated user on a node that has the toolkit installed. On shared build/dev nodes and on container images built from the Intel toolkits, that is a broad set of people.","remediation":"Upgrade the affected component and, just as importantly, audit directory permissions on already-provisioned nodes and container images - upgrading the package does not remove a writable directory an earlier install created. Userspace only: no reboot, no BIOS, no microcode. Rebuild base images rather than patching running nodes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-47795","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01243.html"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2025-05-13"},{"id":"CVE-2024-47976","cve":"CVE-2024-47976","aliases":["Solidigm SA-000563","improper access removal handling"],"title":"Solidigm DC SSDs with TCG Opal (DC P4510/P4511/P4610 Opal, D5-P4320/P4326 Opal, D5-P5316 Opal, D7-P5510/P5520/P5620","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Solidigm DC SSDs with TCG Opal (DC P4510/P4511/P4610 Opal, D5-P4320/P4326 Opal, D5-P5316 Opal, D7-P5510/P5520/P5620…","year":"2024","cvss_score":6.7,"severity":"medium","kev":false,"impact":"The firmware mishandles the REMOVAL of access - that is, revocation does not fully take effect. An authorization that should have ended when you deprovisioned the drive can still be usable. BREAKS TENANT HANDOFF precisely at the moment of handoff: the step where you revoke the previous tenant's access to a locking range is the step that fails, so credentials or authority granted to the departing customer keep working against the drive. Distinct from the access-control-validation bug on the same SKUs, and worth tracking separately because the failure is in de-provisioning rather than in the initial check - it is invisible to any test that only verifies that locking works when you set it up.","attack_vector":"An attacker with physical access to the drive plus some low-level privilege - realistically a departing tenant who retained drive credentials from their tenancy, or anyone who later obtains the physical media through RMA or decommission.","remediation":"Firmware flash, per SKU, drive offline via Solidigm Storage Tool - same firmware trains as the sibling Opal issue (VEV10294/VDV10194/VCV10394 for P4510/P4511/P4610 Opal, 3DV10132 for D5-P4320 Opal, 8DV10564 for D5-P4326 Opal, ACV10340 for D5-P5316 Opal, JCV10404 for D7-P5510, 9CV10410 for D7-P5520/P5620 Opal). Operationally, add a verification step your reclaim pipeline probably lacks: after revoking a locking range, actively re-test that the old credential is rejected, rather than assuming revocation succeeded because the command returned success. And do not let Opal credentials be the only thing standing between two customers - a per-tenant LUKS key you destroy at reclaim is revocation you can actually prove.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-47976","https://www.solidigm.com/support-page/support-security.html","https://www.solidigm.com/content/dam/solidigm/en/site/support/support-community/cve-(security)/documents/public-security-advisory-v2.pdf"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-10-07"},{"id":"CVE-2024-53681","cve":"CVE-2024-53681","aliases":["nvmet subsysnqn overflow","NVMe-oF discovery NQN buffer handling"],"title":"Linux kernel - NVMe-oF target configfs, drivers/nvme/target/configfs.c: Nvmet_root_discovery_nqn_store() treated the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel - NVMe-oF target configfs, drivers/nvme/target/configfs.c","year":"2024","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Nvmet_root_discovery_nqn_store() treated the subsystem NQN string as a fixed-size buffer even though it is dynamically allocated to the length of the existing string, so writing a longer discovery NQN overflows the allocation. The direct exposure is local to whoever administers the target's configfs, but in an operator context that includes storage orchestration and CSI drivers running with elevated privileges - so a compromised control-plane component turns a configuration write into kernel memory corruption on the node holding every tenant's namespaces. It also matters because NQN handling is exactly the surface an attacker probes when attempting NQN spoofing.","attack_vector":"Write an over-long NQN string to the discovery subsystem's configfs attribute on the target. Requires privileged access to the target host's configfs, which orchestration and storage-management agents routinely have. Not remotely reachable on its own, but a natural second stage after compromising a storage control-plane component.","remediation":"Host reboot / kernel upgrade on nvmet target nodes; fold into the same window as the other nvmet findings rather than scheduling separately. Config-side hardening that pays off regardless: keep the target's configfs off any path an untrusted orchestration component can reach, and audit which service accounts can write nvmet configuration.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2024/CVE-2024-53681.json","https://nvd.nist.gov/vuln/detail/CVE-2024-53681"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-01-15"},{"id":"CVE-2025-20017","cve":"CVE-2025-20017","aliases":[],"title":"Intel oneAPI toolkit and component installers: An uncontrolled library search path: the component loads a shared","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel oneAPI toolkit and component installers","year":"2025","cvss_score":6.7,"severity":"medium","kev":false,"impact":"An uncontrolled library search path: the component loads a shared library by name from a directory a non-root user can write. Anyone who can drop a file in that directory gets code execution in the context of whoever next runs the tool - which on an AI node is usually a privileged installer, a service account, or root.","attack_vector":"A local authenticated user on a node that has the toolkit installed. On shared build/dev nodes and on container images built from the Intel toolkits, that is a broad set of people.","remediation":"Upgrade the affected component and, just as importantly, audit directory permissions on already-provisioned nodes and container images - upgrading the package does not remove a writable directory an earlier install created. Userspace only: no reboot, no BIOS, no microcode. Rebuild base images rather than patching running nodes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-20017","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01285.html"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2025-08-12"},{"id":"CVE-2025-20087","cve":"CVE-2025-20087","aliases":[],"title":"Intel oneAPI DPC++/C++ compiler installer: The compiler installer sets permissions that let a local user modify","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel oneAPI DPC++/C++ compiler installer","year":"2025","cvss_score":6.7,"severity":"medium","kev":false,"impact":"The compiler installer sets permissions that let a local user modify installed files that later execute with higher privilege. Same practical outcome as the search-path family: local privilege escalation on shared build and training nodes.","attack_vector":"A local authenticated user on a node that has the toolkit installed. On shared build/dev nodes and on container images built from the Intel toolkits, that is a broad set of people.","remediation":"Upgrade the affected component and, just as importantly, audit directory permissions on already-provisioned nodes and container images - upgrading the package does not remove a writable directory an earlier install created. Userspace only: no reboot, no BIOS, no microcode. Rebuild base images rather than patching running nodes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-20087","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01285.html"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2025-08-12"},{"id":"CVE-2025-20627","cve":"CVE-2025-20627","aliases":[],"title":"Intel oneAPI DPC++/C++ compiler: An uncontrolled library search path: the component loads a shared library by name","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel oneAPI DPC++/C++ compiler","year":"2025","cvss_score":6.7,"severity":"medium","kev":false,"impact":"An uncontrolled library search path: the component loads a shared library by name from a directory a non-root user can write. Anyone who can drop a file in that directory gets code execution in the context of whoever next runs the tool - which on an AI node is usually a privileged installer, a service account, or root.","attack_vector":"A local authenticated user on a node that has the toolkit installed. On shared build/dev nodes and on container images built from the Intel toolkits, that is a broad set of people.","remediation":"Upgrade the affected component and, just as importantly, audit directory permissions on already-provisioned nodes and container images - upgrading the package does not remove a writable directory an earlier install created. Userspace only: no reboot, no BIOS, no microcode. Rebuild base images rather than patching running nodes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-20627","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01285.html"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2025-08-12"},{"id":"CVE-2025-20629","cve":"CVE-2025-20629","aliases":[],"title":"Intel E810 NVM Update Utility: Insecure inherited permissions in the NVM update utility","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel E810 NVM Update Utility","year":"2025","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Insecure inherited permissions in the NVM update utility - the tool you run to fix the NIC firmware issues above - let a local authenticated user escalate. Worth noting because the remediation tool being the vulnerability is a genuine operational trap when you push it fleet-wide under automation.","attack_vector":"Authenticated local user on a node where the utility is staged.","remediation":"Use NVM Update Utility 4.60 or later, and check permissions on the staging directory your automation copies it into. Userspace tool, no reboot for the tool update itself.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-20629","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01295.html"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2025-05-13"},{"id":"CVE-2025-22397","cve":"CVE-2025-22397","aliases":[],"title":"Dell iDRAC9 / iDRAC10 (path traversal): A high-privileged remote attacker traverses paths on the BMC filesystem","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC9 / iDRAC10 (path traversal)","year":"2025","cvss_score":6.7,"severity":"medium","kev":false,"impact":"A high-privileged remote attacker traverses paths on the BMC filesystem, reaching high confidentiality and availability impact. Post-compromise this is how an attacker persists in the BMC.","attack_vector":"Authenticated high-privilege access to iDRAC9 (14G/15G/16G) or iDRAC10 (17G).","remediation":"Update to iDRAC9 7.00.00.181 / past 7.20.10.50, or iDRAC10 1.20.25.00. BMC flash. Because it needs high privilege, the compensating control is tight iDRAC RBAC and no shared admin credentials across the fleet.","references":["https://www.dell.com/support/kbdoc/en-us/000384516/dsa-2025-376-security-update-for-dell-idrac9-and-idrac10-vulnerabilities"],"status":"curated"},{"id":"CVE-2025-23299","cve":"CVE-2025-23299","aliases":[],"title":"NVIDIA ConnectX / BlueField: Privilege escalation leading to arbitrary code execution via the management interface","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA ConnectX / BlueField","year":"2025","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Privilege escalation leading to arbitrary code execution via the management interface","attack_vector":"Local","remediation":"NIC/DPU firmware update","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23299"],"status":"curated","fleet":{"ubiquity":"Universal on RDMA/InfiniBand fleets - ConnectX is the standard NIC in GPU clusters","remediation_pain":"`firmware-flash` on every NIC; requires a node reboot and in some fleets a maintenance window per rack","pain_class":"firmware-flash","why_fleet_wide":"Arbitrary code execution in the NIC/DPU management interface: the NIC sees all inter-node RDMA traffic for training jobs, so a compromise reads or corrupts gradients across tenants"},"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787"],"published":"2025-10-22"},{"id":"CVE-2025-23337","cve":"CVE-2025-23337","aliases":[],"title":"HGX / DGX GB200, GB300, B300 (BMC -> HMC): BMC admin can pivot to the HMC as administrator","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"HGX / DGX GB200, GB300, B300 (BMC -> HMC)","year":"2025","cvss_score":6.7,"severity":"medium","kev":false,"impact":"BMC admin can pivot to the HMC as administrator -> code execution and privesc on the GPU baseboard controller","attack_vector":"Operator or attacker holding BMC admin on the mgmt network","remediation":"Flash BMC + HMC firmware out-of-band; treat BMC admin as a tier-0 credential and rotate","references":["https://github.com/NVIDIA/product-security/tree/main/2025/5692","https://nvd.nist.gov/vuln/detail/CVE-2025-23337"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-1244"],"published":"2025-09-17"},{"cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-427"],"fleet":{"pain_class":"hot-patch"},"id":"CVE-2025-23355","cve":"CVE-2025-23355","aliases":[],"title":"NVIDIA Nsight Graphics (ngfx component, Windows): An ngfx component resolves a DLL from an untrusted search path, so","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Nsight Graphics (ngfx component, Windows)","year":"2025","cvss_score":6.7,"severity":"medium","kev":false,"impact":"An ngfx component resolves a DLL from an untrusted search path, so an attacker who can drop a file in the right directory gets their code loaded the next time a developer launches Nsight Graphics. The consequence is code execution and privilege escalation on the machine of someone who profiles GPU workloads - exactly the account that tends to hold source access, signing material and cluster credentials. Low fleet-wide reach, high value per host.","attack_vector":"A local, low-privileged account that can write into a directory in the loader's search path, plus a real user actually starting Nsight Graphics. NVIDIA rates the complexity high, so this is targeted rather than opportunistic - but the required write is often available on shared dev boxes with a loose system PATH.","remediation":"Upgrade Nsight Graphics to 2025.3 or later. Independently, sweep the system PATH on developer machines for entries that non-admin users can write to and remove them - that removes this whole class of hijack, not just this instance.","references":["https://github.com/NVIDIA/product-security/tree/main/2025/5704","https://nvd.nist.gov/vuln/detail/CVE-2025-23355"],"status":"curated"},{"id":"CVE-2025-33190","cve":"CVE-2025-33190","aliases":[],"title":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware: A second out-of-bounds write in SROOT firmware","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware","year":"2025","cvss_score":6.7,"severity":"medium","kev":false,"impact":"A second out-of-bounds write in SROOT firmware, reachable from a privileged local account, reaching firmware code execution. These sit in the GB10 root-of-trust chain, so a successful exploit undermines the platform's own attestation and secure-boot story rather than just the OS above it.","attack_vector":"Local access to the DGX Spark. Several of the set need no privileges at all; the rest need host root. This is a desk-side developer box, so physical and local access assumptions are much weaker than for a racked DGX.","remediation":"Apply the DGX Spark firmware update from bulletin 5720. Cost: flash plus reboot, low drain cost given the form factor, but not live-patchable and root-of-trust firmware cannot be rolled back once applied.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33190","https://github.com/NVIDIA/product-security/tree/main/2025/5720"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"firmware-flash"},"published":"2025-11-25"},{"id":"CVE-2025-33231","cve":"CVE-2025-33231","aliases":[],"title":"CUDA Toolkit: Code exec via path manipulation on library load","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2025","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Code exec via path manipulation on library load","attack_vector":"Local user","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33231","https://github.com/NVIDIA/product-security/tree/main/2026/5755"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-427"],"published":"2026-01-20"},{"id":"CVE-2025-46366","cve":"CVE-2025-46366","aliases":[],"title":"Dell CloudLink (privilege escalation to database): A privileged user escalates laterally or reads the CloudLink","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Dell CloudLink (privilege escalation to database)","year":"2025","cvss_score":6.7,"severity":"medium","kev":false,"impact":"A privileged user escalates laterally or reads the CloudLink database directly, obtaining confidential information - which for a KMS means key metadata and policy.","attack_vector":"Local high-privilege access on the appliance.","remediation":"Upgrade CloudLink to 8.1.1 or later.","references":["https://www.dell.com/support/kbdoc/en-us/000384363/dsa-2025-374-security-update-for-dell-cloudlink-multiple-security-vulnerabilities"],"status":"curated"},{"id":"CVE-2025-46424","cve":"CVE-2025-46424","aliases":[],"title":"Dell CloudLink (risky cryptographic primitive): Use of a cryptographic primitive with a risky implementation","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Dell CloudLink (risky cryptographic primitive)","year":"2025","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Use of a cryptographic primitive with a risky implementation, exploitable by a high-privileged attacker for denial of service of the key manager. If CloudLink is down, encrypted volumes do not unlock - this is an availability risk to the storage tier.","attack_vector":"Local high-privilege access on the appliance.","remediation":"Upgrade CloudLink to 8.2. Make sure you have tested the KMS-unavailable failure mode for your storage estate.","references":["https://www.dell.com/support/kbdoc/en-us/000384363/dsa-2025-374-security-update-for-dell-cloudlink-multiple-security-vulnerabilities"],"status":"curated"},{"id":"CVE-2025-5187","cve":"CVE-2025-5187","aliases":[],"title":"Kubernetes (kube-apiserver): A node can delete itself, and cascade-delete other objects, by adding an OwnerReference","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-apiserver)","year":"2025","cvss_score":6.7,"severity":"medium","kev":false,"impact":"A node can delete itself, and cascade-delete other objects, by adding an OwnerReference","attack_vector":"A compromised node / kubelet credential","remediation":"Rolling control-plane upgrade; tighten the NodeRestriction admission plugin","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated","published":"2025-08-27"},{"cvss_vector":"CVSS:3.0/AV:A/AC:L/PR:H/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-288"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2018-10841","cve":"CVE-2018-10841","aliases":[],"title":"GlusterFS (glusterd management): An authenticated TLS client can use gluster cli --remote-host to add itself to the","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"GlusterFS (glusterd management)","year":"2018","cvss_score":6.6,"severity":"medium","kev":false,"impact":"An authenticated TLS client can use gluster cli --remote-host to add itself to the trusted storage pool, at which point it runs privileged management operations. A tenant with a client certificate becomes a storage administrator and can reconfigure or destroy other tenants' volumes.","attack_vector":"Any client holding a valid TLS credential that can reach glusterd on a server node.","remediation":"Upgrade glusterfs and restart glusterd on every server. Separate the management network from the client data network so tenant nodes cannot reach glusterd's management port at all, and review the trusted pool membership for hosts you did not add.","references":["https://access.redhat.com/security/cve/CVE-2018-10841","https://nvd.nist.gov/vuln/detail/CVE-2018-10841"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2018-7113","cve":"CVE-2018-7113","aliases":["HPESBHF03894"],"title":"HPE iLO 5 (firmware update security restriction bypass): Bypass of the security restrictions that guard iLO 5 firmware","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HPE iLO 5 (firmware update security restriction bypass)","year":"2018","cvss_score":6.6,"severity":"medium","kev":false,"impact":"Bypass of the security restrictions that guard iLO 5 firmware updates. The severity number undersells what this is: the firmware-update gate is the control that stops an attacker from writing their own image to the service processor. Defeat it and the attacker installs persistent BMC firmware of their choosing - the deepest and most durable implant available on the node, below the hypervisor and untouched by any host reimage. On a bare-metal cloud, a node whose iLO firmware was replaced by a previous tenant never comes clean again through normal reprovisioning.","attack_vector":"Local exploitation - an attacker who already has a foothold on the node or its iLO context, not an unauthenticated remote attacker. The realistic path in a bare-metal fleet is a tenant with host-level access using it during their tenancy to leave something behind for the next one.","remediation":"Flash iLO 5 to v1.37 or later - out-of-band, per-node, no host reboot and no job drain. Complementary control that matters more than the version number: enable and actually check the iLO firmware integrity/attestation features HPE exposes, and verify iLO firmware version and measurement as part of node reprovisioning between tenants rather than trusting that a wipe covered it.","references":["https://support.hpe.com/hpsc/doc/public/display?docLocale=en_US&docId=emr_na-hpesbhf03894en_us","https://nvd.nist.gov/vuln/detail/CVE-2018-7113"],"status":"curated","published":"2018-12-03"},{"id":"CVE-2021-0060","cve":"CVE-2021-0060","aliases":[],"title":"Intel SPS (HECI subsystem compartmentalisation): Insufficient compartmentalisation in the HECI interface","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel SPS (HECI subsystem compartmentalisation)","year":"2021","cvss_score":6.6,"severity":"medium","kev":false,"impact":"Insufficient compartmentalisation in the HECI interface - the host-to-management-engine channel - on Server Platform Services firmware. HECI is the door between the OS and the management engine, so weak compartmentalisation there means host-side code reaches further into the engine than it should.","attack_vector":"Local access on the host with the ability to talk to the HECI device.","remediation":"Fixed in Intel CSME/SPS firmware, which reaches you as an OEM BIOS or firmware package - not as a microcode or OS update. That means: wait for your server vendor to ship it, drain the node, flash, and reboot. OEM availability is the long pole and routinely lags the Intel advisory by one or more quarters on server platforms. Track it per platform SKU, because vendors ship these unevenly across their own product lines.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-0060","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00470.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2022-02-09"},{"id":"CVE-2021-1076","cve":"CVE-2021-1076","aliases":[],"title":"NVIDIA GPU Display Driver (Windows nvlddmkm.sys + Linux nvidia.ko): Improper access control in the kernel-mode layer","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver (Windows nvlddmkm.sys + Linux nvidia.ko)","year":"2021","cvss_score":6.6,"severity":"medium","kev":false,"impact":"Improper access control in the kernel-mode layer on both Windows and Linux, producing disclosure or data corruption. Data corruption from a driver-level access control gap is the quiet kind of failure - silently wrong training runs rather than a visible crash. Debian and Gentoo shipped it as a security update.","attack_vector":"Any local user or GPU container on the host.","remediation":"Install the fixed GPU Display Driver branch on both Windows and Linux nodes. The kernel component (nvlddmkm.sys / nvidia.ko) cannot be hot-swapped under load, so this is a node drain and reboot per host; restart the container runtime afterwards so mounted driver libraries match the kernel module. No VBIOS or BMC flash.","references":["https://lists.debian.org/debian-lts-announce/2022/01/msg00013.html","https://security.gentoo.org/glsa/202310-02","https://nvd.nist.gov/vuln/detail/CVE-2021-1076"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-04-21"},{"id":"CVE-2021-1077","cve":"CVE-2021-1077","aliases":[],"title":"NVIDIA GPU Display Driver (Windows nvlddmkm.sys + Linux nvidia.ko): Reference-count mishandling on a driver resource","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver (Windows nvlddmkm.sys + Linux nvidia.ko)","year":"2021","cvss_score":6.6,"severity":"medium","kev":false,"impact":"Reference-count mishandling on a driver resource in the R450 and R460 branches. Refcount bugs of this shape typically become use-after-free with effort; NVIDIA rates the confirmed impact as denial of service.","attack_vector":"Any local user or GPU container with device access.","remediation":"Install the fixed GPU Display Driver branch on both Windows and Linux nodes. The kernel component (nvlddmkm.sys / nvidia.ko) cannot be hot-swapped under load, so this is a node drain and reboot per host; restart the container runtime afterwards so mounted driver libraries match the kernel module. No VBIOS or BMC flash.","references":["https://security.gentoo.org/glsa/202310-02","https://nvd.nist.gov/vuln/detail/CVE-2021-1077"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-04-21"},{"id":"CVE-2022-3294","cve":"CVE-2022-3294","aliases":[],"title":"Kubernetes (kube-apiserver): Node address not verified when proxying","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-apiserver)","year":"2022","cvss_score":6.6,"severity":"medium","kev":false,"impact":"Node address not verified when proxying; a user who can modify Node objects reaches control-plane-only endpoints","attack_vector":"Cluster user able to patch Node objects","remediation":"Rolling control-plane upgrade; restrict Node patch RBAC","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-3294"],"status":"curated","published":"2023-03-01"},{"id":"CVE-2023-0198","cve":"CVE-2023-0198","aliases":[],"title":"GPU Display Driver: Local privesc (kernel buffer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2023","cvss_score":6.6,"severity":"medium","kev":false,"impact":"Local privesc (kernel buffer overflow)","attack_vector":"Any tenant with a container; vGPU guest","remediation":"Driver + vGPU Manager upgrade; drain + reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0198","https://github.com/NVIDIA/product-security/tree/main/2023/5452"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:L/A:H","cwe":["CWE-119"],"fleet":{"pain_class":"node-reboot"},"published":"2023-04-01"},{"id":"CVE-2023-31015","cve":"CVE-2023-31015","aliases":[],"title":"DGX H100 BMC (REST): Privesc via auth flaw","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX H100 BMC (REST)","year":"2023","cvss_score":6.6,"severity":"medium","kev":false,"impact":"Privesc via auth flaw","attack_vector":"Network-adjacent BMC REST client","remediation":"Flash BMC 23.08.18 out-of-band","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5473/5473.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:H/A:L","cwe":["CWE-287"],"published":"2023-09-20"},{"id":"CVE-2023-31034","cve":"CVE-2023-31034","aliases":[],"title":"DGX A100 SBIOS: Integer overflow in SBIOS","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX A100 SBIOS","year":"2023","cvss_score":6.6,"severity":"medium","kev":false,"impact":"Integer overflow in SBIOS","attack_vector":"Local operator","remediation":"Flash SBIOS 1.25+","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31034","https://github.com/NVIDIA/product-security/tree/main/2024/5513"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:R/S:C/C:L/I:L/A:H","cwe":["CWE-190"],"fleet":{"pain_class":"firmware-flash"},"published":"2024-01-12"},{"id":"CVE-2023-34473","cve":"CVE-2023-34473","aliases":["AMI-SA-2023006","Nozomi Labs BMC audit"],"title":"AMI MegaRAC SPx (BMC hard-coded credentials): Hard-coded credentials inside the BMC firmware","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx (BMC hard-coded credentials)","year":"2023","cvss_score":6.6,"severity":"medium","kev":false,"impact":"Hard-coded credentials inside the BMC firmware. Once extracted from a publicly downloadable image, they work on every BMC running that build regardless of what passwords the operator set, so your entire fleet shares an authentication backdoor you cannot rotate. That converts a per-node authentication story into a single fleet-wide key, and BMC access means power control, console, virtual media and firmware.","attack_vector":"Adjacent network reachability to the BMC plus a valid user session and some interaction, per AMI's vector. The credential itself is obtained offline by unpacking a firmware image - no access to your systems is needed for that half.","remediation":"Firmware flash to SPx_12.2 / SPx_13.0 or later. This one has been fixed since early SPx builds, so the operator task is an audit: enumerate the actual running BMC firmware version across the fleet and find the SKUs still on a pre-fix ODM image - typically older or white-box nodes whose vendor stopped publishing BMC updates. There is no config-only fix, because the credential is baked into the image; the only compensating control is hard network isolation of the BMC plane so the credential has nothing to authenticate against.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2023006.pdf","https://nvd.nist.gov/vuln/detail/CVE-2023-34473"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-07-05"},{"id":"CVE-2023-4622","cve":"CVE-2023-4622","aliases":[],"title":"Linux kernel (AF_UNIX): Use-after-free in unix_stream_sendpage - local privilege escalation, no capabilities needed","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (AF_UNIX)","year":"2023","cvss_score":6.6,"severity":"medium","kev":false,"impact":"Use-after-free in unix_stream_sendpage - local privilege escalation, no capabilities needed","attack_vector":"Any tenant process in a container","remediation":"Livepatchable; otherwise drain + reboot. Note this one needs no CAP_NET_ADMIN, so userns hardening does not help","references":["https://access.redhat.com/security/cve/CVE-2023-4622"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2023-09-06"},{"id":"CVE-2023-52819","cve":"CVE-2023-52819","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd): An out-of-bounds access in the amdgpu power management","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd)","year":"2023","cvss_score":6.6,"severity":"medium","kev":false,"impact":"An out-of-bounds access in the amdgpu power management (SMU/powerplay) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd: Fix UBSAN array-index-out-of-bounds for Polaris and Tonga","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52819","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-21"},{"id":"CVE-2024-20294","cve":"CVE-2024-20294","aliases":[],"title":"Cisco FXOS / NX-OS (LLDP frame handling denial of service): An unauthenticated adjacent attacker sends crafted LLDP","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco FXOS / NX-OS (LLDP frame handling denial of service)","year":"2024","cvss_score":6.6,"severity":"medium","kev":false,"impact":"An unauthenticated adjacent attacker sends crafted LLDP frames and takes the switch down, with scope change. Any compromised host NIC plugged into the fabric can do this.","attack_vector":"Layer-2 adjacency - i.e. anything connected to the switch.","remediation":"Upgrade NX-OS/FXOS per cisco-sa-nxos-lldp-dos-z7PncTgt. Switch reboot required. Interim: disable LLDP on host-facing ports where you do not depend on it for topology discovery.","references":["https://sec.cloudapps.cisco.com/security/center/content/CiscoSecurityAdvisory/cisco-sa-nxos-lldp-dos-z7PncTgt"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-28224","cve":"CVE-2024-28224","aliases":[],"title":"Ollama: DNS rebinding grants a remote page full API access","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ollama","year":"2024","cvss_score":6.6,"severity":"medium","kev":false,"impact":"DNS rebinding grants a remote page full API access","attack_vector":"Browser of anyone on a network with an Ollama host","remediation":"Upgrade past 0.1.29; bind loopback","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-28224"],"status":"curated","published":"2024-04-08"},{"id":"CVE-2024-38482","cve":"CVE-2024-38482","aliases":[],"title":"Dell CloudLink (cluster component exception handling): A highly privileged remote attacker performs unauthorized","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Dell CloudLink (cluster component exception handling)","year":"2024","cvss_score":6.6,"severity":"medium","kev":false,"impact":"A highly privileged remote attacker performs unauthorized actions and reads sensitive information from the CloudLink database via the cluster component.","attack_vector":"Remote, high-privilege, against CloudLink 7.1.x/8.x.","remediation":"Apply the DSA-2024-343 CloudLink update. Appliance upgrade.","references":["https://www.dell.com/support/kbdoc/en-us/000227493/dsa-2024-343-security-update-for-dell-cloudlink-vulnerability"],"status":"curated"},{"id":"CVE-2025-0037","cve":"CVE-2025-0037","aliases":[],"title":"AMD Versal Adaptive SoC - PLM runtime services address validation: The Platform Loader and Manager firmware on AMD","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD Versal Adaptive SoC - PLM runtime services address validation","year":"2025","cvss_score":6.6,"severity":"medium","kev":false,"impact":"The Platform Loader and Manager firmware on AMD Versal devices does not validate addresses when executing runtime services, so a caller reaches isolated or protected memory spaces. On Versal parts the PLM is the root of trust and the isolation enforcer for the device's partitions - bypassing its address checks means crossing whatever partition boundary the design relies on, which on a multi-tenant SmartNIC or accelerator card is the tenant boundary.","attack_vector":"Local to the device, via PLM runtime service calls.","remediation":"Fixed in updated PLM firmware from AMD/Xilinx, applied as a device firmware image plus a card reset. Reaches you through whoever built the card, so expect integrator lag on top of AMD's release. Track which Versal-based cards are in your fleet and who owns their firmware pipeline.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0037","https://www.amd.com/en/resources/product-security.html"],"status":"curated","tags":["tenant-isolation"],"published":"2025-06-10"},{"id":"CVE-2025-0038","cve":"CVE-2025-0038","aliases":[],"title":"AMD Zynq UltraScale+ - CSU runtime service address validation in PMU firmware: The PMU firmware on Zynq UltraScale+","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Zynq UltraScale+ - CSU runtime service address validation in PMU firmware","year":"2025","cvss_score":6.6,"severity":"medium","kev":false,"impact":"The PMU firmware on Zynq UltraScale+ devices does not validate addresses when executing Configuration Security Unit runtime services, giving access to isolated or protected memory. Companion to the Versal PLM issue and the same shape: the firmware component that enforces isolation on the device can be steered outside its own boundaries.","attack_vector":"Local to the device, via CSU runtime service calls through the PMU firmware.","remediation":"Fixed in updated PMU firmware from AMD/Xilinx. Device firmware update plus card reset, gated on your card integrator shipping it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0038","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2025-10-06"},{"id":"CVE-2025-21692","cve":"CVE-2025-21692","aliases":[],"title":"Linux kernel (net/sched ETS): Out-of-bounds indexing in the ETS qdisc - memory corruption from CAP_NET_ADMIN","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/sched ETS)","year":"2025","cvss_score":6.6,"severity":"medium","kev":false,"impact":"Out-of-bounds indexing in the ETS qdisc - memory corruption from CAP_NET_ADMIN","attack_vector":"Any tenant process in a container with CAP_NET_ADMIN in a userns","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2025-21692"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-02-10"},{"id":"CVE-2025-2884","cve":"CVE-2025-2884","aliases":[],"title":"TCG TPM 2.0 reference implementation (CryptHmacSign): Out-of-bounds read in the reference implementation's HMAC signing","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"TCG TPM 2.0 reference implementation (CryptHmacSign)","year":"2025","cvss_score":6.6,"severity":"medium","kev":false,"impact":"Out-of-bounds read in the reference implementation's HMAC signing helper because the signature scheme is not validated against the key's algorithm. Reads past the buffer can disclose TPM-internal memory - the one place in the system that is supposed to be opaque. Because this is the reference code, the same bug propagates into every fTPM, software TPM and vendor TPM derived from it, which is most of them.","attack_vector":"Local user able to issue TPM commands. On a shared or bare-metal node, that is any tenant.","remediation":"Update to TPM 2.0 reference implementation 1.83 or later - in practice that arrives as a platform firmware/BIOS update, a swtpm/libtpms package update for virtualised TPMs, or nothing at all if your TPM vendor has not rebased. Virtualised TPMs are the easy case (package update plus VM restart); silicon and fTPM are a per-node firmware flash. Check both paths separately; fleets usually have some of each.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-2884","https://trustedcomputinggroup.org/trusted-computing-group-vulnerability-response/"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2025-06-10"},{"id":"CVE-2025-51481","cve":"CVE-2025-51481","aliases":[],"title":"Dagster (gRPC `get_notebook_data`): Local file inclusion — read arbitrary files","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Dagster (gRPC `get_notebook_data`)","year":"2025","cvss_score":6.6,"severity":"medium","kev":false,"impact":"Local file inclusion — read arbitrary files","attack_vector":"Attacker with access to the Dagster gRPC server, i.e. a co-tenant in a flat network","remediation":"Upgrade past 1.10.14","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-51481"],"status":"curated","published":"2025-07-22"},{"id":"CVE-2026-47619","cve":"CVE-2026-47619","aliases":[],"title":"NVIDIA Dynamo: Improper access control on privileged operations","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Dynamo","year":"2026","cvss_score":6.6,"severity":"medium","kev":false,"impact":"Improper access control on privileged operations","attack_vector":"Tenant with cluster network access","remediation":"Bump Dynamo; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47619","https://github.com/NVIDIA/product-security/tree/main/2026/5842"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-1357"],"published":"2026-08-04"},{"cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-53115","cve":"CVE-2026-53115","aliases":[],"title":"Linux kernel (drivers/bus/fsl-mc): The fsl-mc bus read its driver_override string without holding the device lock, so","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/bus/fsl-mc)","year":"2026","cvss_score":6.6,"severity":"medium","kev":false,"impact":"The fsl-mc bus read its driver_override string without holding the device lock, so the string can be freed underneath a concurrent probe - a use-after-free in the exact mechanism used to force a device onto vfio-fsl-mc for passthrough. A corrupted override can also land a device on the wrong driver, meaning a device the operator intended to hand to a tenant instead binds to a host driver, or the reverse.","attack_vector":"Needs host root: writing /sys/bus/fsl-mc/devices/*/driver_override while a driver bind is in flight. This is the operator's own passthrough-provisioning path, not a tenant surface - the risk is an automation race in node preparation, plus anything that gets root on the host. Hardware-conditional: NXP DPAA2 / fsl-mc platforms only, not x86 or standard Arm server nodes.","remediation":"Update to a stable kernel carrying commits 4911b836 / 8139ce66 on fsl-mc platforms. Interim: serialize driver_override writes against bind/unbind in your node-provisioning tooling. No action on x86 or Arm server fleets.","references":["https://git.kernel.org/stable/c/4911b836f35c034c36f102db4ecbe339b38e7d1d","https://git.kernel.org/stable/c/8139ce66b52a4a5638bfb445b037c07d4abeb08e","https://nvd.nist.gov/vuln/detail/CVE-2026-53115"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-53120","cve":"CVE-2026-53120","aliases":[],"title":"Linux kernel (drivers/pci): The PCI bus match callback read driver_override without the device lock, so the override","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/pci)","year":"2026","cvss_score":6.6,"severity":"medium","kev":false,"impact":"The PCI bus match callback read driver_override without the device lock, so the override string can be freed while a probe is reading it - a use-after-free in the single mechanism a GPU cloud uses to force a card onto vfio-pci for passthrough. Beyond the memory-safety bug, a torn override read means a device can bind to the wrong driver: a card meant for a tenant lands on a host driver, or a host device lands on vfio-pci and becomes reachable from a container.","attack_vector":"Needs host root: writing /sys/bus/pci/devices/<bdf>/driver_override concurrently with a bind, which is exactly what node-provisioning automation does when it flips GPUs and NICs between host drivers and vfio-pci. Not tenant-reachable, but it sits on the provisioning path that decides who owns which device, and the xen-pciback stub is affected the same way.","remediation":"Update to a stable kernel carrying commits dfe950d9 / 58a42be0. Interim: serialize driver_override writes against driver bind/unbind in your provisioning tooling - write the override, then bind, never concurrently - and verify the resulting driver binding before offering a device to a tenant.","references":["https://git.kernel.org/stable/c/dfe950d9464cad609f3b118c6203e2708055bc61","https://git.kernel.org/stable/c/58a42be0d70307d765594fc581f5f5e5ef059712","https://nvd.nist.gov/vuln/detail/CVE-2026-53120"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-125"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-74346","cve":"CVE-2026-74346","aliases":[],"title":"Linux kernel (drivers/infiniband/hw/irdma): A stale flag caused the CQ memory-registration path to read one element","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/hw/irdma)","year":"2026","cvss_score":6.6,"severity":"medium","kev":false,"impact":"A stale flag caused the CQ memory-registration path to read one element past the end of the page-address array on current-generation hardware, and to program whatever it read as the CQ shadow DMA address. That is both a kernel out-of-bounds read and an adapter pointed at an address derived from adjacent heap contents.","attack_vector":"A tenant container holding /dev/infiniband/uverbs* on an Intel irdma node registers CQ memory where the declared page count equals the region's page count - the normal case, not a crafted one. Local to the tenant; no fabric peer needed.","remediation":"No fixed version is listed in the record - take the stable kernel carrying a80b3b13786e (or 3159c6fac43d / ad360a31092a) and reboot. Interim: drop /dev/infiniband/* from untrusted containers on irdma nodes.","references":["https://git.kernel.org/stable/c/a80b3b13786e9ab1c52b31a1f16c7d6708fa9220","https://git.kernel.org/stable/c/3159c6fac43dc24b34d31971884d98a7a1bf4c4b","https://nvd.nist.gov/vuln/detail/CVE-2026-74346"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:H/I:H/A:N","cwe":["CWE-532"],"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2021-016-argo-cd-helm-oci-repository-cred","cve":null,"aliases":["GHSA-6w87-g839-9wv7"],"title":"Argo CD (Helm OCI repository credential logging): CREDENTIAL DISCLOSURE THROUGH THE LOG PIPELINE: Argo CD wrote the","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD (Helm OCI repository credential logging)","year":"2021","cvss_score":6.6,"severity":"medium","kev":false,"impact":"CREDENTIAL DISCLOSURE THROUGH THE LOG PIPELINE: Argo CD wrote the credentials for authenticated Helm OCI repositories into its pod logs. Anyone with read access to those logs gets the repository credentials — and in a real cluster that population is much larger than the set of people trusted with registry writes, because logs are aggregated into a centralized platform where operators, SREs and often tenant-facing observability tooling can read them. Repository credentials for an OCI registry mean the ability to publish charts that Argo CD will then reconcile onto the cluster, which turns a log-read into a deployment primitive against everything that GitOps controller manages. The exposure is retroactive and lives in log retention rather than in the running system.","attack_vector":"Local / log access: anyone with permission to read Argo CD pod logs via the Kubernetes control plane, or any downstream log aggregation system into which those logs were shipped. Affects all versions before 1.7.14 and 1.8.7 that connect to Helm OCI repositories with authentication enabled.","remediation":"Upgrade Argo CD to 1.7.14 or 1.8.7 or later. Then rotate the Helm OCI repository credentials — patching stops new writes but does nothing about what is already sitting in log retention and downstream indexes. Purge or expire the affected log ranges in your aggregation platform, and review who has read access to Argo CD logs.","references":["https://github.com/argoproj/argo-cd/security/advisories/GHSA-6w87-g839-9wv7"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:L/A:H","cwe":["CWE-125"],"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2025-020-llama-cpp-gguf-vocabulary-parsin","cve":null,"aliases":["GHSA-g4cc-763q-h9h6"],"title":"llama.cpp (GGUF vocabulary parsing, llama_vocab::impl::print_info): MALICIOUS MODEL FILE CRASHES THE SERVER: the GGUF","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"llama.cpp (GGUF vocabulary parsing, llama_vocab::impl::print_info)","year":"2025","cvss_score":6.6,"severity":"medium","kev":false,"impact":"MALICIOUS MODEL FILE CRASHES THE SERVER: the GGUF parser reads special token IDs without bounds-checking them against the vocabulary. special_bos_id defaults to 1, so a model whose vocabulary contains a single token makes print_info index id_to_token[1] out of bounds on the heap and segfault. Any deployment that loads user-supplied or third-party GGUF files — a model-hosting service, a shared inference gateway that lets tenants bring their own weights, a self-service fine-tune endpoint — can be taken down by uploading a small crafted file. It is a heap over-read rather than a write, so the realistic ceiling is denial of service, but on a GPU serving node that means the accelerator sitting idle and every co-tenant on that replica losing service until the process is recycled.","attack_vector":"Local to the loading process, reached remotely wherever the model path is attacker-influenced: the victim runs llama.cpp against a crafted GGUF file. Any bring-your-own-model or shared model-cache workflow puts this within reach of a tenant.","remediation":"Update llama.cpp past the fixed commit and restart the servers. Validate GGUF files before loading — reject models whose declared vocabulary size is smaller than the special token IDs they reference — and load only from model stores whose writers you trust. Set pod restart policies so a crashed serving process is recycled rather than leaving the GPU stranded.","references":["https://github.com/ggml-org/llama.cpp/security/advisories/GHSA-g4cc-763q-h9h6"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-362","CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2011-0695","cve":"CVE-2011-0695","aliases":[],"title":"Linux kernel InfiniBand connection manager drivers/infiniband/core/cma.c and cm.c - cm_work_handler race: A race in the","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel InfiniBand connection manager drivers/infiniband/core/cma.c and cm.c - cm_work_handler race","year":"2011","cvss_score":6.5,"severity":"medium","kev":false,"impact":"A race in the InfiniBand connection-manager work handler. Sending a CM request while other request handlers are still running drives an invalid pointer dereference and panics the node - no account on the target, no authentication step, just frames on the fabric. The fix is two commits, and the second one takes a reference on the cm_id before invoking the callback, which means the underlying defect is a live object being used without a reference held. That is use-after-free shaped, so treating the impact ceiling as 'panic' is the optimistic reading. Either way it is one tenant crashing other tenants' nodes across a shared IB fabric.","attack_vector":"Adjacent network, pre-auth. Any host that can send InfiniBand CM requests to the target - i.e. any node or tenant on the same fabric partition.","remediation":"Kernel upgrade or vendor backport of both commits 25ae21a10112875763c18b385624df713a288a05 (RDMA/cma: fix crash in request handlers) and 29963437a48475036353b95ab142bf199adb909e (IB/cm: bump reference count on cm_id before invoking callback) - applying only the first leaves the refcount defect in place. Rolling reboot. The structural control, shared with every other pre-auth RDMA-CM issue in this set, is fabric partitioning: IB P_Keys or RoCE VLAN separation so tenants cannot address each other's connection managers at all, because the CM itself has no authentication to enable.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=25ae21a10112875763c18b385624df713a288a05","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=29963437a48475036353b95ab142bf199adb909e","https://access.redhat.com/security/cve/CVE-2011-0695","https://www.openwall.com/lists/oss-security/2011/03/11/1"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:N/I:H/A:N","cwe":["CWE-732","CWE-276"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2012-4518","cve":"CVE-2012-4518","aliases":[],"title":"ibacm 1.0.7 (InfiniBand Communication Manager Assistant daemon) - world-writable log and ibacm.port files: The RDMA","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ibacm 1.0.7 (InfiniBand Communication Manager Assistant daemon) - world-writable log and ibacm.port files","year":"2012","cvss_score":6.5,"severity":"medium","kev":false,"impact":"The RDMA fabric-assistance daemon creates its files world-writable, so any local user rewrites the ibacm.port file - which is precisely the input librdmacm consults to decide which ib_acm service to trust for address resolution. On its own it is a permissions bug; chained with CVE-2012-4516 it lets an unprivileged tenant on a shared node redirect every other process's RDMA address resolution to a service the tenant controls. Also lets a tenant tamper with the daemon's log, which removes the record of them having done it.","attack_vector":"Local, unprivileged - any user on a node running ibacm. On a shared or multi-tenant management/compute node this is any tenant process.","remediation":"Update ibacm past 1.0.7 (upstream commit d204fca2b6298d7799e918141ea8e11e7ad43cec; Red Hat shipped it in RHSA-2013-0509) and restart the daemon. No reboot needed. Independently verifiable and worth checking right now regardless of package version: stat the ibacm files on a live node and confirm they are not group- or world-writable, since a permissions defect can be reintroduced by packaging, by a container image, or by an operator's own configuration-management run long after the upstream fix.","references":["https://www.openwall.com/lists/oss-security/2012/10/11/6","https://access.redhat.com/errata/RHSA-2013-0509","https://access.redhat.com/security/cve/CVE-2012-4518"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-20","CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2014-2739","cve":"CVE-2014-2739","aliases":[],"title":"Linux kernel RDMA connection manager drivers/infiniband/core/cma.c - cma_req_handler (RoCE): Pre-authentication remote","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel RDMA connection manager drivers/infiniband/core/cma.c - cma_req_handler (RoCE)","year":"2014","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Pre-authentication remote crash of any node running an RDMA-CM listener. Crafted RoCE connection-request traffic drives cma_req_handler down an address-resolution path owned by a different module, the pointer is wrong, and the kernel panics. The isolation lesson matters more than the crash: the RDMA connection manager accepts and parses attacker-controlled connection requests before any authentication exists, so on a shared fabric every node's kernel is parsing untrusted input from every tenant by design. One tenant reboots a rack's worth of training nodes and takes out everyone's in-flight jobs and their unsaved optimizer state.","attack_vector":"Adjacent network, pre-auth. Any host on the RoCE/IB fabric that can reach the node's RDMA-CM listener. No credentials and no account on the target.","remediation":"Kernel upgrade past 3.14.1 or a vendor backport; rolling reboot. The structural fix that outlives this one CVE is fabric segmentation - keep the RDMA control plane on a dedicated, tenant-unreachable L2 domain, and where the fabric must be shared, use partition keys (IB P_Keys) or RoCE VLAN isolation so tenants cannot address each other's CM listeners at all. RDMA-CM itself has no authentication to turn on.","references":["https://access.redhat.com/security/cve/CVE-2014-2739","https://bugzilla.redhat.com/show_bug.cgi?id=1088078","https://nvd.nist.gov/vuln/detail/CVE-2014-2739"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:N/I:N/A:H","cwe":["CWE-399","CWE-248"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2015-5307","cve":"CVE-2015-5307","aliases":["XSA-156"],"title":"Linux KVM (arch/x86/kvm/svm.c, vmx.c) and Xen 4.3.x-4.6.x - #AC exception handling: A guest raises alignment-check","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux KVM (arch/x86/kvm/svm.c, vmx.c) and Xen 4.3.x-4.6.x - #AC exception handling","year":"2015","cvss_score":6.5,"severity":"medium","kev":false,"impact":"A guest raises alignment-check exceptions in a tight loop and the host panics or hangs. No exploit chain, no memory corruption, no privilege needed inside the guest beyond running instructions - a few lines of assembly from an unprivileged process in any tenant VM takes down the whole physical machine and every co-tenant's workload with it. For a GPU cloud where one node carries eight accelerators and potentially several tenants, this is the cheapest possible cross-tenant availability attack and it hits KVM and Xen alike.","attack_vector":"Guest OS user - not even guest administrator - in any VM on the host. Unprivileged code inside the tenant's own guest is sufficient.","remediation":"Kernel update for KVM hosts, XSA-156 patches for Xen; both need a host reboot with tenants evacuated. There is no configuration workaround - you cannot disable #AC delivery - so unpatched hosts simply have this exposure. Given how trivially triggerable it is, treat it as a gating check before a host is allowed to accept multi-tenant placement, rather than something to schedule into a routine patch cycle.","references":["https://xenbits.xen.org/xsa/advisory-156.html","https://access.redhat.com/security/cve/CVE-2015-5307","https://nvd.nist.gov/vuln/detail/CVE-2015-5307"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:N/A:N","cwe":["CWE-200","CWE-908"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2015-8553","cve":"CVE-2015-8553","aliases":["XSA-120"],"title":"Xen PCI passthrough - device memory/IO decoding and host memory initialisation: With memory and I/O decoding left","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen PCI passthrough - device memory/IO decoding and host memory initialisation","year":"2015","cvss_score":6.5,"severity":"medium","kev":false,"impact":"With memory and I/O decoding left disabled on an assigned device, reads that should have gone to the device instead return uninitialised host kernel memory to the guest. A tenant with a passed-through GPU sweeps its own BAR window and harvests whatever the host had in those pages - a pure cross-boundary confidentiality leak, silent, with no crash and no error counter to trip. It is the incomplete-fix follow-on to CVE-2015-0777 and the reason the XSA-120 family should be treated as an information-disclosure issue rather than only a DoS.","attack_vector":"Guest user with an assigned PCI device; reads its own device BARs while decoding is disabled.","remediation":"Xen update per the XSA-120 family plus the corrected fix for CVE-2015-0777; host reboot. Because the leak is read-only and silent, there is no detection to fall back on - an operator cannot tell after the fact whether a tenant harvested host memory, which is the argument for treating this as patch-now rather than accepting it until the next maintenance window.","references":["https://xenbits.xen.org/xsa/advisory-120.html","https://access.redhat.com/security/cve/CVE-2015-8553","https://nvd.nist.gov/vuln/detail/CVE-2015-8553"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.0/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-20"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2017-16816","cve":"CVE-2017-16816","aliases":["HTCONDOR-2017-0001"],"title":"HTCondor (condor_schedd, GSI/VOMS extension parsing): An authenticated user crashes the schedd by feeding it malformed","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HTCondor (condor_schedd, GSI/VOMS extension parsing)","year":"2017","cvss_score":6.5,"severity":"medium","kev":false,"impact":"An authenticated user crashes the schedd by feeding it malformed GSI/VOMS extensions. The schedd owns the job queue, so while it is down nobody in the pool can submit, query or reap jobs and the GPUs behind it drain.","attack_vector":"A remote authenticated user of the pool, on a schedd configured to use GSI with VOMS extensions.","remediation":"Upgrade to HTCondor 8.6.8 or 8.7.5 and restart condor_schedd. GSI is deprecated in modern HTCondor - moving the pool to IDTOKENS or SSL removes this code path entirely.","references":["https://htcondor.org/security/vulnerabilities/HTCONDOR-2017-0001.html","https://nvd.nist.gov/vuln/detail/CVE-2017-16816"],"status":"curated"},{"cvss_vector":"CVSS:3.0/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-502","CWE-190","CWE-200"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2018-10911","cve":"CVE-2018-10911","aliases":[],"title":"GlusterFS (dict_unserialize): A negative key length in a serialized dict makes the server read memory from elsewhere in","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"GlusterFS (dict_unserialize)","year":"2018","cvss_score":6.5,"severity":"medium","kev":false,"impact":"A negative key length in a serialized dict makes the server read memory from elsewhere in the process into a returned value. The client gets back chunks of brick process memory, which on a shared brick can contain other tenants' file data and credentials.","attack_vector":"Any authenticated gluster client able to send a crafted RPC to a brick.","remediation":"Upgrade glusterfs to 4.1.4 / 3.12.x-fixed or later and restart the bricks. Rotate any secrets that lived in the brick process address space if you believe the flaw was exercised.","references":["https://access.redhat.com/security/cve/CVE-2018-10911","https://nvd.nist.gov/vuln/detail/CVE-2018-10911"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.0/AV:A/AC:L/PR:N/UI:N/S:U/C:N/I:H/A:N","cwe":["CWE-284"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2018-1129","cve":"CVE-2018-1129","aliases":[],"title":"Ceph CephX authentication protocol: The CephX signature calculation can be bypassed, so an on-path attacker can alter","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ceph CephX authentication protocol","year":"2018","cvss_score":6.5,"severity":"medium","kev":false,"impact":"The CephX signature calculation can be bypassed, so an on-path attacker can alter the payload of an authenticated Ceph message without the signature failing. In practice that means tampering with another tenant's in-flight RADOS operations - silent data corruption on a shared pool with no authentication failure logged.","attack_vector":"An on-path attacker on the Ceph public or cluster network who can modify frames in transit.","remediation":"Upgrade to a fixed Ceph release and restart all daemons. Turn on msgr2 secure mode so integrity is protected by real AEAD rather than the legacy signature, and keep the cluster network physically or logically separate from tenant traffic.","references":["https://access.redhat.com/security/cve/CVE-2018-1129","https://nvd.nist.gov/vuln/detail/CVE-2018-1129"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2018-12207","cve":"CVE-2018-12207","aliases":["iTLB multihit","Machine Check Error on Page Size Change","MCEPSC","No eXcuses"],"title":"Intel Core and Xeon CPUs - INTEL-SA-00210: This one is availability, not confidentiality, and it is the most","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel Core and Xeon CPUs - INTEL-SA-00210","year":"2018","cvss_score":6.5,"severity":"medium","kev":false,"impact":"This one is availability, not confidentiality, and it is the most operationally brutal item in the family. A malicious guest changes page size in its own page tables so that the instruction TLB holds multiple hits for one address; the CPU raises an unrecoverable machine check and the ENTIRE HOST hangs. One tenant VM, using only its own memory mappings, hard-kills the physical machine - and every other tenant on it. On a GPU host that means every co-resident training job loses its in-flight state back to the last checkpoint, and the node needs a physical or BMC-driven power cycle. It is a one-guest, no-privilege denial of service against the whole box.","attack_vector":"A privileged user inside any guest VM - i.e. root in a tenant's own VM, which on a bare-metal or VM-rental product is something every customer legitimately has. Does not require SMT or core sharing. Container tenants cannot reach it (no control of page tables); VM tenants can.","remediation":"Hypervisor patch, plus microcode on some parts. KVM's fix restricts large pages to non-executable mappings and splits huge pages down to 4K the moment a guest executes from them. That is the expensive part: losing 2MB/1GB EPT pages for guest code raises TLB pressure and costs memory-access performance on large-footprint guests, and it burns extra host memory on page tables. Linux exposes kvm.nx_huge_pages=off to reclaim it - do not set that on a multi-tenant host. If you rent whole machines to one tenant at a time, the blast radius is only that tenant and you may reasonably leave huge pages on. Recent Xeon generations are not affected.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-12207","https://docs.kernel.org/admin-guide/hw-vuln/multihit.html","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00210.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"published":"2019-11-14"},{"cvss_vector":"CVSS:3.0/AV:L/AC:L/PR:L/UI:N/S:C/C:N/I:N/A:H","cwe":["CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2018-1782","cve":"CVE-2018-1782","aliases":[],"title":"IBM GPFS kernel module (mmap path): An unprivileged user panics the kernel on a GPFS node just by mmap-ing a file on","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"IBM GPFS kernel module (mmap path)","year":"2018","cvss_score":6.5,"severity":"medium","kev":false,"impact":"An unprivileged user panics the kernel on a GPFS node just by mmap-ing a file on the filesystem or running a crafted binary stored there. Every job on that node dies with it, and the node needs a reboot.","attack_vector":"Local account with read access to the GPFS filesystem on a node running Spectrum Scale 5.0.1.0 or 5.0.1.1. Any tenant who can place a file on shared storage can trigger it on any node that opens it.","remediation":"Upgrade to 5.0.1.2 or later. Because the fault is in the kernel module, the fix needs the portability layer rebuilt and the node drained and rebooted, not just a daemon bounce.","references":["https://www.ibm.com/support/docview.wss?uid=ibm10730967","https://nvd.nist.gov/vuln/detail/CVE-2018-1782"],"status":"curated"},{"id":"CVE-2018-3979","cve":"CVE-2018-3979","aliases":[],"title":"Nouveau display driver (in-tree Linux nouveau, NV117): Remote denial of service against a workstation or node running","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Nouveau display driver (in-tree Linux nouveau, NV117)","year":"2018","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Remote denial of service against a workstation or node running the open-source Nouveau driver: a crafted pixel shader delivered through a web page wedges the GPU driver and takes the machine's graphics stack down. No code execution, but on a shared render or VDI host it is a free reboot for anyone who can get a browser to load their page.","attack_vector":"Anyone who can get a user on the host to open a web page - so effectively internet-reachable. No local account needed.","remediation":"This is the in-tree open-source Nouveau driver, not NVIDIA's proprietary stack. Update the distribution kernel (Ubuntu 18.04 shipped the vulnerable NV117 code) or, on GPU nodes, blacklist nouveau entirely and run the proprietary NVIDIA driver, which is what a compute fleet should be doing anyway. Kernel update means a node reboot.","references":["https://talosintelligence.com/vulnerability_reports/TALOS-2018-0647","https://nvd.nist.gov/vuln/detail/CVE-2018-3979"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2019-04-01"},{"id":"CVE-2019-1000008","cve":"CVE-2019-1000008","aliases":[],"title":"Helm: Path traversal in `helm fetch --untar` writes outside the target directory","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Helm","year":"2019","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Path traversal in `helm fetch --untar` writes outside the target directory","attack_vector":"Malicious chart","remediation":"Upgrade Helm on all CI and operator machines","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-1000008"],"status":"curated","published":"2019-02-04"},{"id":"CVE-2019-11135","cve":"CVE-2019-11135","aliases":["TAA","TSX Asynchronous Abort","ZombieLoad v2"],"title":"Intel CPUs supporting TSX, including Cascade Lake Xeon Scalable - INTEL-SA-00270: Same class of in-flight data leak","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel CPUs supporting TSX, including Cascade Lake Xeon Scalable - INTEL-SA-00270","year":"2019","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Same class of in-flight data leak as MDS, but reached through a TSX transaction abort - and crucially it works on Cascade Lake, the generation Intel shipped with the MDS_NO silicon fix. That is the operator lesson: the parts you bought specifically because they were 'the fixed ones' were still vulnerable. An attacker samples data another tenant or the host kernel is moving through the fill buffers. Only the attacker needs to use TSX; the victim does not. Cross-hyperthread, so on a shared-SMT fleet it is a direct tenant-boundary break yielding keys, tokens and credentials.","attack_vector":"Unprivileged local code on a TSX-capable CPU, in any guest or container. Cross-hyperthread attacks work because the fill buffers are shared between siblings.","remediation":"Microcode/BIOS update plus kernel patch - firmware flash, host reboot, job drain. Then a real choice: (a) tsx=off disables TSX entirely and fully closes it including cross-thread - this is the Linux default and the right call for a multi-tenant operator, and the only workloads that notice are the rare ones using hardware lock elision; or (b) keep TSX and rely on VERW buffer clearing, which leaves cross-hyperthread attacks possible unless you also disable SMT. Choose tsx=off. It costs almost nothing on AI infrastructure workloads and removes both the mitigation overhead and the SMT dilemma.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-11135","https://docs.kernel.org/admin-guide/hw-vuln/tsx_async_abort.html","https://zombieloadattack.com/"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2019-11-14"},{"id":"CVE-2019-11246","cve":"CVE-2019-11246","aliases":[],"title":"Kubernetes (kubectl): `kubectl cp` path traversal from a malicious container tar overwrites files on the operator's","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kubectl)","year":"2019","cvss_score":6.5,"severity":"medium","kev":false,"impact":"`kubectl cp` path traversal from a malicious container tar overwrites files on the operator's workstation","attack_vector":"Malicious image, triggered when an operator runs kubectl cp","remediation":"Upgrade kubectl on all operator and CI machines; not a cluster change","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-11246"],"status":"curated","published":"2019-08-29"},{"id":"CVE-2019-11249","cve":"CVE-2019-11249","aliases":[],"title":"Kubernetes (kubectl): Follow-up incomplete fix for the kubectl cp traversal","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kubectl)","year":"2019","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Follow-up incomplete fix for the kubectl cp traversal","attack_vector":"Malicious image","remediation":"Upgrade kubectl on operator and CI machines","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-11249"],"status":"curated","published":"2019-08-29"},{"id":"CVE-2019-11250","cve":"CVE-2019-11250","aliases":[],"title":"Kubernetes (client-go): Bearer tokens logged at verbosity 7+","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (client-go)","year":"2019","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Bearer tokens logged at verbosity 7+; credential disclosure via logs","attack_vector":"Anyone with log-pipeline read access","remediation":"Lower component verbosity; rotate service-account tokens; scrub log store","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-11250"],"status":"curated","published":"2019-08-29"},{"id":"CVE-2019-11254","cve":"CVE-2019-11254","aliases":[],"title":"Kubernetes (kube-apiserver): YAML parsing CPU exhaustion in the apiserver","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-apiserver)","year":"2019","cvss_score":6.5,"severity":"medium","kev":false,"impact":"YAML parsing CPU exhaustion in the apiserver","attack_vector":"Any authorized cluster user","remediation":"Rolling control-plane upgrade; no GPU drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-11254"],"status":"curated","published":"2020-04-01"},{"id":"CVE-2019-16097","cve":"CVE-2019-16097","aliases":[],"title":"Harbor: Non-admin users create admin accounts via POST /api/users","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Harbor","year":"2019","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Non-admin users create admin accounts via POST /api/users; full registry takeover","attack_vector":"Any registered registry user","remediation":"Upgrade Harbor; audit the admin user list","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-16097"],"status":"curated","published":"2019-09-08"},{"id":"CVE-2019-1890","cve":"CVE-2019-1890","aliases":[],"title":"Cisco Nexus 9000 ACI Mode Switch Software (fabric infrastructure VLAN): The earlier instance of the same ACI class of","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco Nexus 9000 ACI Mode Switch Software (fabric infrastructure VLAN)","year":"2019","cvss_score":6.5,"severity":"medium","kev":false,"impact":"The earlier instance of the same ACI class of bug — security validation on infrastructure-VLAN connection setup can be bypassed, letting an adjacent device join the fabric's internal VLAN. Worth tracking separately because the fixed releases differ and clusters running older ACI trains are exposed to this one and not the 2021 variant.","attack_vector":"Unauthenticated, adjacent — a device connected to a leaf port.","remediation":"ACI fabric software upgrade (APIC + switches). Staged, multi-hour. No config-only workaround beyond locking down unused ports.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-1890"],"status":"curated","tags":["tenant-isolation"],"published":"2019-07-04"},{"id":"CVE-2019-9901","cve":"CVE-2019-9901","aliases":[],"title":"Envoy: No URL path normalization, so `something/../admin` bypasses access control","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy","year":"2019","cvss_score":6.5,"severity":"medium","kev":false,"impact":"No URL path normalization, so `something/../admin` bypasses access control","attack_vector":"Unauthenticated network","remediation":"Upgrade Envoy; enable path normalization","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-9901"],"status":"curated","published":"2019-04-25"},{"id":"CVE-2020-15136","cve":"CVE-2020-15136","aliases":[],"title":"etcd: Gateway TLS authentication applied only to endpoints found in DNS SRV records","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"etcd","year":"2020","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Gateway TLS authentication applied only to endpoints found in DNS SRV records; unauthenticated etcd access","attack_vector":"Unauthenticated network reaching etcd","remediation":"Rolling etcd upgrade; enforce mutual TLS on all etcd peers and clients","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-15136"],"status":"curated","published":"2020-08-06"},{"id":"CVE-2020-24501","cve":"CVE-2020-24501","aliases":[],"title":"Intel E810 Ethernet Controller firmware: Buffer overflow in early E810 firmware, triggerable by an unauthenticated","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel E810 Ethernet Controller firmware","year":"2020","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Buffer overflow in early E810 firmware, triggerable by an unauthenticated adjacent attacker for denial of service. Worth carrying in an operator database because E810 cards that shipped in 2020-2021 server generations and were never NVM-updated are still in production fleets — NIC firmware is the layer operators most reliably forget to patch.","attack_vector":"Unauthenticated, adjacent — same L2 segment.","remediation":"Flash E810 firmware to 1.4.1.13 or later. Cold power cycle. Practically: audit your fleet's NVM versions first (`ethtool -i` reports the firmware-version string) — most operators discover a wide spread of versions and should batch the whole update rather than chase individual CVEs.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-24501"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-02-17"},{"id":"CVE-2020-24511","cve":"CVE-2020-24511","aliases":[],"title":"Intel processors (shared resource isolation): Improper isolation of shared processor resources allowing information","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (shared resource isolation)","year":"2020","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Improper isolation of shared processor resources allowing information disclosure to a local authenticated user. Fixed in the same microcode drop as the associated Atom domain-bypass issue; relevant to any multi-tenant host.","attack_vector":"Local authenticated code on the host.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-24511","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00464.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"tags":["tenant-isolation"],"published":"2021-06-09"},{"id":"CVE-2020-24513","cve":"CVE-2020-24513","aliases":["Vector Register Sampling on Atom","SRBDS-family"],"title":"Intel Atom processors (domain-bypass transient execution): A domain-bypass transient execution flaw on Atom parts","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Atom processors (domain-bypass transient execution)","year":"2020","cvss_score":6.5,"severity":"medium","kev":false,"impact":"A domain-bypass transient execution flaw on Atom parts leaking information across privilege domains. Matters for edge inference boxes and storage/management appliances built on Atom silicon rather than for Xeon compute nodes - but those appliances often sit inside the trusted network of an AI datacenter.","attack_vector":"Local authenticated code on an affected Atom platform.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-24513","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00465.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"tags":["tenant-isolation"],"published":"2021-06-09"},{"id":"CVE-2020-8766","cve":"CVE-2020-8766","aliases":[],"title":"Intel SGX DCAP (datacenter attestation primitives): An improper conditions check in DCAP lets an unauthenticated","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SGX DCAP (datacenter attestation primitives)","year":"2020","cvss_score":6.5,"severity":"medium","kev":false,"impact":"An improper conditions check in DCAP lets an unauthenticated adjacent attacker deny service to the attestation path. In a confidential-compute fleet, killing attestation means new workloads cannot start and existing ones cannot renew - an availability failure that looks like a control-plane outage.","attack_vector":"Unauthenticated attacker with adjacent network access to the attestation service - so anything on the same network segment as your PCCS/quote-generation service.","remediation":"Upgrade SGX DCAP to 1.6 or later and keep the caching service off flat networks. Userspace service update and restart; no node reboot or firmware.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-8766","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00398"],"status":"curated","published":"2020-11-12"},{"id":"CVE-2021-0009","cve":"CVE-2021-0009","aliases":[],"title":"Intel Ethernet 800 Series Controller firmware: Out-of-bounds read in 800-series (E810 family) adapter firmware","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Ethernet 800 Series Controller firmware","year":"2021","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Out-of-bounds read in 800-series (E810 family) adapter firmware, reachable by an unauthenticated user, causing denial of service. Listed separately from the later E810 issues because the fixed version target is different (1.5.3.0), so a fleet standardised on a mid-2021 NVM image is exposed to this one even if it is patched for the 2023-2025 batch.","attack_vector":"Unauthenticated, over the network to the adapter.","remediation":"Flash 800-series firmware to 1.5.3.0 or later; cold power cycle. In practice, set a single fleet-wide minimum NVM version well above all of these and enforce it in provisioning, rather than tracking each CVE's individual threshold.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-0009"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-08-11"},{"id":"CVE-2021-1115","cve":"CVE-2021-1115","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): NULL dereference reachable through private IOCTLs, with NVIDIA noting","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2021","cvss_score":6.5,"severity":"medium","kev":false,"impact":"NULL dereference reachable through private IOCTLs, with NVIDIA noting the denial of service lands in a component beyond the driver itself - the failure propagates past the GPU stack.","attack_vector":"Any local unprivileged user with GPU device access on a Windows host.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1115"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-10-27"},{"id":"CVE-2021-21285","cve":"CVE-2021-21285","aliases":[],"title":"Docker / moby: Malformed image manifest crashes dockerd","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2021","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Malformed image manifest crashes dockerd; node-level DoS","attack_vector":"Malicious image","remediation":"Upgrade Docker Engine; daemon restart","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-21285"],"status":"curated","published":"2021-02-02"},{"id":"CVE-2021-25735","cve":"CVE-2021-25735","aliases":[],"title":"Kubernetes (kube-apiserver): Node updates bypass a validating admission webhook, defeating node-level policy","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-apiserver)","year":"2021","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Node updates bypass a validating admission webhook, defeating node-level policy","attack_vector":"Cluster user able to update Node objects","remediation":"Rolling control-plane upgrade; no GPU drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-25735"],"status":"curated","published":"2021-09-06"},{"id":"CVE-2021-26341","cve":"CVE-2021-26341","aliases":[],"title":"AMD processors - transient execution beyond unconditional direct branches: Some AMD CPUs transiently execute","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD processors - transient execution beyond unconditional direct branches","year":"2021","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Some AMD CPUs transiently execute instructions past an unconditional direct branch - code that should never run, running speculatively and leaving traces in the cache. That gives an attacker speculative gadgets in places the compiler and the kernel's own Spectre auditing assume are unreachable, so hardened code can still leak. The practical outcome is data disclosure across privilege and guest boundaries.","attack_vector":"Local, from an unprivileged process or a guest VM.","remediation":"Mitigated by kernel-side changes that insert INT3 speculation barriers after unconditional branches in sensitive paths. Take the distro kernel update and reboot; no firmware step for the kernel mitigation itself. Mitigated by AMD microcode plus, on most of these, a kernel-side change - and the durable delivery vehicle is the OEM SBIOS/AGESA package, which carries **one to six months of OEM lag** and needs a drained node and a full power cycle. The linux-firmware amd-ucode blobs get you the microcode sooner via initramfs early-load and a reboot, but AMD does not support late-loading microcode on a running EPYC host, so either way this is reboot-required, not a live patch.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26341","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2022-03-11"},{"id":"CVE-2021-26921","cve":"CVE-2021-26921","aliases":[],"title":"Argo CD: Tokens keep working after the user account is disabled","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2021","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Tokens keep working after the user account is disabled; offboarding does not revoke access","attack_vector":"A former user with a cached token","remediation":"Rolling Argo CD upgrade; force-rotate all tokens","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26921"],"status":"curated","published":"2021-02-09"},{"id":"CVE-2021-31920","cve":"CVE-2021-31920","aliases":[],"title":"Istio: Multiple or escaped slashes bypass an Istio authorization policy","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Istio","year":"2021","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Multiple or escaped slashes bypass an Istio authorization policy","attack_vector":"Unauthenticated network","remediation":"Rolling istiod upgrade plus sidecar restart","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-31920"],"status":"curated","published":"2021-05-27"},{"id":"CVE-2021-3524","cve":"CVE-2021-3524","aliases":[],"title":"Ceph RGW: HTTP header injection via a newline in the CORS ExposeHeader tag","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ceph RGW","year":"2021","cvss_score":6.5,"severity":"medium","kev":false,"impact":"HTTP header injection via a newline in the CORS ExposeHeader tag","attack_vector":"Network (remote)","remediation":"Control-plane: RGW upgrade only; no OSD disruption","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-3524"],"status":"curated","published":"2021-05-17"},{"id":"CVE-2021-3979","cve":"CVE-2021-3979","aliases":[],"title":"Ceph: Key length incorrectly passed to the encryption algorithm","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ceph","year":"2021","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Key length incorrectly passed to the encryption algorithm -> non-random, weak key on encrypted disks","attack_vector":"Network (remote)","remediation":"Data-plane: storage node upgrade; re-encrypt affected RBD volumes","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-3979"],"status":"curated","published":"2022-08-25"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-863"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2021-43337","cve":"CVE-2021-43337","aliases":[],"title":"Slurm (slurmdbd, AccountingStoreFlags=job_script / job_env): When the site turns on job-script and job-environment","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Slurm (slurmdbd, AccountingStoreFlags=job_script / job_env)","year":"2021","cvss_score":6.5,"severity":"medium","kev":false,"impact":"When the site turns on job-script and job-environment archival, SlurmDBD's access-control rules let a user fetch other users' job scripts and environments. Those routinely contain API keys, model-registry tokens, S3 credentials and dataset paths, so this is a credential harvest across tenants, not just metadata leakage.","attack_vector":"Any user with a Slurm account on a cluster where AccountingStoreFlags includes job_script or job_env. 21.08.0 through 21.08.3 only - the feature did not exist before 21.08.","remediation":"Upgrade to Slurm 21.08.4 and restart slurmdbd. If you cannot upgrade now, remove job_script and job_env from AccountingStoreFlags and reconfigure - that removes the exposure immediately. Treat any secret that appeared in a job script or job env during the exposure window as burned and rotate it.","references":["https://lists.schedmd.com/pipermail/slurm-announce/2021/000068.html","https://nvd.nist.gov/vuln/detail/CVE-2021-43337"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2021-44776","cve":"CVE-2021-44776","aliases":[],"title":"Lanner IAC-AST2500A BMC firmware: The attacker rewrites who is permitted to use KVM and virtual media on the BMC","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lanner IAC-AST2500A BMC firmware","year":"2021","cvss_score":6.5,"severity":"medium","kev":false,"impact":"The attacker rewrites who is permitted to use KVM and virtual media on the BMC - which is to say they grant themselves console access and the ability to attach a boot image. Even at a medium score this is a direct route to the two BMC capabilities that matter most to an operator: watching a tenant's console, and booting the node into attacker-supplied media. It is also a quieter attack than the overflow bugs, because it leaves the BMC running normally with an altered permission set rather than crashing anything. Broken access control in the SubNet_handler_func function of spx_restservice, allowing an attacker to change the security access rights governing KVM and virtual media.","attack_vector":"Network access to the BMC REST service with sufficient standing to invoke the subnet handler. Reachable from the out-of-band management network.","remediation":"Firmware flash, subject to the same Lanner sourcing problem as the rest of the cluster. Because this bug manipulates configuration rather than corrupting memory, it leaves auditable state: check the KVM and virtual-media permission settings on every IAC-AST2500A BMC against your intended baseline, and alert on changes. Where fixed firmware is unobtainable, disabling virtual media and KVM outright on the BMC - where the platform permits it - removes the capability the attacker is trying to grant themselves.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-44776","https://www.nozominetworks.com/labs/vulnerability-advisories/cve-2021-44776/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-46744","cve":"CVE-2021-46744","aliases":["CipherLeaks","ciphertext side channel"],"title":"AMD SEV / SEV-ES / SEV-SNP - ciphertext observability: SEV encrypts guest memory deterministically per physical","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV / SEV-ES / SEV-SNP - ciphertext observability","year":"2021","cvss_score":6.5,"severity":"medium","kev":false,"impact":"SEV encrypts guest memory deterministically per physical address, so the same plaintext at the same address always produces the same ciphertext. A malicious hypervisor that can read the encrypted pages watches ciphertext blocks change over time and infers the plaintext values underneath - the CipherLeaks researchers recovered full RSA and ECDSA private keys out of the SEV-protected VMSA this way. This is the deepest structural problem in the SEV family: it is a consequence of the encryption mode, not a coding bug, so it cannot simply be patched away.","attack_vector":"Requires a malicious or compromised hypervisor with the ability to read guest ciphertext - the exact adversary SEV exists to defeat. No guest bug needed.","remediation":"Partially mitigated: AMD added ciphertext-hiding for the VMSA in SEV-SNP firmware, delivered through AGESA/SEV firmware as an OEM SBIOS package with the usual **one to six month lag** and a drained-node power cycle, plus a TCB bump that requires refreshing VCEK certificates. On newer parts, SEV-SNP Ciphertext Hiding can be enabled to block host reads of guest ciphertext outright - check whether your platform and firmware support it and turn it on. The residual risk on older EPYC generations is **not fully fixable**: guest software must avoid keeping secrets in memory patterns an observer can correlate, which is a burden you cannot impose on a tenant's workload. Be honest with confidential-computing customers about which EPYC generation their VM lands on.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-46744","https://arxiv.org/abs/2204.11669","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2022-05-11"},{"cwe":["CWE-843","CWE-787"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2021-46931","cve":"CVE-2021-46931","aliases":[],"title":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core/en): The transmit health reporter's dump callback casts its","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core/en)","year":"2021","cvss_score":6.5,"severity":"medium","kev":false,"impact":"The transmit health reporter's dump callback casts its argument to the wrong structure type on the TX-timeout recovery path. The type confusion walks bogus pointers, overflows the kernel stack past the guard page and panics the machine - a single stuck transmit queue turns into a fatal crash that takes every tenant on the node with it.","attack_vector":"The trigger is a TX timeout on an mlx5 queue, which is reachable through fabric-level conditions (severe congestion, link flap, a queue wedged by a heavy neighbour) rather than through a tenant device node. No privileges are required to be the workload that causes the stalled queue, but there is no direct attacker-controlled input - treat this as a shared-node availability and memory-safety defect, not a targeted exploit path.","remediation":"Update to a patched kernel on your stream. There is no configuration workaround short of disabling the devlink TX health reporter's dump on affected kernels; plan node reboots.","references":["https://git.kernel.org/stable/c/73665165b64a8f3c5b3534009a69be55bb744f05","https://git.kernel.org/stable/c/07f13d58a8ecc3baf9a488588fb38c5cb0db484f","https://nvd.nist.gov/vuln/detail/CVE-2021-46931"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-755","CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2021-47076","cve":"CVE-2021-47076","aliases":[],"title":"Linux kernel Soft-RoCE completer (rdma_rxe, invalid lkey handling in atomic operations): The local key is the RDMA","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel Soft-RoCE completer (rdma_rxe, invalid lkey handling in atomic operations)","year":"2021","cvss_score":6.5,"severity":"medium","kev":false,"impact":"The local key is the RDMA access-control token - it is what stops one queue pair touching another's registered memory. Soft-RoCE failed to record the WQE status when a LOCAL_WRITE failed, so an atomic operation submitted with a deliberately wrong lkey walked into a WARN and kernel panic instead of returning a completion error. The reachability is the point: supplying a bad lkey is the most basic thing an attacker probing RDMA isolation does, and on this driver it crashed the node rather than being rejected cleanly.","attack_vector":"Local, unprivileged - a tenant posts an atomic work request with an invalid lkey on a Soft-RoCE device.","remediation":"Kernel update returning a CQE error instead of falling through. Blacklist rdma_rxe on nodes with hardware RDMA that do not need software RoCE.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=a81e72d49b0648f68ab201163dbfc7841056d7a2","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2021/CVE-2021-47076.json"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-125","CWE-200"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2021-47255","cve":"CVE-2021-47255","aliases":[],"title":"Linux kernel (arch/x86/kvm): The guard against accessing bytes 4-15 of an emulated APIC register was dropped, and","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (arch/x86/kvm)","year":"2021","cvss_score":6.5,"severity":"medium","kev":false,"impact":"The guard against accessing bytes 4-15 of an emulated APIC register was dropped, and reading those offsets leaks host kernel stack contents straight back to the guest. A tenant VM gets a repeatable read primitive into host kernel memory - useful for defeating KASLR and for harvesting whatever else sits on that stack.","attack_vector":"Issued from inside the guest with no privileges beyond guest ring 0: read an emulated local-APIC register at a misaligned offset. Applies to guests running with the in-kernel LAPIC and without APIC virtualization handling the access.","remediation":"Update to a kernel with the referenced stable commits (no fixed release string published - match by commit). There is no practical interim control short of patching; the LAPIC is not something you can take away from a guest.","references":["https://git.kernel.org/stable/c/bf99ea52970caeb4583bdba1192c1f9b53b12c84","https://git.kernel.org/stable/c/018685461a5b9a9a70e664ac77aef0d7415a3fd5","https://nvd.nist.gov/vuln/detail/CVE-2021-47255"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2021-47265","cve":"CVE-2021-47265","aliases":[],"title":"Linux kernel RDMA core + mlx5_ib (ib_uverbs_ex_create_flow, flow steering rule creation): The port number a tenant","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel RDMA core + mlx5_ib (ib_uverbs_ex_create_flow, flow steering rule creation)","year":"2021","cvss_score":6.5,"severity":"medium","kev":false,"impact":"The port number a tenant supplies when creating an RDMA flow steering rule was never validated against the device's real port count before being handed to the driver. Flow steering is the mechanism that decides which packets land in which tenant's queue pair, so an unvalidated port index in the rule-creation path is both a kernel crash primitive (the mlx5_ib oops in the report) and a reason to distrust the boundary that is supposed to keep one tenant's traffic out of another's receive queues.","attack_vector":"Local ioctl on /dev/infiniband/uverbs* by any process allowed to create flow rules - which is any RDMA-capable tenant container. Unprivileged.","remediation":"Kernel update moving port validation into the core create_flow handler. No configuration workaround; RDMA flow steering cannot be selectively disabled without breaking RoCE traffic classification.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=2adcb4c5a52a2623cd2b43efa7041e74d19f3a5e","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2021/CVE-2021-47265.json"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2022-0001","cve":"CVE-2022-0001","aliases":["BHI","Spectre-BHB","Branch History Injection"],"title":"Intel processors (branch history injection): BHI / Spectre-BHB: even with eIBRS enabled, the branch history buffer is","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (branch history injection)","year":"2022","cvss_score":6.5,"severity":"medium","kev":false,"impact":"BHI / Spectre-BHB: even with eIBRS enabled, the branch history buffer is shared across privilege levels, so unprivileged code can steer kernel-side speculation and read kernel memory. This is the attack that showed hardware Spectre-v2 mitigations were not the end of the story, and it is directly a container-to-host and guest-to-host read primitive.","attack_vector":"Local unprivileged code - any container or VM on the node.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect. Linux additionally offers unprivileged-eBPF disabling and BHB-clearing sequences; check the spectre_v2 sysfs file after patching to see which mitigation actually engaged.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-0001","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00598.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"tags":["tenant-isolation"],"published":"2022-03-11"},{"id":"CVE-2022-0002","cve":"CVE-2022-0002","aliases":["Intra-mode BTI"],"title":"Intel processors (intra-mode branch target injection): The intra-mode sibling of BHI: branch predictor state is shared","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (intra-mode branch target injection)","year":"2022","cvss_score":6.5,"severity":"medium","kev":false,"impact":"The intra-mode sibling of BHI: branch predictor state is shared within a privilege level, so one sandboxed context can steer another's speculation without crossing rings. The concern is sandbox escape inside a single process - JIT tenants sharing a runtime.","attack_vector":"Local code inside the same privilege level as the victim, e.g. a sandboxed JIT tenant.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-0002","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00598.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"published":"2022-03-11"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:H/A:N","cwe":["CWE-732"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2022-22411","cve":"CVE-2022-22411","aliases":[],"title":"IBM Spectrum Scale Data Access Services (DAS): An authenticated DAS user inserts code that manipulates cluster","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"IBM Spectrum Scale Data Access Services (DAS)","year":"2022","cvss_score":6.5,"severity":"medium","kev":false,"impact":"An authenticated DAS user inserts code that manipulates cluster resources, because DAS runs with more permission than the caller should inherit. The user ends up changing shared cluster state rather than just their own data path.","attack_vector":"Any authenticated user of the Data Access Services layer in Spectrum Scale DAS 5.1.3.1 - typically the S3/object front end offered to tenants.","remediation":"Upgrade DAS to the fixed level in IBM's bulletin and restart the service. Review which service account DAS runs as and tighten it so an escape yields less.","references":["https://www.ibm.com/support/pages/node/6610277","https://nvd.nist.gov/vuln/detail/CVE-2022-22411"],"status":"curated"},{"id":"CVE-2022-23816","cve":"CVE-2022-23816","aliases":["RetBleed (AMD)","Branch Type Confusion"],"title":"AMD processors - branch predictor aliasing causing wrong branch type prediction (AMD-SB-1037): Aliases in the branch","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD processors - branch predictor aliasing causing wrong branch type prediction (AMD-SB-1037)","year":"2022","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Aliases in the branch predictor cause some AMD processors to predict the wrong branch type, which is AMD's form of the RetBleed / Branch Type Confusion problem. An attacker in one security domain trains the predictor so the victim's return or branch speculates to an attacker-chosen target, leaking data across process, kernel and guest boundaries. This is the AMD analogue people look for when they ask about Spectre-BHB, and it is the real answer.","attack_vector":"Local, cross-privilege and cross-guest. Reachable from any tenant workload on affected silicon.","remediation":"Needs both halves: AGESA/microcode from the OEM SBIOS package (**one to six months of lag**, drained node, power cycle) **and** an OS update that issues IBPB on context switch. Zen 1 and Zen 2 additionally depend on the LFENCE/JMP construct, which was itself found insufficient - so check that your kernel is using retpoline or IBRS rather than LFENCE/JMP by reading /sys/devices/system/cpu/vulnerabilities/spectre_v2 on the fleet. Patch alongside the other AMD-SB-1037 CVEs; they are one disclosure.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-23816","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-1037.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2023-01-17"},{"id":"CVE-2022-28199","cve":"CVE-2022-28199","aliases":[],"title":"NVIDIA MLNX_DPDK: Improper error recovery in NVIDIA's DPDK distribution lets a remote attacker cause denial of service","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA MLNX_DPDK","year":"2022","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Improper error recovery in NVIDIA's DPDK distribution lets a remote attacker cause denial of service with some integrity and confidentiality impact. Relevant on DPDK-based storage and network dataplanes fronting GPU clusters.","attack_vector":"Network - the attacker sends traffic that the DPDK dataplane mishandles. No authentication involved; this is packet-level reachability.","remediation":"Update MLNX_DPDK per bulletin 5389 and restart the dataplane application. Cost: a dataplane restart drops in-flight connections; on a storage path that means an I/O stall visible to running jobs, so drain or fail over first.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28199","https://github.com/NVIDIA/product-security/tree/main/2022/5389"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-1284"],"fleet":{"pain_class":"node-drain"},"published":"2022-09-01"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:H/A:N","cwe":["CWE-298","CWE-613"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2022-31145","cve":"CVE-2022-31145","aliases":["GHSA-qwrj-9hmp-gpxh"],"title":"FlyteAdmin (external IdP access token / ID token expiration check): FlyteAdmin does not enforce expiry on access and ID","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"FlyteAdmin (external IdP access token / ID token expiration check)","year":"2022","cvss_score":6.5,"severity":"medium","kev":false,"impact":"FlyteAdmin does not enforce expiry on access and ID tokens issued by an external identity provider, so a token stays usable after the user's session should have ended. Offboarded users and stolen tokens keep working, letting someone launch and inspect workflows on the cluster long after their access was supposed to be revoked.","attack_vector":"Anyone holding an expired but otherwise valid token from the external IdP, including a token pulled from a browser, a CI log or a laptop after offboarding. Deployments using flyteadmin itself as the OAuth2 authorization server are unaffected.","remediation":"Upgrade FlyteAdmin to 1.1.30 or later and restart it. Until the upgrade lands, rotate the signing keys repeatedly - each rotation invalidates all open sessions and forces re-authentication, which is the only way to expire the outstanding tokens.","references":["https://github.com/flyteorg/flyteadmin/security/advisories/GHSA-qwrj-9hmp-gpxh","https://nvd.nist.gov/vuln/detail/CVE-2022-31145"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2022-3162","cve":"CVE-2022-3162","aliases":[],"title":"Kubernetes (kube-apiserver): Users authorized to list/watch one namespaced CR type can read other CR types in the same","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-apiserver)","year":"2022","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Users authorized to list/watch one namespaced CR type can read other CR types in the same API group cluster-wide; cross-tenant data exposure","attack_vector":"Cluster user with namespace access","remediation":"Rolling control-plane upgrade; no GPU drain. Audit CRD RBAC","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-3162"],"status":"curated","published":"2023-03-01"},{"id":"CVE-2022-3287","cve":"CVE-2022-3287","aliases":[],"title":"fwupd's Redfish plugin: Any unprivileged local user on the host can read a working BMC credential out of a config file","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"fwupd's Redfish plugin","year":"2022","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Any unprivileged local user on the host can read a working BMC credential out of a config file. That is a direct escalation from 'has a shell on the node' to 'has an account on the node's out-of-band controller', which is the boundary a bare-metal operator is selling. From there the attacker reaches the BMC's Redfish surface with a legitimate operator account and can chain into any of the authenticated BMC bugs elsewhere in this list. If your provisioning tooling deploys fwupd with the Redfish plugin across the fleet, the same class of credential is sitting on every node. When it creates an OPERATOR account on the BMC, it writes the auto-generated password into /etc/fwupd/redfish.conf without restricting the file's permissions, leaving BMC credentials world-readable on the host.","attack_vector":"Any local unprivileged account on a host running fwupd with the Redfish plugin enabled. No root, no network position on the management VLAN, no exploit - just file read. On rented bare metal, the tenant is that local account.","remediation":"Update fwupd to a version carrying the permissions fix. Because the credential has already been written in the clear on existing installs, updating the package is not sufficient: you must also rotate the BMC operator account fwupd created on every affected node, and check the permissions of /etc/fwupd/redfish.conf directly rather than trusting the package version. This is config-and-credential work rather than a firmware flash, so it is cheap to remediate but easy to leave half-done.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-3287","https://github.com/fwupd/fwupd/commit/ea676855f2119e36d433fbd2ed604039f53b2091"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-34665","cve":"CVE-2022-34665","aliases":[],"title":"NVIDIA GPU Display Driver - kernel mode layer (Windows nvlddmkm.sys and Linux nvidia.ko): A null-pointer dereference","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - kernel mode layer (Windows nvlddmkm.sys and Linux nvidia.ko)","year":"2022","cvss_score":6.5,"severity":"medium","kev":false,"impact":"A null-pointer dereference in the kernel mode layer with a changed CVSS scope crashes the node from an unprivileged account. Both the Windows and Linux datacenter drivers are affected, so a mixed fleet needs two separate rollouts.","attack_vector":"Local and unprivileged on either OS. On Linux it is reachable from any GPU container via /dev/nvidia*; on Windows from any session holding a GPU handle.","remediation":"Upgrade both the Linux and the Windows datacenter driver branches listed in bulletin 5383. Cost: Linux needs a drain and nvidia.ko reload per node; Windows needs a reboot per node. Two change windows unless your fleet is homogeneous.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34665","https://github.com/NVIDIA/product-security/tree/main/2022/5383"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"published":"2022-11-19"},{"id":"CVE-2022-34666","cve":"CVE-2022-34666","aliases":[],"title":"NVIDIA GPU Display Driver - kernel mode layer (Windows nvlddmkm.sys and Linux nvidia.ko): A second null-pointer","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - kernel mode layer (Windows nvlddmkm.sys and Linux nvidia.ko)","year":"2022","cvss_score":6.5,"severity":"medium","kev":false,"impact":"A second null-pointer dereference path in the kernel mode layer, same unprivileged local reach. Both the Windows and Linux datacenter drivers are affected, so a mixed fleet needs two separate rollouts.","attack_vector":"Local and unprivileged on either OS. On Linux it is reachable from any GPU container via /dev/nvidia*; on Windows from any session holding a GPU handle.","remediation":"Upgrade both the Linux and the Windows datacenter driver branches listed in bulletin 5383. Cost: Linux needs a drain and nvidia.ko reload per node; Windows needs a reboot per node. Two change windows unless your fleet is homogeneous.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34666","https://github.com/NVIDIA/product-security/tree/main/2022/5383"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"published":"2022-11-10"},{"id":"CVE-2022-34678","cve":"CVE-2022-34678","aliases":[],"title":"NVIDIA GPU Display Driver - kernel mode layer (Windows nvlddmkm.sys and Linux nvidia.ko): An unprivileged user","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - kernel mode layer (Windows nvlddmkm.sys and Linux nvidia.ko)","year":"2022","cvss_score":6.5,"severity":"medium","kev":false,"impact":"An unprivileged user null-pointer-dereferences the kernel mode layer, scope-changed, and panics the node. Both the Windows and Linux datacenter drivers are affected, so a mixed fleet needs two separate rollouts.","attack_vector":"Local and unprivileged on either OS. On Linux it is reachable from any GPU container via /dev/nvidia*; on Windows from any session holding a GPU handle.","remediation":"Upgrade both the Linux and the Windows datacenter driver branches listed in bulletin 5415. Cost: Linux needs a drain and nvidia.ko reload per node; Windows needs a reboot per node. Two change windows unless your fleet is homogeneous.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34678","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"published":"2022-12-30"},{"id":"CVE-2022-36055","cve":"CVE-2022-36055","aliases":[],"title":"Helm: OOM panic in the strvals package from crafted `--set` input","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Helm","year":"2022","cvss_score":6.5,"severity":"medium","kev":false,"impact":"OOM panic in the strvals package from crafted `--set` input","attack_vector":"Anyone who can supply chart values, e.g. through a self-service portal","remediation":"Upgrade Helm; validate tenant-supplied values","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-36055"],"status":"curated","published":"2022-09-01"},{"id":"CVE-2022-40716","cve":"CVE-2022-40716","aliases":[],"title":"HashiCorp Consul: Internal RPC endpoint does not check multiple SAN URIs in a CSR","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HashiCorp Consul","year":"2022","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Internal RPC endpoint does not check multiple SAN URIs in a CSR -> bypass service-mesh intentions","attack_vector":"Network (remote)","remediation":"Control-plane: Consul server upgrade then clients; re-verify mesh intentions","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-40716"],"status":"curated","published":"2022-09-23"},{"id":"CVE-2022-40982","cve":"CVE-2022-40982","aliases":[],"title":"Intel CPU (Downfall / GDS): Downfall: Gather Data Sampling leaks AVX gather-instruction data across SMT siblings","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel CPU (Downfall / GDS)","year":"2022","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Downfall: Gather Data Sampling leaks AVX gather-instruction data across SMT siblings, containers and VMs - directly breaks multi-tenant isolation","attack_vector":"Any tenant process in a container; tenant VM guest","remediation":"Microcode update + reboot, standing perf cost (reported up to ~50% on gather-heavy vector code). Alternative is disabling AVX gather, which is worse for AI workloads. On shared-GPU nodes with co-tenanted CPUs this is a must-fix","references":["https://access.redhat.com/security/cve/CVE-2022-40982"],"status":"curated","fleet":{"ubiquity":"very common - Skylake through Tiger Lake era Xeons, still the host CPU under a large installed base of GPU nodes","remediation_pain":"microcode+reboot - microcode is loaded at boot, so every node drains and reboots; the mitigation carries a measurable AVX2/AVX-512 gather slowdown, i.e. a permanent throughput tax on the fleet","pain_class":"microcode + reboot","why_fleet_wide":"Cross-tenant data leakage from stale vector registers on shared hardware - exactly the isolation property a multi-tenant GPU cloud sells - so it forces a fleet-wide reboot campaign regardless of workload."},"published":"2023-08-11"},{"id":"CVE-2022-42282","cve":"CVE-2022-42282","aliases":[],"title":"DGX-2 BMC: Info disclosure via path traversal","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX-2 BMC","year":"2022","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Info disclosure via path traversal","attack_vector":"Network-adjacent authenticated","remediation":"Flash DGX-2 BMC firmware","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42282","https://github.com/NVIDIA/product-security/tree/main/2024/5513"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-22"],"fleet":{"pain_class":"firmware-flash"},"published":"2023-01-13"},{"cwe":["CWE-843","CWE-674"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-49732","cve":"CVE-2022-49732","aliases":[],"title":"Linux kernel (net/tls): A BPF sockmap psock could be attached to a socket that already had the kTLS ULP installed. The","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/tls)","year":"2022","cvss_score":6.5,"severity":"medium","kev":false,"impact":"A BPF sockmap psock could be attached to a socket that already had the kTLS ULP installed. The psock teardown path then unwinds the ULP and calls tcp_update_ulp() with the ULP's own protocol ops, so the ULP loops its own callbacks - a protocol-ops confusion between two layers that both think they own sk->sk_prot, ending in runaway recursion on the node.","attack_vector":"Needs the ability to add a socket to a BPF sockmap (CAP_BPF/CAP_NET_ADMIN - in practice the node's CNI, a service-mesh sidecar, or a tenant granted BPF), against a socket that already has kTLS enabled. Both halves are normal in an encrypted service mesh, so the collision arises without an attacker if kTLS and sockmap are used together; a tenant with BPF access can force it deliberately.","remediation":"Boot a kernel carrying the linked stable commits. Interim: do not grant CAP_BPF to tenant containers, and do not run sockmap/sk_msg policy over kTLS sockets on the same node.","references":["https://git.kernel.org/stable/c/72fa0f65b56605b8a9ae9fba2082f2123f7fe017","https://git.kernel.org/stable/c/922309e50befb0cfa5cb65e4989b7706d6578846","https://nvd.nist.gov/vuln/detail/CVE-2022-49732"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2023-0204","cve":"CVE-2023-0204","aliases":[],"title":"ConnectX-5/6/6-DX NIC firmware: NIC DoS (improper exception handling)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"ConnectX-5/6/6-DX NIC firmware","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"NIC DoS (improper exception handling)","attack_vector":"Any unprivileged tenant with the NIC exposed (SR-IOV VF)","remediation":"Flash NIC firmware to 35.1012+; requires node reboot, evict tenants","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5459/5459.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:N/I:N/A:H","cwe":["CWE-703"],"fleet":{"pain_class":"node-reboot"},"published":"2023-04-22"},{"id":"CVE-2023-20575","cve":"CVE-2023-20575","aliases":[],"title":"AMD processors - power reporting side channel against SEV VMs: An authenticated attacker uses the platform's power","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD processors - power reporting side channel against SEV VMs","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"An authenticated attacker uses the platform's power reporting functionality to monitor execution inside an AMD SEV VM. The whole promise of SEV is that the host cannot see what the confidential guest is doing; power telemetry is a channel the memory encryption does not cover, so a host operator watches the guest's execution profile through the power meter. For anyone selling confidential computing on EPYC this is a direct hole in the product claim.","attack_vector":"Local, authenticated, with access to power reporting interfaces on a host running SEV guests.","remediation":"Mitigated by restricting access to power reporting interfaces and by AGESA-level changes to reduce telemetry resolution. Mitigated by AMD microcode plus, on most of these, a kernel-side change - and the durable delivery vehicle is the OEM SBIOS/AGESA package, which carries **one to six months of OEM lag** and needs a drained node and a full power cycle. The linux-firmware amd-ucode blobs get you the microcode sooner via initramfs early-load and a reboot, but AMD does not support late-loading microcode on a running EPYC host, so either way this is reboot-required, not a live patch. The immediately actionable step is to lock down who can read host power telemetry on confidential-computing nodes - that is a permissions change, not a maintenance window.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-20575","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2023-07-11"},{"id":"CVE-2023-20591","cve":"CVE-2023-20591","aliases":[],"title":"AMD IOMMU - not re-initialized during DRTM (AMD-SB-3003): The IOMMU is not re-initialized during a Dynamic Root of","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD IOMMU - not re-initialized during DRTM (AMD-SB-3003)","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"The IOMMU is not re-initialized during a Dynamic Root of Trust for Measurement launch, so stale DMA mappings survive into the measured environment and an attacker can read or modify **hypervisor** memory via a device. DRTM exists precisely to establish a clean, measured starting state; leaving the IOMMU stale means the thing you are measuring can already be under device-level attack. On a GPU host, where accelerators and RDMA NICs are DMA-capable and numerous, that is a large set of devices to leave pointed at hypervisor memory.","attack_vector":"Local, requires the ability to set up DMA mappings before a DRTM launch and to drive a device afterwards.","remediation":"Fixed in AMD PI/AGESA firmware and delivered only as an OEM SBIOS package - AMD ships the PI drop to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before releasing BIOS. **Budget one to six months of OEM lag**, and note that several CVEs in this batch are marked 'no fix planned' on Naples (EPYC 7001) - for those the only remediation is retiring the hardware. Applying it means cordon, drain and a full power cycle per node; there is no driver reload, no live patch and no VBIOS step.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-20591","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3003.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"tags":["tenant-isolation"],"published":"2024-08-13"},{"id":"CVE-2023-20593","cve":"CVE-2023-20593","aliases":[],"title":"AMD CPU (Zenbleed): Zenbleed: cross-process/cross-VM register-file data leak on Zen 2 at ~30 kB/s per core, no special","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"AMD CPU (Zenbleed)","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Zenbleed: cross-process/cross-VM register-file data leak on Zen 2 at ~30 kB/s per core, no special privileges","attack_vector":"Any tenant process in a container; tenant VM guest","remediation":"AMD microcode (AGESA) update + reboot; kernel chicken-bit workaround (DE_CFG[9]) available with a measured perf cost. Zen 2 EPYC is still common as the CPU side of A100/L40S nodes","references":["https://access.redhat.com/security/cve/CVE-2023-20593"],"status":"curated","fleet":{"ubiquity":"common - Zen 2 EPYC (Rome) still hosts a large installed base of GPU nodes and rental fleets","remediation_pain":"microcode+reboot for the real fix; the interim DE_CFG MSR chicken-bit workaround is a kernel change with a measurable FP/vector performance cost - so the fleet either reboots for microcode or eats a permanent tax","pain_class":"microcode + reboot","why_fleet_wide":"Leaks ~30 KB/s/core of stale vector-register data across any privilege boundary including cross-process and cross-VM, i.e. exactly the co-tenancy isolation a GPU cloud sells, on every Rome host at once."},"published":"2023-07-24"},{"id":"CVE-2023-22276","cve":"CVE-2023-22276","aliases":[],"title":"Intel Ethernet Controller E810 Series firmware: A race condition in E810 firmware lets an authenticated local user","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Ethernet Controller E810 Series firmware","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"A race condition in E810 firmware lets an authenticated local user cause denial of service. In a bare-metal multi-tenant cluster 'authenticated local' is the tenant, so this is a tenant able to kill its own node's NIC — and on shared-NIC designs, potentially the NIC serving other functions on the same host.","attack_vector":"Authenticated local access to the host. In a bare-metal GPU rental model, that is the customer.","remediation":"Flash E810 firmware to 1.7.2.4 or later; cold power cycle. This one is a good argument for reflashing NIC firmware as part of tenant handoff rather than only on a CVE-driven schedule.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-22276"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2023-08-11"},{"id":"CVE-2023-22497","cve":"CVE-2023-22497","aliases":[],"title":"Netdata: Agent MACHINE GUID is readable and reusable","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Netdata","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Agent MACHINE GUID is readable and reusable -> impersonate an agent to the registry/parent","attack_vector":"Network (remote)","remediation":"Data-plane: Netdata agents run on GPU nodes - fleet-wide agent upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-22497"],"status":"curated","published":"2023-01-14"},{"id":"CVE-2023-25526","cve":"CVE-2023-25526","aliases":[],"title":"Cumulus Linux (neighmgrd/nlmanager): Switch DoS via crafted packet","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Cumulus Linux (neighmgrd/nlmanager)","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Switch DoS via crafted packet","attack_vector":"Network-adjacent attacker (any tenant on the L2 domain)","remediation":"Upgrade Cumulus Linux to 5.5.0+; rolling switch upgrade","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5480/5480.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-248"],"published":"2023-09-20"},{"id":"CVE-2023-25532","cve":"CVE-2023-25532","aliases":[],"title":"DGX H100 BMC (IPMI): Info disclosure of credentials","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX H100 BMC (IPMI)","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Info disclosure of credentials","attack_vector":"Network-adjacent IPMI client","remediation":"Flash BMC 23.08.18; rotate credentials","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5473/5473.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-522"],"published":"2023-09-20"},{"id":"CVE-2023-26054","cve":"CVE-2023-26054","aliases":[],"title":"BuildKit: Git URL credentials in a build request are persisted into the build cache and can be read by other builds","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"BuildKit","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Git URL credentials in a build request are persisted into the build cache and can be read by other builds","attack_vector":"Anyone sharing a multi-tenant builder","remediation":"Upgrade BuildKit; purge shared cache; rotate git credentials","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-26054"],"status":"curated","published":"2023-03-06"},{"id":"CVE-2023-2727","cve":"CVE-2023-2727","aliases":[],"title":"Kubernetes (kube-apiserver): Ephemeral containers bypass the ImagePolicyWebhook, so unapproved images run","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-apiserver)","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Ephemeral containers bypass the ImagePolicyWebhook, so unapproved images run","attack_vector":"Cluster user with namespace access","remediation":"Rolling control-plane upgrade; no GPU drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-2727"],"status":"curated","published":"2023-07-03"},{"id":"CVE-2023-2728","cve":"CVE-2023-2728","aliases":[],"title":"Kubernetes (kube-apiserver): Ephemeral containers bypass the ServiceAccount mountable-secrets policy","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-apiserver)","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Ephemeral containers bypass the ServiceAccount mountable-secrets policy; a tenant mounts secrets they were denied","attack_vector":"Cluster user with namespace access","remediation":"Rolling control-plane upgrade; no GPU drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-2728"],"status":"curated","published":"2023-07-03"},{"id":"CVE-2023-27595","cve":"CVE-2023-27595","aliases":[],"title":"Cilium: On agent start, eBPF programs are briefly detached, so traffic bypasses NetworkPolicy","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"On agent start, eBPF programs are briefly detached, so traffic bypasses NetworkPolicy","attack_vector":"Any pod on the cluster network during an agent restart","remediation":"Upgrade Cilium; note that every agent restart, including your own upgrades, is a policy-gap window","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-27595"],"status":"curated","published":"2023-03-17"},{"id":"CVE-2023-28376","cve":"CVE-2023-28376","aliases":["INTEL-SA-00869"],"title":"Intel E810 Ethernet Controller firmware: Out-of-bounds read in E810 firmware reachable from an adjacent","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel E810 Ethernet Controller firmware","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Out-of-bounds read in E810 firmware reachable from an adjacent, unauthenticated attacker, causing denial of service. Adjacent means the same L2 segment — in a leaf/spine cluster that is every other node in the rack, including other tenants' nodes if you share a VLAN. One compromised tenant machine can walk the rack knocking NICs offline.","attack_vector":"Unauthenticated, adjacent — a host on the same layer-2 segment as the target NIC.","remediation":"Flash E810 firmware to 1.7.1 or later. Cold power cycle required for the NVM image to activate; plan a per-node drain. If you cannot patch immediately, keeping tenants on separate VLANs limits who is 'adjacent' — a switch config change that meaningfully shrinks the exposed set.","references":["https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00869.html","https://nvd.nist.gov/vuln/detail/CVE-2023-28376"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2023-11-14"},{"id":"CVE-2023-28746","cve":"CVE-2023-28746","aliases":["RFDS","Register File Data Sampling"],"title":"Intel processors (register file data sampling): RFDS: stale data left in the integer, floating-point and vector","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (register file data sampling)","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"RFDS: stale data left in the integer, floating-point and vector register files after transient execution can be sampled by a local attacker, crossing the process, VM and enclave boundaries. Vector register files are where model activations and weights live during compute, so on an AI host this leaks the workload's actual data, not just pointers.","attack_vector":"Local authenticated code on an affected processor, including a co-tenant VM or container.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect. The kernel-side mitigation reuses the VERW buffer-clearing path, so a current kernel plus current microcode is the whole story; check the reg_file_data_sampling sysfs entry after reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-28746","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00898.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"tags":["tenant-isolation"],"published":"2024-03-14"},{"id":"CVE-2023-2878","cve":"CVE-2023-2878","aliases":[],"title":"secrets-store-csi-driver: Service account tokens written to driver logs","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"secrets-store-csi-driver","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Service account tokens written to driver logs","attack_vector":"Anyone with log-pipeline read access","remediation":"Upgrade the driver via DaemonSet rollout; rotate the exposed service-account tokens; scrub logs","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated","published":"2023-06-07"},{"id":"CVE-2023-30575","cve":"CVE-2023-30575","aliases":[],"title":"Apache Guacamole: Miscalculated instruction lengths during the Guacamole handshake","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Apache Guacamole","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Miscalculated instruction lengths during the Guacamole handshake -> protocol instruction injection","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade guacd and the webapp; console gateways are high-value","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-30575"],"status":"curated","published":"2023-06-07"},{"id":"CVE-2023-31018","cve":"CVE-2023-31018","aliases":[],"title":"GPU Display Driver (Linux + Windows KMD): Host DoS (null deref from unprivileged user)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver (Linux + Windows KMD)","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Host DoS (null deref from unprivileged user)","attack_vector":"Any tenant with a container holding /dev/nvidia*","remediation":"Driver upgrade; rolling node reboot","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5491/5491.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"published":"2023-11-02"},{"id":"CVE-2023-31025","cve":"CVE-2023-31025","aliases":[],"title":"DGX A100 BMC: LDAP injection in BMC auth","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX A100 BMC","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"LDAP injection in BMC auth","attack_vector":"Network-adjacent attacker","remediation":"Flash BMC 00.22.05+","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31025","https://github.com/NVIDIA/product-security/tree/main/2024/5510"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-90"],"published":"2024-01-12"},{"id":"CVE-2023-31419","cve":"CVE-2023-31419","aliases":[],"title":"Elasticsearch: Crafted _search query string","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Elasticsearch","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Crafted _search query string -> stack overflow and node denial of service","attack_vector":"Network (remote)","remediation":"Control-plane: rolling upgrade of the log cluster","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31419"],"status":"curated","published":"2023-10-26"},{"id":"CVE-2023-34345","cve":"CVE-2023-34345","aliases":["AMI-SA-2023005","NVIDIA OSR review"],"title":"AMI MegaRAC SPx (SPX REST API): Path traversal in the BMC REST API letting a low-privilege user read arbitrary files","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx (SPX REST API)","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Path traversal in the BMC REST API letting a low-privilege user read arbitrary files off the BMC filesystem. That is where the interesting material lives: the shadow file, TLS private keys, SSH host keys, IPMI user database and stored configuration. It is a credential-harvesting step - the attacker takes what they read here and uses it to escalate on this BMC and, because BMC credentials are usually cloned across a fleet, on every sibling node.","attack_vector":"Network-reachable REST API with only a low-privilege BMC account. That is a much lower bar than the admin-required bugs in the same advisory - a read-only telemetry or monitoring account is sufficient, and those are exactly the accounts operators hand out widely and rarely rotate.","remediation":"Firmware flash to SPx_12.5 / SPx_13.3 or later, out-of-band per node, ODM-gated. Because the exploit path starts from a low-privilege account, the immediate config-only step is to inventory BMC accounts and delete or rotate every stale, shared or monitoring credential - and to assume any credential material stored on an unpatched BMC has already been read, and rotate it rather than leaving it in place.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2023005.pdf","https://nvd.nist.gov/vuln/detail/CVE-2023-34345"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-06-12"},{"id":"CVE-2023-40285","cve":"CVE-2023-40285","aliases":[],"title":"Supermicro BMC (IPMI web interface, XSS): Another injection point in the BMC web interface, lower-impact than","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC (IPMI web interface, XSS)","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Another injection point in the BMC web interface, lower-impact than its siblings but usable for the same session-hijack chain toward virtual media and firmware flash.","attack_vector":"Network reach to the BMC web UI plus an operator loading the affected page.","remediation":"BMC firmware flash per board; same batch as the rest of the 2023 Supermicro BMC advisories, so fix them together rather than one at a time.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-40285"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-03-27"},{"id":"CVE-2023-40584","cve":"CVE-2023-40584","aliases":[],"title":"Argo CD: repo-server extracts a user-controlled tar.gz without size validation","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Argo CD","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"repo-server extracts a user-controlled tar.gz without size validation -> denial of service","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; isolate and resource-cap the repo-server","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-40584"],"status":"curated","published":"2023-09-07"},{"id":"CVE-2023-4091","cve":"CVE-2023-4091","aliases":[],"title":"Samba: SMB client can truncate files despite read-only permissions when acl_xattr ignores system ACLs","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Samba","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"SMB client can truncate files despite read-only permissions when acl_xattr ignores system ACLs","attack_vector":"Network (remote)","remediation":"Data-plane: smbd upgrade + review acl_xattr share config","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-4091"],"status":"curated","published":"2023-11-03"},{"id":"CVE-2023-43040","cve":"CVE-2023-43040","aliases":[],"title":"Ceph RGW (IBM Spectrum Fusion HCI): Improper bucket access lets an actor perform unauthorized actions in RGW","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ceph RGW (IBM Spectrum Fusion HCI)","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Improper bucket access lets an actor perform unauthorized actions in RGW","attack_vector":"Network (remote)","remediation":"Control-plane: RGW/appliance firmware upgrade; re-audit bucket policies","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-43040"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-05-14"},{"id":"CVE-2023-45229","cve":"CVE-2023-45229","aliases":["PixieFail","VU#132380"],"title":"EDK II NetworkPkg (DHCPv6 Advertise, IA_NA/IA_TA option parsing): An integer underflow when parsing","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"EDK II NetworkPkg (DHCPv6 Advertise, IA_NA/IA_TA option parsing)","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"An integer underflow when parsing the identity-association options in a DHCPv6 Advertise causes the firmware to read outside its buffer. On its own this leaks firmware memory contents or crashes the boot; chained with the overflow bugs in the same advertisement path it is the information-leak half of a reliable pre-OS exploit (defeating whatever address-layout guesswork the attacker would otherwise need).","attack_vector":"Anyone able to send DHCPv6 Advertise messages on the segment the node PXE-boots from. Unauthenticated, pre-OS.","remediation":"Firmware flash via the server OEM's BIOS package - the fix is in upstream edk2 but only reaches you after the IBV rebase and the OEM's own validation cycle. Reboot per node. Config-only stopgap: disable IPv6 network boot, or PXE entirely, and treat the provisioning VLAN as a trust boundary that tenant workloads must never reach.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-45229","https://kb.cert.org/vuls/id/132380","https://github.com/tianocore/edk2/security/advisories/GHSA-hc6x-cw6p-gj7h"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-01-16"},{"id":"CVE-2023-45231","cve":"CVE-2023-45231","aliases":["PixieFail","VU#132380"],"title":"EDK II NetworkPkg (IPv6 Neighbor Discovery Redirect handling): A truncated ND Redirect message drives an out-of-bounds","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"EDK II NetworkPkg (IPv6 Neighbor Discovery Redirect handling)","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"A truncated ND Redirect message drives an out-of-bounds read in the firmware IPv6 stack. Practical outcome is firmware memory disclosure or a wedged boot; the more interesting operational consequence is that the Redirect path itself lets an on-link attacker steer where the booting node sends its traffic, so this is both a leak and a foothold for redirecting the netboot fetch.","attack_vector":"On-link IPv6 attacker on the provisioning segment - any host that can emit ICMPv6 Neighbor Discovery to the booting node. Unauthenticated, pre-OS.","remediation":"OEM BIOS update, flash + reboot per node; the IBV-to-OEM rebase lag applies. Interim: enable IPv6 RA Guard / ND inspection on the provisioning switches, and disable the UEFI IPv6 network stack on nodes that boot locally. No OS-level or config-in-firmware toggle short of turning network boot off.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-45231","https://kb.cert.org/vuls/id/132380","https://github.com/tianocore/edk2/security/advisories/GHSA-hc6x-cw6p-gj7h"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-01-16"},{"id":"CVE-2023-4969","cve":"CVE-2023-4969","aliases":["LeftoverLocals","VU#446598"],"title":"GPU local/shared memory not cleared between kernels (AMD, Apple, Qualcomm, Imagination): A GPU kernel reads whatever","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"GPU local/shared memory not cleared between kernels (AMD, Apple, Qualcomm, Imagination)","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"A GPU kernel reads whatever the previous kernel left in local (shared/scratchpad) memory, including a kernel belonging to a different user, process or container. Trail of Bits recovered another process's LLM inference output token by token - on an AMD Radeon RX 7900 XT the leak was around 5.5 MB per GPU invocation, roughly 181 MB per query against a 7B model on llama.cpp. That is enough to reconstruct prompts, activations and responses, not just fragments. Affected vendors are AMD, Apple, Qualcomm and Imagination; NVIDIA, Intel and Arm tested clean and told CERT/CC they were not impacted. If you run AMD Instinct or ROCm, this is your bug and AMD was still investigating mitigations at disclosure.","attack_vector":"A co-tenant. The attacker needs only to run an ordinary OpenCL/Vulkan/Metal compute kernel on the same physical GPU - no privileges, no kernel exploit, no driver bug in the usual sense. Time-sliced sharing, MPS-style sharing and sequential job scheduling on the same device all qualify.","remediation":"Qualcomm shipped firmware v2.07 (January 2024) and Imagination fixed it in DDK 23.3 (December 2023); Apple fixed it in silicon from A17/M3 onward, leaving older Apple GPUs UNPATCHABLE; ChromeOS shipped AMD and Qualcomm mitigations in stable 120 / LTS 114. AMD's position at disclosure was that devices remained vulnerable pending mitigation work - track AMD-SB-6010 for your specific Instinct parts rather than assuming a fix exists. Cost where a driver fix does exist: driver upgrade plus node drain to reload the kernel module. Where it does not, the only controls are refusing to share a GPU across trust boundaries, or having your runtime explicitly zero local memory at kernel entry - which costs measurable throughput on small kernels and has to be done by whoever compiles the kernels, not by you.","references":["https://kb.cert.org/vuls/id/446598","https://blog.trailofbits.com/2024/01/16/leftoverlocals-listening-to-llm-responses-through-leaked-gpu-local-memory/","https://nvd.nist.gov/vuln/detail/CVE-2023-4969","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-6010"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"tags":["tenant-isolation"],"published":"2024-01-16"},{"cwe":["CWE-362","CWE-835"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-52587","cve":"CVE-2023-52587","aliases":[],"title":"Linux kernel (drivers/infiniband/ulp/ipoib): The IPoIB multicast join task drops its lock mid-iteration, letting a","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/ulp/ipoib)","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"The IPoIB multicast join task drops its lock mid-iteration, letting a concurrent device flush move entries off the list it is walking. The loop then spins forever and the node hard-locks - observed in production on RHEL kernels. One fabric event during multicast activity takes the entire shared node offline, killing every tenant's workload on it.","attack_vector":"Driven by fabric and link events, not by a tenant device node: an IB port event, subnet-manager sweep, or link flap runs ipoib_ib_dev_flush_light concurrently with a multicast join. Any node running IPoIB with multicast groups is exposed; a tenant generating multicast join churn on the IPoIB interface widens the window. Conditional on the ib_ipoib module being in use.","remediation":"No fixed version is recorded in this entry; boot a stable kernel carrying the mcast list locking fix (commits 4c8922ae8eb8 / 615e3adc2042). Interim: reduce IPoIB multicast usage on shared nodes and stabilise the fabric (avoid unnecessary port flaps / SM re-sweeps) until patched.","references":["https://git.kernel.org/stable/c/4c8922ae8eb8dcc1e4b7d1059d97a8334288d825","https://git.kernel.org/stable/c/615e3adc2042b7be4ad122a043fc9135e6342c90","https://nvd.nist.gov/vuln/detail/CVE-2023-52587"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-53147","cve":"CVE-2023-53147","aliases":[],"title":"Linux kernel (net/xfrm): XFRM_MSG_NEWAE lets a caller update replay-window state on a state that never had replay_esn","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"XFRM_MSG_NEWAE lets a caller update replay-window state on a state that never had replay_esn allocated, so the kernel memcpys into a NULL pointer. The commit is explicit that a malicious user can crash the kernel this way - a container-triggerable node panic straight through the replay-state update path.","attack_vector":"xfrm_new_ae over xfrm netlink, requiring CAP_NET_ADMIN in the network namespace - satisfied by any container granted NET_ADMIN with its own netns, and by the node's IKE daemon. No fabric access or device node needed; the SA does not need to be a valid ESN state, which is the whole point.","remediation":"Boot a kernel carrying the linked stable commits. Interim: drop CAP_NET_ADMIN from tenant containers so xfrm netlink is not reachable from tenant workloads.","references":["https://git.kernel.org/stable/c/ed1cba039309c80b49719fcff3e3d7cdddb73d96","https://git.kernel.org/stable/c/44f69c96f8a147413c23c68cda4d6fb5e23137cd","https://nvd.nist.gov/vuln/detail/CVE-2023-53147"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-191","CWE-770"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-53171","cve":"CVE-2023-53171","aliases":[],"title":"Linux kernel (drivers/vfio): Pinned-memory accounting for a VFIO container is lost across exec(), then underflows to a","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/vfio)","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Pinned-memory accounting for a VFIO container is lost across exec(), then underflows to a huge unsigned value on unmap. A tenant that maps DMA, execs, and repeats sheds its RLIMIT_MEMLOCK charge each round and can pin host RAM without bound until the node runs out of memory - a noisy-neighbour outage for everyone on the box. The underflow also permanently wedges further DMA maps for that container.","attack_vector":"Any container holding /dev/vfio/vfio and a group fd, using the legacy type1 container. The sequence is map DMA, exec() (the container fd survives), repeat. No host root, no special hardware, no race window to win.","remediation":"The record lists no fixed release; boot a kernel carrying the stable fix commits below. Interim controls: put a hard memory cgroup limit on tenants holding VFIO containers rather than relying on RLIMIT_MEMLOCK alone, and drop /dev/vfio from containers that do not need passthrough.","references":["https://git.kernel.org/stable/c/5a271242716846cc016736fb76be2b40ee49b0c3","https://git.kernel.org/stable/c/eafb81c50da899dd80b340c841277acc4a1945b7","https://nvd.nist.gov/vuln/detail/CVE-2023-53171"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-908","CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-53525","cve":"CVE-2023-53525","aliases":[],"title":"Linux kernel (drivers/infiniband/core): Rdma_join_multicast accepted queue-pair types other than UD and built the","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/core)","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Rdma_join_multicast accepted queue-pair types other than UD and built the multicast event with an uninitialized qkey, so kernel stack contents leaked into fabric-visible multicast state and non-UD QPs could be attached to multicast groups they have no business joining. That is both an information disclosure and a queue-pair type confusion on the shared fabric.","attack_vector":"Local and unprivileged: reached through the userspace RDMA CM character device - a tenant container holding /dev/infiniband/rdma_cm issues the multicast join (the syzkaller path is ucma_write -> ucma_join_multicast). No hardware RDMA adapter is required if soft-RoCE (rxe) or another software provider is present.","remediation":"No fixed version is recorded in this entry; boot a stable kernel carrying the UD-only multicast restriction (commits ae1149885142 / 48e8e7851dc0). Interim: remove /dev/infiniband/rdma_cm from tenant containers and blacklist rdma_ucm where tenants do not need userspace connection management.","references":["https://git.kernel.org/stable/c/ae11498851423d6de27aebfe12a5ee85060ab1d5","https://git.kernel.org/stable/c/48e8e7851dc0b1584d83817a78fc7108c8904b54","https://nvd.nist.gov/vuln/detail/CVE-2023-53525"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-190","CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-54239","cve":"CVE-2023-54239","aliases":[],"title":"Linux kernel (drivers/iommu/iommufd): Iommufd accepts a user address plus length that wraps past zero, then asks the mm","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/iommufd)","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Iommufd accepts a user address plus length that wraps past zero, then asks the mm to pin pages for a range that cannot exist. The DMA-map path proceeds on arithmetic the kernel never validated, which is the class of bug that ends with a mapping describing the wrong physical pages.","attack_vector":"A holder of /dev/iommu calling IOMMU_IOAS_MAP with a user pointer near UINTPTR_MAX so uptr + length overflows. Plain ioctl, no race, no root. Conditional on iommufd being the passthrough path in use.","remediation":"The record lists no fixed release; boot a kernel carrying the stable fix commits below. Interim control: do not expose /dev/iommu to tenants.","references":["https://git.kernel.org/stable/c/800963e7eb001ada8cf2418f159fb649694467f1","https://git.kernel.org/stable/c/e4395701330fc4aee530905039516fe770b81417","https://nvd.nist.gov/vuln/detail/CVE-2023-54239"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2023-6570","cve":"CVE-2023-6570","aliases":[],"title":"Kubeflow: SSRF","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Kubeflow","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"SSRF","attack_vector":"Authenticated notebook/pipeline user","remediation":"Upgrade; block IMDS egress from Kubeflow pods","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-6570"],"status":"curated","published":"2023-12-14"},{"id":"CVE-2024-0078","cve":"CVE-2024-0078","aliases":[],"title":"GPU Display Driver: DoS (null deref)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"DoS (null deref)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; rolling reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0078","https://github.com/NVIDIA/product-security/tree/main/2024/5520"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"published":"2024-03-27"},{"id":"CVE-2024-0079","cve":"CVE-2024-0079","aliases":[],"title":"vGPU Manager: Guest-triggered host DoS","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU Manager","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Guest-triggered host DoS","attack_vector":"Tenant VM guest","remediation":"vGPU Manager upgrade; evacuate VMs","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0079","https://github.com/NVIDIA/product-security/tree/main/2024/5520"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:N/I:N/A:H","cwe":["CWE-476"],"published":"2024-03-27"},{"id":"CVE-2024-0083","cve":"CVE-2024-0083","aliases":[],"title":"ChatRTX: XSS","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"ChatRTX","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"XSS","attack_vector":"Web user","remediation":"Consumer app; no DC action","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0083","https://github.com/NVIDIA/product-security/tree/main/2024/5532"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:L/I:N/A:L","cwe":["CWE-79"],"published":"2024-04-08"},{"id":"CVE-2024-0093","cve":"CVE-2024-0093","aliases":[],"title":"vGPU Manager: Cross-tenant info disclosure","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU Manager","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Cross-tenant info disclosure","attack_vector":"Tenant VM guest","remediation":"vGPU Manager upgrade; evacuate VMs","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0093","https://github.com/NVIDIA/product-security/tree/main/2024/5551"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:N/A:N","cwe":["CWE-200"],"published":"2024-06-13"},{"id":"CVE-2024-0100","cve":"CVE-2024-0100","aliases":[],"title":"Triton Inference Server: Info disclosure via path traversal","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Info disclosure via path traversal","attack_vector":"Client of the inference endpoint","remediation":"Upgrade Triton; redeploy serving images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0100","https://github.com/NVIDIA/product-security/tree/main/2024/5535"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:H/UI:N/S:U/C:N/I:H/A:H","cwe":["CWE-73"],"published":"2024-05-14"},{"id":"CVE-2024-10270","cve":"CVE-2024-10270","aliases":[],"title":"Keycloak: Regex complexity in SearchQueryUtils","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Keycloak","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Regex complexity in SearchQueryUtils -> resource exhaustion DoS of the auth service","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; rate-limit the token/search endpoints","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-10270"],"status":"curated","published":"2024-11-25"},{"id":"CVE-2024-11185","cve":"CVE-2024-11185","aliases":["Arista Security Advisory 0118"],"title":"Arista EOS (L2 forwarding / VLAN isolation): Ingress traffic on a layer-2 port is forwarded out ports belonging to a","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (L2 forwarding / VLAN isolation)","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Ingress traffic on a layer-2 port is forwarded out ports belonging to a different VLAN. VLAN separation is the primary tenant boundary in most GPU-cluster builds, so this is one tenant's frames landing in another tenant's broadcast domain. CVSS 6.5 understates the operator consequence — for a neocloud selling isolated tenancy this is a contractual failure, not a medium-severity bug.","attack_vector":"An attacker on any L2 port under the conditions the advisory describes. No credentials — this is a forwarding-plane defect, not an access-control one.","remediation":"EOS upgrade plus switch reload on every affected leaf. No config workaround — you cannot ACL your way out of a forwarding-plane leak. Plan a rolling upgrade across the leaf layer; on MLAG pairs you can do one side at a time and keep the rack up.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-11185"],"status":"curated","tags":["tenant-isolation"],"published":"2025-05-27"},{"id":"CVE-2024-1725","cve":"CVE-2024-1725","aliases":[],"title":"KubeVirt: kubevirt-csi in OpenShift Virtualization HCP grants access to the root HCP worker node's volume","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"KubeVirt","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"kubevirt-csi in OpenShift Virtualization HCP grants access to the root HCP worker node's volume via a crafted PV","attack_vector":"Authenticated cluster user","remediation":"Upgrade the kubevirt-csi component","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-1725"],"status":"curated","published":"2024-03-07"},{"id":"CVE-2024-20365","cve":"CVE-2024-20365","aliases":[],"title":"Redfish API implementation on Cisco UCS B-Series, UCS Managed C-Series and UCS X-Series servers: An administrator-level","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Redfish API implementation on Cisco UCS B-Series, UCS Managed C-Series and UCS X-Series servers","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"An administrator-level Redfish user escapes the API's intended command surface and executes commands on the underlying management controller. The operator-relevant point is the boundary being crossed: Redfish administrative privilege is supposed to mean 'can configure the server', not 'can run code on the controller', and that distinction is what lets an operator delegate Redfish admin to automation without handing over the box. When it collapses, every service account your provisioning system holds becomes a shell on every BMC it manages. Command injection reachable through the Redfish interface by a user who already holds administrative privileges in the management software.","attack_vector":"An authenticated remote user with administrative privileges on the UCS management software, reaching the Redfish API. In practice this is your automation's own credentials - Terraform providers, Ansible modules and inventory collectors all hold Redfish admin, and any compromise of the automation host inherits it.","remediation":"Firmware/software update to the fixed UCS release per Cisco's advisory; this is a controller firmware update, so plan a per-chassis maintenance window rather than a rolling config push. The durable lesson is architectural: stop treating Redfish administrative accounts as a safe delegation boundary. Scope automation credentials to the narrowest Redfish role that works, keep them out of shared secret stores that tenant-facing systems can read, and log Redfish administrative calls centrally so an anomalous command sequence is visible.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-20365","https://sec.cloudapps.cisco.com/security/center/content/CiscoSecurityAdvisory/cisco-sa-cimc-redfish-cominj-sbkv5ZZ"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2024-21121","cve":"CVE-2024-21121","aliases":[],"title":"Oracle VirtualBox: Easily exploitable Core flaw allowing unauthorised access to VirtualBox-accessible data","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Oracle VirtualBox","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Easily exploitable Core flaw allowing unauthorised access to VirtualBox-accessible data","attack_vector":"Tenant VM guest","remediation":"VirtualBox update + VM restart","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21121"],"status":"curated","published":"2024-04-16"},{"id":"CVE-2024-2206","cve":"CVE-2024-2206","aliases":[],"title":"Gradio: SSRF in the `/proxy` route","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Gradio","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"SSRF in the `/proxy` route","attack_vector":"Unauthenticated network","remediation":"Upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-2206"],"status":"curated","published":"2024-03-27"},{"id":"CVE-2024-23499","cve":"CVE-2024-23499","aliases":[],"title":"Intel ice driver (Ethernet 800 Series, Linux kernel mode): A protection-mechanism failure in the E810 Linux kernel","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel ice driver (Ethernet 800 Series, Linux kernel mode)","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"A protection-mechanism failure in the E810 Linux kernel driver reachable by an unauthenticated attacker. Driver-side counterpart to the E810 firmware protection failures in the same advisory.","attack_vector":"Unauthenticated attacker able to present traffic to the interface.","remediation":"Fixed in the Intel out-of-tree ice driver (or the equivalent in-kernel version). Updating the driver requires unloading and reloading the module, which drops every link on that NIC - on a node whose RDMA fabric carries collective traffic, that is a job-killing event, so drain first. If you take it via a distro kernel update instead, it is a reboot. No firmware flash for the driver-side fixes. Target ice 28.3 or later.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-23499","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00918.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-08-14"},{"id":"CVE-2024-24580","cve":"CVE-2024-24580","aliases":[],"title":"Intel Data Center GPU Max Series 1100 / 1550: A second improper conditions check in the Max Series allowing","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel Data Center GPU Max Series 1100 / 1550","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"A second improper conditions check in the Max Series allowing a privileged local user to cause denial of service on the accelerator.","attack_vector":"Local, privileged. Host root on the node.","remediation":"Apply the Intel update for the Max Series. Cost: driver reload, or drain and reboot if the fix is in firmware.","references":["https://www.intel.com/content/www/us/en/security-center/default.html","https://nvd.nist.gov/vuln/detail/CVE-2024-24580"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-08-14"},{"id":"CVE-2024-24983","cve":"CVE-2024-24983","aliases":["INTEL-SA-00918"],"title":"Intel Ethernet Controller E810 firmware: An unauthenticated attacker on the network can take an E810 NIC out of service","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Ethernet Controller E810 firmware","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"An unauthenticated attacker on the network can take an E810 NIC out of service through a protection-mechanism failure in the adapter firmware. E810 is the standard 100/200GbE front-end NIC on a large share of AI servers, and it is often also the storage-network NIC — so knocking it down removes the node from both the job and its dataset. Because the fault is in firmware, the host OS sees a dead link, not a driver problem, and normal remediation (restart the driver) does not recover it.","attack_vector":"Unauthenticated, remote — traffic arriving at the NIC over the network. No host credentials and no adjacency requirement.","remediation":"Flash E810 adapter firmware to 4.4 or later using Intel's NVM Update Tool (or the OEM-repackaged version — Dell DSA-2025-236, HPE and Lenovo ship their own). NVM updates require a **cold power cycle**, not a warm reboot, for the new image to take effect — so this is a full node drain per server, which across a GPU fleet is the dominant cost. Sequence it with your normal node-maintenance rotation rather than as an emergency.","references":["https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00918.html","https://nvd.nist.gov/vuln/detail/CVE-2024-24983"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-08-14"},{"id":"CVE-2024-25742","cve":"CVE-2024-25742","aliases":["#VC injection","WeSee-class"],"title":"SEV-ES / SEV-SNP guest kernel - unsolicited #VC (vector 29) injection: An untrusted hypervisor can inject the #VC","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"SEV-ES / SEV-SNP guest kernel - unsolicited #VC (vector 29) injection","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"An untrusted hypervisor can inject the #VC exception (vector 29) into an SEV-ES or SEV-SNP guest at a moment of its choosing and force the guest's #VC handler to run in an unexpected context. The handler is privileged guest code that touches the GHCB shared page, so a host that controls when it fires gains a lever on guest execution - the WeSee research turned this class into reading and writing confidential guest memory and executing code inside the VM. The guest's own hardware isolation is what gets used against it.","attack_vector":"Malicious or compromised hypervisor against its own guest. No guest bug required and no tenant cooperation - the host simply injects.","remediation":"Fixed in the **guest** kernel, not the host - the hardening lives in the SEV-ES/SNP guest's #VC handler and interrupt entry code. That inverts the usual rollout: you can patch every hypervisor you own and still be exposed, because the protection has to be in the tenant's own VM image. As an operator your job is to ship updated confidential-guest images (or tell tenants which minimum kernel to run) and, where you can, enforce it as an admission requirement. Each guest picks the fix up on its next boot; no host reboot, no firmware update. Fixed in Linux 6.9 and backported; guests must run a kernel that hardens the #VC entry path. As the operator you cannot fix this for a tenant who brings their own image - the honest control is to document a minimum guest kernel and enforce it at admission for confidential workloads.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-25742","https://ahoi-attacks.github.io/wesee/","https://www.amd.com/en/resources/product-security.html"],"status":"curated","tags":["tenant-isolation"],"published":"2024-05-17"},{"cwe":["CWE-476","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-26615","cve":"CVE-2024-26615","aliases":[],"title":"Linux kernel SMC-D diagnostics (smc_diag, rmb_desc access during connection dump): Dumping SMC-D connections while","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel SMC-D diagnostics (smc_diag, rmb_desc access during connection dump)","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Dumping SMC-D connections while connections are churning dereferences a remote memory buffer descriptor that has already gone away, crashing the node. The reproducer is nothing exotic - run a web benchmark under smc_run and poll `smcss -D` in a loop. That means routine monitoring can kill a node, and it also means a tenant able to trigger the diag dump path can do so deliberately while generating connection churn.","attack_vector":"Local. Requires the ability to issue SMC diag netlink dumps while SMC connections are being torn down; monitoring agents do this on a timer.","remediation":"Kernel update guarding the rmb_desc access. Until patched, stop polling SMC diagnostics on nodes carrying live SMC traffic - the monitoring is the trigger.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=1fea9969b81c67d0cb1611d1b8b7d19049d937be","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2024/CVE-2024-26615.json"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-26808","cve":"CVE-2024-26808","aliases":[],"title":"Linux kernel (netfilter): nft_chain_filter NETDEV_UNREGISTER mishandling for inet/ingress basechains - UAF","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (netfilter)","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"nft_chain_filter NETDEV_UNREGISTER mishandling for inet/ingress basechains - UAF","attack_vector":"Any tenant process in a container with CAP_NET_ADMIN","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2024-26808"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-04-04"},{"cwe":["CWE-476","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-26858","cve":"CVE-2024-26858","aliases":[],"title":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core): Nothing orders the PTP send-queue tracking list against","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core)","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Nothing orders the PTP send-queue tracking list against population of the timestamp metadata map, so the compiler or the CPU can publish a queue entry before its metadata exists. The NAPI poll path on another CPU then walks the list and dereferences NULL, panicking the node and taking down every co-tenant workload with it.","attack_vector":"Reachable from ordinary unprivileged socket traffic: any local process that requests hardware TX timestamping (SO_TIMESTAMPING) routes packets through the mlx5e PTP send queue, and the race is between that transmit and NAPI completion on a different CPU. A tenant container needs no device node and no capability - just a socket. Requires the mlx5e PTP TX queue to be active (hardware timestamping enabled on the interface, which is normal on clusters running PTP time sync).","remediation":"Update to 6.6.22 or later on the 6.6.x branch, or a mainline kernel from 6.6 onward carrying the memory-barrier fix. Interim control: disable hardware TX timestamping on the mlx5 interface where PTP time sync is not required, which retires the vulnerable send queue.","references":["https://git.kernel.org/stable/c/936ef086161ab89a7f38f7a0761d6a3063c3277e","https://git.kernel.org/stable/c/b7cf07586c40f926063d4d09f7de28ff82f62b2a","https://nvd.nist.gov/vuln/detail/CVE-2024-26858"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-835"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-26891","cve":"CVE-2024-26891","aliases":[],"title":"Linux kernel (drivers/iommu/intel): The whole node hangs. VT-d keeps re-issuing an ATS device-TLB invalidation to a","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/intel)","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"The whole node hangs. VT-d keeps re-issuing an ATS device-TLB invalidation to a device that is already gone, retrying forever inside interrupt context after the invalidation-timeout fault. The NMI watchdog reports a hard lockup and the box stops serving every tenant on it.","attack_vector":"Needs a device behind a hotplug-capable PCIe port to vanish or its link to flap: an operator hot-unplugging a GPU or NIC, a flaky link, or a tenant-initiated device/bus reset on a passthrough function that ends in link-down. Conditional on Intel VT-d with ATS enabled, which is the normal state for passthrough-capable accelerators and SmartNICs. Not reachable from the network fabric; the trigger is a PCIe topology event, not a crafted request.","remediation":"The record lists no fixed release; boot a kernel carrying the stable fix commits below. Interim controls: keep ATS-capable endpoints out of hotplug-capable slots, disable PCIe hotplug on slots holding passthrough devices, and do not let tenants initiate bus resets on assigned functions.","references":["https://git.kernel.org/stable/c/f873b85ec762c5a6abe94a7ddb31df5d3ba07d85","https://git.kernel.org/stable/c/d70f1c85113cd8c2aa8373f491ca5d1b22ec0554","https://nvd.nist.gov/vuln/detail/CVE-2024-26891"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-29893","cve":"CVE-2024-29893","aliases":[],"title":"Argo CD: Repo-server DoS, halting all GitOps reconciliation","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Repo-server DoS, halting all GitOps reconciliation","attack_vector":"Any authenticated Argo CD user","remediation":"Rolling Argo CD upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-29893"],"status":"curated","published":"2024-03-29"},{"id":"CVE-2024-31141","cve":"CVE-2024-31141","aliases":[],"title":"Apache Kafka (client): ConfigProvider plugins let an untrusted app read files/env of the Kafka client host","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Apache Kafka (client)","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"ConfigProvider plugins let an untrusted app read files/env of the Kafka client host","attack_vector":"Network (remote)","remediation":"Control-plane: dependency upgrade in telemetry/billing pipelines","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-31141"],"status":"curated","published":"2024-11-19"},{"id":"CVE-2024-32048","cve":"CVE-2024-32048","aliases":[],"title":"Intel Distribution of OpenVINO Model Server: An unauthenticated user can reach an input-validation flaw in OpenVINO","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel Distribution of OpenVINO Model Server","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"An unauthenticated user can reach an input-validation flaw in OpenVINO Model Server. Model Server is a production inference endpoint, so this is on the request path of a live service.","attack_vector":"Anything that can send a request to the model server - which for most deployments is the whole cluster network, and sometimes the internet.","remediation":"Upgrade OpenVINO Model Server to 2024.0 or later. Container image swap and rolling restart; no node reboot or firmware.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-32048","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01158.html"],"status":"curated","fleet":{"pain_class":"daemon-restart"},"published":"2024-11-13"},{"id":"CVE-2024-36293","cve":"CVE-2024-36293","aliases":[],"title":"Intel processors with SGX (EDECCSSA leaf): Improper access control on the EDECCSSA user leaf function lets","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors with SGX (EDECCSSA leaf)","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Improper access control on the EDECCSSA user leaf function lets an authenticated local user deny SGX service. EDECCSSA is the SGX2 instruction used for in-enclave exception handling, so this is reachable from ordinary enclave-adjacent code.","attack_vector":"Local authenticated user on an SGX-enabled host.","remediation":"Microcode update and reboot; late-loadable at boot without an OEM BIOS release. Re-attest afterwards.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36293","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01213.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"published":"2025-02-12"},{"id":"CVE-2024-36620","cve":"CVE-2024-36620","aliases":[],"title":"Docker / moby: NULL pointer dereference in image_history crashes the daemon","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"NULL pointer dereference in image_history crashes the daemon","attack_vector":"Any tenant workload / API client","remediation":"Upgrade moby","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36620"],"status":"curated","published":"2024-11-29"},{"id":"CVE-2024-36621","cve":"CVE-2024-36621","aliases":[],"title":"Docker / moby: Race in the buildkit snapshot adapter","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Race in the buildkit snapshot adapter; concurrent builds leak resources and exhaust the host","attack_vector":"Anyone who can submit builds to a shared builder","remediation":"Upgrade moby; isolate tenant builders","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36621"],"status":"curated","published":"2024-11-29"},{"id":"CVE-2024-3744","cve":"CVE-2024-3744","aliases":[],"title":"azure-file-csi-driver: Service account tokens disclosed in driver logs","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"azure-file-csi-driver","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Service account tokens disclosed in driver logs","attack_vector":"Anyone with log-pipeline read access","remediation":"DaemonSet rollout; rotate tokens; scrub logs","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated","published":"2024-05-15"},{"id":"CVE-2024-45806","cve":"CVE-2024-45806","aliases":[],"title":"Envoy: External clients manipulate Envoy internal headers, reaching unauthorized behaviour","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"External clients manipulate Envoy internal headers, reaching unauthorized behaviour","attack_vector":"Unauthenticated network","remediation":"Upgrade Envoy","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-45806"],"status":"curated","published":"2024-09-20"},{"cwe":["CWE-476","CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-46824","cve":"CVE-2024-46824","aliases":[],"title":"Linux kernel (drivers/iommu/iommufd): The cache-invalidation ioctl calls a driver operation that may not exist, jumping","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/iommufd)","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"The cache-invalidation ioctl calls a driver operation that may not exist, jumping through a NULL function pointer on a guest VMM's invalidation request. The host takes an unhandled fault at address zero inside the very path that is supposed to make a guest's IOMMU invalidations real, so the failure mode is both a node crash and an invalidation that never happened.","attack_vector":"A VMM holding /dev/iommu issuing IOMMU_HWPT_INVALIDATE - the nested-translation path qemu uses to forward a guest's IOTLB invalidations to hardware. No host root. Conditional on running nested translation on an IOMMU driver that never implemented the user-invalidation op; the upstream trace is qemu on arm64.","remediation":"The record lists no fixed release; boot a kernel carrying the stable fix commits below. Interim control: do not enable nested translation / vIOMMU for tenant VMs on drivers that lack user cache invalidation, and keep /dev/iommu out of tenant containers.","references":["https://git.kernel.org/stable/c/89827a4de802765b1ebb401fc1e73a90108c7520","https://git.kernel.org/stable/c/a11dda723c6493bb1853bbc61c093377f96e2d47","https://nvd.nist.gov/vuln/detail/CVE-2024-46824"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-49953","cve":"CVE-2024-49953","aliases":[],"title":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core/en_accel): The IPsec offload worker does not check the xfrm","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core/en_accel)","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"The IPsec offload worker does not check the xfrm state's lifecycle before expiring it, so an SA that is already dead gets deleted a second time and the kernel dereferences poison list pointers. On a node where the fabric encryption terminates, this is a general protection fault that panics the host and takes every tenant on it offline.","attack_vector":"Requires mlx5 IPsec packet/full offload to be configured on the node (fabric-level encryption). The racing worker fires on SA soft/hard packet-count limits, so the timing window is opened by traffic volume across the tunnel - a fabric peer pushing traffic influences when it triggers - while SA teardown comes from the local IKE daemon. This is not a tenant-controlled primitive, but it is a peer-influenced panic on shared infrastructure.","remediation":"Boot a kernel with the mlx5e IPsec state-check fix (stable commits below; no fixed-version list published for this ID). Interim control: if IPsec offload is not actually needed on a given node, run the fabric without mlx5 IPsec offload configured - the vulnerable worker is not scheduled when no offloaded SAs exist.","references":["https://git.kernel.org/stable/c/0b1672834634df9ac9cedf856db9fc36d92c50ef","https://git.kernel.org/stable/c/151e7dead1f5399a73c19c4b50307ea48aff1dc0","https://nvd.nist.gov/vuln/detail/CVE-2024-49953"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-20","CWE-125"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-50142","cve":"CVE-2024-50142","aliases":[],"title":"Linux kernel (net/xfrm): An SA created with an AF_UNSPEC selector escaped prefix-length validation, and the kernel then","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"An SA created with an AF_UNSPEC selector escaped prefix-length validation, and the kernel then copied the SA's own family (AF_INET) into the selector - leaving an IPv4 selector claiming a 128-bit prefix. Selector matching afterwards compares more address bits than exist, so which traffic that SA claims is undefined: the wrong flows can be steered into, or excluded from, an encrypted association.","attack_vector":"XFRM_MSG_NEWSA over xfrm netlink, which needs CAP_NET_ADMIN in the network namespace - satisfied by the node's IKE daemon and by any container granted NET_ADMIN with its own netns. A tenant with that capability can install selectors the validator was supposed to reject.","remediation":"Boot a kernel carrying the linked stable commits. Interim: drop CAP_NET_ADMIN from tenant containers so only the node's IKE daemon can install SAs, and audit installed selectors for prefix lengths that exceed the address family.","references":["https://git.kernel.org/stable/c/f31398570acf0f0804c644006f7bfa9067106b0a","https://git.kernel.org/stable/c/401ad99a5ae7180dd9449eac104cb755f442e7f3","https://nvd.nist.gov/vuln/detail/CVE-2024-50142"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-476","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-56668","cve":"CVE-2024-56668","aliases":[],"title":"Linux kernel (drivers/iommu/intel): Attaching a nested parent domain skips allocating the invalidation batch structure","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/intel)","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Attaching a nested parent domain skips allocating the invalidation batch structure, so the first DMA map into that domain dereferences NULL inside the cache-flush path and oopses the host. The same allocation is done without a lock, so it can also be raced and leaked. The crash arrives from a guest VMM's ordinary map operation.","attack_vector":"A VMM holding /dev/iommu that has created a nested parent domain (vIOMMU / nested translation for a passthrough device) and then calls IOMMU_IOAS_MAP. The upstream trace is qemu-system-x86 on Intel VT-d. No host root. Conditional on nested translation being enabled - plain single-level passthrough does not take this path.","remediation":"The record lists no fixed release; boot a kernel carrying the stable fix commits below. Interim control: do not expose nested translation / vIOMMU to tenant VMs on unpatched VT-d hosts.","references":["https://git.kernel.org/stable/c/ffd774c34774fd4cc0e9cf2976595623a6c3a077","https://git.kernel.org/stable/c/74536f91962d5f6af0a42414773ce61e653c10ee","https://nvd.nist.gov/vuln/detail/CVE-2024-56668"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-58093","cve":"CVE-2024-58093","aliases":[],"title":"Linux kernel (drivers/pci/pcie): The ASPM link state of a PCIe switch is freed as soon as ANY function on the upstream","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/pci/pcie)","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"The ASPM link state of a PCIe switch is freed as soon as ANY function on the upstream port is removed, while downstream ports still hold it as their parent link. Every subsequent reference is a use-after-free, and the reported symptom is a general protection fault that takes the whole node down - all tenants on it, not just the one whose device went away.","attack_vector":"Triggered by device removal under a multi-function PCIe switch upstream port, which is exactly the topology in a GPU box where a Broadcom/PLX switch fans out to eight accelerators or a bank of NVMe. The upstream note says the faults are ESPECIALLY frequent during hot-unplug, because pciehp removes devices on the link bus in reverse order and therefore hits the non-zero function before function 0. No tenant credentials are needed - a surprise removal, a device that drops off the link, or a maintenance pull supplies the event. Requires ASPM enabled and a switch with a multi-function upstream port.","remediation":"Update to 5.4.292 / 5.10.236 / 5.15.180 / 6.1.134 / 6.4 / 6.5 or later. Interim: drain the node before any planned PCIe removal under a switch, and consider disabling ASPM (pcie_aspm=off) on nodes where hot-removal under a switch is routine.","references":["https://git.kernel.org/stable/c/0a0f9aecf66b98959aab7fb5764b4b3e522f4f5b","https://git.kernel.org/stable/c/62db339ecc3d58d8fd83a9e4d80061cd943bb6e2","https://nvd.nist.gov/vuln/detail/CVE-2024-58093"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-14269","cve":"CVE-2025-14269","aliases":[],"title":"Headlamp: Credential caching in Headlamp when Helm is enabled","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Headlamp","year":"2025","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Credential caching in Headlamp when Helm is enabled","attack_vector":"Cluster user with dashboard access","remediation":"Upgrade Headlamp; rotate cached credentials","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated"},{"id":"CVE-2025-1767","cve":"CVE-2025-1767","aliases":[],"title":"Kubernetes (kubelet): gitRepo volume grants inadvertent access to local repositories on the node","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kubelet)","year":"2025","cvss_score":6.5,"severity":"medium","kev":false,"impact":"gitRepo volume grants inadvertent access to local repositories on the node","attack_vector":"Cluster user able to create a gitRepo volume","remediation":"Rolling kubelet upgrade; block gitRepo volumes","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated","published":"2025-03-13"},{"id":"CVE-2025-1944","cve":"CVE-2025-1944","aliases":[],"title":"picklescan: ZIP manipulation crashes the scanner (scan bypass by DoS)","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"picklescan","year":"2025","cvss_score":6.5,"severity":"medium","kev":false,"impact":"ZIP manipulation crashes the scanner (scan bypass by DoS)","attack_vector":"Customer-supplied model archive","remediation":"Upgrade to 0.0.23+; fail-closed on scanner crash","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-1944"],"status":"curated","published":"2025-03-10"},{"cwe":["CWE-457","CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-21996","cve":"CVE-2025-21996","aliases":[],"title":"Linux kernel (drivers/gpu/drm/radeon): The radeon video-encode command-stream parser used an uninitialised stack value","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/radeon)","year":"2025","cvss_score":6.5,"severity":"medium","kev":false,"impact":"The radeon video-encode command-stream parser used an uninitialised stack value as the size argument when validating a relocation, so a tenant that crafts an encode command stream whose first opcode is the encode command gets the bounds check performed against garbage. That turns the CS parser's relocation validation into a coin flip and opens the door to out-of-bounds buffer access driven entirely from the ioctl.","attack_vector":"A container holding /dev/dri/renderD* on a host with a radeon-class AMD GPU submits a hand-built VCE command stream through the CS ioctl. No privilege beyond the render node; conditional on the radeon driver being loaded and the GPU exposing VCE, so it does not apply to amdgpu-era or NVIDIA nodes.","remediation":"Update to a kernel with the fix commits below. Interim: blacklist the radeon module on nodes that do not need it, or drop /dev/dri from untrusted containers on radeon hosts.","references":["https://git.kernel.org/stable/c/0effb378ebce52b897f85cd7f828854b8c7cb636","https://git.kernel.org/stable/c/5b4d9d20fd455a97920cf158dd19163b879cf65d","https://nvd.nist.gov/vuln/detail/CVE-2025-21996"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-22892","cve":"CVE-2025-22892","aliases":[],"title":"OpenVINO Model Server: An unauthenticated request can drive OpenVINO Model Server into unbounded resource consumption","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"OpenVINO Model Server","year":"2025","cvss_score":6.5,"severity":"medium","kev":false,"impact":"An unauthenticated request can drive OpenVINO Model Server into unbounded resource consumption and take the endpoint down. Cheap remote DoS against a serving tier.","attack_vector":"Anyone who can send a request to the model server.","remediation":"Upgrade OpenVINO Model Server to 2024.4 or later, and put request-size and rate limits in front of the endpoint. Rolling container restart.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-22892","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01272.html"],"status":"curated","published":"2025-05-13"},{"id":"CVE-2025-23047","cve":"CVE-2025-23047","aliases":[],"title":"Cilium: Insecure default Access-Control-Allow-Origin in Hubble UI exposes sensitive observability data","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2025","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Insecure default Access-Control-Allow-Origin in Hubble UI exposes sensitive observability data","attack_vector":"Anyone who can get an operator's browser to a hostile page","remediation":"Upgrade Cilium; put Hubble UI behind auth","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23047"],"status":"curated","published":"2025-01-22"},{"id":"CVE-2025-23243","cve":"CVE-2025-23243","aliases":[],"title":"NVIDIA Riva: Weak authentication","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Riva","year":"2025","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Weak authentication -> unauthorized use / info disclosure","attack_vector":"Network client","remediation":"Upgrade Riva containers; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23243","https://github.com/NVIDIA/product-security/tree/main/2025/5625"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:L/A:L","cwe":["CWE-284"],"published":"2025-03-11"},{"id":"CVE-2025-23259","cve":"CVE-2025-23259","aliases":[],"title":"Mellanox DPDK: DoS / data tampering via race condition","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Mellanox DPDK","year":"2025","cvss_score":6.5,"severity":"medium","kev":false,"impact":"DoS / data tampering via race condition","attack_vector":"Tenant with a DPDK-attached VF","remediation":"Bump DPDK packages; rebuild dataplane images; restart dataplane","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23259","https://github.com/NVIDIA/product-security/tree/main/2025/5655"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:L/I:N/A:H","cwe":["CWE-362"],"published":"2025-09-04"},{"id":"CVE-2025-24323","cve":"CVE-2025-24323","aliases":["INTEL-SA-01339"],"title":"Intel PCIe Switch firmware package and LED mode toggle tool before version MR4_1.0b1: Improper access control","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel PCIe Switch firmware package and LED mode toggle tool before version MR4_1.0b1","year":"2025","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Improper access control in the PCIe switch firmware package and its management tool lets a privileged local user escalate. The PCIe switch is the fabric between the CPU complex and the GPUs, NVMe and NICs inside a node - it decides which device can DMA where. Firmware-level control of the switch means an attacker can re-map or observe traffic between accelerators and the host, which is below every isolation boundary the OS or hypervisor enforces, and it persists in the switch's own flash across any host reimage and across tenant handoff. This is the kind of component operators rarely inventory at all, so exposure tends to be unmeasured rather than accepted.","attack_vector":"A privileged local user on the host running the vendor firmware/management tooling against the switch.","remediation":"Update the PCIe switch firmware package to MR4_1.0b1 or later. In practice this is delivered by the system builder (Supermicro, Gigabyte, Quanta, Wiwynn and the GPU-system ODMs) rather than by Intel directly, and it is one of the least reliably shipped firmware components in the stack - you will often have to ask the integrator for it by name. Requires the node quiesced and power-cycled. Remove the LED mode toggle tool and other switch management binaries from tenant-visible host images.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-24323","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01339.html"],"status":"curated"},{"id":"CVE-2025-32003","cve":"CVE-2025-32003","aliases":[],"title":"Intel Ethernet Network Adapter E810 (100GbE) firmware: Out-of-bounds read in 100GbE E810 firmware reachable","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Ethernet Network Adapter E810 (100GbE) firmware","year":"2025","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Out-of-bounds read in 100GbE E810 firmware reachable from privileged host software, causing denial of service. The version threshold here (cvl fw 1.7.6, cpk 1.3.7) is different again from the other E810 advisories — the E810 firmware CVE trail is long enough that version-by-CVE tracking is not workable and operators should treat it as a rolling minimum-version policy.","attack_vector":"Privileged local software on the host (Ring 0).","remediation":"Flash E810 firmware to at least cvl fw 1.7.6 / cpk 1.3.7 — and given the later advisories, go straight to 1.7.8.x or newer. Cold power cycle, per node.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-32003"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-02-10"},{"id":"CVE-2025-32386","cve":"CVE-2025-32386","aliases":[],"title":"Helm: Decompression bomb chart exhausts memory on the rendering host","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Helm","year":"2025","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Decompression bomb chart exhausts memory on the rendering host","attack_vector":"Malicious chart","remediation":"Upgrade Helm","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-32386"],"status":"curated","published":"2025-04-09"},{"id":"CVE-2025-32387","cve":"CVE-2025-32387","aliases":[],"title":"Helm: Deeply nested JSON Schema references cause stack overflow","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Helm","year":"2025","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Deeply nested JSON Schema references cause stack overflow","attack_vector":"Malicious chart","remediation":"Upgrade Helm","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-32387"],"status":"curated","published":"2025-04-09"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-277","CWE-732"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2025-36104","cve":"CVE-2025-36104","aliases":[],"title":"IBM Storage Scale SMB protocol stack (inherited ACL handling): Files created or modified over SMB inherit permissions","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"IBM Storage Scale SMB protocol stack (inherited ACL handling)","year":"2025","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Files created or modified over SMB inherit permissions that are wider than intended, so an authenticated user reads data they were never granted. On a mixed-protocol cluster this quietly opens one tenant's directories to another.","attack_vector":"Any authenticated SMB client of a Storage Scale 5.2.3.0 or 5.2.3.1 cluster where directories use inherited ACLs.","remediation":"Apply the fix from IBM's bulletin, then audit effective ACLs on every directory that was created or modified through SMB while the affected versions ran - the upgrade corrects behaviour going forward but does not repair permissions already written.","references":["https://www.ibm.com/support/pages/node/7239562","https://nvd.nist.gov/vuln/detail/CVE-2025-36104"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-36608","cve":"CVE-2025-36608","aliases":[],"title":"Dell SmartFabric OS10 (XML external entity): XXE in SmartFabric OS10 before 10.6.0.5, reachable remotely","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell SmartFabric OS10 (XML external entity)","year":"2025","cvss_score":6.5,"severity":"medium","kev":false,"impact":"XXE in SmartFabric OS10 before 10.6.0.5, reachable remotely by a low-privileged attacker. XXE on a switch typically yields file read from the switch's filesystem — which holds the running configuration, and therefore the fabric's secrets and its full topology.","attack_vector":"Low-privileged attacker with remote access to the OS10 management interface.","remediation":"Upgrade OS10 to 10.6.0.5 or later plus reload. Related file-exposure issue in the same release: CVE-2025-30103.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-36608","https://nvd.nist.gov/vuln/detail/CVE-2025-30103"],"status":"curated","published":"2025-07-30"},{"cwe":["CWE-362","CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-37772","cve":"CVE-2025-37772","aliases":[],"title":"Linux kernel (drivers/infiniband/core): Bursts of network neighbour updates crash the node. Each event re-initializes a","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/core)","year":"2025","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Bursts of network neighbour updates crash the node. Each event re-initializes a work item that is already queued for the same connection id, corrupting the kernel workqueue and faulting in the worker thread - a whole-host outage triggered by traffic conditions rather than by anything the operator does.","attack_vector":"Reachable from the fabric/IP network with no host credentials: a peer that provokes back-to-back neighbour (ARP/ND) updates for an address that has a live rdma_cm connection id - route churn, address flapping, or a deliberately noisy neighbour on the same segment. Any RDMA provider using rdma_cm is affected, hardware or soft-RoCE.","remediation":"No fixed release is listed in this record - apply the referenced stable fix commits or run a current stable kernel (the follow-up fix for the same code is in 6.1.142 / 6.6.94 / 6.12.34 / 6.15, so land both). Interim: none reliable, since the trigger is ordinary network event handling.","references":["https://git.kernel.org/stable/c/51003b2c872c63d28bcf5fbcc52cf7b05615f7b7","https://git.kernel.org/stable/c/c2b169fc7a12665d8a675c1ff14bca1b9c63fb9a","https://nvd.nist.gov/vuln/detail/CVE-2025-37772"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-833","CWE-667"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-37868","cve":"CVE-2025-37868","aliases":[],"title":"Linux kernel (drivers/gpu/drm/xe): The Xe userptr path takes folio locks while holding the MMU notifier lock, which","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/xe)","year":"2025","cvss_score":6.5,"severity":"medium","kev":false,"impact":"The Xe userptr path takes folio locks while holding the MMU notifier lock, which inverts against core-kernel page migration doing the reverse. A tenant using userptr GPU mappings can deadlock the notifier against memory compaction/migration, hanging tasks that have nothing to do with the GPU and taking the node out for every workload on it.","attack_vector":"Any container with /dev/dri/renderD* on an Intel Xe host that registers userptr GPU mappings over its own anonymous memory; the collision happens when the kernel's page migration batches those folios at the same time as the driver marks them accessed/dirty. Needs no privilege - just userptr use plus normal memory pressure or THP compaction on the node.","remediation":"Update to 6.12.25 / 6.14 or later, which drops the unnecessary mark-accessed/dirty under the notifier lock. Interim: deny render-node access to untrusted tenants on xe hosts; nothing in the driver disables userptr independently.","references":["https://git.kernel.org/stable/c/65dc4e3d5b01db0179fc95c1f0bdb87194c28ab5","https://git.kernel.org/stable/c/90574ecf6052be83971d91d16600c5cf07003bbb","https://nvd.nist.gov/vuln/detail/CVE-2025-37868"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-48942","cve":"CVE-2025-48942","aliases":[],"title":"vLLM (`/v1/completions` guided decoding): Invalid `json_schema` kills the server","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (`/v1/completions` guided decoding)","year":"2025","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Invalid `json_schema` kills the server","attack_vector":"Unauthenticated network to an exposed serving port","remediation":"Upgrade to 0.9.0+; single-request DoS against a shared serving tier","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-48942"],"status":"curated","published":"2025-05-30"},{"cvss_vector":"CVSS:3.1/AV:A/AC:H/PR:L/UI:N/S:C/C:H/I:L/A:N","cwe":["CWE-269"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2025-52555","cve":"CVE-2025-52555","aliases":[],"title":"CephFS (ceph-fuse client): A tenant with an ordinary unprivileged UID on a node that has a CephFS volume mounted via","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"CephFS (ceph-fuse client)","year":"2025","cvss_score":6.5,"severity":"medium","kev":false,"impact":"A tenant with an ordinary unprivileged UID on a node that has a CephFS volume mounted via ceph-fuse can chmod its way to running code as root on that node, which in turn hands it whatever CephX key the mount was made with. On a shared GPU node that means the tenant inherits the mount's view of the whole file system, not just its own subtree.","attack_vector":"Any local unprivileged user on a compute node where ceph-fuse has mounted CephFS. No cluster network access is needed and no CephX credential of the attacker's own is required.","remediation":"Upgrade ceph-fuse to 17.2.8 / 18.2.5 / 19.2.3 or later on every node that fuse-mounts CephFS, then remount. Where possible prefer the kernel CephFS client for tenant nodes and scope each mount's CephX key to a single subtree so a compromised mount cannot read the rest of the tree.","references":["https://github.com/ceph/ceph/security/advisories/GHSA-89hm-qq33-2fjm","https://nvd.nist.gov/vuln/detail/CVE-2025-52555"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-55199","cve":"CVE-2025-55199","aliases":[],"title":"Helm: Crafted JSON Schema causes OOM termination of the renderer","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Helm","year":"2025","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Crafted JSON Schema causes OOM termination of the renderer","attack_vector":"Malicious chart","remediation":"Upgrade Helm to 3.18.5+","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-55199"],"status":"curated","published":"2025-08-14"},{"id":"CVE-2025-64433","cve":"CVE-2025-64433","aliases":[],"title":"KubeVirt: A VM reads arbitrary files from the virt-launcher pod filesystem","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"KubeVirt","year":"2025","cvss_score":6.5,"severity":"medium","kev":false,"impact":"A VM reads arbitrary files from the virt-launcher pod filesystem","attack_vector":"Any tenant VM","remediation":"Upgrade KubeVirt to 1.5.3/1.6.1+","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-64433"],"status":"curated","published":"2025-11-07"},{"cwe":["CWE-400","CWE-190"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-71104","cve":"CVE-2025-71104","aliases":[],"title":"Linux kernel (arch/x86/kvm): A guest using its APIC timer in periodic mode can leave KVM programming an already-expired","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (arch/x86/kvm)","year":"2025","cvss_score":6.5,"severity":"medium","kev":false,"impact":"A guest using its APIC timer in periodic mode can leave KVM programming an already-expired hypervisor timer over and over, producing a practically unbounded storm of host hrtimer interrupts. The upstream fix states this can hard-lock the host - a whole-node outage that takes every other tenant on the box down with it, so the impact is DoS rather than escape.","attack_vector":"The guest side is trivial and unprivileged: any tenant programs a periodic LAPIC timer. The trigger is a long gap in vCPU execution - VM pause/suspend, a live-migration blackout window, or severe host oversubscription - after which the expiration delta goes negative and overflows what the VMX preemption timer can encode. Intel hosts using the hypervisor timer, which is the default.","remediation":"Update to a stable kernel carrying the linked fix; the record's version list does not name a usable fixed release, so track the branch containing commit 786ed625c125. Interim controls: avoid long pause/suspend of running tenant VMs and, at a performance cost, disable the VMX preemption timer for KVM (kvm_intel.preemption_timer=0) so the software hrtimer path is used.","references":["https://git.kernel.org/stable/c/786ed625c125c5cd180d6aaa37e653e3e4ffb8d9","https://git.kernel.org/stable/c/807dbe8f3862fa7c164155857550ce94b36a11b9","https://nvd.nist.gov/vuln/detail/CVE-2025-71104"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-7445","cve":"CVE-2025-7445","aliases":[],"title":"secrets-store-sync-controller: Service account tokens disclosed in controller logs","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"secrets-store-sync-controller","year":"2025","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Service account tokens disclosed in controller logs","attack_vector":"Anyone with log-pipeline read access","remediation":"Controller rollout, no GPU drain; rotate exposed tokens; scrub logs","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated","published":"2025-09-05"},{"id":"CVE-2026-13208","cve":"CVE-2026-13208","aliases":[],"title":"KubeVirt: virt-handler notify server derives VMI identity from the request body without validating the connection","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"KubeVirt","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"virt-handler notify server derives VMI identity from the request body without validating the connection; cross-VM event spoofing","attack_vector":"An attacker with virt-launcher access","remediation":"Upgrade KubeVirt","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-13208"],"status":"curated","published":"2026-06-24"},{"id":"CVE-2026-18727","cve":"CVE-2026-18727","aliases":[],"title":"open-iscsi iscsiuio (DHCPv6 handling): Integer underflow and out-of-bounds read in iscsiuio's DHCPv6 handling","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"open-iscsi iscsiuio (DHCPv6 handling)","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Integer underflow and out-of-bounds read in iscsiuio's DHCPv6 handling. iscsiuio is the userspace daemon that drives iSCSI offload on Broadcom and QLogic adapters, and it processes DHCPv6 from the network to bring up the offload interface — so this is unauthenticated network input reaching a privileged storage daemon. Relevant to GPU clusters that boot or mount datasets over iSCSI, which is still common on the cheaper storage tiers.","attack_vector":"Unauthenticated, adjacent — a rogue DHCPv6 responder on the storage network. DHCPv6 has no authentication and responds fastest-wins, so this needs only presence on the segment.","remediation":"Upgrade open-iscsi and restart iscsiuio — package upgrade with a service restart; iSCSI sessions may briefly drop, so drain storage-dependent workloads first. Independently: disable IPv6 on storage networks that do not need it, or enforce DHCPv6 guard on the storage VLAN at the switch, both live config changes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-18727"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2026-08-12"},{"cwe":["CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-23441","cve":"CVE-2026-23441","aliases":[],"title":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core/en_accel): All IPsec offload objects on a physical function share","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core/en_accel)","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"All IPsec offload objects on a physical function share one DMA-mapped hardware context, and the driver drops the lock before the hardware has finished writing it. A second operation overwrites the buffer, so the first one reads another security association's state - the offloaded IPsec engine acts on cross-SA data (replay counters, lifetimes, soft/hard limits) where the fabric encryption terminates.","attack_vector":"Reachable wherever mlx5 IPsec crypto offload is enabled on the node. Concurrency is supplied by normal traffic across multiple SAs, so a peer on the encrypted fabric can time SA queries and updates against each other; a tenant that owns any IPsec SA on the shared PF contributes operations to the same context. Conditional on IPsec full/crypto offload being configured - not reachable if IPsec is done in software.","remediation":"Update to a kernel carrying the fix on your stream. Interim: disable mlx5 IPsec crypto/full offload and terminate IPsec in software until nodes are rebooted onto a fixed kernel, especially where SAs belong to different tenants on the same physical function.","references":["https://git.kernel.org/stable/c/99aaee927800ea00b441b607737f9f67b1899755","https://git.kernel.org/stable/c/c3db55dc0f3344b62da25b025a8396d78763b5fa","https://nvd.nist.gov/vuln/detail/CVE-2026-23441"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-24182","cve":"CVE-2026-24182","aliases":[],"title":"GPU Display Driver: DoS / privesc (race in GPU memory management)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"DoS / privesc (race in GPU memory management)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24182","https://github.com/NVIDIA/product-security/tree/main/2026/5821"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:N/I:N/A:H","cwe":["CWE-667"],"fleet":{"pain_class":"node-reboot"},"published":"2026-05-26"},{"id":"CVE-2026-24197","cve":"CVE-2026-24197","aliases":[],"title":"GPU Display Driver: Cross-tenant info disclosure via GPU memory leakage","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Cross-tenant info disclosure via GPU memory leakage","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24197","https://github.com/NVIDIA/product-security/tree/main/2026/5821"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:N/I:N/A:H","cwe":["CWE-1188"],"fleet":{"pain_class":"node-reboot"},"published":"2026-05-26"},{"id":"CVE-2026-24204","cve":"CVE-2026-24204","aliases":[],"title":"NVIDIA FLARE SDK: Auth bypass via insufficient input validation","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA FLARE SDK","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Auth bypass via insufficient input validation","attack_vector":"Network peer","remediation":"Upgrade FLARE; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24204","https://github.com/NVIDIA/product-security/tree/main/2026/5819"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-20"],"published":"2026-04-28"},{"id":"CVE-2026-24514","cve":"CVE-2026-24514","aliases":[],"title":"ingress-nginx: Admission controller denial of service","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Admission controller denial of service","attack_vector":"Any pod on the cluster network","remediation":"Rolling controller upgrade","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated","published":"2026-02-03"},{"cwe":["CWE-476","CWE-696"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-31425","cve":"CVE-2026-31425","aliases":[],"title":"Linux kernel RDS over InfiniBand (FRMR registration before connection establishment): An RDS sendmsg carrying","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel RDS over InfiniBand (FRMR registration before connection establishment)","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"An RDS sendmsg carrying RDS_CMSG_RDMA_MAP parses the control message and starts fast-registration memory-region work before the RDMA connection exists, dereferencing a NULL rdma_cm_id to reach its queue pair. The existing guard only checked whether the connection object was present, not whether it had a cm_id. Memory registration executing against a connection that has not been established is registration outside any established security context - and the immediate effect is that an unprivileged tenant crashes the node with one sendmsg.","attack_vector":"Local, unprivileged. One sendmsg with RDS_CMSG_RDMA_MAP on a freshly created outgoing RDS connection.","remediation":"Kernel update strengthening the guard in rds_ib_reg_frmr() to check the cm_id. Blacklist rds/rds_rdma - the fastest fix for a module you almost certainly do not need.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=23e07c340c445f0ebff7757ba15434cb447eb662","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-31425.json"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-476","CWE-665"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-31601","cve":"CVE-2026-31601","aliases":[],"title":"Linux kernel (drivers/vfio/pci/xe): Resetting a passed-through Intel GPU virtual function that does not support","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/vfio/pci/xe)","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Resetting a passed-through Intel GPU virtual function that does not support migration makes the host dereference a NULL migration context and oops in the reset-done callback. The host kernel goes down on a GPU-passthrough node, so one tenant resetting its own card is an outage for every tenant sharing that box.","attack_vector":"The reproducer runs through the sysfs reset attribute, which needs host root, but the same pci_reset_function() path is what runs on VFIO_DEVICE_RESET and on device-fd release - both of which a tenant holding /dev/vfio/<group> drives directly. Conditional on the xe_vfio_pci variant driver being bound to an Intel Xe SR-IOV VF that lacks migration support. Not reachable on NVIDIA-only fleets.","remediation":"Update to a stable kernel carrying commits 8fa4113f / 73e53ff1. Interim: on Intel GPU SR-IOV nodes, bind VFs to plain vfio-pci rather than xe_vfio_pci where live migration is not needed, and block tenant-initiated device reset in the VMM.","references":["https://git.kernel.org/stable/c/8fa4113fc65b8b29a30fbbca5fd82221dc6e146e","https://git.kernel.org/stable/c/73e53ff144a538f1843b3dea1e2740a755031cdc","https://nvd.nist.gov/vuln/detail/CVE-2026-31601"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-3864","cve":"CVE-2026-3864","aliases":[],"title":"CSI Driver NFS: Path traversal via `subDir` lets a tenant delete unintended directories on the shared NFS server","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"CSI Driver NFS","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Path traversal via `subDir` lets a tenant delete unintended directories on the shared NFS server; cross-tenant data destruction","attack_vector":"Cluster user able to create a PV/PVC with a crafted subDir","remediation":"DaemonSet/controller rollout; add admission validation on subDir. High priority for neoclouds sharing one NFS backend across tenants","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated","published":"2026-03-20"},{"id":"CVE-2026-3865","cve":"CVE-2026-3865","aliases":[],"title":"CSI Driver SMB: Same `subDir` path traversal against a shared SMB server","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"CSI Driver SMB","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Same `subDir` path traversal against a shared SMB server","attack_vector":"Cluster user able to create a PV/PVC with a crafted subDir","remediation":"DaemonSet/controller rollout; validate subDir at admission","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated"},{"cwe":["CWE-617","CWE-131"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-43107","cve":"CVE-2026-43107","aliases":[],"title":"Linux kernel (net/xfrm): The async-event reply buffer was sized without accounting for the interface-ID attribute, so","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"The async-event reply buffer was sized without accounting for the interface-ID attribute, so querying an SA that has an if_id set overflows the reply skb, returns -EMSGSIZE, and lands on a BUG_ON - a deterministic kernel panic from a single netlink query. One tenant with namespace-scoped network privilege takes the whole node down and every co-tenant loses its GPUs.","attack_vector":"Needs an xfrm netlink socket and an SA carrying an if_id, then one XFRM_MSG_GETAE request. That is host root or, more interestingly, a tenant container holding CAP_NET_ADMIN in its own user+network namespace: it can create its own SA with an if_id and immediately query it, so it does not depend on the host's SA configuration at all. No fabric access needed.","remediation":"Boot a kernel carrying the fix commits below (no fixed stable version published). Interim control: remove CAP_NET_ADMIN from tenant user namespaces, or block the XFRM netlink family for tenant workloads - there is no way to make the panic survivable once the request reaches the kernel.","references":["https://git.kernel.org/stable/c/2c41283d94af943a05f7f2cc1a01f0c872f3cf43","https://git.kernel.org/stable/c/e62e322ea20be78e346e4b49f9a6b9f03313af4c","https://nvd.nist.gov/vuln/detail/CVE-2026-43107"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-362","CWE-459"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-43220","cve":"CVE-2026-43220","aliases":[],"title":"Linux kernel (drivers/iommu/amd): AMD-Vi hands out the completion-wait sequence number outside the IOMMU lock, so","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/amd)","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"AMD-Vi hands out the completion-wait sequence number outside the IOMMU lock, so completion commands get queued out of order and the driver's wait for 'invalidation finished' matches the wrong command or times out. The driver then continues as though an IOTLB flush completed when it did not, leaving stale translations usable by devices; the timeout storms themselves stall the invalidation path for every tenant on the node.","attack_vector":"Concurrent TLB invalidations - on a shared node that means several tenants issuing DMA map/unmap through vfio or iommufd at the same time, or heavy device DMA with SVA/PASID. This is load, not a crafted request: no host root and no special ioctl sequence, just enough parallel IOMMU traffic. AMD-Vi (EPYC) hosts only.","remediation":"Update to 6.6.140 or 6.12.88 or later. No safe interim control on AMD hosts other than reducing concurrent passthrough DMA churn per node; do not treat unmap-then-reuse as a hard boundary until patched.","references":["https://git.kernel.org/stable/c/d51bf43193b1e95dc4e34e540dc76e19def2ae5a","https://git.kernel.org/stable/c/fca7aa0264ae99e5ff287d0ced5af0b82b121c4f","https://nvd.nist.gov/vuln/detail/CVE-2026-43220"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-46433","cve":"CVE-2026-46433","aliases":[],"title":"lldpd (802.1Q VLAN tag stripping in lldpd_decode): lldpd strips 802.1Q VLAN tags by memmove-ing the frame payload four","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"lldpd (802.1Q VLAN tag stripping in lldpd_decode)","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"lldpd strips 802.1Q VLAN tags by memmove-ing the frame payload four bytes left, and the byte count is not correctly bounded — so a crafted tagged frame drives a bad copy. Worth noting for anyone running lldpd on trunk ports, which in a leaf/spine fabric is most of them.","attack_vector":"Unauthenticated, adjacent — a crafted VLAN-tagged Ethernet frame on a port where lldpd is listening.","remediation":"Upgrade lldpd to 1.0.22 or later and restart the daemon. Package upgrade plus service restart; on switch NOSes it comes with the image update.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-46433"],"status":"curated","fleet":{"pain_class":"daemon-restart"},"published":"2026-06-09"},{"id":"CVE-2026-47155","cve":"CVE-2026-47155","aliases":[],"title":"vLLM (revision pinning): Revision pinning does not apply to all model artifacts","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (revision pinning)","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Revision pinning does not apply to all model artifacts → supply-chain drift","attack_vector":"Poisoned Hub repo where a pinned revision is silently not enforced","remediation":"Upgrade to 0.22.0+. Undermines model-supply-chain controls the provider may be advertising","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47155"],"status":"curated","published":"2026-06-22"},{"id":"CVE-2026-47481","cve":"CVE-2026-47481","aliases":[],"title":"Triton Inference Server: MITM via weak TLS verification","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"MITM via weak TLS verification","attack_vector":"Network attacker on the serving path","remediation":"Upgrade Triton; verify TLS config","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47481","https://github.com/NVIDIA/product-security/tree/main/2026/5853"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:L/I:L/A:N","cwe":["CWE-288"],"published":"2026-07-14"},{"id":"CVE-2026-47606","cve":"CVE-2026-47606","aliases":[],"title":"NVIDIA Triton Inference Server: An absolute path traversal reaches code execution and information disclosure","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"An absolute path traversal reaches code execution and information disclosure - the attacker reads or writes outside the model repository. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5865. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47606","https://github.com/NVIDIA/product-security/tree/main/2026/5865"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:L/I:L/A:N","cwe":["CWE-36"],"published":"2026-08-18"},{"id":"CVE-2026-47620","cve":"CVE-2026-47620","aliases":[],"title":"NVIDIA Dynamo: DoS / corruption via race in resource sync","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Dynamo","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"DoS / corruption via race in resource sync","attack_vector":"Concurrent inference clients","remediation":"Bump Dynamo; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47620","https://github.com/NVIDIA/product-security/tree/main/2026/5842"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:N/I:L/A:H","cwe":["CWE-362"],"published":"2026-08-04"},{"id":"CVE-2026-47621","cve":"CVE-2026-47621","aliases":[],"title":"NVIDIA Dynamo: TOCTOU race","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Dynamo","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"TOCTOU race","attack_vector":"Local/network client","remediation":"Bump Dynamo; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47621","https://github.com/NVIDIA/product-security/tree/main/2026/5842"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:N/I:L/A:H","cwe":["CWE-367"],"published":"2026-08-04"},{"cwe":["CWE-833","CWE-667"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-53197","cve":"CVE-2026-53197","aliases":[],"title":"Linux kernel (net/xfrm): Tearing down an IPTFS security association cancels its hrtimers while holding the very locks","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Tearing down an IPTFS security association cancels its hrtimers while holding the very locks those timer callbacks take. On SMP that closes an ABBA cycle - one CPU spins on the softirq expiry lock while the softirq holding it waits for the spinlock the canceller owns. Neither side ever progresses, so a lock held on the IPsec datapath is never released and packet processing on the node stops. Every tenant on that host loses fabric connectivity until it is power-cycled.","attack_vector":"Requires the ability to delete an IPTFS xfrm state while its output or drop timer is firing - host root, or a tenant container holding CAP_NET_ADMIN in its own user+network namespace, which can create and delete IPTFS SAs in a loop while pushing traffic through them to keep the timers hot. Conditional on IP-TFS mode being available. Found by source-code audit rather than a fuzzer, so treat the deadlock as reachable but not yet weaponised in public.","remediation":"Boot a kernel carrying the fix commits below (no fixed stable version published). Interim control: remove CAP_NET_ADMIN from tenant user namespaces so tenants cannot churn IPTFS SAs, or blacklist xfrm_iptfs where IP-TFS is not deliberately in use.","references":["https://git.kernel.org/stable/c/a13ca53e47e500854a3b9ec18b5dc83acfec863e","https://git.kernel.org/stable/c/822b98d354e63e8249e85473c5f3c519f3c9cecc","https://nvd.nist.gov/vuln/detail/CVE-2026-53197"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-667","CWE-400"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-53274","cve":"CVE-2026-53274","aliases":[],"title":"Linux kernel (net/smc): Setsockopt() on an SMC socket copies the option value from user memory while holding the socket","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/smc)","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Setsockopt() on an SMC socket copies the option value from user memory while holding the socket lock, so a tenant pointing optval at a userfaultfd- or FUSE-backed page can hold that lock forever. Combined with asynchronous shutdown work, this drains the kernel worker pool and trips the hung-task watchdog - one tenant's process wedges kworkers for every workload on the node.","attack_vector":"Local and fully unprivileged: register a userfaultfd region (or mmap a FUSE-backed file where unprivileged userfaultfd is disabled), pass it as optval to setsockopt() on an AF_SMC socket, then call shutdown() from another thread. socket(AF_SMC, ...) autoloads the module with no capability check, so any tenant container reaches it. This is a noisy-neighbour outage, not a corruption bug - but it is a whole-node one.","remediation":"Boot a kernel carrying the fix commits (moves the user copy outside lock_sock). Interim: blacklist the smc module or deny socket family 43 in tenant seccomp profiles; disabling unprivileged userfaultfd (vm.unprivileged_userfaultfd=0) narrows but does not close the path, since FUSE-backed memory works too.","references":["https://git.kernel.org/stable/c/35a22117839602bb52283de08894c5a7dde92420","https://git.kernel.org/stable/c/89f6fbe0033c942cb790ffd53ca93a45eeaf1c91","https://nvd.nist.gov/vuln/detail/CVE-2026-53274"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-54235","cve":"CVE-2026-54235","aliases":[],"title":"vLLM - sampling parameter validation: Temperature validation uses strict comparison operators, so boundary values slip","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM - sampling parameter validation","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Temperature validation uses strict comparison operators, so boundary values slip through the gate silently. In a multi-tenant serving deployment this is a request-parameter validation gap rather than memory corruption, but it lets a caller push the engine into a sampling state the operator believed was fenced off - which matters where sampling parameters are part of a tenant-facing safety or cost control.","attack_vector":"Network. Any client able to submit a request with sampling parameters to the vLLM endpoint.","remediation":"Upgrade to vLLM 0.23.1rc0 or later and roll the serving deployment. Cost: rolling restart of the inference tier; no driver or firmware change.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-54235"],"status":"curated","fleet":{"pain_class":"daemon-restart"},"published":"2026-06-22"},{"id":"CVE-2026-56794","cve":"CVE-2026-56794","aliases":[],"title":"Dell OpenManage Server Administrator (relative path traversal): A low-privileged remote attacker reads arbitrary files","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Dell OpenManage Server Administrator (relative path traversal)","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"A low-privileged remote attacker reads arbitrary files from the server hosting OMSA.","attack_vector":"Authenticated low-privilege network access to OMSA below 11.1.0.2.","remediation":"Upgrade OMSA to 11.1.0.2, or remove the agent if unused.","references":["https://www.dell.com/support/kbdoc/en-us/000494958/dsa-2026-326-security-update-for-dell-openmanage-server-administrator-omsa-network-access-vulnerabilities"],"status":"curated"},{"id":"CVE-2026-63457","cve":"CVE-2026-63457","aliases":["HPESBHF05090"],"title":"HPE iLO 6 (denial of service): An unauthenticated attacker on an adjacent network can knock out iLO 6 availability","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HPE iLO 6 (denial of service)","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"An unauthenticated attacker on an adjacent network can knock out iLO 6 availability. No data is read or written - the damage is losing the out-of-band plane itself. For a GPU fleet that is a real operational hit rather than a cosmetic one: when iLO is down you cannot power-cycle a hung training node, cannot console into it to see why the job died, and cannot push firmware or re-provision it. A coordinated version of this against a rack turns a recoverable incident into a physical dispatch. Affects iLO 6 before v1.78, so it covers the Gen11 fleet.","attack_vector":"Adjacent network, unauthenticated - anything sharing the management segment with the iLOs. No account, no host access, no user interaction.","remediation":"Flash iLO 6 to v1.78 or later. Out-of-band, per-node, no host reboot and no job drain - which makes this a cheap fix relative to the availability risk it removes. Because the vector is adjacent-network and unauthenticated, network segmentation is the meaningful compensating control while the rollout runs: keep BMCs off any segment shared with general-purpose hosts.","references":["https://support.hpe.com/hpesc/public/docDisplay?docId=hpesbhf05090en_us&docLocale=en_US","https://nvd.nist.gov/vuln/detail/CVE-2026-63457"],"status":"curated","published":"2026-08-05"},{"cwe":["CWE-834","CWE-1284"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-64289","cve":"CVE-2026-64289","aliases":[],"title":"Linux kernel (drivers/iommu/iommufd): A tenant using nested translation can ask iommufd to process a cache-invalidation","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/iommufd)","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"A tenant using nested translation can ask iommufd to process a cache-invalidation batch of essentially unbounded size, pinning a CPU in a non-preemptible loop until the soft-lockup watchdog fires. On a node booted with softlockup_panic that is a panic; otherwise it is a burnt core plus a wedged IOMMU invalidation path that every other tenant on the box shares.","attack_vector":"One IOMMUFD_CMD_HWPT_INVALIDATE ioctl from a tenant container holding /dev/iommu, with entry_num or entry_len set near U32_MAX. No host root and no special hardware beyond a nested/vIOMMU setup (VT-d nested is the case called out in the fix). Trivially repeatable, so a tenant can burn cores one at a time.","remediation":"Update to a stable kernel carrying commits d2bd041e / 32ca4aed. Interim: do not expose /dev/iommu to untrusted tenants directly - run passthrough through an operator-controlled VMM - and do not boot tenant nodes with softlockup_panic, which converts this from a stalled core into a full node panic.","references":["https://git.kernel.org/stable/c/d2bd041e0efaf7d81789779b135279d18b33d6d5","https://git.kernel.org/stable/c/32ca4aed2a66205b072fcfecabe220289a8149ff","https://nvd.nist.gov/vuln/detail/CVE-2026-64289"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-770","CWE-789"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-64291","cve":"CVE-2026-64291","aliases":[],"title":"Linux kernel (drivers/iommu/iommufd): Iommufd accepted any non-zero virtual event queue depth up to U32_MAX, so a","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/iommufd)","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Iommufd accepted any non-zero virtual event queue depth up to U32_MAX, so a tenant can ask the kernel to allocate an arbitrarily large queue and drive the node into memory exhaustion. That is an OOM kill storm or a hung box for every tenant on the host, from one ioctl.","attack_vector":"A single IOMMUFD_CMD_VEVENTQ_ALLOC from a container holding /dev/iommu, with veventq_depth set enormous. No host root, no hardware prerequisite beyond a vIOMMU/nested setup that exposes virtual event queues.","remediation":"Update to a stable kernel carrying commits f565297e / e7b5e556, which caps the depth. Interim: keep /dev/iommu out of tenant containers and set per-container memory cgroup limits that the kernel allocation is charged against where your kernel supports it.","references":["https://git.kernel.org/stable/c/f565297edf316016be4a1a9e2eb9f39359313f43","https://git.kernel.org/stable/c/e7b5e55652746b1221b9c10ff80eae8a154101ba","https://nvd.nist.gov/vuln/detail/CVE-2026-64291"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-825","CWE-754"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-64579","cve":"CVE-2026-64579","aliases":[],"title":"Linux kernel (net/xfrm): The policy-hash rebuild preallocates for exactly the wrong half of the policy set - the guard","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"The policy-hash rebuild preallocates for exactly the wrong half of the policy set - the guard is inverted, so the inexact policies (the ones whose reinsert actually allocates) are skipped. When the GFP_ATOMIC allocation then fails mid-rebuild the code only warns and continues, leaving a policy with a poisoned list node: it is neither properly on the inexact chain nor off it. Two consequences for an operator - a policy that should have matched can stop matching, so traffic that was supposed to be tunnelled leaves in clear, and the next rebuild dereferences LIST_POISON2 and panics the node.","attack_vector":"The rebuild is queued by XFRM_MSG_NEWSPDINFO or a policy-threshold change, which needs CAP_NET_ADMIN - host root or a tenant container holding it in its own user+network namespace. The second ingredient is an atomic allocation failure during the rebuild, which is deterministic under failslab and realistically reachable under genuine memory pressure; on a shared node a co-tenant can supply that pressure. Requires a non-trivial number of inexact policies installed.","remediation":"Boot a kernel carrying the fix commits below (no fixed stable version published). Interim control: deny CAP_NET_ADMIN in tenant user namespaces so tenants cannot trigger rebuilds, and keep memory headroom on nodes running IPsec so atomic allocations do not fail.","references":["https://git.kernel.org/stable/c/e48f4c3e3df35b34be719d72d737bbeaca77cf0c","https://git.kernel.org/stable/c/1cdeed9df1306f1a277e715600772640d63defa9","https://nvd.nist.gov/vuln/detail/CVE-2026-64579"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-772"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-72128","cve":"CVE-2026-72128","aliases":[],"title":"Linux kernel (drivers/nvme/target): A client that asks the target to create a submission queue with an invalid queue ID","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/nvme/target)","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"A client that asks the target to create a submission queue with an invalid queue ID leaks a controller reference on every attempt. The controller object and everything hanging off it (queues, buffers, AER state) is then never freed even after the client disconnects, so a peer can pin target memory indefinitely and eventually starve the shared storage node.","attack_vector":"Reachable by any peer that has established a connection to the nvmet subsystem and can issue an admin Create SQ command with an out-of-range sqid - a one-field change in a command the client fully controls. No target-side privilege and no tenant device node required. Conditional on nvmet being configured and exporting a subsystem the peer can reach.","remediation":"No fixed version is listed on this record - boot a kernel carrying the linked stable commits. Interim: restrict which hosts may connect to the subsystem (host NQN allow-list rather than allow-any-host), and watch for controllers that never disappear after disconnect.","references":["https://git.kernel.org/stable/c/26355295ce21cb046546085c3a81abe68160a784","https://git.kernel.org/stable/c/fcef60ed5f714a24104eb021d6397a67955ebeff","https://nvd.nist.gov/vuln/detail/CVE-2026-72128"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-8045","cve":"CVE-2026-8045","aliases":[],"title":"Schneider Electric Data Center Expert - SOAP service endpoints: XML external entity processing on DCE SOAP endpoints","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Schneider Electric Data Center Expert - SOAP service endpoints","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"XML external entity processing on DCE SOAP endpoints lets an authenticated user read server-side files. On a DCIM appliance that means configuration and credential material for the power and cooling estate - the recurring theme with DCE is that any read primitive is a facility-wide credential leak.","attack_vector":"Any user with a DCE account submitting crafted XML to the SOAP endpoints.","remediation":"Apply the Schneider fix for DCE. Audit DCE account holders in the same pass. Cheap software update; the credential rotation afterwards is the expensive half.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-8045"],"status":"curated","published":"2026-06-09"},{"id":"NCVD-2019-003-rnic-on-board-sram-metadata-cach","cve":null,"aliases":["Pythia","RDMA remote side channel","RNIC SRAM/PTE cache timing attack","Tsai, Payer, Zhang - USENIX Security 2019"],"title":"RNIC on-board SRAM metadata cache (page table entries, QP context) - most widely deployed RDMA NIC: RNICs cache","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"RNIC on-board SRAM metadata cache (page table entries, QP context) - most widely deployed RDMA NIC","year":"2019","cvss_score":6.5,"severity":"medium","kev":false,"impact":"RNICs cache page-table entries and connection context in a small on-board SRAM and spill to host memory over PCIe when it overflows. Pythia turns the resulting timing difference into a remote side channel: an attacker on one client machine learns the memory access patterns of victims on other client machines against a shared in-memory data service. No memory contents are read directly, but access patterns over a key-value store or a shared embedding table are often enough to recover which records a victim touched. In an AI cluster the same primitive applies to shared parameter servers and RDMA-backed caches, where access pattern equals query content.","attack_vector":"The attacker is an ordinary RDMA client of the same server - no special privilege, no injection needed. They issue their own RDMA reads to addresses chosen to contend for specific RNIC SRAM cache sets, then measure completion latency to infer whether a victim's access evicted their entry. The authors reverse-engineered the memory architecture of the most widely deployed RNIC to make the eviction sets precise, raising the channel's efficiency substantially.","remediation":"No patch. Mitigations are all structural: do not let mutually untrusted tenants share an RNIC or a server-side RDMA data service; partition the server's registered memory so different tenants' regions do not share cache sets; or add deliberate noise/padding to server-side access patterns (application change with a throughput cost). Where a DPU fronts the fabric, terminating tenant connections on separate DPU cores reduces sharing. Scheduling policy - not co-locating untrusted tenants on the same RDMA service - is the realistic control and costs bin-packing efficiency, not downtime.","references":["https://www.usenix.org/conference/usenixsecurity19/presentation/tsai","https://www.usenix.org/conference/usenixsecurity22/presentation/xing"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"tags":["tenant-isolation"]},{"id":"NCVD-2019-006-rnic-on-board-sram-metadata-cach","cve":null,"aliases":["Pythia","RDMA remote side channel","RNIC SRAM/PTE cache timing attack","Tsai, Payer, Zhang - USENIX Security 2019"],"title":"RNIC on-board SRAM metadata cache (page table entries, QP context) - most widely deployed RDMA NIC: RNICs cache","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"RNIC on-board SRAM metadata cache (page table entries, QP context) - most widely deployed RDMA NIC","year":"2019","cvss_score":6.5,"severity":"medium","kev":false,"impact":"RNICs cache page-table entries and connection context in a small on-board SRAM and spill to host memory over PCIe when it overflows. Pythia turns the resulting timing difference into a remote side channel: an attacker on one client machine learns the memory access patterns of victims on other client machines against a shared in-memory data service. No memory contents are read directly, but access patterns over a key-value store or a shared embedding table are often enough to recover which records a victim touched. In an AI cluster the same primitive applies to shared parameter servers and RDMA-backed caches, where access pattern equals query content.","attack_vector":"The attacker is an ordinary RDMA client of the same server - no special privilege, no injection needed. They issue their own RDMA reads to addresses chosen to contend for specific RNIC SRAM cache sets, then measure completion latency to infer whether a victim's access evicted their entry. The authors reverse-engineered the memory architecture of the most widely deployed RNIC to make the eviction sets precise, raising the channel's efficiency substantially.","remediation":"No patch. Mitigations are all structural: do not let mutually untrusted tenants share an RNIC or a server-side RDMA data service; partition the server's registered memory so different tenants' regions do not share cache sets; or add deliberate noise/padding to server-side access patterns (application change with a throughput cost). Where a DPU fronts the fabric, terminating tenant connections on separate DPU cores reduces sharing. Scheduling policy - not co-locating untrusted tenants on the same RDMA service - is the realistic control and costs bin-packing efficiency, not downtime.","references":["https://www.usenix.org/conference/usenixsecurity19/presentation/tsai","https://www.usenix.org/conference/usenixsecurity22/presentation/xing"],"status":"curated","tags":["tenant-isolation"]},{"id":"NCVD-2022-003-rocev2-congestion-control-dcqcn","cve":null,"aliases":["DCQCN manipulation","RoCE congestion-control abuse","CNP spoofing","arXiv:2207.10898"],"title":"RoCEv2 congestion control - DCQCN, ECN marking and Congestion Notification Packets: DCQCN reacts to ECN marks by having","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"RoCEv2 congestion control - DCQCN, ECN marking and Congestion Notification Packets","year":"2022","cvss_score":6.5,"severity":"medium","kev":false,"impact":"DCQCN reacts to ECN marks by having the receiver send Congestion Notification Packets that make the sender cut its rate. CNPs are unauthenticated RoCE packets like any other, so an attacker who can spoof onto the fabric can forge CNPs at a victim's sender and drive its rate to the floor while their own traffic is unaffected - a targeted bandwidth-theft and starvation primitive. Conversely, a tenant whose NIC simply ignores CNPs takes an unfair share and pushes everyone else into PFC. Published measurement shows the native PFC-based scheme already suffers unfairness and head-of-line blocking, and that congestion-control choice materially changes distributed DNN training time, so the manipulation lands directly on job completion times in a GPU cluster.","attack_vector":"Forged CNPs need only a spoofed source GID and the victim's QP number - the same predictability that makes packet injection work. Rate-ignoring is even simpler: run a NIC configuration or a custom firmware/driver that under-responds to congestion notifications, which looks like a tuning choice rather than an attack. Both are invisible to host-level monitoring; they show up only as unexplained throughput asymmetry between tenants.","remediation":"Config change: enforce switch-side per-tenant rate limiting and ECN marking policy rather than trusting endpoint congestion response, apply source-address filtering so CNPs cannot be spoofed across tenants, and keep tenants in separate traffic classes so an unresponsive one cannot starve others. Standardise and lock the DCQCN parameter set through the NIC driver configuration (mlxconfig / sysfs) so tenants cannot retune their own NICs - a driver-level config change applied at provisioning, no reboot. Per-tenant switch queues are the durable fix and may require a QoS profile change plus a switch reload on constrained platforms. Monitor per-QP CNP counts as a detection signal.","references":["https://arxiv.org/abs/2207.10898","https://arxiv.org/abs/1806.08159"],"status":"curated","fleet":{"pain_class":"hot-patch"},"tags":["fabric-dos"]},{"id":"NCVD-2023-006-rnic-microarchitectural-resource","cve":null,"aliases":["Husky","RDMA performance isolation test suite","RNIC microarchitecture resource contention","Kong et al., NSDI 2023"],"title":"RNIC microarchitectural resources (NIC cache, processing units) under multi-tenant RDMA: This is the paper that","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"RNIC microarchitectural resources (NIC cache, processing units) under multi-tenant RDMA","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"This is the paper that established RDMA performance isolation in the cloud is not solved. The authors built a test suite (released as host-bench/husky) that models how RDMA operations consume RNIC microarchitecture resources, and report that it breaks every existing performance-isolation solution in various scenarios - a result acknowledged and reproduced by one of the largest RDMA NIC vendors. For an operator selling RDMA into guest VMs or containers, this means a neighbouring tenant can degrade your customer's collective-communication bandwidth at will, and no NIC-level QoS knob currently on the market reliably prevents it. Jobs miss their step-time SLOs with no attributable cause in host metrics.","attack_vector":"A co-resident tenant issues RDMA verb patterns chosen to thrash specific RNIC resources - for example many small operations across many queue pairs and memory regions to blow out the NIC's address-translation cache, or operation mixes that monopolise particular NIC processing stages. The attacker needs nothing more than normal RDMA access on a shared NIC; the damage is done inside the NIC where host-side rate limiters and cgroups have no visibility.","remediation":"No patch. Run the Husky suite against your own NIC/firmware/isolation configuration before promising RDMA SLAs - that is a test-harness exercise, not a change window. Practical controls: give each tenant a dedicated VF with vendor rate limiters plus caps on QP and MR counts (driver config change, applied at VF creation), keep per-tenant working sets small enough to stay resident in NIC cache, and where the risk is unacceptable, dedicate physical NICs. Upgrading to newer RNIC generations with larger caches and better per-VF quotas helps but does not close it - firmware flash plus driver upgrade, rolling host reboots.","references":["https://www.usenix.org/conference/nsdi23/presentation/kong","https://github.com/host-bench/husky"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"tags":["tenant-isolation"]},{"id":"NCVD-2023-008-rnic-microarchitectural-resource","cve":null,"aliases":["Husky","RDMA performance isolation test suite","RNIC microarchitecture resource contention","Kong et al., NSDI 2023"],"title":"RNIC microarchitectural resources (NIC cache, processing units) under multi-tenant RDMA: This is the paper that","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"RNIC microarchitectural resources (NIC cache, processing units) under multi-tenant RDMA","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"This is the paper that established RDMA performance isolation in the cloud is not solved. The authors built a test suite (released as host-bench/husky) that models how RDMA operations consume RNIC microarchitecture resources, and report that it breaks every existing performance-isolation solution in various scenarios - a result acknowledged and reproduced by one of the largest RDMA NIC vendors. For an operator selling RDMA into guest VMs or containers, this means a neighbouring tenant can degrade your customer's collective-communication bandwidth at will, and no NIC-level QoS knob currently on the market reliably prevents it. Jobs miss their step-time SLOs with no attributable cause in host metrics.","attack_vector":"A co-resident tenant issues RDMA verb patterns chosen to thrash specific RNIC resources - for example many small operations across many queue pairs and memory regions to blow out the NIC's address-translation cache, or operation mixes that monopolise particular NIC processing stages. The attacker needs nothing more than normal RDMA access on a shared NIC; the damage is done inside the NIC where host-side rate limiters and cgroups have no visibility.","remediation":"No patch. Run the Husky suite against your own NIC/firmware/isolation configuration before promising RDMA SLAs - that is a test-harness exercise, not a change window. Practical controls: give each tenant a dedicated VF with vendor rate limiters plus caps on QP and MR counts (driver config change, applied at VF creation), keep per-tenant working sets small enough to stay resident in NIC cache, and where the risk is unacceptable, dedicate physical NICs. Upgrading to newer RNIC generations with larger caches and better per-VF quotas helps but does not close it - firmware flash plus driver upgrade, rolling host reboots.","references":["https://www.usenix.org/conference/nsdi23/presentation/kong","https://github.com/host-bench/husky"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.0/AV:N/AC:H/PR:N/UI:N/S:U/C:H/I:L/A:N","cwe":["CWE-295"],"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2024-010-ceph-python-bindings-imap4-ssl-s","cve":null,"aliases":["GHSA-xj9f-7g59-m4jx","CVE-2024-31884 (reserved)"],"title":"Ceph (Python bindings, IMAP4_SSL/SMTP_SSL TLS clients): Ceph's Python code constructs imaplib.IMAP4_SSL and","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ceph (Python bindings, IMAP4_SSL/SMTP_SSL TLS clients)","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Ceph's Python code constructs imaplib.IMAP4_SSL and smtplib.SMTP_SSL without passing an SSL context, so no X.509 validation happens and any certificate is accepted. An attacker positioned on the path between a Ceph manager node and the mail server it uses for alerting can present their own certificate, harvest the mail credentials Ceph authenticates with, and read or tamper with cluster alert traffic. For a cluster operator the practical loss is twofold: reusable SMTP/IMAP credentials (frequently shared with other infrastructure) and the ability to suppress or forge the alert channel that is supposed to tell you a storage node is failing or a tenant is misbehaving.","attack_vector":"Network, machine-in-the-middle between the Ceph mgr node and its configured mail server. No credentials and no user interaction needed — the attacker only needs a position on the path, which is realistic when alerting egresses over a shared or upstream-provider network.","remediation":"Upgrade to a build carrying the fix (20.2.1, 19.2.4, 18.2.9 or later) and restart the mgr modules. Rotate the SMTP/IMAP credentials Ceph was configured with, since a MITM window means they may already be captured. Where possible route alert egress over a path you control rather than shared transit.","references":["https://github.com/ceph/ceph/security/advisories/GHSA-xj9f-7g59-m4jx","https://github.com/ceph/ceph/pull/66089"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-20","CWE-400","CWE-770","CWE-789"],"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2025-018-vllm-openai-compatible-server-ch","cve":null,"aliases":["GHSA-6fvq-23cw-5628","CVE-2025-61620 (reserved)"],"title":"vLLM OpenAI-compatible server (chat_template / chat_template_kwargs): NOISY-NEIGHBOUR DENIAL OF SERVICE: one tenant","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM OpenAI-compatible server (chat_template / chat_template_kwargs)","year":"2025","cvss_score":6.5,"severity":"medium","kev":false,"impact":"NOISY-NEIGHBOUR DENIAL OF SERVICE: one tenant takes down a shared inference endpoint, and the GPU sits idle while it happens. The OpenAI-compatible server lets callers supply a Jinja template through chat_template. Jinja has loops and nesting, so a small request body can pin CPU and balloon memory until the server stops answering anyone. The operationally important detail is that blocking the chat_template parameter is not enough: the server builds its kwargs dict and then calls dict.update() with the caller's chat_template_kwargs, so an attacker simply nests a chat_template key inside chat_template_kwargs and overwrites the template anyway. Any filtering you wrote against the obvious parameter name misses the real one. On a GPU fleet the cost is the expensive resource stranded behind a wedged Python process, plus every co-tenant on that replica losing service.","attack_vector":"Network, any client authenticated to the OpenAI-compatible endpoint, if that endpoint accepts chat_template or chat_template_kwargs from untrusted callers. A single ordinary-looking chat completion request is sufficient.","remediation":"Upgrade vLLM to 0.11.0 or later and restart the servers. If you front vLLM with a gateway, strip both chat_template and chat_template_kwargs from inbound requests — stripping only chat_template does not close it. Set per-pod CPU and memory limits with a restart policy so a wedged replica is recycled rather than dragging the node, and rate-limit per tenant.","references":["https://github.com/vllm-project/vllm/security/advisories/GHSA-6fvq-23cw-5628"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:L/I:L/A:N","cwe":["CWE-22"],"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2026-042-rclone-serve-s3","cve":null,"aliases":["GHSA-8v25-v8p6-qf7v"],"title":"rclone (serve s3): Path traversal in rclone's S3 gateway lets a caller read and overwrite files above the served root.","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"rclone (serve s3)","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Path traversal in rclone's S3 gateway lets a caller read and overwrite files above the served root. Sites use rclone serve s3 to put an S3 front end on a scratch or dataset directory for training jobs; this turns that gateway into arbitrary read/write on the host filesystem outside the intended prefix.","attack_vector":"Any client that can reach the rclone serve s3 endpoint, with no authentication required in the advisory's rating.","remediation":"Upgrade rclone to the fixed release and restart the serve process. Until then, run rclone serve s3 as an unprivileged user in a container or with a bind-mounted root so traversal cannot escape into anything that matters. No CVE ID has been assigned; track it by the GHSA.","references":["https://github.com/rclone/rclone/security/advisories/GHSA-8v25-v8p6-qf7v"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-789"],"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2026-053-kubeedge-cloudhub-viaduct-packer","cve":null,"aliases":["GHSA-gfw4-49f9-cp25","CVE-2026-62370 (reserved)"],"title":"KubeEdge CloudHub (viaduct packer, pkg/viaduct/pkg/packer): ONE COMPROMISED EDGE NODE TAKES DOWN CLOUD-EDGE","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"KubeEdge CloudHub (viaduct packer, pkg/viaduct/pkg/packer)","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"ONE COMPROMISED EDGE NODE TAKES DOWN CLOUD-EDGE COMMUNICATION FOR THE FLEET. The viaduct packer read a 32-bit payload length from the message header and allocated a buffer that size with no upper bound, so an authenticated peer sends a header declaring an enormous payload and CloudHub allocates accordingly. Repeat it and CloudHub climbs into OOM termination and restart loops, cutting the control channel to every edge node, not just the attacker's. The asymmetry is what matters for an operator: a single stolen or decommissioned-but-not-revoked node credential — the credential class hardest to keep clean across a distributed estate — is enough to deny the whole cloud-edge plane.","attack_vector":"Network, authenticated: any peer that can establish a viaduct connection to CloudHub, i.e. a malicious or compromised enrolled edge node. No unauthenticated access and no code execution.","remediation":"Upgrade to KubeEdge 1.23.1, 1.22.2 or 1.21.2, which enforce a 32 MiB maximum viaduct payload in both reader and writer and reject oversized lengths before allocation. Before that: restrict the CloudHub endpoint to trusted edge networks, rotate edge-node credentials and actively revoke those of decommissioned nodes, set memory limits and a restart policy on the CloudHub workload, and alert on CloudHub memory growth and unusual connection activity.","references":["https://github.com/kubeedge/kubeedge/security/advisories/GHSA-gfw4-49f9-cp25"],"status":"curated"},{"id":"CVE-2018-6622","cve":"CVE-2018-6622","aliases":[],"title":"TPM 2.0 (S3 sleep PCR reset): Platform Configuration Registers can be reset without a full platform restart by abusing","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"TPM 2.0 (S3 sleep PCR reset)","year":"2018","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Platform Configuration Registers can be reset without a full platform restart by abusing the S3 sleep path, letting an attacker replay chosen measurements into a TPM that should only ever accumulate them. The effect is that a machine which booted a tampered image can present PCR values identical to a clean boot, defeating sealed-storage unlock policies and remote attestation. Any control you built on 'the PCRs cannot lie' stops holding.","attack_vector":"Local attacker with the ability to put the system into and out of S3 sleep - so a tenant with root on a bare-metal node, or anyone with console access.","remediation":"Platform firmware/BIOS update from the OEM (per node, reboot required) that correctly re-establishes the static root of trust across sleep. Practical compensating control on servers: disable S3 suspend entirely in BIOS, which most datacenter nodes never use anyway - a config-only change that eliminates the trigger. Do that first, then patch on the normal cycle.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-6622","https://kb.cert.org/vuls/id/922681"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2018-08-17"},{"id":"CVE-2020-14308","cve":"CVE-2020-14308","aliases":["BootHole family"],"title":"GRUB2 (grub_malloc allocator): GRUB's allocator never checks the requested size for arithmetic overflow, so a tenant","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (grub_malloc allocator)","year":"2020","cvss_score":6.4,"severity":"medium","kev":false,"impact":"GRUB's allocator never checks the requested size for arithmetic overflow, so a tenant who can influence any parsed structure gets a heap that is smaller than GRUB believes it is. The payoff is arbitrary write inside the bootloader, which means loading an unsigned kernel with Secure Boot still reporting green. On a bare-metal GPU node that is the difference between 'the next tenant gets a clean box' and 'the next tenant gets the last tenant's rootkit'.","attack_vector":"A tenant with root on a node they rented, or anyone who can write the EFI System Partition (including via BMC virtual media). Not remote on its own - it is the persistence half of a two-stage attack.","remediation":"grub2 package update plus reboot on every node. The package update alone does not close it: the old signed GRUB binary stays trusted until the UEFI revocation list (dbx) is updated, and pushing dbx before every node is on the new shim/GRUB will brick nodes at next boot. Plan it as two passes with a verification gate between them.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-14308","https://access.redhat.com/security/cve/CVE-2020-14308"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2020-07-29"},{"id":"CVE-2020-14309","cve":"CVE-2020-14309","aliases":["BootHole family"],"title":"GRUB2 (squashfs symlink parser): Integer overflow in grub_squash_read_symlink lets a crafted squashfs image drive","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (squashfs symlink parser)","year":"2020","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Integer overflow in grub_squash_read_symlink lets a crafted squashfs image drive a heap overflow inside GRUB. Attacker-chosen code runs before the kernel and before any measured-boot evidence the operator would trust, so an implant planted here is invisible to every agent running in the tenant OS.","attack_vector":"Requires control of a filesystem image GRUB will read - the boot partition on a node the attacker already had, or an image served over the provisioning path.","remediation":"grub2 package update + reboot per node. Real closure needs the dbx revocation of the old signed GRUB, which is a separate and riskier rollout. On GPU nodes the reboot means draining running training jobs, so batch it with an existing maintenance window rather than doing it alone.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-14309","https://access.redhat.com/security/cve/CVE-2020-14309"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2020-07-30"},{"id":"CVE-2020-14310","cve":"CVE-2020-14310","aliases":["BootHole family"],"title":"GRUB2 (read_section_from_string): Integer overflow while reading a section string overflows the heap and gives control","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (read_section_from_string)","year":"2020","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Integer overflow while reading a section string overflows the heap and gives control of GRUB before the kernel loads. Practical outcome is a Secure Boot bypass that survives an OS reinstall, because the compromise lives in the boot partition rather than the root filesystem.","attack_vector":"Local write access to boot-time data on the node - a previous tenant, an operator with remote-hands, or a BMC-mounted virtual disk.","remediation":"grub2 package update + reboot. Follow with a dbx update, sequenced after every node is confirmed on the fixed binary. Config-only mitigation does not exist for this class.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-14310","https://access.redhat.com/security/cve/CVE-2020-14310"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2020-07-31"},{"id":"CVE-2020-14311","cve":"CVE-2020-14311","aliases":["BootHole family"],"title":"GRUB2 (ext2/ext4 symlink reader): Integer overflow in grub_ext2_read_link on a crafted ext filesystem yields a heap","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (ext2/ext4 symlink reader)","year":"2020","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Integer overflow in grub_ext2_read_link on a crafted ext filesystem yields a heap overflow in the bootloader. Because ext4 is what almost every Linux GPU node actually boots from, this one is reachable on stock images rather than exotic filesystems.","attack_vector":"Attacker controls the boot filesystem - realistically a tenant who had root on the box, or anyone who can attach media over the BMC.","remediation":"grub2 package update + reboot per node; then the dbx revocation pass. Note that a node that PXE-boots a fresh image every provisioning cycle is not automatically safe - the vulnerable GRUB is in the image you serve, so fix the golden image too, not just running nodes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-14311","https://access.redhat.com/security/cve/CVE-2020-14311"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2020-07-31"},{"id":"CVE-2020-15706","cve":"CVE-2020-15706","aliases":["BootHole family"],"title":"GRUB2 (script function redefinition): Use-after-free when a GRUB script redefines a function while that function","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (script function redefinition)","year":"2020","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Use-after-free when a GRUB script redefines a function while that function is executing. Gives arbitrary code execution in the bootloader from nothing more than a modified grub.cfg, which on most distros is not itself signature-checked.","attack_vector":"Anyone who can write grub.cfg - local root, or a previous bare-metal tenant. This is the classic BootHole shape: config file trusted more than it deserves.","remediation":"grub2 package update + reboot per node. Also worth checking that grub.cfg is not writable from a tenant-reachable partition on your image layout.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-15706","https://ubuntu.com/security/CVE-2020-15706"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2020-07-29"},{"id":"CVE-2020-8559","cve":"CVE-2020-8559","aliases":[],"title":"Kubernetes (kube-apiserver): Unvalidated redirect on proxied upgrade requests lets a compromised node escalate to other","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-apiserver)","year":"2020","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Unvalidated redirect on proxied upgrade requests lets a compromised node escalate to other nodes","attack_vector":"An attacker who already owns one node, pivoting cluster-wide","remediation":"Rolling control-plane upgrade; no GPU drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-8559"],"status":"curated","published":"2020-07-22"},{"id":"CVE-2021-3418","cve":"CVE-2021-3418","aliases":[],"title":"GRUB2 (grub-install shim_lock regression): GRUB 2.06~rc1 reintroduced the earlier direct-boot flaw: grub-install could","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (grub-install shim_lock regression)","year":"2021","cvss_score":6.4,"severity":"medium","kev":false,"impact":"GRUB 2.06~rc1 reintroduced the earlier direct-boot flaw: grub-install could produce an installation that skips shim and therefore skips kernel signature verification. Nodes you believed you had already fixed silently regress when they are rebuilt with a newer GRUB.","attack_vector":"Not directly attacker-triggered - it is a build/provisioning regression that reopens the earlier bypass. The attacker then needs only local root.","remediation":"grub2 package update + reboot, and re-verify the boot chain on any node reimaged between the original BootHole fix and this one. Worth a fleet-wide audit script that asserts shim is in the chain, not a one-off check.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-3418","https://access.redhat.com/security/cve/CVE-2021-3418"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-03-15"},{"id":"CVE-2022-26362","cve":"CVE-2022-26362","aliases":["XSA-401"],"title":"Xen (x86 PV): Race condition in typeref acquisition - PV guest escalates to host privilege","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen (x86 PV)","year":"2022","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Race condition in typeref acquisition - PV guest escalates to host privilege","attack_vector":"Tenant VM guest (PV)","remediation":"Hypervisor patch + host reboot with guest evacuation, or use Xen livepatch if the deployment supports it. Simplest structural fix: stop offering PV guests, run PVH/HVM only","references":["https://xenbits.xen.org/xsa/advisory-401.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2022-06-09"},{"id":"CVE-2022-30773","cve":"CVE-2022-30773","aliases":["INSYDE-SA-2022042"],"title":"Insyde InsydeH2O (IhisiSmm parameter buffer, DMA TOCTOU): IHISI is Insyde's own firmware-services interface","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (IhisiSmm parameter buffer, DMA TOCTOU)","year":"2022","cvss_score":6.4,"severity":"medium","kev":false,"impact":"IHISI is Insyde's own firmware-services interface - the channel BIOS update and configuration tooling talks to. Its parameter buffer sits outside SMRAM, so a device can rewrite the parameters after SMM has validated them and before SMM uses them. What the attacker reaches through this particular driver is the firmware update path itself, which is the shortest route from a peripheral to a permanent SPI implant on a GPU node.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in kernel 5.4 / 05.44.23 and 5.5 / 05.52.23.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-30773","https://www.insyde.com/security-pledge/SA-2022042"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-11-14"},{"id":"CVE-2022-30774","cve":"CVE-2022-30774","aliases":["INSYDE-SA-2022043"],"title":"Insyde InsydeH2O (PnpSmm parameter buffer, DMA TOCTOU): The plug-and-play SMI handler's parameters can be swapped","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (PnpSmm parameter buffer, DMA TOCTOU)","year":"2022","cvss_score":6.4,"severity":"medium","kev":false,"impact":"The plug-and-play SMI handler's parameters can be swapped by DMA between the check and the use. PnpSmm is the driver that writes SMBIOS/platform-description data, so corrupting it lets an attacker both corrupt SMRAM and poison the hardware inventory the OS and your fleet-management tooling read back - a node can be made to misreport what it is.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in kernel 5.2 / 05.27.29, 5.3 / 05.36.25, 5.4 / 05.44.25, 5.5 / 05.52.25.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-30774","https://www.insyde.com/security-pledge/SA-2022043"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-11-15"},{"id":"CVE-2022-31243","cve":"CVE-2022-31243","aliases":["INSYDE-SA-2022044"],"title":"Insyde InsydeH2O (FvbServicesRuntimeDxe input buffer, DMA TOCTOU): Firmware Volume Block services are the abstraction","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (FvbServicesRuntimeDxe input buffer, DMA TOCTOU)","year":"2022","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Firmware Volume Block services are the abstraction SMM uses to read and write the SPI flash. A DMA race on this handler's input buffer means an attacker influences what gets written to the boot flash - the most direct path in this whole batch to firmware that survives OS reinstall, disk replacement and reimaging between tenants.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in kernel 5.2 / 05.27.21, 5.3 / 05.36.21, 5.4 / 05.44.21, 5.5 / 05.52.21. Verify SPI write protection (BIOS Lock Enable, protected range registers) is actually set - it blunts the primitive even before the flash lands. The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31243","https://www.insyde.com/security-pledge/SA-2022044"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-11-15"},{"id":"CVE-2022-31602","cve":"CVE-2022-31602","aliases":[],"title":"NVIDIA DGX A100 - SBIOS / SMM firmware: An out-of-bounds write in IpSecDxe, exploitable against a preconditioned heap","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX A100 - SBIOS / SMM firmware","year":"2022","cvss_score":6.4,"severity":"medium","kev":false,"impact":"An out-of-bounds write in IpSecDxe, exploitable against a preconditioned heap, reaches firmware code execution. This is firmware-level persistence: it survives OS reinstall, image re-flash and tenant handoff, and it is invisible to anything running above it. On a bare-metal GPU rental business it is the difference between wiping a node between tenants and not actually being able to.","attack_vector":"Local and already privileged - host root, or code that has reached the platform firmware/SMM path. It is not a first foothold; it is what turns a one-time root compromise into something you cannot remediate by reimaging.","remediation":"Flash the fixed SBIOS from bulletin 5367. Cost: not live-patchable. Full node drain, host power cycle, and on DGX the SBIOS ships inside a firmware bundle alongside BMC and CPLD components, so budget 30-60 minutes of node downtime plus a post-flash health check. Firmware rollback protection means you cannot cleanly revert - stage on one node before the fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31602","https://github.com/NVIDIA/product-security/tree/main/2022/5367"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"firmware-flash"},"published":"2022-07-04"},{"id":"CVE-2022-31603","cve":"CVE-2022-31603","aliases":[],"title":"NVIDIA DGX A100 - SBIOS / SMM firmware: Improper array index validation in IpSecDxe gives firmware-phase code execution","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX A100 - SBIOS / SMM firmware","year":"2022","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Improper array index validation in IpSecDxe gives firmware-phase code execution against preconditioned global data. This is firmware-level persistence: it survives OS reinstall, image re-flash and tenant handoff, and it is invisible to anything running above it. On a bare-metal GPU rental business it is the difference between wiping a node between tenants and not actually being able to.","attack_vector":"Local and already privileged - host root, or code that has reached the platform firmware/SMM path. It is not a first foothold; it is what turns a one-time root compromise into something you cannot remediate by reimaging.","remediation":"Flash the fixed SBIOS from bulletin 5367. Cost: not live-patchable. Full node drain, host power cycle, and on DGX the SBIOS ships inside a firmware bundle alongside BMC and CPLD components, so budget 30-60 minutes of node downtime plus a post-flash health check. Firmware rollback protection means you cannot cleanly revert - stage on one node before the fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31603","https://github.com/NVIDIA/product-security/tree/main/2022/5367"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-129"],"fleet":{"pain_class":"firmware-flash"},"published":"2022-07-04"},{"id":"CVE-2022-31667","cve":"CVE-2022-31667","aliases":[],"title":"Harbor: Robot accounts in other projects can be updated","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Harbor","year":"2022","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Robot accounts in other projects can be updated","attack_vector":"Authenticated registry user","remediation":"Upgrade Harbor","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31667"],"status":"curated","published":"2024-11-14"},{"id":"CVE-2022-32266","cve":"CVE-2022-32266","aliases":["INSYDE-SA-2022045"],"title":"Insyde InsydeH2O (PcdSmmDxe parameter buffer, DMA TOCTOU): A DMA race against the Platform Configuration Database SMI","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (PcdSmmDxe parameter buffer, DMA TOCTOU)","year":"2022","cvss_score":6.4,"severity":"medium","kev":false,"impact":"A DMA race against the Platform Configuration Database SMI handler corrupts ACPI fields and adjacent memory. ACPI tables are consumed by the OS after boot, so this driver is the one that reaches OS-visible platform description - an attacker can corrupt what the kernel believes about the hardware, not just SMRAM. Insyde notes exploitation needs detailed knowledge of the PCD contents on the specific platform, which raises the bar but does not close the race.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in the kernel releases named in the advisory (Insyde does not enumerate per-kernel versions for this one).  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-32266","https://www.insyde.com/security-pledge/SA-2022045"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-11-14"},{"id":"CVE-2022-32267","cve":"CVE-2022-32267","aliases":["INSYDE-SA-2022046"],"title":"Insyde InsydeH2O (SmmResourceCheckDxe input buffer, DMA TOCTOU): The sharpest irony in the batch: SmmResourceCheckDxe","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (SmmResourceCheckDxe input buffer, DMA TOCTOU)","year":"2022","cvss_score":6.4,"severity":"medium","kev":false,"impact":"The sharpest irony in the batch: SmmResourceCheckDxe is the driver whose job is to validate SMM resource access, and its own input buffer is racy. An attacker who wins this race corrupts SMRAM through the guard rather than around it, which means the check other handlers rely on can be made to approve what it should reject.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in kernel 5.2 / 05.27.23, 5.3 / 05.36.23, 5.4 / 05.44.23, 5.5 / 05.52.23.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-32267","https://www.insyde.com/security-pledge/SA-2022046"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-11-15"},{"id":"CVE-2022-33906","cve":"CVE-2022-33906","aliases":["INSYDE-SA-2022048"],"title":"Insyde InsydeH2O (FwBlockServiceSmm input buffer, DMA TOCTOU): The firmware block service is the SMM-side writer","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (FwBlockServiceSmm input buffer, DMA TOCTOU)","year":"2022","cvss_score":6.4,"severity":"medium","kev":false,"impact":"The firmware block service is the SMM-side writer for the boot flash. Racing its input buffer with DMA corrupts SMRAM and puts the attacker on the flash-write path - a persistent implant that no reimage between tenant leases will remove, sitting underneath Secure Boot and underneath whatever the node later attests.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in kernel 5.2 / 05.27.23, 5.3 / 05.36.23, 5.4 / 05.44.23, 5.5 / 05.52.23. Confirm SPI flash write protection is enforced as a stopgap. The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-33906","https://www.insyde.com/security-pledge/SA-2022048"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-11-15"},{"id":"CVE-2022-33907","cve":"CVE-2022-33907","aliases":["INSYDE-SA-2022049"],"title":"Insyde InsydeH2O (IdeBusDxe SMI input buffer, DMA TOCTOU): SMRAM corruption via a DMA race on the legacy IDE/ATA bus","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (IdeBusDxe SMI input buffer, DMA TOCTOU)","year":"2022","cvss_score":6.4,"severity":"medium","kev":false,"impact":"SMRAM corruption via a DMA race on the legacy IDE/ATA bus driver. Worth noting for fleet operators that this driver is usually only live when CSM / legacy storage compatibility is enabled - so unlike most of this batch, there is a real chance the attack surface is simply not present on a modern UEFI-only server profile.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in kernel 5.2 / 05.27.25, 5.3 / 05.36.25, 5.4 / 05.44.25. Genuine config lever on this one: disable CSM / legacy storage support in BIOS on UEFI-only nodes, which removes the driver rather than just patching it. The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-33907","https://www.insyde.com/security-pledge/SA-2022049"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-11-14"},{"id":"CVE-2022-33982","cve":"CVE-2022-33982","aliases":["INSYDE-SA-2022052"],"title":"Insyde InsydeH2O (Int15ServiceSmm parameter buffer, DMA TOCTOU): DMA race against the legacy INT15 services SMI handler","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (Int15ServiceSmm parameter buffer, DMA TOCTOU)","year":"2022","cvss_score":6.4,"severity":"medium","kev":false,"impact":"DMA race against the legacy INT15 services SMI handler corrupts SMRAM. This is legacy BIOS callback plumbing that most operators do not know is still resident on a modern server image - it is, and it is reachable, which is the general lesson of this batch: the attack surface is the union of every driver the IBV compiled in, not the subset you actually use.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in kernel 5.2 / 05.27.23, 5.3 / 05.36.23, 5.4 / 05.44.23, 5.5 / 05.52.23.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-33982","https://www.insyde.com/security-pledge/SA-2022052"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-11-14"},{"id":"CVE-2022-33986","cve":"CVE-2022-33986","aliases":["INSYDE-SA-2022056"],"title":"Insyde InsydeH2O (VariableRuntimeDxe parameter buffer, DMA TOCTOU): The UEFI variable store is where the Secure Boot","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (VariableRuntimeDxe parameter buffer, DMA TOCTOU)","year":"2022","cvss_score":6.4,"severity":"medium","kev":false,"impact":"The UEFI variable store is where the Secure Boot key databases live - PK, KEK, db and dbx. A DMA race on this handler's parameter buffer lets an attacker corrupt SMRAM through the driver that guards the keys deciding what firmware and bootloaders are allowed to run. Of everything in this batch, this is the driver whose compromise most directly invalidates the boot-integrity story you sell to tenants.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in kernel 5.4 / 05.44.23 and 5.5 / 05.52.23.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-33986","https://www.insyde.com/security-pledge/SA-2022056"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-11-15"},{"id":"CVE-2022-42283","cve":"CVE-2022-42283","aliases":[],"title":"DGX-2 BMC: Buffer overflow","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX-2 BMC","year":"2022","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Buffer overflow -> code exec on BMC","attack_vector":"Network-adjacent authenticated","remediation":"Flash DGX-2 BMC firmware","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42283","https://github.com/NVIDIA/product-security/tree/main/2024/5513"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-120"],"fleet":{"pain_class":"firmware-flash"},"published":"2023-01-13"},{"id":"CVE-2023-22745","cve":"CVE-2023-22745","aliases":[],"title":"tpm2-tss (Tss2_RC_Decode / Tss2_RC_SetHandler): An 8-bit layer number indexes an array with far fewer entries, so a TPM","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"tpm2-tss (Tss2_RC_Decode / Tss2_RC_SetHandler)","year":"2023","cvss_score":6.4,"severity":"medium","kev":false,"impact":"An 8-bit layer number indexes an array with far fewer entries, so a TPM response code the library did not expect reads or writes outside the buffer - and the path to arbitrary code execution runs through the userspace component that every attestation and key-sealing tool on the node depends on. The disclosed trigger is a man-in-the-middle on the TPM bus returning 0xFFFFFFFF, which ties this directly to the physical-interposer threat model.","attack_vector":"Local, privileged - or an attacker sitting on the TPM's LPC/SPI bus, which is a hardware-interposer attack that a colo tenant or remote-hands contractor can mount.","remediation":"Package update to tpm2-tss 4.0.1 / 3.2.2 or later and restart anything linked against it - no reboot, no firmware flash. One of the genuinely cheap fixes in this cluster, so there is no reason to carry it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-22745","https://github.com/tpm2-software/tpm2-tss/security/advisories"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2023-01-19"},{"id":"CVE-2023-5088","cve":"CVE-2023-5088","aliases":[],"title":"QEMU (IDE/ATAPI): Improper IDE controller reset lets a guest overwrite the host MBR of an attached device","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"QEMU (IDE/ATAPI)","year":"2023","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Improper IDE controller reset lets a guest overwrite the host MBR of an attached device","attack_vector":"Tenant VM guest","remediation":"QEMU update + VM restart/live-migration. Also stop exposing raw block devices as IDE to tenant VMs","references":["https://access.redhat.com/security/cve/CVE-2023-5088"],"status":"curated","published":"2023-11-03"},{"id":"CVE-2024-21823","cve":"CVE-2024-21823","aliases":[],"title":"Intel DSA/IAA (idxd): Hardware erratum: direct access to Intel DSA/IAA accelerators by an untrusted application allows","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel DSA/IAA (idxd)","year":"2024","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Hardware erratum: direct access to Intel DSA/IAA accelerators by an untrusted application allows privilege escalation","attack_vector":"Tenant process granted direct accelerator access; tenant VM guest with passthrough","remediation":"Microcode + kernel driver update + reboot. Relevant wherever DSA/IAA is exposed to tenants alongside GPUs","references":["https://access.redhat.com/security/cve/CVE-2024-21823"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"published":"2024-05-16"},{"id":"CVE-2024-22278","cve":"CVE-2024-22278","aliases":[],"title":"Harbor: Incorrect permission validation lets authenticated users modify Harbor configuration","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Harbor","year":"2024","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Incorrect permission validation lets authenticated users modify Harbor configuration","attack_vector":"Authenticated registry user","remediation":"Upgrade Harbor","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-22278"],"status":"curated","published":"2024-08-02"},{"id":"CVE-2024-25620","cve":"CVE-2024-25620","aliases":[],"title":"Helm: Relative path in a chart name writes the chart outside the intended directory","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Helm","year":"2024","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Relative path in a chart name writes the chart outside the intended directory","attack_vector":"Malicious chart","remediation":"Upgrade Helm on CI and operator machines","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-25620"],"status":"curated","published":"2024-02-15"},{"cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-35809","cve":"CVE-2024-35809","aliases":[],"title":"Linux kernel (drivers/pci): Pm_runtime_get_sync() does not wait for an already-running .runtime_idle() callback, so a","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/pci)","year":"2024","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Pm_runtime_get_sync() does not wait for an already-running .runtime_idle() callback, so a driver can be torn out from under a callback that is still executing and touching driver data. The result is a use-after-free and an unhandled page fault in the host kernel during driver unbind.","attack_vector":"The unbind path is not an exotic operation in a GPU cloud - it is the passthrough provisioning flow. Every time the control plane detaches a GPU, NIC or NVMe from its host driver to bind it to vfio-pci for a tenant, pci_device_remove() runs, and if the device's .runtime_idle() callback happens to be in flight the two race. Needs host root to initiate (the operator's own automation), and a driver that implements runtime_idle; the reported crash was rtsx_pcr, but the defect is in the generic PCI driver core so any such driver qualifies. Frequent rebind cycles - i.e. fast tenant turnover - widen the window.","remediation":"Update to a kernel carrying the fix (no fixed_in published; take the stable commits below into your branch). Interim: quiesce and idle a device before unbinding it - disable runtime PM on the device (power/control = on) prior to unbind/rebind in your passthrough provisioning scripts.","references":["https://git.kernel.org/stable/c/9a87375bb586515c0af63d5dcdcd58ec4acf20a6","https://git.kernel.org/stable/c/47d8aafcfe313511a98f165a54d0adceb34e54b1","https://nvd.nist.gov/vuln/detail/CVE-2024-35809"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-8105","cve":"CVE-2024-8105","aliases":["PKfail"],"title":"UEFI Secure Boot Platform Key: ~791 firmware releases across Acer, Dell, Fujitsu, Gigabyte, HP, Intel, Lenovo","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"UEFI Secure Boot Platform Key","year":"2024","cvss_score":6.4,"severity":"medium","kev":false,"impact":"~791 firmware releases across Acer, Dell, Fujitsu, Gigabyte, HP, Intel, Lenovo, Supermicro shipped with AMI's *test* Platform Key, whose private half is public on GitHub. Anyone can sign a bootloader that the platform trusts — complete Secure Boot bypass with a below-OS, reimage-surviving implant","attack_vector":"Local, or supply-chain","remediation":"Requires generating and enrolling a real per-vendor Platform Key, which is a BIOS-level key-enrollment operation, not a patch. Many affected models never received a fixed firmware, so for those the only remediation is hardware replacement or accepting that Secure Boot is decorative","references":["https://www.binarly.io/advisories/brly-2024-005"],"status":"curated","fleet":{"ubiquity":"very common - ~900 device models across Dell, HP, Lenovo, Gigabyte, Supermicro, Intel, Fujitsu, spanning 2012-2024","remediation_pain":"firmware-flash + key re-provisioning - each node needs a genuinely secret PK enrolled and the KEK/db chain re-signed; a leaked private key cannot be patched, only rotated","pain_class":"firmware-flash","why_fleet_wide":"The private Platform Key is public, so Secure Boot is decorative on affected nodes: anyone who can write the ESP can sign a bootkit the firmware trusts, below the OS, across the whole affected SKU population."},"published":"2024-08-26"},{"id":"CVE-2025-0622","cve":"CVE-2025-0622","aliases":["GRUB2 2025 batch"],"title":"GRUB2 (commands/gpg): Module unload leaves registered hooks behind, so GRUB later calls through freed function pointers","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (commands/gpg)","year":"2025","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Module unload leaves registered hooks behind, so GRUB later calls through freed function pointers. Notable because the affected module is the one meant to verify signatures - the bug is in the verification machinery itself.","attack_vector":"Local, via GRUB command sequences or grub.cfg.","remediation":"grub2 package update + reboot. Part of the same February 2025 distro update as the rest of the batch.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0622","https://www.openwall.com/lists/oss-security/2025/02/18/3"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-02-18"},{"id":"CVE-2025-0677","cve":"CVE-2025-0677","aliases":["GRUB2 2025 batch"],"title":"GRUB2 (UFS symlink handling): Integer overflow on symlink handling in UFS gives a heap out-of-bounds write and a path","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (UFS symlink handling)","year":"2025","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Integer overflow on symlink handling in UFS gives a heap out-of-bounds write and a path to loading unsigned code with Secure Boot on.","attack_vector":"Attacker-supplied UFS filesystem image.","remediation":"grub2 package update + reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0677","https://www.openwall.com/lists/oss-security/2025/02/18/3"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-02-19"},{"id":"CVE-2025-0684","cve":"CVE-2025-0684","aliases":["GRUB2 2025 batch"],"title":"GRUB2 (ReiserFS symlink handling): Same symlink integer-overflow pattern in the ReiserFS parser","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (ReiserFS symlink handling)","year":"2025","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Same symlink integer-overflow pattern in the ReiserFS parser - heap out-of-bounds write, pre-boot execution.","attack_vector":"Attacker-supplied ReiserFS image.","remediation":"grub2 package update + reboot; or build without the module, since nothing in a modern GPU fleet boots ReiserFS.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0684","https://www.openwall.com/lists/oss-security/2025/02/18/3"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-03-03"},{"id":"CVE-2025-0685","cve":"CVE-2025-0685","aliases":["GRUB2 2025 batch"],"title":"GRUB2 (JFS symlink handling): Symlink integer overflow in the JFS parser producing a heap out-of-bounds write","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (JFS symlink handling)","year":"2025","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Symlink integer overflow in the JFS parser producing a heap out-of-bounds write.","attack_vector":"Attacker-supplied JFS image.","remediation":"grub2 package update + reboot; or drop the unused module from the build.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0685","https://www.openwall.com/lists/oss-security/2025/02/18/3"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-03-03"},{"id":"CVE-2025-0686","cve":"CVE-2025-0686","aliases":["GRUB2 2025 batch"],"title":"GRUB2 (romfs symlink handling): Symlink integer overflow in the romfs parser producing a heap out-of-bounds write","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (romfs symlink handling)","year":"2025","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Symlink integer overflow in the romfs parser producing a heap out-of-bounds write.","attack_vector":"Attacker-supplied romfs image.","remediation":"grub2 package update + reboot; or drop the module.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0686","https://www.openwall.com/lists/oss-security/2025/02/18/3"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-03-03"},{"id":"CVE-2025-0689","cve":"CVE-2025-0689","aliases":["GRUB2 2025 batch"],"title":"GRUB2 (UDF filesystem parser): Heap buffer overflow in grub_udf_read_block","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (UDF filesystem parser)","year":"2025","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Heap buffer overflow in grub_udf_read_block. UDF is the optical/ISO filesystem, which is precisely what BMC virtual media presents - so this one is reachable by anyone with BMC credentials, not only by a local tenant.","attack_vector":"Crafted UDF/ISO image, including one mounted remotely through iDRAC/iLO/XCC virtual media.","remediation":"grub2 package update + reboot. Disable BMC virtual media where you do not need it - that closes the most convenient remote path to this and several other parser bugs at once.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0689","https://www.openwall.com/lists/oss-security/2025/02/18/3"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-03-03"},{"id":"CVE-2025-0690","cve":"CVE-2025-0690","aliases":["GRUB2 2025 batch"],"title":"GRUB2 (read command): Integer overflow in the read command's accumulator writes out of bounds","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (read command)","year":"2025","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Integer overflow in the read command's accumulator writes out of bounds. Straightforward pre-boot memory corruption from GRUB script.","attack_vector":"Local, via grub.cfg or the GRUB shell.","remediation":"grub2 package update + reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0690","https://www.openwall.com/lists/oss-security/2025/02/18/3"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-02-24"},{"id":"CVE-2025-1125","cve":"CVE-2025-1125","aliases":["GRUB2 2025 batch"],"title":"GRUB2 (HFS filesystem parser): Integer overflow computing internal buffer sizes from HFS metadata, leading to a heap","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (HFS filesystem parser)","year":"2025","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Integer overflow computing internal buffer sizes from HFS metadata, leading to a heap out-of-bounds write.","attack_vector":"Attacker-supplied HFS volume, physical or virtual media.","remediation":"grub2 package update + reboot; or build GRUB without HFS support.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-1125","https://www.openwall.com/lists/oss-security/2025/02/18/3"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-03-03"},{"cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:C/C:H/I:L/A:N","cwe":["CWE-532"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2025-1979","cve":"CVE-2025-1979","aliases":["GHSA-w4rh-fgx7-q63m"],"title":"Ray (GCS Redis credential handling / logging): When the Redis password is passed on the Ray command line it gets","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ray (GCS Redis credential handling / logging)","year":"2025","cvss_score":6.4,"severity":"medium","kev":false,"impact":"When the Redis password is passed on the Ray command line it gets written into standard logs. Anyone who can read those logs - a log aggregator, a sidecar, a co-tenant with pod-log RBAC - recovers the password to the Global Control Store, which is the cluster's single source of truth for actor placement and job state. From there they can read or tamper with every tenant's workload metadata.","attack_vector":"A local or in-cluster reader of Ray's log output who also has network reach to the Redis instance. Requires that Redis auth is enabled and the password was supplied as an argument.","remediation":"Upgrade Ray to 2.43.0 or later, then rotate the Redis password - the upgrade stops new leakage but does nothing about passwords already sitting in retained logs. Purge or re-key any log archives that captured the old value.","references":["https://github.com/advisories/GHSA-w4rh-fgx7-q63m","https://nvd.nist.gov/vuln/detail/CVE-2025-1979"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-416","CWE-911"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-37946","cve":"CVE-2025-37946","aliases":[],"title":"Linux kernel (drivers/pci/hotplug): Powering off a physical function that still has child virtual functions drops the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/pci/hotplug)","year":"2025","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Powering off a physical function that still has child virtual functions drops the pci_dev reference twice. The refcount underflows, the struct pci_dev can be freed while still in use, and the host then operates on a released device object - a use-after-free reachable through the PF/VF lifecycle, which is the same lifecycle that hands VFs to tenants.","attack_vector":"s390 only - the bug is in the s390 PCI hotplug slot driver, so it applies to Linux in an IBM Z LPAR or z/VM guest, not to x86 or ARM GPU nodes. Reached by writing to the slot power attribute in sysfs (host root) for a PF that still has VFs attached; the double put is on the path that was supposed to REFUSE that operation. Include it in your inventory only if s390 is part of your estate; on a conventional GPU fleet it is inert.","remediation":"Update to a kernel carrying the fix (no fixed_in published; stable commits below). Interim on s390: disable all VFs on a PF before powering its slot off, and keep slot power control out of any automation that runs while VFs are assigned.","references":["https://git.kernel.org/stable/c/c488f8b53e156d6dcc0514ef0afa3a33376b8f9e","https://git.kernel.org/stable/c/957529baef142d95e0d1b1bea786675bd47dbe53","https://nvd.nist.gov/vuln/detail/CVE-2025-37946"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-57852","cve":"CVE-2025-57852","aliases":[],"title":"KServe ModelMesh: Group-writable `/etc/passwd` in the container image","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"KServe ModelMesh","year":"2025","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Group-writable `/etc/passwd` in the container image → privilege escalation inside the container","attack_vector":"Tenant with code execution in a ModelMesh pod","remediation":"Rebuild the container images; no host patch. Provider owns the images if it ships a managed KServe","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-57852"],"status":"curated","published":"2025-09-30"},{"id":"CVE-2026-13318","cve":"CVE-2026-13318","aliases":[],"title":"KubeVirt: SSRF in the virt-api port-forward handler via attacker-influenced VMI status IP","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"KubeVirt","year":"2026","cvss_score":6.4,"severity":"medium","kev":false,"impact":"SSRF in the virt-api port-forward handler via attacker-influenced VMI status IP","attack_vector":"Cluster user with namespace access","remediation":"Upgrade KubeVirt","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-13318"],"status":"curated","published":"2026-06-26"},{"id":"CVE-2026-24205","cve":"CVE-2026-24205","aliases":[],"title":"NVIDIA TensorRT-LLM: Concurrent requests race inside TensorRT-LLM and reach data tampering with a changed scope. On a","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA TensorRT-LLM","year":"2026","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Concurrent requests race inside TensorRT-LLM and reach data tampering with a changed scope. On a shared LLM serving tier, a race between concurrent requests means one tenant's request can affect another's - the mechanism by which response bleed-through happens.","attack_vector":"Network, low privileges. Any client able to issue concurrent requests to the serving endpoint, which is every client.","remediation":"Upgrade TensorRT-LLM per bulletin 5805 and roll the serving deployment. Cost: rolling restart. Until patched, the compensating control is reducing concurrency or dedicating an engine per tenant, both of which cost throughput.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24205","https://github.com/NVIDIA/product-security/tree/main/2026/5805"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:C/C:L/I:L/A:N","cwe":["CWE-362"],"fleet":{"pain_class":"daemon-restart"},"tags":["tenant-isolation"]},{"id":"CVE-2026-24220","cve":"CVE-2026-24220","aliases":[],"title":"TensorRT-LLM: Code exec via insecure deserialization on model load","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT-LLM","year":"2026","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Code exec via insecure deserialization on model load","attack_vector":"Malicious model artifact","remediation":"Bump TensorRT-LLM; rebuild serving images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24220","https://github.com/NVIDIA/product-security/tree/main/2026/5840"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"],"published":"2026-07-14"},{"id":"CVE-2026-24259","cve":"CVE-2026-24259","aliases":[],"title":"TensorRT-LLM: Missing authentication in configuration processing","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT-LLM","year":"2026","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Missing authentication in configuration processing","attack_vector":"Network client","remediation":"Bump TensorRT-LLM; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24259","https://github.com/NVIDIA/product-security/tree/main/2026/5840"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-306"],"published":"2026-07-14"},{"id":"CVE-2026-44774","cve":"CVE-2026-44774","aliases":[],"title":"Traefik: A tenant with HTTPRoute creation rights exposes the REST provider handler, bypassing provider isolation","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Traefik","year":"2026","cvss_score":6.4,"severity":"medium","kev":false,"impact":"A tenant with HTTPRoute creation rights exposes the REST provider handler, bypassing provider isolation","attack_vector":"Cluster user with namespace access","remediation":"Rolling Traefik upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-44774"],"status":"curated","published":"2026-05-15"},{"id":"CVE-2019-9508","cve":"CVE-2019-9508","aliases":[],"title":"Vertiv Avocent UMG-4000 universal management gateway: An authenticated admin can plant a maliciously named file","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Vertiv Avocent UMG-4000 universal management gateway","year":"2019","cvss_score":6.3,"severity":"medium","kev":false,"impact":"An authenticated admin can plant a maliciously named file in the web application that executes JavaScript every time any user (including a higher-privileged one) browses to the page listing it — a stepping stone to hijacking another operator's session on the KVM gateway.","attack_vector":"Requires an authenticated administrator account to upload/name the malicious file; the payload then fires against any user who later views that page.","remediation":"Same fixed firmware/software build as the UMG-4000 command-injection issue (CVE-2019-9507) — apply both in the same maintenance window since they land in the same release. One flash per gateway.","references":["https://www.vertiv.com/en-us/support/software-download/it-management/avocent-universal-management-gateway-appliance--software-downloads/"],"status":"curated","published":"2020-03-30"},{"id":"CVE-2020-5969","cve":"CVE-2020-5969","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): Time-of-check to time-of-use on a shared resource between guest and host plugin. A","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2020","cvss_score":6.3,"severity":"medium","kev":false,"impact":"Time-of-check to time-of-use on a shared resource between guest and host plugin. A tenant that wins the race gets host-side information disclosure or crashes the shared GPU. vGPU 8.x before 8.4, 9.x before 9.4, 10.x before 10.3.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5969"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2020-06-30"},{"id":"CVE-2020-8554","cve":"CVE-2020-8554","aliases":[],"title":"Kubernetes (kube-apiserver): Any user who can create a Service with externalIPs (or patch LB status) intercepts cluster","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-apiserver)","year":"2020","cvss_score":6.3,"severity":"medium","kev":false,"impact":"Any user who can create a Service with externalIPs (or patch LB status) intercepts cluster traffic to that IP; MITM","attack_vector":"Cluster user with namespace access","remediation":"No upstream code fix; deploy an admission policy denying externalIPs and status.loadBalancer patches for tenants","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-8554"],"status":"curated","published":"2021-01-21"},{"id":"CVE-2020-8555","cve":"CVE-2020-8555","aliases":[],"title":"Kubernetes (kube-controller-manager): Half-blind SSRF from the controller manager into the cloud metadata service","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-controller-manager)","year":"2020","cvss_score":6.3,"severity":"medium","kev":false,"impact":"Half-blind SSRF from the controller manager into the cloud metadata service and internal network","attack_vector":"Cluster user able to create storage objects","remediation":"Rolling control-plane upgrade; block link-local metadata from control-plane nodes","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-8555"],"status":"curated","published":"2020-06-05"},{"id":"CVE-2021-1061","cve":"CVE-2021-1061","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): The plugin keeps using a resource it validated after the guest has changed it - a","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":6.3,"severity":"medium","kev":false,"impact":"The plugin keeps using a resource it validated after the guest has changed it - a double-fetch race that yields host information disclosure or a shared-GPU crash. vGPU 8.x before 8.6, 11.0 before 11.3.","attack_vector":"Any unprivileged user inside a guest VM able to race the host's validation window.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1061"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2021-01-08"},{"id":"CVE-2021-21334","cve":"CVE-2021-21334","aliases":[],"title":"containerd: Environment variables from an unrelated image leak into a container, exposing another tenant's secrets","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2021","cvss_score":6.3,"severity":"medium","kev":false,"impact":"Environment variables from an unrelated image leak into a container, exposing another tenant's secrets","attack_vector":"Any tenant workload on a shared node","remediation":"Rolling containerd upgrade with node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-21334"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2021-03-10"},{"id":"CVE-2021-41091","cve":"CVE-2021-41091","aliases":[],"title":"Docker / moby: /var/lib/docker subdirectories world-traversable","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2021","cvss_score":6.3,"severity":"medium","kev":false,"impact":"/var/lib/docker subdirectories world-traversable; unprivileged host user reaches container filesystems","attack_vector":"Any local user on the node","remediation":"Upgrade Docker Engine and fix directory modes","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-41091"],"status":"curated","published":"2021-10-04"},{"id":"CVE-2022-21820","cve":"CVE-2022-21820","aliases":[],"title":"NVIDIA DCGM - nv-hostengine: A network-reachable caller drives nv-hostengine into an unhandled error condition","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DCGM - nv-hostengine","year":"2022","cvss_score":6.3,"severity":"medium","kev":false,"impact":"A network-reachable caller drives nv-hostengine into an unhandled error condition, reaching limited code execution and privilege escalation. DCGM runs as root on every GPU node and holds the fleet's telemetry, so it is a high-value target sitting on an open port.","attack_vector":"Network, with low privileges. nv-hostengine listens on TCP 5555 by default and many operators leave it bound beyond localhost so a central collector can scrape it - that binding is the exposure.","remediation":"Update DCGM per bulletin 5328 and restart nv-hostengine. Cost: restarting the host engine briefly interrupts telemetry but does not touch running GPU jobs - no drain needed. While you are there, bind nv-hostengine to localhost and scrape via a local exporter instead of exposing 5555.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-21820","https://github.com/NVIDIA/product-security/tree/main/2022/5328"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:L/I:L/A:L","cwe":["CWE-20"],"published":"2022-03-24"},{"id":"CVE-2022-23823","cve":"CVE-2022-23823","aliases":["Hertzbleed (AMD)"],"title":"AMD processors - frequency scaling / power management: A remote or local attacker times operations and infers secret","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD processors - frequency scaling / power management","year":"2022","cvss_score":6.3,"severity":"medium","kev":false,"impact":"A remote or local attacker times operations and infers secret data from how DVFS frequency scaling responds to the data being processed - turning a power side channel into a timing side channel that works over the network. On a GPU host node the exposure is the CPU-side crypto: TLS termination for your API, key material in the control plane, tenant secrets handled by the host. It does not read GPU memory.","attack_vector":"An authenticated attacker able to time operations on the target, including remotely for network-facing crypto. Co-tenancy is not required, which is what made Hertzbleed notable.","remediation":"AMD's guidance is not a microcode patch: the fix is constant-time or blinded implementations in the affected cryptographic software, and optionally disabling frequency boost - which costs you real performance on every workload on the node. Practically: update OpenSSL/libcrypto and any SIKE-like primitives, and treat disabling boost as a last resort. Effectively UNPATCHABLE at the silicon level.","references":["https://www.amd.com/en/resources/product-security/bulletin/amd-sb-1038","https://www.hertzbleed.com/","https://nvd.nist.gov/vuln/detail/CVE-2022-23823"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"published":"2022-06-15"},{"id":"CVE-2022-24436","cve":"CVE-2022-24436","aliases":["Hertzbleed (Intel)","INTEL-SA-00698"],"title":"Intel processors - power management throttling: The Intel half of Hertzbleed: observable behaviour in power-management","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors - power management throttling","year":"2022","cvss_score":6.3,"severity":"medium","kev":false,"impact":"The Intel half of Hertzbleed: observable behaviour in power-management throttling lets an authenticated user infer processed data over a network path. Same operator exposure as the AMD variant - host-side cryptography on your GPU nodes and control plane, not GPU memory.","attack_vector":"Authenticated user, remotely exploitable via timing of network-facing crypto operations. No co-tenancy required.","remediation":"Intel likewise did not ship a microcode fix. Mitigation is constant-time cryptographic software, or disabling Turbo Boost / SpeedStep which costs substantial throughput on a GPU host that is already CPU-bound in the data loader. Treat as UNPATCHABLE in hardware; patch the crypto libraries instead.","references":["https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00698.html","https://www.hertzbleed.com/","https://nvd.nist.gov/vuln/detail/CVE-2022-24436"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"published":"2022-06-15"},{"cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:H/I:N/A:H","cwe":["CWE-276"],"fleet":{"pain_class":"hot-patch"},"id":"CVE-2022-31251","cve":"CVE-2022-31251","aliases":[],"title":"Slurm (openSUSE slurm-testsuite packaging): The openSUSE slurm testsuite package ships files with permissive default","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Slurm (openSUSE slurm-testsuite packaging)","year":"2022","cvss_score":6.3,"severity":"medium","kev":false,"impact":"The openSUSE slurm testsuite package ships files with permissive default ownership, so an attacker who already controls the slurm service account can pivot to root on that host. Since slurmctld runs as SlurmUser, a compromise of the controller daemon becomes a compromise of the controller machine.","attack_vector":"Local, requires already having control of the slurm user - so it is a privilege-escalation chain step after a slurmctld or slurmd compromise, not an entry point.","remediation":"Only affects hosts where the openSUSE or SLES slurm-testsuite package is installed. Uninstall it from production controllers - a test suite has no business on a scheduler that runs tenant workloads - or update to the fixed package.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31251","https://bugzilla.suse.com/show_bug.cgi?id=1201674"],"status":"curated"},{"id":"CVE-2022-35888","cve":"CVE-2022-35888","aliases":["Hertzbleed (Ampere Altra)"],"title":"Ampere Altra / Altra Max processors: The Arm-server variant of Hertzbleed","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Ampere Altra / Altra Max processors","year":"2022","cvss_score":6.3,"severity":"medium","kev":false,"impact":"The Arm-server variant of Hertzbleed. Relevant because Ampere Altra is a common host CPU under GPU nodes in Arm-based AI racks and in several neocloud fleets - operators who assumed the Hertzbleed story was x86-only still have it.","attack_vector":"Authenticated user able to time operations, including over the network against host crypto.","remediation":"Ampere published a security bulletin rather than a firmware fix. Mitigation is constant-time crypto and, at high cost, disabling frequency scaling. Treat as UNPATCHABLE in hardware.","references":["https://amperecomputing.com/products/security-bulletins/hertzbleed.html","https://nvd.nist.gov/vuln/detail/CVE-2022-35888"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"published":"2022-09-29"},{"id":"CVE-2023-25163","cve":"CVE-2023-25163","aliases":[],"title":"Argo CD: Repository access credentials leaked in error messages surfaced in the UI and logs","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2023","cvss_score":6.3,"severity":"medium","kev":false,"impact":"Repository access credentials leaked in error messages surfaced in the UI and logs","attack_vector":"Any Argo CD user who can trigger a sync error","remediation":"Rolling Argo CD upgrade; rotate repo credentials","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25163"],"status":"curated","published":"2023-02-08"},{"id":"CVE-2023-26604","cve":"CVE-2023-26604","aliases":[],"title":"systemd: Privilege escalation via the systemctl `less` pager when sudo-granted systemctl is available","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"systemd","year":"2023","cvss_score":6.3,"severity":"medium","kev":false,"impact":"Privilege escalation via the systemctl `less` pager when sudo-granted systemctl is available","attack_vector":"Local user with any sudo systemctl grant","remediation":"Package update; audit sudoers for systemctl grants. No reboot","references":["https://access.redhat.com/security/cve/CVE-2023-26604"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2023-03-03"},{"id":"CVE-2023-34471","cve":"CVE-2023-34471","aliases":["AMI-SA-2023006","Nozomi Labs BMC audit"],"title":"AMI MegaRAC SPx (BMC cryptography / HMAC): A step is missing when the BMC generates its HMAC, so the authentication tag","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx (BMC cryptography / HMAC)","year":"2023","cvss_score":6.3,"severity":"medium","kev":false,"impact":"A step is missing when the BMC generates its HMAC, so the authentication tag it produces is weaker than intended and can be forged. The result is that an attacker can present traffic the BMC accepts as authentic - loss of authentication as well as confidentiality and integrity. Practically it degrades whatever assurance you thought you had that a management command came from your orchestration system rather than from something else on the wire.","attack_vector":"Adjacent network, requires an existing high-privilege position and user interaction, at high attack complexity. This is a chaining bug rather than a standalone break-in - it matters mostly as the thing that lets an attacker who already has partial management-plane access forge their way further.","remediation":"Firmware flash to SPx_12.2 / SPx_13.0 or later; fixed in early SPx branches, so the real work is verifying that the ODM build actually running on each node is from a fixed branch rather than trusting AMI's fix version. No config-only remediation. Compensate by treating the management network as untrusted transit: bastion-only access, mutual TLS with your own CA, and alerting on BMC sessions that do not originate from your management hosts.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2023006.pdf","https://nvd.nist.gov/vuln/detail/CVE-2023-34471"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-07-05"},{"id":"CVE-2024-0085","cve":"CVE-2024-0085","aliases":[],"title":"vGPU Manager: Improper permission management","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU Manager","year":"2024","cvss_score":6.3,"severity":"medium","kev":false,"impact":"Improper permission management","attack_vector":"Tenant VM guest","remediation":"vGPU Manager upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0085","https://github.com/NVIDIA/product-security/tree/main/2024/5551"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:N/I:H/A:H","cwe":["CWE-266"],"published":"2024-06-13"},{"id":"CVE-2024-0129","cve":"CVE-2024-0129","aliases":[],"title":"NVIDIA NeMo: SaveRestoreConnector extracts .tar archives unsafely, so a crafted archive writes files outside","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NeMo","year":"2024","cvss_score":6.3,"severity":"medium","kev":false,"impact":"SaveRestoreConnector extracts .tar archives unsafely, so a crafted archive writes files outside the extraction directory and reaches code execution. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5580 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0129","https://github.com/NVIDIA/product-security/tree/main/2024/5580"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:L/I:L/A:L","cwe":["CWE-22"],"published":"2024-10-15"},{"id":"CVE-2024-10026","cve":"CVE-2024-10026","aliases":[],"title":"gVisor: Weak hashing and small seeds let a remote attacker derive a local IP and per-boot identifier","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"gVisor","year":"2024","cvss_score":6.3,"severity":"medium","kev":false,"impact":"Weak hashing and small seeds let a remote attacker derive a local IP and per-boot identifier; sandbox fingerprinting","attack_vector":"Unauthenticated network","remediation":"Upgrade runsc","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-10026"],"status":"curated","published":"2025-01-30"},{"id":"CVE-2024-10603","cve":"CVE-2024-10603","aliases":[],"title":"gVisor: Predictable TCP/UDP source ports and header values enable off-path attacks","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"gVisor","year":"2024","cvss_score":6.3,"severity":"medium","kev":false,"impact":"Predictable TCP/UDP source ports and header values enable off-path attacks","attack_vector":"Unauthenticated network","remediation":"Upgrade runsc","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-10603"],"status":"curated","published":"2025-01-30"},{"id":"CVE-2024-20280","cve":"CVE-2024-20280","aliases":[],"title":"Cisco UCS Central Software (weak backup encryption): Weak encryption on full-state and configuration backups means","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Cisco UCS Central Software (weak backup encryption)","year":"2024","cvss_score":6.3,"severity":"medium","kev":false,"impact":"Weak encryption on full-state and configuration backups means anyone who obtains a backup file recovers the sensitive information inside it - including the credentials UCS Central uses across the estate.","attack_vector":"Access to a UCS Central backup file. Backups routinely sit on shared file servers and in ticket attachments, which is what makes this practical.","remediation":"Upgrade UCS Central per cisco-sa-ucsc-bkpsky-TgJ5f73J, then re-take backups and securely destroy old ones. Rotate credentials contained in previously-generated backups - the patch does not protect files already created.","references":["https://sec.cloudapps.cisco.com/security/center/content/CiscoSecurityAdvisory/cisco-sa-ucsc-bkpsky-TgJ5f73J"],"status":"curated"},{"id":"CVE-2024-28961","cve":"CVE-2024-28961","aliases":[],"title":"Dell OpenManage Enterprise (credential disclosure): A low-privileged local user obtains stored credentials from OME","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Dell OpenManage Enterprise (credential disclosure)","year":"2024","cvss_score":6.3,"severity":"medium","kev":false,"impact":"A low-privileged local user obtains stored credentials from OME, leading to elevated unauthorized access - in practice, the BMC credentials OME uses to manage the fleet.","attack_vector":"Local low-privilege access to the OME appliance (versions 4.0.0/4.0.1).","remediation":"Apply the DSA-2024-184 update, then rotate the discovery/management credentials OME holds. The rotation matters more than the patch.","references":["https://www.dell.com/support/kbdoc/en-us/000224251/dsa-2024-184-security-update-for-dell-openmanage-enterprise-vulnerability"],"status":"curated"},{"id":"CVE-2024-36319","cve":"CVE-2024-36319","aliases":[],"title":"AMD Video Decoder Engine Firmware (VCN FW) - debug code left active: Debug code was shipped active in AMD's Video Core","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Video Decoder Engine Firmware (VCN FW) - debug code left active","year":"2024","cvss_score":6.3,"severity":"medium","kev":false,"impact":"Debug code was shipped active in AMD's Video Core Next firmware, so a maliciously crafted command makes the VCN firmware read and write hardware registers. This is the shipped-debug-hooks failure applied to a GPU IP block: an attacker who can submit VCN commands - which any workload with the GPU device node can - gets arbitrary hardware register access, and hardware registers are how you reconfigure memory apertures, power state and access control on the device. On MI-series parts the VCN block is present whether or not your AI workload uses video decode, so 'we do not do video' is not a mitigation.","attack_vector":"Local, by submitting a crafted command to the video decode engine - reachable from any process holding /dev/dri/renderD*, i.e. an unprivileged tenant container.","remediation":"Fixed in AMD GPU firmware, which on Instinct parts is delivered as a firmware bundle through the ROCm/amdgpu driver package (the PSP loads the signed blobs at driver init) rather than through the server BIOS. Practically: update the AMD GPU driver/firmware package, then **drain the node and reboot** - the firmware is loaded once at driver init, so a reload of the module with no process holding /dev/kfd is the minimum, and a reboot is what you will actually schedule. Some fixes at this layer also require a **GPU VBIOS flash** via AMD's amdvbflash/amdfwtool, which is an offline, per-card operation with real bricking risk - check the AMD bulletin for whether a VBIOS update is called out before assuming a driver package covers it. If your workloads genuinely never touch video decode, consider whether the VCN block can be gated off in your deployment as a stopgap - but verify rather than assume, since the ROCm stack initialises IP blocks it does not use.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36319","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2026-02-12"},{"id":"CVE-2024-38805","cve":"CVE-2024-38805","aliases":["GHSA-p7wp-52j7-6r5x"],"title":"EDK II NetworkPkg (IScsiDxe, iSCSI login response processing): A hostile iSCSI target answers the firmware initiator","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"EDK II NetworkPkg (IScsiDxe, iSCSI login response processing)","year":"2024","cvss_score":6.3,"severity":"medium","kev":false,"impact":"A hostile iSCSI target answers the firmware initiator with a malformed login response and gets out-of-bounds reads and writes in the pre-OS network stack. Realistic outcome per the upstream advisory is a hung or crashed boot rather than clean code execution, but for a fleet that boots from SAN this is an attacker holding nodes down from the storage side, and the write primitive is a corruption bug that has not been proven unexploitable so much as judged unlikely.","attack_vector":"Whoever controls or can impersonate the iSCSI target the node boots from - a compromised storage appliance, an attacker on the storage VLAN, or a rogue target answering discovery. Unauthenticated from the firmware's point of view, pre-OS.","remediation":"OEM BIOS update; flash + reboot per node. Effective config workaround exists and is cheap: disable the UEFI iSCSI initiator on nodes that do not boot from SAN (a BIOS setting, no flash), and where you do boot from iSCSI, enable mutual CHAP so a rogue target cannot complete the login, and keep the storage network on its own VLAN.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-38805","https://github.com/tianocore/edk2/security/advisories/GHSA-p7wp-52j7-6r5x"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2025-08-12"},{"id":"CVE-2024-45104","cve":"CVE-2024-45104","aliases":[],"title":"Lenovo XClarity Administrator (LXCA, insufficient authorization): An authenticated LXCA user without sufficient","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lenovo XClarity Administrator (LXCA, insufficient authorization)","year":"2024","cvss_score":6.3,"severity":"medium","kev":false,"impact":"An authenticated LXCA user without sufficient privileges modifies a managed device through a crafted web API call using the device identifier - a horizontal privilege bypass over the fleet management API.","attack_vector":"Authenticated low-privilege LXCA user.","remediation":"Apply the LXCA update per LEN-154748. Appliance upgrade; review LXCA role assignments while you are in there.","references":["https://support.lenovo.com/us/en/product_security/LEN-154748"],"status":"curated"},{"cvss_vector":"CVSS:4.0/AV:N/AC:L/AT:P/PR:N/UI:N/VC:L/VI:N/VA:N/SC:L/SI:N/SA:N","cwe":["CWE-200"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2024-53862","cve":"CVE-2024-53862","aliases":["GHSA-h36c-m3rf-34h9"],"title":"Argo Workflows (Argo Server, archived workflow retrieval under client/sso auth mode): With --auth-mode=client the","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo Workflows (Argo Server, archived workflow retrieval under client/sso auth mode)","year":"2024","cvss_score":6.3,"severity":"medium","kev":false,"impact":"With --auth-mode=client the archived-workflow GET endpoint accepts a forged token, and with --auth-mode=sso any valid token reads any archived workflow regardless of namespace. An attacker or a low-privilege tenant reads other tenants' completed run history - parameters, artifact locations, and whatever those specs embed.","attack_vector":"Anyone who can reach the Argo Server API. In client mode a fabricated token is enough; in SSO mode any legitimate account, including the lowest-privilege one, reads everything.","remediation":"Upgrade Argo Server to 3.5.13 or 3.6.2 and restart. Prefer --auth-mode=server or sso with the fix in place over client mode, and treat archived workflow contents as having been readable by every authenticated user before the upgrade.","references":["https://github.com/argoproj/argo-workflows/security/advisories/GHSA-h36c-m3rf-34h9","https://nvd.nist.gov/vuln/detail/CVE-2024-53862"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-9594","cve":"CVE-2024-9594","aliases":[],"title":"Kubernetes Image Builder: Default credentials present during the build window for several providers","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes Image Builder","year":"2024","cvss_score":6.3,"severity":"medium","kev":false,"impact":"Default credentials present during the build window for several providers","attack_vector":"Attacker present during the image build","remediation":"Rebuild images; isolate the build network","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated","published":"2024-10-15"},{"id":"CVE-2025-23262","cve":"CVE-2025-23262","aliases":[],"title":"ConnectX: Access-control flaw in NIC firmware","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"ConnectX","year":"2025","cvss_score":6.3,"severity":"medium","kev":false,"impact":"Access-control flaw in NIC firmware","attack_vector":"Tenant with a VF","remediation":"Flash NIC firmware; node reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23262","https://github.com/NVIDIA/product-security/tree/main/2025/5655"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:L/I:H/A:H","cwe":["CWE-863"],"fleet":{"pain_class":"node-reboot"},"published":"2025-09-04"},{"id":"CVE-2025-37752","cve":"CVE-2025-37752","aliases":[],"title":"Linux kernel (net/sched SFQ): Missing limit validation in sch_sfq - out-of-bounds write","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/sched SFQ)","year":"2025","cvss_score":6.3,"severity":"medium","kev":false,"impact":"Missing limit validation in sch_sfq - out-of-bounds write","attack_vector":"Any tenant process in a container with CAP_NET_ADMIN in a userns","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2025-37752"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-05-01"},{"id":"CVE-2026-15044","cve":"CVE-2026-15044","aliases":[],"title":"TrustyAI Service Operator: unauthenticated access to AI guardrail and orchestrator APIs","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"TrustyAI Service Operator (Red Hat OpenShift AI)","year":"2026","cvss_score":6.3,"severity":"medium","kev":false,"impact":"Guardrail and orchestrator APIs are reachable without authentication from anywhere on the cluster network. An attacker who can reach them reads or reconfigures the guardrails meant to constrain model behaviour, so the safety layer a deployment depends on can be turned off without touching the model itself.","attack_vector":"Any workload with cluster network access to the operator's services.","remediation":"Upgrade the TrustyAI Service Operator to the fixed release from Red Hat, then put a NetworkPolicy in front of the guardrail and orchestrator services and require mTLS between them. The operator restart is a rolling one; no node work.","references":["https://access.redhat.com/security/cve/CVE-2026-15044","https://bugzilla.redhat.com/show_bug.cgi?id=2498039","https://nvd.nist.gov/vuln/detail/CVE-2026-15044"],"status":"curated","fleet":{"pain_class":"daemon-restart"},"published":"2026-07-08"},{"id":"CVE-2026-24142","cve":"CVE-2026-24142","aliases":[],"title":"TensorRT-LLM: Code exec via insecure deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT-LLM","year":"2026","cvss_score":6.3,"severity":"medium","kev":false,"impact":"Code exec via insecure deserialization","attack_vector":"Malicious model artifact","remediation":"Bump TensorRT-LLM; rebuild serving images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24142","https://github.com/NVIDIA/product-security/tree/main/2026/5805"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:L/I:L/A:L","cwe":["CWE-502"],"published":"2026-05-20"},{"id":"CVE-2026-24226","cve":"CVE-2026-24226","aliases":[],"title":"TensorRT-LLM: Unsafe external file loading","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT-LLM","year":"2026","cvss_score":6.3,"severity":"medium","kev":false,"impact":"Unsafe external file loading -> code exec","attack_vector":"Malicious model reference","remediation":"Bump TensorRT-LLM; rebuild serving images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24226","https://github.com/NVIDIA/product-security/tree/main/2026/5840"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-829"],"published":"2026-07-14"},{"id":"CVE-2026-49943","cve":"CVE-2026-49943","aliases":[],"title":"CZ.NIC BIRD Internet Routing Daemon (BGP AS_PATH mask matching): Stack-based buffer overflow in BIRD's AS_PATH mask","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"CZ.NIC BIRD Internet Routing Daemon (BGP AS_PATH mask matching)","year":"2026","cvss_score":6.3,"severity":"medium","kev":false,"impact":"Stack-based buffer overflow in BIRD's AS_PATH mask matching: `as_path_match()` uses a fixed 2049-entry stack array while the parsed path can exceed it. BIRD is a common choice for route servers, for BGP-to-the-host designs, and inside open networking stacks — including some SONiC and Linux-router-based cluster underlays. A stack overflow in the BGP path-attribute parser is reachable from any peer, and in a route-reflector topology from beyond the direct peer.","attack_vector":"A BGP peer, or anything upstream of one whose AS_PATH propagates, sending a long AS_PATH that hits the mask-matching path. Requires a BGP filter using AS path masks.","remediation":"Upgrade BIRD past 2.19.0 and restart the daemon — package upgrade plus service restart, which briefly drops BGP sessions and reconverges. Interim: apply an inbound AS_PATH length limit on every eBGP session, a live config change and sound policy regardless.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-49943"],"status":"curated","fleet":{"pain_class":"daemon-restart"},"published":"2026-06-02"},{"id":"CVE-2019-3016","cve":"CVE-2019-3016","aliases":[],"title":"Linux KVM - PV TLB shootdown leaks memory between guest processes: In a KVM guest with paravirtualised TLB enabled, one","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux KVM - PV TLB shootdown leaks memory between guest processes","year":"2019","cvss_score":6.2,"severity":"medium","kev":false,"impact":"In a KVM guest with paravirtualised TLB enabled, one process in the guest can read memory belonging to another process in the same guest. The isolation that breaks is inside the VM rather than between VMs - which matters for any operator whose customers run multi-user workloads inside a single VM, and for confidential guests where the tenant assumed process separation held.","attack_vector":"Local, from one process to another inside a KVM guest with PV TLB enabled. The host must be running Linux with KVM.","remediation":"Fixed in the Linux kernel. The fix belongs in the **guest** kernel, so update confidential and tenant VM images, not just hosts. Interim mitigation: disable PV TLB flush in the guest (the kvm.pv_tlb boot option / KVM_FEATURE_PV_TLB_FLUSH), which costs some scheduling efficiency and needs a guest reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-3016"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2020-01-31"},{"id":"CVE-2021-1093","cve":"CVE-2021-1093","aliases":[],"title":"NVIDIA GPU Display Driver, GPU firmware: An attacker-triggerable assert in GPU firmware aborts harder than it needs","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver, GPU firmware","year":"2021","cvss_score":6.2,"severity":"medium","kev":false,"impact":"An attacker-triggerable assert in GPU firmware aborts harder than it needs to, crashing the system. Firmware-level asserts are worth noting because the failure is below the driver: the recovery is a node reset, not a service restart. Debian and Gentoo shipped it as a security update.","attack_vector":"Any local user or GPU container able to drive the firmware into the asserting path.","remediation":"Install the fixed GPU Display Driver branch on both Windows and Linux nodes. The kernel component (nvlddmkm.sys / nvidia.ko) cannot be hot-swapped under load, so this is a node drain and reboot per host; restart the container runtime afterwards so mounted driver libraries match the kernel module. No VBIOS or BMC flash.","references":["https://lists.debian.org/debian-lts-announce/2022/01/msg00013.html","https://security.gentoo.org/glsa/202310-02","https://nvd.nist.gov/vuln/detail/CVE-2021-1093"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-07-22"},{"id":"CVE-2021-1100","cve":"CVE-2021-1100","aliases":[],"title":"NVIDIA vGPU Manager kernel module (nvidia.ko, host): The host vGPU kernel module dereferences an unvalidated user-space","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager kernel module (nvidia.ko, host)","year":"2021","cvss_score":6.2,"severity":"medium","kev":false,"impact":"The host vGPU kernel module dereferences an unvalidated user-space pointer. Guest-reachable crash of the hypervisor's GPU kernel module, taking down every tenant on the card. vGPU 12.x before 12.3, 11.x before 11.5, 8.x before 8.8.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1100"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2021-07-21"},{"cwe":["CWE-835"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2021-47617","cve":"CVE-2021-47617","aliases":[],"title":"Linux kernel (drivers/pci/hotplug): A power fault on a PCIe hotplug slot latches a sticky status bit that the hardirq","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/pci/hotplug)","year":"2021","cvss_score":6.2,"severity":"medium","kev":false,"impact":"A power fault on a PCIe hotplug slot latches a sticky status bit that the hardirq handler cannot clear, so the interrupt handler spins forever inside hardirq context. The CPU never returns to the scheduler - the node hard-hangs and every tenant on it goes down at once, with no clean drain and no way in to diagnose it.","attack_vector":"Not tenant-software reachable. It needs a device or sled in a PCIe hotplug slot to assert a main power fault - a failing or over-current NVMe/GPU carrier, a bad riser, or an add-in device someone with physical or smart-hands access installs. Relevant to GPU clusters because hot-pluggable NVMe and GPU sleds are exactly the population that generates power faults, and one faulty card silently converts into a whole-node outage. Requires pciehp driving the slot.","remediation":"Update to 4.19.233 / 5.4.177 or later (the fix sets the power_fault_detected flag in the hardirq handler so the loop terminates). Interim: there is no software mitigation once the fault latches - track slot power-fault events in your health telemetry and drain nodes reporting them, and restrict physical/smart-hands installation of unvetted cards into hotplug slots.","references":["https://git.kernel.org/stable/c/ff27f7d0333cff89ec85c419f431aca1b38fb16a","https://git.kernel.org/stable/c/464da38ba827f670deac6500a1de9a4f0f44c41d","https://nvd.nist.gov/vuln/detail/CVE-2021-47617"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2022-42284","cve":"CVE-2022-42284","aliases":[],"title":"DGX servers BMC: Sensitive data exposure from BMC","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX servers BMC","year":"2022","cvss_score":6.2,"severity":"medium","kev":false,"impact":"Sensitive data exposure from BMC","attack_vector":"Network-adjacent","remediation":"Flash BMC 2.09.00+; rotate credentials","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42284","https://github.com/NVIDIA/product-security/tree/main/2024/5513"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-312"],"published":"2023-01-13"},{"id":"CVE-2022-46146","cve":"CVE-2022-46146","aliases":[],"title":"Prometheus (exporter-toolkit): Poisoning the built-in auth cache bypasses basic-auth on exporters","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Prometheus (exporter-toolkit)","year":"2022","cvss_score":6.2,"severity":"medium","kev":false,"impact":"Poisoning the built-in auth cache bypasses basic-auth on exporters","attack_vector":"Local","remediation":"Data-plane: rebuild and roll every exporter - node_exporter/DCGM run on GPU nodes","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-46146"],"status":"curated","published":"2022-11-29"},{"id":"CVE-2023-25153","cve":"CVE-2023-25153","aliases":[],"title":"containerd: Unbounded read on OCI image import causes containerd OOM","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2023","cvss_score":6.2,"severity":"medium","kev":false,"impact":"Unbounded read on OCI image import causes containerd OOM; node DoS","attack_vector":"Malicious image","remediation":"Rolling containerd upgrade with node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25153"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2023-02-16"},{"cwe":["CWE-400","CWE-834"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-52484","cve":"CVE-2023-52484","aliases":[],"title":"Linux kernel (drivers/iommu/arm/arm-smmu-v3): A process using SVA that unmaps memory drives a flood of SMMU range","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/arm/arm-smmu-v3)","year":"2023","cvss_score":6.2,"severity":"medium","kev":false,"impact":"A process using SVA that unmaps memory drives a flood of SMMU range invalidations, spinning a CPU in the command queue long enough to trip the soft-lockup watchdog - 26 seconds in the upstream report on a 244-CPU box. One tenant's munmap becomes a multi-second stall of the shared SMMU command queue, which stalls invalidation and DMA setup for every other device and tenant behind that SMMU.","attack_vector":"An unprivileged process using SVA/PASID on an arm64 host with SMMUv3 - the configuration on Grace-based GPU nodes. The trigger is an ordinary large munmap in a process that holds an SVA binding: no host root, no crafted request, no race to win. x86 hosts are unaffected.","remediation":"The record lists no fixed release; boot a kernel carrying the stable fix commits below. Interim control: do not enable SVA/PASID for tenant workloads on unpatched arm64 SMMUv3 nodes.","references":["https://git.kernel.org/stable/c/f5a604757aa8e37ea9c7011dc9da54fa1b30f29b","https://git.kernel.org/stable/c/f90f4c562003ac3d3b135c5a40a5383313f27264","https://nvd.nist.gov/vuln/detail/CVE-2023-52484"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-8982","cve":"CVE-2024-8982","aliases":[],"title":"OpenLLM: Local file inclusion via the web application","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"OpenLLM","year":"2024","cvss_score":6.2,"severity":"medium","kev":false,"impact":"Local file inclusion via the web application","attack_vector":"Network user of the OpenLLM UI","remediation":"Upgrade past 0.6.10","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-8982"],"status":"curated","published":"2025-03-20"},{"id":"CVE-2025-0426","cve":"CVE-2025-0426","aliases":[],"title":"Kubernetes (kubelet): Unauthenticated node DoS through the kubelet checkpoint API filling node disk","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kubelet)","year":"2025","cvss_score":6.2,"severity":"medium","kev":false,"impact":"Unauthenticated node DoS through the kubelet checkpoint API filling node disk","attack_vector":"Any pod on the cluster network reaching the kubelet port","remediation":"Rolling kubelet upgrade with node drain; disable the ContainerCheckpoint feature gate","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2025-02-13"},{"id":"CVE-2025-33176","cve":"CVE-2025-33176","aliases":[],"title":"NVIDIA Run:ai: Improper restriction of communication channels lets an attacker on an adjacent network reach","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Run:ai","year":"2025","cvss_score":6.2,"severity":"medium","kev":false,"impact":"Improper restriction of communication channels lets an attacker on an adjacent network reach privilege escalation, data tampering and information disclosure in the Run:ai scheduler. Run:ai is the component deciding which tenant's job lands on which GPU, so control over it is control over the fleet's allocation and quota model.","attack_vector":"Adjacent network, low privileges, user interaction, high complexity. A tenant workload or a compromised pod inside the cluster network.","remediation":"Upgrade Run:ai per bulletin 5719. Cost: a control-plane upgrade - the scheduler restarts, queued jobs pause, running jobs keep their GPUs. Plan it in a low-submission window rather than draining nodes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33176","https://github.com/NVIDIA/product-security/tree/main/2025/5719"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:H/PR:L/UI:R/S:C/C:L/I:H/A:N","cwe":["CWE-923"],"published":"2025-11-04"},{"cwe":["CWE-681","CWE-190"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-23067","cve":"CVE-2026-23067","aliases":[],"title":"Linux kernel (drivers/iommu): The ARM long-descriptor unmap path returns a negative errno through an unsigned size_t","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu)","year":"2026","cvss_score":6.2,"severity":"medium","kev":false,"impact":"The ARM long-descriptor unmap path returns a negative errno through an unsigned size_t, so the caller is told roughly 2^64 bytes were unmapped. That value is added to the IOVA in the unmap loop, overflowing the address and hitting a BUG_ON - a kernel panic that takes every tenant on the node with it - while the unmap itself walks into addresses that were never part of the request.","attack_vector":"Reached when an unmap request covers a page-table entry that is already absent. On a passthrough node the unmap originates from a tenant's VMM through vfio or iommufd, though the upper layers normally track mapped areas well enough that hitting an absent entry means state is already inconsistent (a WARN fires first). arm64 hosts using io-pgtable-arm, i.e. SMMUv3 on Grace-class GPU nodes. x86 is unaffected.","remediation":"The record lists no fixed release; boot a kernel carrying the stable fix commits below. No practical interim control on affected arm64 nodes beyond patching.","references":["https://git.kernel.org/stable/c/41ec6988547819756fb65e94fc24f3e0dddf84ac","https://git.kernel.org/stable/c/374e7af67d9d9d6103c2cfc8eb32abfecf3a2fd8","https://nvd.nist.gov/vuln/detail/CVE-2026-23067"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-24271","cve":"CVE-2026-24271","aliases":[],"title":"TensorRT-LLM: DoS via large tensor allocation","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT-LLM","year":"2026","cvss_score":6.2,"severity":"medium","kev":false,"impact":"DoS via large tensor allocation","attack_vector":"Any inference client","remediation":"Bump TensorRT-LLM; add request limits","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24271","https://github.com/NVIDIA/product-security/tree/main/2026/5840"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-770"],"published":"2026-07-14"},{"id":"CVE-2026-41187","cve":"CVE-2026-41187","aliases":[],"title":"Calico: DeleteCollection skips AuthorizeTierOperation, so a tenant can delete tiered NetworkPolicies they","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Calico","year":"2026","cvss_score":6.2,"severity":"medium","kev":false,"impact":"DeleteCollection skips AuthorizeTierOperation, so a tenant can delete tiered NetworkPolicies they cannot delete individually","attack_vector":"Cluster user with namespace access","remediation":"Rolling Calico apiserver upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-41187"],"status":"curated","published":"2026-07-30"},{"cwe":["CWE-667","CWE-833"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-43147","cve":"CVE-2026-43147","aliases":[],"title":"Linux kernel (drivers/pci): Tearing down a PF that still has SR-IOV VFs takes pci_rescan_remove_lock recursively and","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/pci)","year":"2026","cvss_score":6.2,"severity":"medium","kev":false,"impact":"Tearing down a PF that still has SR-IOV VFs takes pci_rescan_remove_lock recursively and deadlocks. The thread wedges holding the global PCI rescan/remove lock, so every subsequent PCI enumeration, device bind, VF create and hot-remove on that node blocks behind it forever - the node stops being able to hand devices to anyone, and only a reboot clears it.","attack_vector":"Host-side, needs the ability to write PCI sysfs (root, or a privileged container with writable /sys) on a node with SR-IOV VFs created - the mlx5 trace in the report is a NIC, and the same shape applies to any PF that calls sriov_disable() from its remove path. The trigger is the ordinary fleet operation of removing a PF that still has VFs (echo 1 > .../remove after sriov_numvfs), which is exactly what VF-reclaim and node-reimage automation does between tenants, so this fires from your own control plane as readily as from an attacker.","remediation":"Update to 5.10.252 / 5.15.202 / 6.1.165 / 6.6.128 / 6.12.75 / 6.18 or later (the offending rescan-remove locking commit is reverted). Interim: always set sriov_numvfs to 0 and let VF teardown complete before removing or unbinding a PF, and keep PCI sysfs unwritable from containers.","references":["https://git.kernel.org/stable/c/f61cdd7e9b67bb8961b0a81bf294b78343e5db05","https://git.kernel.org/stable/c/0de341b2365bad430aade0853fe09c2cbe468f59","https://nvd.nist.gov/vuln/detail/CVE-2026-43147"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-47470","cve":"CVE-2026-47470","aliases":[],"title":"TensorRT-LLM: DoS / memory corruption (insufficient tensor validation)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT-LLM","year":"2026","cvss_score":6.2,"severity":"medium","kev":false,"impact":"DoS / memory corruption (insufficient tensor validation)","attack_vector":"Any inference client","remediation":"Bump TensorRT-LLM; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47470","https://github.com/NVIDIA/product-security/tree/main/2026/5840"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-20"],"published":"2026-07-14"},{"id":"CVE-2026-47475","cve":"CVE-2026-47475","aliases":[],"title":"TensorRT-LLM: DoS via assertion failure","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT-LLM","year":"2026","cvss_score":6.2,"severity":"medium","kev":false,"impact":"DoS via assertion failure","attack_vector":"Any inference client","remediation":"Bump TensorRT-LLM; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47475","https://github.com/NVIDIA/product-security/tree/main/2026/5840"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-617"],"published":"2026-07-14"},{"cwe":["CWE-835","CWE-667"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-64290","cve":"CVE-2026-64290","aliases":[],"title":"Linux kernel (drivers/iommu/iommufd): A failed copy_to_user while draining the iommufd fault queue restarts the same","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/iommufd)","year":"2026","cvss_score":6.2,"severity":"medium","kev":false,"impact":"A failed copy_to_user while draining the iommufd fault queue restarts the same failing copy forever, spinning in the kernel at 100% CPU while holding the fault mutex. The task is not killable, the core is gone until reboot, and the I/O page-fault path for that device is wedged - repeat it once per core and the node is finished.","attack_vector":"A tenant container holding /dev/iommu reads the fault fd into a deliberately bad buffer (unmapped or read-only memory). One read syscall, no host privilege, no hardware prerequisite beyond an IOPF-capable device being attached. Trivially repeatable and deterministic - this is the cheapest node-kill in the shard.","remediation":"Update to a stable kernel carrying commits a38e0714 / 5539da12. Interim: do not hand /dev/iommu to tenant containers; run passthrough through a VMM the operator controls so the fault fd is never in tenant hands.","references":["https://git.kernel.org/stable/c/a38e0714affc5c0bbb40cba5a65d6d32a5e72a71","https://git.kernel.org/stable/c/5539da127d03c1f6c2e2a49fdfbe331a0ccbdea8","https://nvd.nist.gov/vuln/detail/CVE-2026-64290"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-835"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-64474","cve":"CVE-2026-64474","aliases":[],"title":"Linux kernel (drivers/vfio): A blocked migration-state transition makes the vfio state machine spin forever while","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/vfio)","year":"2026","cvss_score":6.2,"severity":"medium","kev":false,"impact":"A blocked migration-state transition makes the vfio state machine spin forever while holding the driver's state mutex. The calling thread never returns, a CPU is pinned in kernel mode, and on a node booted with softlockup_panic the box panics outright - one ioctl from one tenant is a noisy-neighbour outage for everyone else on the node.","attack_vector":"A tenant holding /dev/vfio/* for a migration-capable device issues VFIO_DEVICE_FEATURE with MIG_DEVICE_STATE requesting the blocked STOP_COPY to PRE_COPY (or PRE_COPY_P2P) transition. Conditional on a vfio variant driver that advertises precopy - mlx5 VFs, Intel Xe and QAT VFs and similar - being bound to the tenant's device. No host privilege needed.","remediation":"Update to a stable kernel carrying commits 8e872c07 / ed7d5599. Interim: bind tenant devices to plain vfio-pci rather than a migration-capable variant driver where you do not need live migration, and do not boot tenant nodes with softlockup_panic.","references":["https://git.kernel.org/stable/c/8e872c07e40d51a66dee7b280a23a460a2e1e3fa","https://git.kernel.org/stable/c/ed7d5599e6c398da74845767cd1e6a8370a160fc","https://nvd.nist.gov/vuln/detail/CVE-2026-64474"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2020-15157","cve":"CVE-2020-15157","aliases":[],"title":"containerd: \"ContainerDrip\": registry credentials leaked to an attacker-controlled URL referenced in an image manifest","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2020","cvss_score":6.1,"severity":"medium","kev":false,"impact":"\"ContainerDrip\": registry credentials leaked to an attacker-controlled URL referenced in an image manifest","attack_vector":"Malicious image pulled from an untrusted registry","remediation":"Upgrade containerd; rotate any registry pull credentials that may have leaked","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-15157"],"status":"curated","published":"2020-10-16"},{"id":"CVE-2021-1094","cve":"CVE-2021-1094","aliases":[],"title":"NVIDIA GPU Display Driver (Windows nvlddmkm.sys + Linux nvidia.ko): Out-of-bounds array access in the escape handler","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver (Windows nvlddmkm.sys + Linux nvidia.ko)","year":"2021","cvss_score":6.1,"severity":"medium","kev":false,"impact":"Out-of-bounds array access in the escape handler on Windows and Linux, giving disclosure or a crash from unprivileged local code.","attack_vector":"Any local user or GPU container with device access.","remediation":"Install the fixed GPU Display Driver branch on both Windows and Linux nodes. The kernel component (nvlddmkm.sys / nvidia.ko) cannot be hot-swapped under load, so this is a node drain and reboot per host; restart the container runtime afterwards so mounted driver libraries match the kernel module. No VBIOS or BMC flash.","references":["https://lists.debian.org/debian-lts-announce/2022/01/msg00013.html","https://security.gentoo.org/glsa/202310-02","https://nvd.nist.gov/vuln/detail/CVE-2021-1094"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-07-22"},{"id":"CVE-2021-22810","cve":"CVE-2021-22810","aliases":["SEVD-2021-313-03"],"title":"APC Network Management Card 2 (AP9630/AP9631/AP9635) in Smart-UPS, Symmetra and Galaxy 3500: Stored/reflected","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"APC Network Management Card 2 (AP9630/AP9631/AP9635) in Smart-UPS, Symmetra and Galaxy 3500","year":"2021","cvss_score":6.1,"severity":"medium","kev":false,"impact":"Stored/reflected cross-site scripting in the NMC2 policy-file pages. On its own it is a browser bug; in context it is a route to hijack a facility engineer's authenticated session on the card that controls UPS behaviour. An attacker with an NMC session can change shutdown policies, thresholds and outlet-group behaviour - which is a path to a power event, not just a defacement.","attack_vector":"Requires tricking an already-privileged NMC user into clicking a crafted URL. Realistic in a colo where facility staff routinely click links in tickets.","remediation":"Firmware update to NMC2 AOS v6.9.6 or later (SEVD-2021-313-03 covers the whole CVE-2021-22810 through -22815 batch, so treat it as one campaign). Non-disruptive flash. Enforce that NMC admin sessions are only opened from a dedicated management workstation.","references":["https://download.schneider-electric.com/files?p_Doc_Ref=SEVD-2021-313-03"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-01-28"},{"id":"CVE-2021-28509","cve":"CVE-2021-28509","aliases":[],"title":"Arista EOS (TerminAttr / OpenConfig telemetry transport): The streaming-telemetry agent can leak MACsec keys over the","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (TerminAttr / OpenConfig telemetry transport)","year":"2021","cvss_score":6.1,"severity":"medium","kev":false,"impact":"The streaming-telemetry agent can leak MACsec keys over the telemetry transport. Whoever consumes your telemetry stream — often a monitoring platform with far weaker access control than the switches themselves — ends up holding the keys that protect inter-site and inter-pod links. From there an attacker decrypts traffic for every tenant crossing those links.","attack_vector":"An attacker with access to the telemetry stream or to the collector storing it. That is usually a much softer target than the switch.","remediation":"EOS/TerminAttr upgrade plus agent restart. Then rotate every MACsec key that could have been exposed — a fabric-wide key rotation is disruptive and is the real cost here, not the upgrade. Treat telemetry collectors as secret-bearing systems going forward.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-28509"],"status":"curated","tags":["tenant-isolation"],"published":"2022-05-26"},{"id":"CVE-2021-38961","cve":"CVE-2021-38961","aliases":["IBM X-Force 212049"],"title":"IBM OpenBMC OP910 web UI (phosphor-webui lineage): Stored/reflected script injection in the BMC web interface","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"IBM OpenBMC OP910 web UI (phosphor-webui lineage)","year":"2021","cvss_score":6.1,"severity":"medium","kev":false,"impact":"Stored/reflected script injection in the BMC web interface. The victim is your own operator: an admin opens the BMC console and the injected script runs with their authenticated session, which is a session that can power-cycle nodes, mount virtual media and push firmware. On a GPU fleet the practical scenario is an attacker who has read-only or low-privilege access to one BMC planting the payload and waiting for an administrator to visit, converting a foothold into administrative control without ever cracking a password.","attack_vector":"Requires getting attacker-controlled content into a field the BMC web UI renders, plus an administrator subsequently loading that page. Network access to the BMC web interface.","remediation":"Fixed in later OP910 firmware - per-node system firmware update, maintenance window. Cheap compensating control: do not browse BMC web UIs from the same browser profile you use for anything else, and prefer Redfish API calls over the web UI for routine operations. Fleet-scale automation against Redfish rather than humans clicking through per-node web UIs removes the victim this bug needs.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-38961","https://www.ibm.com/support/pages/node/6536720","https://exchange.xforce.ibmcloud.com/vulnerabilities/212049"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2021-12-27"},{"id":"CVE-2021-46759","cve":"CVE-2021-46759","aliases":[],"title":"AMD TEE / ASP bootloader syscall input validation: Insufficient validation of syscall inputs in the AMD trusted","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD TEE / ASP bootloader syscall input validation","year":"2021","cvss_score":6.1,"severity":"medium","kev":false,"impact":"Insufficient validation of syscall inputs in the AMD trusted execution environment lets an attacker who controls a user application running under the ASP bootloader read back ASP bootloader memory - disclosing firmware internals and, more usefully to an attacker, the layout and secrets needed to build a reliable exploit against the secure processor.","attack_vector":"Local plus physical access, and control of a Uapp running under the bootloader. High bar; realistic for an attacker with hands on the hardware (supply chain, colocation insider, returned hardware).","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. Physical-access requirement means datacenter physical controls and tamper-evident handling are a genuine compensating control here.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-46759","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-05-09"},{"id":"CVE-2022-21123","cve":"CVE-2022-21123","aliases":[],"title":"Intel CPU (MMIO Stale Data / SBDR): Incomplete cleanup of multi-core shared buffers - stale data read across domains","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel CPU (MMIO Stale Data / SBDR)","year":"2022","cvss_score":6.1,"severity":"medium","kev":false,"impact":"Incomplete cleanup of multi-core shared buffers - stale data read across domains","attack_vector":"Tenant VM guest; any tenant process in a container","remediation":"Microcode + kernel mitigation + reboot; consider disabling SMT for hard-isolation tenants","references":["https://access.redhat.com/security/cve/CVE-2022-21123"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"published":"2022-06-15"},{"id":"CVE-2022-21813","cve":"CVE-2022-21813","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): An unprivileged local user gets limited write access","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":6.1,"severity":"medium","kev":false,"impact":"An unprivileged local user gets limited write access to memory the driver treats as protected, which is enough to crash the node. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5312. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-21813","https://github.com/NVIDIA/product-security/tree/main/2022/5312"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:L/A:H","cwe":["CWE-280"],"fleet":{"pain_class":"node-drain"},"published":"2022-02-07"},{"id":"CVE-2022-21814","cve":"CVE-2022-21814","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): An unprivileged local user gets limited write access","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":6.1,"severity":"medium","kev":false,"impact":"An unprivileged local user gets limited write access to protected memory through the kernel driver package, ending in a node-wide GPU denial of service. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5312. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-21814","https://github.com/NVIDIA/product-security/tree/main/2022/5312"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:L/A:H","cwe":["CWE-280"],"fleet":{"pain_class":"node-drain"},"published":"2022-02-07"},{"id":"CVE-2022-28186","cve":"CVE-2022-28186","aliases":[],"title":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys): Improper input validation in the DxgkDdiEscape","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys)","year":"2022","cvss_score":6.1,"severity":"medium","kev":false,"impact":"Improper input validation in the DxgkDdiEscape handler lets a local user crash the node or tamper with driver state. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5353. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28186","https://github.com/NVIDIA/product-security/tree/main/2022/5353"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:L/A:H","cwe":["CWE-20"],"fleet":{"pain_class":"node-drain"},"published":"2022-05-17"},{"id":"CVE-2022-31616","cve":"CVE-2022-31616","aliases":[],"title":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys): An out-of-bounds read through DxgkDdiEscape","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys)","year":"2022","cvss_score":6.1,"severity":"medium","kev":false,"impact":"An out-of-bounds read through DxgkDdiEscape yields a crash or kernel information disclosure. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5383. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31616","https://github.com/NVIDIA/product-security/tree/main/2022/5383"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:N/A:H","cwe":["CWE-20"],"fleet":{"pain_class":"node-drain"},"published":"2022-11-19"},{"id":"CVE-2023-0186","cve":"CVE-2023-0186","aliases":[],"title":"GPU Display Driver: DoS (GPU firmware buffer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2023","cvss_score":6.1,"severity":"medium","kev":false,"impact":"DoS (GPU firmware buffer overflow)","attack_vector":"Any tenant with a container","remediation":"Driver + GPU firmware bundle upgrade; reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0186","https://github.com/NVIDIA/product-security/tree/main/2023/5452"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:L/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-reboot"},"published":"2023-04-01"},{"id":"CVE-2023-0187","cve":"CVE-2023-0187","aliases":[],"title":"GPU Display Driver: Info disclosure (OOB read)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2023","cvss_score":6.1,"severity":"medium","kev":false,"impact":"Info disclosure (OOB read)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; rolling node reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0187","https://github.com/NVIDIA/product-security/tree/main/2023/5452"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:N/A:H","cwe":["CWE-125"],"fleet":{"pain_class":"node-reboot"},"published":"2023-04-01"},{"id":"CVE-2023-0199","cve":"CVE-2023-0199","aliases":[],"title":"NVIDIA GPU Display Driver - kernel mode layer (Windows nvlddmkm.sys and Linux nvidia.ko): An out-of-bounds write","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - kernel mode layer (Windows nvlddmkm.sys and Linux nvidia.ko)","year":"2023","cvss_score":6.1,"severity":"medium","kev":false,"impact":"An out-of-bounds write in the kernel mode handler gives a local user denial of service and data tampering across the driver boundary. Both the Windows and Linux datacenter drivers are affected, so a mixed fleet needs two separate rollouts.","attack_vector":"Local and unprivileged on either OS. On Linux it is reachable from any GPU container via /dev/nvidia*; on Windows from any session holding a GPU handle.","remediation":"Upgrade both the Linux and the Windows datacenter driver branches listed in bulletin 5452. Cost: Linux needs a drain and nvidia.ko reload per node; Windows needs a reboot per node. Two change windows unless your fleet is homogeneous.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0199","https://github.com/NVIDIA/product-security/tree/main/2023/5452"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:L/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-reboot"},"published":"2023-04-22"},{"id":"CVE-2023-22655","cve":"CVE-2023-22655","aliases":[],"title":"Intel 3rd/4th Gen Xeon with SGX or TDX (protection mechanism failure): A protection mechanism in 3rd and 4th generation","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel 3rd/4th Gen Xeon with SGX or TDX (protection mechanism failure)","year":"2023","cvss_score":6.1,"severity":"medium","kev":false,"impact":"A protection mechanism in 3rd and 4th generation Xeon fails when SGX or TDX is in use, letting a privileged local user escalate. Affects the exact Xeon generations most AI datacenter hosts were built on between 2021 and 2024.","attack_vector":"Privileged local access on the host.","remediation":"Microcode update plus TCB recovery. Late-loadable microcode, reboot, then re-attest enclaves and trust domains.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-22655","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00960.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"tags":["tenant-isolation"],"published":"2024-03-14"},{"id":"CVE-2023-28642","cve":"CVE-2023-28642","aliases":[],"title":"runc: AppArmor bypass when /proc inside the container is symlinked with a specific mount config","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"runc","year":"2023","cvss_score":6.1,"severity":"medium","kev":false,"impact":"AppArmor bypass when /proc inside the container is symlinked with a specific mount config","attack_vector":"Malicious image or tenant-controlled pod spec","remediation":"Replace runc binary; drain node","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-28642"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2023-03-29"},{"id":"CVE-2023-31012","cve":"CVE-2023-31012","aliases":[],"title":"DGX H100 BMC (REST): DoS / data tampering","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX H100 BMC (REST)","year":"2023","cvss_score":6.1,"severity":"medium","kev":false,"impact":"DoS / data tampering","attack_vector":"Network-adjacent BMC REST client","remediation":"Flash BMC 23.08.18 out-of-band","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5473/5473.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:H/UI:N/S:U/C:H/I:H/A:N","cwe":["CWE-20"],"published":"2023-09-20"},{"id":"CVE-2023-31013","cve":"CVE-2023-31013","aliases":[],"title":"DGX H100 BMC (REST): DoS / data tampering","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX H100 BMC (REST)","year":"2023","cvss_score":6.1,"severity":"medium","kev":false,"impact":"DoS / data tampering","attack_vector":"Network-adjacent BMC REST client","remediation":"Flash BMC 23.08.18 out-of-band","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5473/5473.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:H/UI:N/S:U/C:H/I:H/A:N","cwe":["CWE-20"],"published":"2023-09-20"},{"id":"CVE-2023-31020","cve":"CVE-2023-31020","aliases":[],"title":"GPU Display Driver (Windows): DoS / data tampering","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver (Windows)","year":"2023","cvss_score":6.1,"severity":"medium","kev":false,"impact":"DoS / data tampering","attack_vector":"Local user","remediation":"Upgrade Oct-2023 driver branch","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5491/5491.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:L/A:H","cwe":["CWE-284"],"published":"2023-11-02"},{"id":"CVE-2023-40546","cve":"CVE-2023-40546","aliases":["shim 15.8 batch"],"title":"shim (mok.c mirror_one_esl): NULL pointer dereference while printing an error message stops the node from booting","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"shim (mok.c mirror_one_esl)","year":"2023","cvss_score":6.1,"severity":"medium","kev":false,"impact":"NULL pointer dereference while printing an error message stops the node from booting. On a fleet this is a denial of service you cannot fix over SSH - it needs console or BMC access per affected node, which is expensive at scale.","attack_vector":"Requires attacker-influenced MOK/ESL data on the node.","remediation":"shim package update + reboot. Availability-only, so it can ride a normal maintenance window rather than an emergency one.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-40546","https://www.openwall.com/lists/oss-security/2024/01/26/1"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-01-29"},{"cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-52617","cve":"CVE-2023-52617","aliases":[],"title":"Linux kernel (drivers/pci/switch): If a userspace process is holding the Switchtec management character device open","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/pci/switch)","year":"2023","cvss_score":6.1,"severity":"medium","kev":false,"impact":"If a userspace process is holding the Switchtec management character device open when the switch is surprise-removed, the final release runs long after the driver is gone - the MMIO mapping is already torn down, so the release path's register write takes a fatal page fault, and the DMA teardown that follows hands a stale device pointer to dma_free_coherent(). A userspace file descriptor thereby outlives and then corrupts kernel state belonging to the PCIe switch that fans out the node's GPUs and NVMe.","attack_vector":"Needs the switchtec driver bound to a Microsemi/Microchip PCIe switch - real hardware in GPU and NVMe fabric chassis - plus a process holding /dev/switchtec* open, which is the management or telemetry agent that normally does. The trigger is device-side: a surprise removal, link drop or switch reset while that fd is open. Access to the chardev is root/administrative, so this is not a tenant-initiated exploit; it is a management-plane crash of a shared node that a misbehaving switch can provoke.","remediation":"Boot a kernel that moves the MRPC DMA shutdown into switchtec_pci_remove() after stdev_kill() and takes a counted reference on the pdev. Interim: have management agents close /dev/switchtec* rather than holding it open indefinitely, and keep the chardev out of containers.","references":["https://git.kernel.org/stable/c/d8c293549946ee5078ed0ab77793cec365559355","https://git.kernel.org/stable/c/4a5d0528cf19dbf060313dffbe047bc11c90c24c","https://nvd.nist.gov/vuln/detail/CVE-2023-52617"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-0075","cve":"CVE-2024-0075","aliases":[],"title":"GPU Display Driver: DoS (null deref)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2024","cvss_score":6.1,"severity":"medium","kev":false,"impact":"DoS (null deref)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; rolling reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0075","https://github.com/NVIDIA/product-security/tree/main/2024/5520"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"published":"2024-03-27"},{"id":"CVE-2024-0115","cve":"CVE-2024-0115","aliases":[],"title":"NVIDIA CV-CUDA: A long-running CV-CUDA Python process consumes resources without bound, ending in denial of service","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA CV-CUDA","year":"2024","cvss_score":6.1,"severity":"medium","kev":false,"impact":"A long-running CV-CUDA Python process consumes resources without bound, ending in denial of service and data loss. On a shared inference node this means one tenant's preprocessing job can starve the box.","attack_vector":"Local, low privileges - a user able to submit work through the CV-CUDA Python API.","remediation":"Update CV-CUDA per bulletin 5560 and rebuild affected images. Cost: package update and job restart; no driver or firmware change. Consider cgroup memory limits on preprocessing containers as a standing control.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0115","https://github.com/NVIDIA/product-security/tree/main/2024/5560"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:L/A:H","cwe":["CWE-400"],"published":"2024-08-12"},{"id":"CVE-2024-10086","cve":"CVE-2024-10086","aliases":[],"title":"HashiCorp Consul: Missing Content-Type header lets user input be reinterpreted","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HashiCorp Consul","year":"2024","cvss_score":6.1,"severity":"medium","kev":false,"impact":"Missing Content-Type header lets user input be reinterpreted -> reflected XSS on the Consul UI","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; keep the Consul UI behind the ops VPN","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-10086"],"status":"curated","published":"2024-10-30"},{"id":"CVE-2024-10649","cve":"CVE-2024-10649","aliases":[],"title":"Weights & Biases OpenUI: Unauthenticated endpoints allow file upload and download","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Weights & Biases OpenUI","year":"2024","cvss_score":6.1,"severity":"medium","kev":false,"impact":"Unauthenticated endpoints allow file upload and download","attack_vector":"Unauthenticated network","remediation":"Upgrade; only affects the OpenUI project, not the core W&B SDK","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-10649"],"status":"curated","published":"2025-02-10"},{"id":"CVE-2024-25630","cve":"CVE-2024-25630","aliases":[],"title":"Cilium: WireGuard transparent encryption not applied to some pod traffic","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2024","cvss_score":6.1,"severity":"medium","kev":false,"impact":"WireGuard transparent encryption not applied to some pod traffic","attack_vector":"Anyone on the underlay network","remediation":"Rolling Cilium upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-25630"],"status":"curated","published":"2024-02-20"},{"id":"CVE-2024-25631","cve":"CVE-2024-25631","aliases":[],"title":"Cilium: With an external kvstore and WireGuard, pod-to-pod traffic is unencrypted","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2024","cvss_score":6.1,"severity":"medium","kev":false,"impact":"With an external kvstore and WireGuard, pod-to-pod traffic is unencrypted","attack_vector":"Anyone on the underlay network","remediation":"Rolling Cilium upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-25631"],"status":"curated","published":"2024-02-20"},{"cwe":["CWE-476","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-26813","cve":"CVE-2024-26813","aliases":[],"title":"Linux kernel (drivers/vfio/platform): A tenant holding a vfio-platform device can loopback-trigger an interrupt before","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/vfio/platform)","year":"2024","cvss_score":6.1,"severity":"medium","kev":false,"impact":"A tenant holding a vfio-platform device can loopback-trigger an interrupt before any signaling eventfd has been configured, dereferencing a NULL trigger from the interrupt path and taking the host kernel down. Same defect class as the vfio-pci INTx bugs - the device fd's interrupt state and its eventfd lifetime were not tied together.","attack_vector":"A container or VM holding a vfio-platform device fd, calling VFIO_DEVICE_SET_IRQS with the loopback trigger flags before ever registering an eventfd. No host root where it applies. Conditional on the vfio-platform driver being loaded with a non-PCI platform device assigned - an ARM/embedded SoC configuration, so a standard x86 GPU node is not exposed.","remediation":"The record lists no fixed release; boot a kernel carrying the stable fix commits below. Interim control: blacklist vfio-platform and vfio-amba on fleets that never assign platform devices, which is the normal case for x86 GPU nodes.","references":["https://git.kernel.org/stable/c/07afdfd8a68f9eea8db0ddc4626c874f29d2ac5e","https://git.kernel.org/stable/c/09452c8fcbd7817c06e8e3212d99b45917e603a5","https://nvd.nist.gov/vuln/detail/CVE-2024-26813"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-26814","cve":"CVE-2024-26814","aliases":[],"title":"Linux kernel (drivers/vfio/fsl-mc): The eventfd trigger for a vfio-fsl-mc interrupt starts out NULL and becomes NULL","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/vfio/fsl-mc)","year":"2024","cvss_score":6.1,"severity":"medium","kev":false,"impact":"The eventfd trigger for a vfio-fsl-mc interrupt starts out NULL and becomes NULL again if the tenant sets it to -1, but the loopback test path fires the handler without checking. A tenant holding the device fd dereferences NULL in the kernel interrupt path and downs the host for everyone on the node.","attack_vector":"A container or VM holding a vfio-fsl-mc device fd, invoking the loopback interrupt trigger through VFIO_DEVICE_SET_IRQS before setting an eventfd or after clearing it to -1. No host root. Conditional on the NXP DPAA2 fsl-mc bus and its vfio driver - NXP SoC hardware, not present in x86 or ARM GPU fleets.","remediation":"The record lists no fixed release; boot a kernel carrying the stable fix commits below. Interim control: blacklist vfio-fsl-mc on any fleet that does not assign DPAA2 objects.","references":["https://git.kernel.org/stable/c/a563fc18583ca4f42e2fdd0c70c7c618288e7ede","https://git.kernel.org/stable/c/250219c6a556f8c69c5910fca05a59037e24147d","https://nvd.nist.gov/vuln/detail/CVE-2024-26814"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-28249","cve":"CVE-2024-28249","aliases":[],"title":"Cilium: IPsec-eligible traffic matching L7 policy is sent unencrypted","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2024","cvss_score":6.1,"severity":"medium","kev":false,"impact":"IPsec-eligible traffic matching L7 policy is sent unencrypted","attack_vector":"Anyone on the underlay network","remediation":"Rolling Cilium upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-28249"],"status":"curated","published":"2024-03-18"},{"id":"CVE-2024-28250","cve":"CVE-2024-28250","aliases":[],"title":"Cilium: WireGuard-eligible traffic matching L7 policy is sent unencrypted between nodes","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2024","cvss_score":6.1,"severity":"medium","kev":false,"impact":"WireGuard-eligible traffic matching L7 policy is sent unencrypted between nodes","attack_vector":"Anyone on the underlay network","remediation":"Rolling Cilium upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-28250"],"status":"curated","published":"2024-03-18"},{"id":"CVE-2024-34923","cve":"CVE-2024-34923","aliases":[],"title":"Avocent DSR2030 / SVIP1020 KVM-over-IP appliance: A reflected XSS in the appliance's web interface lets an attacker","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Avocent DSR2030 / SVIP1020 KVM-over-IP appliance","year":"2024","cvss_score":6.1,"severity":"medium","kev":false,"impact":"A reflected XSS in the appliance's web interface lets an attacker who can get an operator to click a crafted link run JavaScript in that operator's browser session — enough to steal their session cookie and act as them on the KVM appliance.","attack_vector":"Requires social engineering: the victim (an operator with legitimate access to the KVM appliance) has to click a link the attacker controls while authenticated to the device.","remediation":"Software upgrade — DSR2030 to firmware 03.07.01.23 or later, SVIP1020 to 01.07.00.00 or later. Standard firmware flash per unit; no serial/KVM downtime beyond the reboot itself.","references":["https://ka1ne1.github.io/avocent_xss.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-05-27"},{"id":"CVE-2024-48869","cve":"CVE-2024-48869","aliases":[],"title":"Intel Xeon 6 E-core with TDX or SGX: Improper restriction of software interfaces to hardware features on Xeon 6 E-core","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Xeon 6 E-core with TDX or SGX","year":"2024","cvss_score":6.1,"severity":"medium","kev":false,"impact":"Improper restriction of software interfaces to hardware features on Xeon 6 E-core parts when TDX or SGX is in use. Reaches both confidential-compute technologies on the same silicon, so a single platform update covers both boundaries.","attack_vector":"Local access on an affected Xeon 6 platform with TDX or SGX enabled.","remediation":"OEM platform firmware/BIOS update plus TCB recovery for whichever technology you use. Drain, reboot, re-attest.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-48869","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01268.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2025-05-13"},{"id":"CVE-2024-50302","cve":"CVE-2024-50302","aliases":[],"title":"Linux kernel (HID): Uninitialised HID report buffer leaks kernel memory","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (HID)","year":"2024","cvss_score":6.1,"severity":"medium","kev":true,"impact":"Uninitialised HID report buffer leaks kernel memory [KEV]","attack_vector":"Local user with USB device access","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2024-50302"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-11-19"},{"id":"CVE-2024-5321","cve":"CVE-2024-5321","aliases":[],"title":"Kubernetes (kubelet): Incorrect permissions on Windows container log directories allow privilege escalation","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kubelet)","year":"2024","cvss_score":6.1,"severity":"medium","kev":false,"impact":"Incorrect permissions on Windows container log directories allow privilege escalation","attack_vector":"Any tenant workload on a Windows node","remediation":"Rolling kubelet upgrade; Windows node drain","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2024-07-18"},{"cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-56561","cve":"CVE-2024-56561","aliases":[],"title":"Linux kernel (drivers/pci/endpoint): Pci_epc_destroy() releases the PCI domain ID using a device object it has already","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/pci/endpoint)","year":"2024","cvss_score":6.1,"severity":"medium","kev":false,"impact":"Pci_epc_destroy() releases the PCI domain ID using a device object it has already unregistered and freed, so the domain-ID release runs against freed memory. Beyond the use-after-free itself, it releases the WRONG domain ID (the EPC's rather than its parent's), corrupting the kernel's domain-number allocator for every controller that follows.","attack_vector":"Endpoint mode required: the machine must run a PCIe endpoint controller with the EPC core registered. Reached when the EPC is destroyed - controller driver unbind or module unload, which is host root. Not tenant-reachable and inert on a normal GPU host; it matters on DPU/smartNIC style devices that run Linux as the endpoint and whose controller drivers get reloaded.","remediation":"Update to 6.12 or later, or take the stable commits below. Interim: avoid unbinding or unloading PCIe endpoint controller drivers on a running system; reboot instead.","references":["https://git.kernel.org/stable/c/c74a1df6c2a2df7dd45c3fc1a5edc29a075dcf22","https://git.kernel.org/stable/c/4acc902ed3743edd4ac2d3846604a99d17104359","https://nvd.nist.gov/vuln/detail/CVE-2024-56561"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-39783","cve":"CVE-2025-39783","aliases":[],"title":"Linux kernel (drivers/pci/endpoint): The endpoint function core calls list_del() on a structure that is a list HEAD","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/pci/endpoint)","year":"2025","cvss_score":6.1,"severity":"medium","kev":false,"impact":"The endpoint function core calls list_del() on a structure that is a list HEAD, not a list entry, so tearing down an endpoint function driver writes two pointers into memory that has already been freed. KASAN reports it as a slab use-after-free write - an attacker-useful primitive, not just a crash, if the freed slab has been reallocated.","attack_vector":"Endpoint mode required. Triggered by unloading an endpoint function driver that registered a configfs attribute group; the upstream KASAN splat came from unloading nvmet_pci_epf, the NVMe-over-PCIe endpoint target. That is host root on the endpoint machine, not a tenant action - but nvmet_pci_epf is exactly the kind of tenant-facing storage target a provider would run on an endpoint device, and its module lifecycle is part of normal service management.","remediation":"Update to a kernel carrying the fix (no fixed_in published; stable commits below). Interim: do not unload endpoint function driver modules on a live system - especially nvmet_pci_epf - and reboot to change endpoint function configuration.","references":["https://git.kernel.org/stable/c/80ea6e6904fb2ba4ccb5d909579988466ec65358","https://git.kernel.org/stable/c/d5aecddc3452371d9da82cdbb0c715812524b54b","https://nvd.nist.gov/vuln/detail/CVE-2025-39783"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-23528","cve":"CVE-2026-23528","aliases":[],"title":"Dask distributed (+ Jupyter proxy): Exposure when Dask, JupyterLab and jupyter-server-proxy are combined","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Dask distributed (+ Jupyter proxy)","year":"2026","cvss_score":6.1,"severity":"medium","kev":false,"impact":"Exposure when Dask, JupyterLab and jupyter-server-proxy are combined","attack_vector":"Notebook user / network attacker","remediation":"Upgrade to 2026.1.0+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-23528"],"status":"curated","published":"2026-01-16"},{"id":"CVE-2026-26963","cve":"CVE-2026-26963","aliases":[],"title":"Cilium: With native routing plus WireGuard node encryption, traffic from pods on other nodes is wrongly permitted","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2026","cvss_score":6.1,"severity":"medium","kev":false,"impact":"With native routing plus WireGuard node encryption, traffic from pods on other nodes is wrongly permitted","attack_vector":"Any pod on the cluster network","remediation":"Rolling Cilium upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-26963"],"status":"curated","published":"2026-02-20"},{"id":"CVE-2026-29777","cve":"CVE-2026-29777","aliases":[],"title":"Traefik: A tenant with HTTPRoute write access injects backtick-delimited rule tokens into Traefik's router","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Traefik","year":"2026","cvss_score":6.1,"severity":"medium","kev":false,"impact":"A tenant with HTTPRoute write access injects backtick-delimited rule tokens into Traefik's router rule language; cross-tenant route hijack","attack_vector":"Cluster user with namespace access","remediation":"Rolling Traefik upgrade to 3.6.10+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-29777"],"status":"curated","published":"2026-03-11"},{"id":"CVE-2026-41568","cve":"CVE-2026-41568","aliases":[],"title":"Docker / moby: Companion `docker cp` mount-setup race","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2026","cvss_score":6.1,"severity":"medium","kev":false,"impact":"Companion `docker cp` mount-setup race","attack_vector":"Any tenant workload","remediation":"Upgrade Docker Engine to 29.5.1+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-41568"],"status":"curated","published":"2026-06-12"},{"cwe":["CWE-476","CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-46034","cve":"CVE-2026-46034","aliases":[],"title":"Linux kernel (drivers/vfio/cdx): A tenant can call the interrupt-configuration ioctl with the trigger flags before MSI","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/vfio/cdx)","year":"2026","cvss_score":6.1,"severity":"medium","kev":false,"impact":"A tenant can call the interrupt-configuration ioctl with the trigger flags before MSI has ever been configured, so the driver walks an IRQ array that was never allocated and dereferences NULL. A straight userspace-to-kernel NULL dereference through the passthrough device fd - the same call-ordering assumption that had to be fixed in vfio-pci years earlier.","attack_vector":"A container or VM holding a vfio-cdx device fd, calling VFIO_DEVICE_SET_IRQS with VFIO_IRQ_SET_DATA_BOOL or DATA_NONE before the EVENTFD path has allocated the IRQ array. No host root, no race. Conditional on the AMD/Xilinx CDX bus and the vfio-cdx driver being present - FPGA/embedded hardware rather than a standard GPU node.","remediation":"The record lists no fixed release; boot a kernel carrying the stable fix commits below. Interim control: blacklist vfio-cdx on fleets that do not assign CDX devices.","references":["https://git.kernel.org/stable/c/51bf7638f33aece41cb3f4cbeb942cc52950e329","https://git.kernel.org/stable/c/5d6c349c9823eb819fed8b537b088cf38126018c","https://nvd.nist.gov/vuln/detail/CVE-2026-46034"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.0/AV:N/AC:L/PR:N/UI:R/S:C/C:L/I:L/A:N","cwe":["CWE-601"],"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2025-017-bentoml-1-3-9-open-redirect-in-t","cve":null,"aliases":["GHSA-564p-rx2q-4c8v"],"title":"BentoML 1.3.9 (open redirect in the serving UI): A crafted URL against the BentoML server bounces the visitor to an","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"BentoML 1.3.9 (open redirect in the serving UI)","year":"2025","cvss_score":6.1,"severity":"medium","kev":false,"impact":"A crafted URL against the BentoML server bounces the visitor to an attacker-chosen site while the link still looks like it points at your model endpoint. Useful for credential phishing against the people who operate or consume the endpoint. No CVE was assigned to the BentoML advisory; the underlying Gradio issue is CVE-2024-4940.","attack_vector":"A remote unauthenticated attacker who gets a user to click a link pointing at the BentoML host.","remediation":"Upgrade BentoML past 1.3.9 and restart the serving pods. If you cannot upgrade, block off-host redirect targets at the ingress in front of the endpoint.","references":["https://github.com/advisories/GHSA-564p-rx2q-4c8v","https://huntr.com/bounties/2a284ff6-cc6c-4a10-b72e-1bb31c842bca"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:C/C:N/I:N/A:H","cwe":["CWE-264"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2013-3495","cve":"CVE-2013-3495","aliases":["XSA-59"],"title":"Intel VT-d interrupt remapping engine as used by Xen 3.3.x-4.3.x: Proof that interrupt remapping is not a complete","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel VT-d interrupt remapping engine as used by Xen 3.3.x-4.3.x","year":"2013","cvss_score":6,"severity":"medium","kev":false,"impact":"Proof that interrupt remapping is not a complete containment boundary. A malformed MSI emitted by a bus-mastering device the tenant controls provokes a System Error Reporting NMI, and the NMI is delivered natively - it goes around the VT-d IR engine rather than through it - panicking the host. Every tenant on that node loses their work. The reason this belongs in a modern catalogue is architectural rather than historical: operators reason about passthrough safety as 'IOMMU plus IR equals isolated', and this is the canonical counterexample showing that error-reporting paths in the platform bypass the remapping engine entirely.","attack_vector":"A guest with a passed-through, bus-mastering-capable PCI device on an Intel VT-d host. Attack is on the host and therefore on every co-tenant.","remediation":"Apply XSA-59 and reboot the hypervisor. Note the advisory's own framing - part of the mitigation is platform configuration of SERR/NMI handling, not just hypervisor code, so the fix needs to be validated per server model rather than assumed fleet-wide. For a GPU rental fleet the pragmatic control is to make host NMI-triggered panics a monitored, attributable event: if you cannot prevent a tenant from crashing the node, you should at least be able to tell which tenant did it and stop selling to them.","references":["https://xenbits.xen.org/xsa/advisory-59.html","https://nvd.nist.gov/vuln/detail/CVE-2013-3495"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:C/C:N/I:N/A:H","cwe":["CWE-264"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2015-2150","cve":"CVE-2015-2150","aliases":["XSA-120"],"title":"Xen 3.3.x-4.5.x and Linux kernel through 3.19.1 - PCI command register access for assigned devices: A tenant clears the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen 3.3.x-4.5.x and Linux kernel through 3.19.1 - PCI command register access for assigned devices","year":"2015","cvss_score":6,"severity":"medium","kev":false,"impact":"A tenant clears the memory- or I/O-decode bit in the PCI command register of its assigned device and then touches the device's BAR. The transaction gets an Unsupported Request completion, the platform raises a fatal NMI, and the host dies with every co-tenant on it. What makes this one worth carrying is the scope line - it is not a Xen-only bug. The same unmediated command-register write existed in the Linux kernel's own device-assignment path through 3.19.1, so KVM/VFIO GPU passthrough hosts were affected too, and that is the configuration most GPU clouds actually run.","attack_vector":"Guest administrator with any assigned PCI device, on either Xen or KVM/VFIO. Two config-space writes and one MMIO read.","remediation":"Update Xen per XSA-120 and the Linux kernel past 3.19.1 (the VFIO side gained emulation of the command register rather than passing writes through), then reboot the hosts. Firmware matters here as much as software: whether a UR turns into a fatal NMI or is logged and swallowed is a platform/BIOS AER configuration choice, so validate the behaviour per server model. Track CVE-2015-8553 with it - same advisory family, and that one leaks uninitialised host memory rather than crashing.","references":["https://xenbits.xen.org/xsa/advisory-120.html","https://access.redhat.com/security/cve/CVE-2015-2150","https://nvd.nist.gov/vuln/detail/CVE-2015-2150"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2017-5703","cve":"CVE-2017-5703","aliases":["INTEL-SA-00087"],"title":"SPI flash configuration (flash descriptor / protected range registers) across multiple Intel platforms","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"SPI flash configuration (flash descriptor / protected range registers) across multiple Intel platforms","year":"2017","cvss_score":6,"severity":"medium","kev":false,"impact":"Misconfiguration of SPI flash protection lets a local attacker change how the SPI flash behaves, up to and including bricking the node. In a GPU fleet the operator-facing outcome is a hard denial of service that no reimage fixes: the node will not POST and needs a physical flash recovery (external programmer or OEM RMA), so it is an unplanned rack visit and a node out of revenue for days. Where write protection is incomplete rather than merely unstable, the same weakness is the standard route to a persistent BIOS implant that survives every reimage and every tenant handoff.","attack_vector":"Local privileged code on the host writing to the SPI controller's configuration and protected-range registers. Reachable by any tenant with root on a bare-metal node.","remediation":"BIOS/platform firmware update from the OEM that sets the flash descriptor and protected-range/BIOS-lock registers correctly - HPE, Dell, Supermicro and Lenovo all shipped these; reboot and drain required. Independently of the patch, audit the flash-protection state on your actual fleet (BIOSWE/BLE/SMM_BWP/PRx and descriptor lock) with a tool like CHIPSEC as part of node acceptance and node reclaim, because these bits are set by the OEM's BIOS build and vary by SKU and by BIOS version even within one OEM.","references":["https://nvd.nist.gov/vuln/detail/CVE-2017-5703","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00087.html","https://support.hpe.com/hpsc/doc/public/display?docLocale=en_US&docId=emr_na-hpesbhf03867en_us"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-43784","cve":"CVE-2021-43784","aliases":[],"title":"runc: Netlink bytemsg length integer overflow in libcontainer allows config injection / partial escape","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"runc","year":"2021","cvss_score":6,"severity":"medium","kev":false,"impact":"Netlink bytemsg length integer overflow in libcontainer allows config injection / partial escape","attack_vector":"Any tenant workload with control over container config","remediation":"Replace runc binary; drain node","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-43784"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2021-12-06"},{"id":"CVE-2022-21233","cve":"CVE-2022-21233","aliases":[],"title":"Intel CPU (AEPIC Leak): Stale data read from the legacy xAPIC MMIO page - leaks SGX enclave and cross-domain data","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel CPU (AEPIC Leak)","year":"2022","cvss_score":6,"severity":"medium","kev":false,"impact":"Stale data read from the legacy xAPIC MMIO page - leaks SGX enclave and cross-domain data","attack_vector":"Local user; tenant VM guest","remediation":"Microcode + reboot. Kills naive SGX-based confidential-compute claims on affected parts","references":["https://access.redhat.com/security/cve/CVE-2022-21233"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"published":"2022-08-18"},{"id":"CVE-2022-36382","cve":"CVE-2022-36382","aliases":[],"title":"Intel Ethernet E810 Series and Ethernet 700 Series firmware: Out-of-bounds write in firmware across both the E810 line","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Ethernet E810 Series and Ethernet 700 Series firmware","year":"2022","cvss_score":6,"severity":"medium","kev":false,"impact":"Out-of-bounds write in firmware across both the E810 line and the older 700 Series (X710/XL710/XXV710), triggerable by a privileged host user. Notable because it spans two adapter generations — if you have a mixed fleet, the version target differs per family (E810 before 1.7.0.8, 700 Series before 9.101) and a single blanket firmware policy will miss half of it.","attack_vector":"Privileged local user on the host.","remediation":"Flash E810 to 1.7.0.8+ and 700-Series adapters to 9.101+. Cold power cycle on both. Because the two families need different images, build the NVM update into your provisioning pipeline keyed on device ID rather than doing it by hand.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-36382"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2023-02-16"},{"id":"CVE-2022-38090","cve":"CVE-2022-38090","aliases":[],"title":"Intel processors with SGX (shared resource isolation): Improper isolation of shared microarchitectural resources lets a","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors with SGX (shared resource isolation)","year":"2022","cvss_score":6,"severity":"medium","kev":false,"impact":"Improper isolation of shared microarchitectural resources lets a privileged user extract information from SGX enclaves. Another TCB-recovery event for anyone selling enclave-backed confidentiality.","attack_vector":"Privileged local access on the host.","remediation":"Microcode update and re-attestation. Late-loadable microcode plus reboot; no OEM BIOS strictly required for the microcode component.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-38090","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00767.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"tags":["tenant-isolation"],"published":"2023-02-16"},{"id":"CVE-2022-42285","cve":"CVE-2022-42285","aliases":[],"title":"NVIDIA DGX A100 - SBIOS / SMM firmware: A privileged user can disable SPI flash write protection during the Pre-EFI","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX A100 - SBIOS / SMM firmware","year":"2022","cvss_score":6,"severity":"medium","kev":false,"impact":"A privileged user can disable SPI flash write protection during the Pre-EFI Initialization phase, which removes the platform's own guard against firmware overwrite. This is firmware-level persistence: it survives OS reinstall, image re-flash and tenant handoff, and it is invisible to anything running above it. On a bare-metal GPU rental business it is the difference between wiping a node between tenants and not actually being able to.","attack_vector":"Local and already privileged - host root, or code that has reached the platform firmware/SMM path. It is not a first foothold; it is what turns a one-time root compromise into something you cannot remediate by reimaging.","remediation":"Flash the fixed SBIOS from bulletin 5435. Cost: not live-patchable. Full node drain, host power cycle, and on DGX the SBIOS ships inside a firmware bundle alongside BMC and CPLD components, so budget 30-60 minutes of node downtime plus a post-flash health check. Firmware rollback protection means you cannot cleanly revert - stage on one node before the fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42285","https://github.com/NVIDIA/product-security/tree/main/2022/5435"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:N/I:H/A:H","cwe":["CWE-1231"],"fleet":{"pain_class":"firmware-flash"},"published":"2023-01-13"},{"id":"CVE-2022-42286","cve":"CVE-2022-42286","aliases":[],"title":"DGX-2 SBIOS: OOB write in BIOS","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX-2 SBIOS","year":"2022","cvss_score":6,"severity":"medium","kev":false,"impact":"OOB write in BIOS","attack_vector":"Local operator","remediation":"Flash SBIOS out-of-band; node power cycle","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42286","https://github.com/NVIDIA/product-security/tree/main/2023/5449"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:N/I:H/A:H","cwe":["CWE-119"],"fleet":{"pain_class":"firmware-flash"},"published":"2023-01-13"},{"id":"CVE-2022-42287","cve":"CVE-2022-42287","aliases":[],"title":"DGX-2 BMC: Path traversal","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX-2 BMC","year":"2022","cvss_score":6,"severity":"medium","kev":false,"impact":"Path traversal","attack_vector":"Network-adjacent authenticated","remediation":"Flash DGX-2 BMC firmware","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42287","https://github.com/NVIDIA/product-security/tree/main/2024/5513"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:H/I:H/A:N","cwe":["CWE-22"],"fleet":{"pain_class":"firmware-flash"},"published":"2023-01-13"},{"cwe":["CWE-191"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-49199","cve":"CVE-2022-49199","aliases":[],"title":"Linux kernel RDMA core netlink (nldev_stat_set_counter_dynamic_doit): The dynamic-counter netlink setter bounded its","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel RDMA core netlink (nldev_stat_set_counter_dynamic_doit)","year":"2022","cvss_score":6,"severity":"medium","kev":false,"impact":"The dynamic-counter netlink setter bounded its index from above but not from below, so a negative index underflowed past the start of the counter array. It is reachable from the RDMA netlink admin surface, which in clusters that hand namespaced RDMA administration to tenant operators (or to a fabric-management sidecar) is not as far from the tenant as it looks.","attack_vector":"Local RDMA netlink message; requires the privilege to configure RDMA statistics counters in the namespace.","remediation":"Kernel update changing the index to an unsigned type. Audit which workloads hold CAP_NET_ADMIN in their netns - on many clusters that capability is granted far more broadly than intended, and it is the gate on this and a family of similar nldev bugs.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=2a495ef04d5f42e6f00eb2bec1ee9075e3d5a771","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2022/CVE-2022-49199.json"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-1303","CWE-200"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-49611","cve":"CVE-2022-49611","aliases":[],"title":"Linux kernel (arch/x86/kvm/vmx): The return stack buffer was not refilled on VM exit when the host used IBRS/eIBRS as","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (arch/x86/kvm/vmx)","year":"2022","cvss_score":6,"severity":"medium","kev":false,"impact":"The return stack buffer was not refilled on VM exit when the host used IBRS/eIBRS as its Spectre-v2 mitigation, so return predictions a guest planted survive into host kernel execution immediately after the exit. A tenant VM can speculatively steer host returns and read host kernel memory through a cache side channel - a cross-boundary information leak, no crash and no trace in dmesg.","attack_vector":"Guest-driven and unprivileged inside the VM: any tenant vCPU can shape the RSB and then force a VM exit. Applies to Intel hosts running the IBRS/eIBRS mitigation path with kvm_intel loaded; nothing has to be exposed to the guest beyond a normal vCPU.","remediation":"Update to a stable kernel carrying this fix together with the companion RSB fixes (the record lists 4.14.297 among the affected lines). No runtime knob substitutes for the fix; switching the host to retpoline-based mitigation changes but does not remove the exposure, so patch and reboot.","references":["https://git.kernel.org/stable/c/3d323b99ff5c8c57005184056d65f6af5b0479d8","https://git.kernel.org/stable/c/17a9fc4a7b91f8599223631bb6ae6416bc0de1c0","https://nvd.nist.gov/vuln/detail/CVE-2022-49611"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2023-20588","cve":"CVE-2023-20588","aliases":[],"title":"AMD CPU (DIV0): Division-by-zero leaves stale quotient data readable across contexts - confidentiality loss on Zen 1","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"AMD CPU (DIV0)","year":"2023","cvss_score":6,"severity":"medium","kev":false,"impact":"Division-by-zero leaves stale quotient data readable across contexts - confidentiality loss on Zen 1","attack_vector":"Any tenant process in a container; tenant VM guest","remediation":"Kernel mitigation + reboot; low perf cost","references":["https://access.redhat.com/security/cve/CVE-2023-20588"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2023-08-08"},{"id":"CVE-2023-25509","cve":"CVE-2023-25509","aliases":[],"title":"NVIDIA DGX-1 - SBIOS / SMM firmware: A flaw in the Bds phase reaches firmware code execution and privilege escalation","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX-1 - SBIOS / SMM firmware","year":"2023","cvss_score":6,"severity":"medium","kev":false,"impact":"A flaw in the Bds phase reaches firmware code execution and privilege escalation. This is firmware-level persistence: it survives OS reinstall, image re-flash and tenant handoff, and it is invisible to anything running above it. On a bare-metal GPU rental business it is the difference between wiping a node between tenants and not actually being able to.","attack_vector":"Local and already privileged - host root, or code that has reached the platform firmware/SMM path. It is not a first foothold; it is what turns a one-time root compromise into something you cannot remediate by reimaging.","remediation":"Flash the fixed SBIOS from bulletin 5458. Cost: not live-patchable. Full node drain, host power cycle, and on DGX the SBIOS ships inside a firmware bundle alongside BMC and CPLD components, so budget 30-60 minutes of node downtime plus a post-flash health check. Firmware rollback protection means you cannot cleanly revert - stage on one node before the fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25509","https://github.com/NVIDIA/product-security/tree/main/2023/5458"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:N/I:H/A:H","cwe":["CWE-119"],"fleet":{"pain_class":"firmware-flash"},"published":"2023-04-22"},{"id":"CVE-2023-30456","cve":"CVE-2023-30456","aliases":[],"title":"KVM (nested VMX): Missing CR0/CR4 consistency checks in nVMX - L2 guest can break nested-virt assumptions / crash host","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"KVM (nested VMX)","year":"2023","cvss_score":6,"severity":"medium","kev":false,"impact":"Missing CR0/CR4 consistency checks in nVMX - L2 guest can break nested-virt assumptions / crash host","attack_vector":"Tenant VM guest running nested virtualisation","remediation":"Kernel patch + reboot. Cheaper interim control: disable nested virtualisation for tenant VMs, which most GPU tenants do not need","references":["https://access.redhat.com/security/cve/CVE-2023-30456"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2023-04-10"},{"id":"CVE-2023-31026","cve":"CVE-2023-31026","aliases":[],"title":"vGPU software (Virtual GPU Manager): Host-side DoS (null deref in vGPU Manager)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU software (Virtual GPU Manager)","year":"2023","cvss_score":6,"severity":"medium","kev":false,"impact":"Host-side DoS (null deref in vGPU Manager)","attack_vector":"Tenant VM guest","remediation":"Upgrade vGPU Manager on hypervisor; migrate guest VMs then reboot host","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5491/5491.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:C/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"published":"2023-11-02"},{"id":"CVE-2023-31355","cve":"CVE-2023-31355","aliases":["decommissioned guest memory disclosure","UMC seed reuse"],"title":"AMD SEV-SNP firmware, guest teardown / UMC key seed handling: TENANT HANDOFF FAILURE","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV-SNP firmware, guest teardown / UMC key seed handling","year":"2023","cvss_score":6,"severity":"medium","kev":false,"impact":"TENANT HANDOFF FAILURE. A malicious hypervisor can overwrite a guest's UMC seed such that memory belonging to an already-decommissioned confidential guest becomes readable. In a rented-GPU business the slot a customer just released is immediately resold; this says the previous tenant's plaintext - checkpoints, weights, prompts, keys still resident in DRAM - can be recovered after their VM is gone. It breaks the one property you cannot buy back with an apology, and it does so on a path your own automation exercises thousands of times a day.","attack_vector":"Malicious or compromised hypervisor / host root, acting after a confidential guest terminates. No access to the victim tenant needed at all - only control of the host they used to be on.","remediation":"Same package as CVE-2024-21980 (AMD-SB-3011): hot-loadable SEV firmware 1.37.14 hex (Milan) / 1.37.24 hex (Genoa) with no reboot, or Platform Initialization firmware MilanPI 1.0.0.D / GenoaPI 1.0.0.C via OEM BIOS with a reboot. Prioritize this one over the rest of the batch. Until it is fixed, do not treat guest teardown as a memory-sanitization boundary - force an explicit scrub or a full host reboot between confidential tenants. The fix bumps TCB[SNP], so re-baseline attestation policies.","references":["https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3011.html","https://nvd.nist.gov/vuln/detail/CVE-2023-31355"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-08-05"},{"id":"CVE-2023-34342","cve":"CVE-2023-34342","aliases":["AMI-SA-2023005","NVIDIA OSR review"],"title":"AMI MegaRAC SPx (IPMI handler): Arbitrary file upload and download through the BMC's IPMI handler","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx (IPMI handler)","year":"2023","cvss_score":6,"severity":"medium","kev":false,"impact":"Arbitrary file upload and download through the BMC's IPMI handler. Download gives the attacker the BMC's stored secrets and configuration; upload gives them a way to drop a payload onto the controller's filesystem and, depending on where it lands, get it executed - which is how a credentialed foothold becomes a persistent BMC implant. Availability damage is also on the table: writing over the wrong file bricks the controller.","attack_vector":"Local access to the BMC with high privileges per AMI's vector - i.e. an attacker who already holds a BMC admin credential or has landed on the controller. Its role in a real chain is post-exploitation persistence, not initial access.","remediation":"Firmware flash to SPx_12.7 / SPx_13.5, out-of-band per node, ODM-gated. Config-only reduction: disable IPMI-over-LAN so the handler is not reachable from the network at all and drive management through Redfish, accepting that this breaks ipmitool-based provisioning and monitoring tooling. Also worth doing regardless: alert on any BMC firmware or filesystem change, because this class of bug is invisible from the host OS.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2023005.pdf","https://nvd.nist.gov/vuln/detail/CVE-2023-34342"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-06-12"},{"id":"CVE-2023-47165","cve":"CVE-2023-47165","aliases":[],"title":"Intel Data Center GPU Max Series 1100 / 1550: An improper conditions check lets a privileged local user take","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel Data Center GPU Max Series 1100 / 1550","year":"2023","cvss_score":6,"severity":"medium","kev":false,"impact":"An improper conditions check lets a privileged local user take the accelerator out of service. Low ceiling - denial of service only - but it is one of the very few CVEs that exist against Intel's Ponte Vecchio datacenter GPUs at all, which is itself worth knowing if you are evaluating them.","attack_vector":"Local, privileged. Host root on the node holding the GPU.","remediation":"Apply the Intel firmware/driver update for the Max Series. Cost: a GPU firmware update requires a drain and reboot; the driver alone needs a module reload.","references":["https://www.intel.com/content/www/us/en/security-center/default.html","https://nvd.nist.gov/vuln/detail/CVE-2023-47165"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-05-16"},{"id":"CVE-2023-47855","cve":"CVE-2023-47855","aliases":[],"title":"Intel TDX module: The TDX module is the software that stands between the host/VMM and every confidential VM on the box","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel TDX module","year":"2023","cvss_score":6,"severity":"medium","kev":false,"impact":"The TDX module is the software that stands between the host/VMM and every confidential VM on the box; a privilege escalation inside it is a break of the boundary that separates a tenant's trust domain from the operator and from other TDs. Specific flaw: a second input-validation gap in the same module version range.","attack_vector":"A privileged user on the host - which in the TDX threat model is the adversary the whole design exists to exclude, so 'requires host privilege' is not a mitigating factor here.","remediation":"Update the Intel TDX module. The TDX module is loaded by the SEAM loader at boot, so the practical rollout is: stage the new module, drain every trust domain off the node, and reboot. It is not a live-patchable component and running TDs cannot be migrated through it. After the update, every TD must re-attest because the TDX module SVN is part of the attestation report - so anything that pinned the old measurement will fail until you update your attestation policy too. No OEM BIOS release needed for the module itself, which makes this materially faster than a platform firmware update.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-47855","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01036.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2024-05-16"},{"id":"CVE-2024-20399","cve":"CVE-2024-20399","aliases":[],"title":"Cisco NX-OS CLI: Command injection giving root on the switch's underlying OS from an admin CLI session","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco NX-OS CLI","year":"2024","cvss_score":6,"severity":"medium","kev":true,"impact":"Command injection giving root on the switch's underlying OS from an admin CLI session; exploited in the wild by the Velvet Ant group, who used it to install persistent malware on the switch","attack_vector":"Network, authenticated administrator","remediation":"NX-OS upgrade with fabric failover — one leaf at a time, relying on the fabric's redundancy; a spine upgrade on a rail-optimised GPU fabric costs measurable job throughput","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-20399"],"status":"curated","published":"2024-07-01"},{"id":"CVE-2024-21850","cve":"CVE-2024-21850","aliases":[],"title":"Intel TDX SEAM loader (Seamldr): Sensitive information is not cleared before a resource is reused in the SEAM loader","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel TDX SEAM loader (Seamldr)","year":"2024","cvss_score":6,"severity":"medium","kev":false,"impact":"Sensitive information is not cleared before a resource is reused in the SEAM loader - the component that loads and measures the TDX module itself. Anything wrong at the SEAM loader layer is below the TDX module in the trust stack, so it undermines the measurement every TD attestation ultimately chains to.","attack_vector":"Privileged host user.","remediation":"Update the TDX SEAM loader (Seamldr) to 1.5.02.00 or later alongside the TDX module. Loaded at boot, so drain all trust domains and reboot the node. Re-attest afterwards; the SEAM loader version feeds the attestation chain. No OEM BIOS dependency for the loader itself.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21850","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01076.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-11-13"},{"id":"CVE-2024-21961","cve":"CVE-2024-21961","aliases":["PCIe link guest-to-host DoS"],"title":"AMD PCIe link handling (memory buffer bounds): A guest VM can drive the PCIe link into an out-of-bounds condition","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD PCIe link handling (memory buffer bounds)","year":"2024","cvss_score":6,"severity":"medium","kev":false,"impact":"A guest VM can drive the PCIe link into an out-of-bounds condition and deny service to the entire host. On an AI node the PCIe fabric is the load-bearing structure - GPUs, NVMe scratch, RDMA NICs all hang off it - so a link-level fault does not degrade one tenant, it takes the box down and kills every job on it. This is a guest-to-host availability break reachable from a normal VM, which is a materially different risk class from the ring 0 firmware issues elsewhere in this set.","attack_vector":"Attacker with access to a guest virtual machine - an ordinary paying tenant. Network-adjacent attack vector per AMD's scoring, low privilege required.","remediation":"Firmware update per AMD-SB-4013 from the OEM; BIOS flash and reboot. In the meantime the practical control is blast-radius management rather than prevention: do not co-locate high-value long-running training jobs with untrusted short-lived tenants on the same PCIe complex, and make sure checkpointing intervals assume the node can vanish. Note this bulletin targets client and embedded platform audits, so confirm applicability to your specific EPYC or Instinct SKUs with your vendor before planning a fleet-wide flash.","references":["https://www.amd.com/en/resources/product-security/bulletin/AMD-SB-4013.html","https://nvd.nist.gov/vuln/detail/CVE-2024-21961"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2026-02-13"},{"id":"CVE-2024-21978","cve":"CVE-2024-21978","aliases":[],"title":"AMD SEV-SNP firmware - input validation: Improper input validation in SEV-SNP lets a malicious hypervisor read or","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV-SNP firmware - input validation","year":"2024","cvss_score":6,"severity":"medium","kev":false,"impact":"Improper input validation in SEV-SNP lets a malicious hypervisor read or overwrite guest memory. That is the whole point of SEV-SNP defeated in one line: the host, which SNP exists to exclude, gets both read and write access to the confidential guest's pages.","attack_vector":"Malicious or compromised hypervisor. The guest need do nothing.","remediation":"Fixed in AMD SEV firmware / AGESA and reaches you as an OEM SBIOS package - AMD hands AGESA to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before shipping BIOS. **Budget one to six months of OEM lag**, longer on older platforms and sometimes never on end-of-support SKUs. Applying it means draining the host and doing a full power cycle. Because the fix moves the platform's reported SEV-SNP TCB version, you must also pull fresh VCEK certificates from AMD's Key Distribution Service and update any attestation policy your tenants pin - otherwise guests will start failing launch validation the moment the BIOS lands. Some SEV firmware can alternatively be staged from linux-firmware (amd/amd_sev_*.sbin) and committed via the ccp driver at boot, which is faster than waiting on BIOS - check whether your platform supports firmware hot-load before assuming the OEM is the only route.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21978","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2024-08-05"},{"id":"CVE-2024-24595","cve":"CVE-2024-24595","aliases":[],"title":"ClearML: Passwords stored in plaintext in MongoDB","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"ClearML","year":"2024","cvss_score":6,"severity":"medium","kev":false,"impact":"Passwords stored in plaintext in MongoDB","attack_vector":"Anyone who compromises the ClearML server","remediation":"Upgrade; rotate all credentials after any exposure","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-24595"],"status":"curated","published":"2024-02-05"},{"cwe":["CWE-476","CWE-459"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-27079","cve":"CVE-2024-27079","aliases":[],"title":"Linux kernel (drivers/iommu/intel): On device release VT-d could dereference a NULL domain and, separately, leave the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/intel)","year":"2024","cvss_score":6,"severity":"medium","kev":false,"impact":"On device release VT-d could dereference a NULL domain and, separately, leave the device's scalable-mode context entry uncleared. A context entry that survives device release is a translation the hardware still honours for whatever occupies that bus/device/function next - the stale-mapping shape of a DMA isolation break - and the NULL dereference itself oopses the host.","attack_vector":"Reached on device release/detach: a device leaving its IOMMU group, which on a GPU node happens on driver unbind, VF teardown, or when a tenant's passthrough function is returned to the host. The upstream reproducer is the kdump kernel, where deferred attach means the domain pointer is not yet assigned. Needs host-side device lifecycle events rather than a tenant ioctl - but the residue it leaves is exactly what the next tenant on that BDF would inherit.","remediation":"The record lists no fixed release; boot a kernel carrying the stable fix commits below. Interim controls: avoid rapid rebind/reassign cycles of passthrough functions between tenants, and force a full device reset and re-probe before handing a function to a new tenant.","references":["https://git.kernel.org/stable/c/333fe86968482ca701c609af590003bcea450e8f","https://git.kernel.org/stable/c/81e921fd321614c2ad8ac333b041aae1da7a1c6d","https://nvd.nist.gov/vuln/detail/CVE-2024-27079"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-36346","cve":"CVE-2024-36346","aliases":[],"title":"AMD Power Management Firmware (PMFW) - guest VM input validation causing GPU reset: Improper input validation in AMD's","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Power Management Firmware (PMFW) - guest VM input validation causing GPU reset","year":"2024","cvss_score":6,"severity":"medium","kev":false,"impact":"Improper input validation in AMD's GPU Power Management Firmware lets a **guest VM** send arbitrary input data that forces a GPU reset. On a virtualised GPU host this is a tenant taking the accelerator out from under everyone sharing it: a GPU reset kills in-flight work on the whole device, so one tenant's malformed PMFW message destroys other tenants' training progress since their last checkpoint. Cheap to trigger, expensive to absorb, and it does not require the attacker to escape their VM at all.","attack_vector":"From inside a guest VM with GPU access - SR-IOV virtual function or passthrough. No host privilege and no escape needed; the guest simply talks to the power management firmware through the interface it is legitimately given.","remediation":"Fixed in AMD GPU firmware, which on Instinct parts is delivered as a firmware bundle through the ROCm/amdgpu driver package (the PSP loads the signed blobs at driver init) rather than through the server BIOS. Practically: update the AMD GPU driver/firmware package, then **drain the node and reboot** - the firmware is loaded once at driver init, so a reload of the module with no process holding /dev/kfd is the minimum, and a reboot is what you will actually schedule. Some fixes at this layer also require a **GPU VBIOS flash** via AMD's amdvbflash/amdfwtool, which is an offline, per-card operation with real bricking risk - check the AMD bulletin for whether a VBIOS update is called out before assuming a driver package covers it. Prioritise this on any GPU virtualisation deployment with untrusted tenants - the attacker prerequisite is just 'has a GPU assigned'. Interim mitigation is thin: you cannot easily filter PMFW messages from a VF, so the practical stopgap is not co-tenanting untrusted guests on a shared physical GPU until the firmware is updated.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36346","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2025-09-06"},{"id":"CVE-2024-39283","cve":"CVE-2024-39283","aliases":[],"title":"Intel TDX module: The TDX module is the software that stands between the host/VMM and every confidential VM on the box","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel TDX module","year":"2024","cvss_score":6,"severity":"medium","kev":false,"impact":"The TDX module is the software that stands between the host/VMM and every confidential VM on the box; a privilege escalation inside it is a break of the boundary that separates a tenant's trust domain from the operator and from other TDs. Specific flaw: incomplete filtering of special elements, reachable by an authenticated user.","attack_vector":"A privileged user on the host - which in the TDX threat model is the adversary the whole design exists to exclude, so 'requires host privilege' is not a mitigating factor here.","remediation":"Update the Intel TDX module. The TDX module is loaded by the SEAM loader at boot, so the practical rollout is: stage the new module, drain every trust domain off the node, and reboot. It is not a live-patchable component and running TDs cannot be migrated through it. After the update, every TD must re-attest because the TDX module SVN is part of the attestation report - so anything that pinned the old measurement will fail until you update your attestation policy too. No OEM BIOS release needed for the module itself, which makes this materially faster than a platform firmware update.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-39283","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01010.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2024-08-14"},{"id":"CVE-2024-45779","cve":"CVE-2024-45779","aliases":["GRUB2 2025 batch"],"title":"GRUB2 (BFS filesystem parser): Integer overflow producing a heap out-of-bounds read in the BeFS parser","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (BFS filesystem parser)","year":"2024","cvss_score":6,"severity":"medium","kev":false,"impact":"Integer overflow producing a heap out-of-bounds read in the BeFS parser - leaks bootloader memory, useful for defeating any layout randomisation before pairing with a write primitive.","attack_vector":"Attacker-supplied BFS image.","remediation":"grub2 package update + reboot; or strip unused filesystem modules from the GRUB build.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-45779","https://www.openwall.com/lists/oss-security/2025/02/18/3"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-03-03"},{"id":"CVE-2024-50264","cve":"CVE-2024-50264","aliases":[],"title":"Linux kernel (vsock/virtio): Dangling pointer in vsk","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (vsock/virtio)","year":"2024","cvss_score":6,"severity":"medium","kev":false,"impact":"Dangling pointer in vsk->trans during AF_VSOCK connect - use-after-free, local root; vsock is the host/guest channel on virtualised nodes","attack_vector":"Any tenant process in a container; tenant VM guest","remediation":"Livepatchable; otherwise drain + reboot. Blacklist vsock modules on nodes that do not use them","references":["https://access.redhat.com/security/cve/CVE-2024-50264"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-11-19"},{"id":"CVE-2025-0033","cve":"CVE-2025-0033","aliases":[],"title":"AMD SEV-SNP - RMP write access during SNP initialization: There is a window during SEV-SNP initialization in which an","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV-SNP - RMP write access during SNP initialization","year":"2025","cvss_score":6,"severity":"medium","kev":false,"impact":"There is a window during SEV-SNP initialization in which an admin-privileged attacker can write to the reverse-map table itself. Corrupting the RMP at init time means the ownership map that governs every subsequent guest page assignment starts out wrong - the attacker can arrange for guest pages to be host-writable from the moment the platform comes up, and no later check catches it because the check is the RMP.","attack_vector":"Local, admin-privileged, and specifically during SNP platform initialization - so an attacker who controls the host boot sequence or the SNP init path.","remediation":"Fixed in AMD SEV firmware / AGESA and reaches you as an OEM SBIOS package - AMD hands AGESA to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before shipping BIOS. **Budget one to six months of OEM lag**, longer on older platforms and sometimes never on end-of-support SKUs. Applying it means draining the host and doing a full power cycle. Because the fix moves the platform's reported SEV-SNP TCB version, you must also pull fresh VCEK certificates from AMD's Key Distribution Service and update any attestation policy your tenants pin - otherwise guests will start failing launch validation the moment the BIOS lands. Some SEV firmware can alternatively be staged from linux-firmware (amd/amd_sev_*.sbin) and committed via the ccp driver at boot, which is faster than waiting on BIOS - check whether your platform supports firmware hot-load before assuming the OEM is the only route. Since the exposure is at SNP init, the practical control is boot integrity: measured boot on the host, and refusing to admit confidential tenants onto nodes whose boot chain you cannot attest.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0033","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2025-10-14"},{"id":"CVE-2025-20067","cve":"CVE-2025-20067","aliases":[],"title":"Intel CSME / SPS firmware (timing side channel): An observable timing discrepancy in CSME/SPS firmware allows","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel CSME / SPS firmware (timing side channel)","year":"2025","cvss_score":6,"severity":"medium","kev":false,"impact":"An observable timing discrepancy in CSME/SPS firmware allows a privileged local user to infer information from the management engine. Low direct impact; useful as a step toward key or state recovery from a component that is supposed to be opaque to the host.","attack_vector":"Privileged local access on the host.","remediation":"Fixed in Intel CSME/SPS firmware, which reaches you as an OEM BIOS or firmware package - not as a microcode or OS update. That means: wait for your server vendor to ship it, drain the node, flash, and reboot. OEM availability is the long pole and routinely lags the Intel advisory by one or more quarters on server platforms. Track it per platform SKU, because vendors ship these unevenly across their own product lines.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-20067","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01280.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-08-12"},{"id":"CVE-2025-24296","cve":"CVE-2025-24296","aliases":[],"title":"Intel E810 Ethernet controller firmware: Improper input validation in E810 firmware lets a privileged local user deny","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel E810 Ethernet controller firmware","year":"2025","cvss_score":6,"severity":"medium","kev":false,"impact":"Improper input validation in E810 firmware lets a privileged local user deny service on the adapter. On a node whose NIC carries collective traffic, denying the NIC denies the job.","attack_vector":"Privileged local access on the host.","remediation":"Fixed in the E810 NVM (adapter firmware) image. Deploy with Intel's NVM Update Utility, which needs a driver reload and a power cycle - not just a warm reboot - for the new image to take effect. Drain the node. Distinct from the ice driver updates: you need both, and they ship on different schedules. Target E810 firmware 4.6 or later.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-24296","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01257.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-08-12"},{"id":"CVE-2025-24851","cve":"CVE-2025-24851","aliases":["INTEL-SA-01171"],"title":"Intel Ethernet Controller E810 (100GbE) firmware: Uncaught exception in 100GbE E810 firmware, reachable from privileged","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Ethernet Controller E810 (100GbE) firmware","year":"2025","cvss_score":6,"severity":"medium","kev":false,"impact":"Uncaught exception in 100GbE E810 firmware, reachable from privileged host software, causing denial of service. Ships in the same advisory as the out-of-bounds write, so a single NVM update fixes both — worth listing separately so operators tracking by CVE do not think they are done after one.","attack_vector":"Privileged local software on the host.","remediation":"Same NVM update as CVE-2025-27243 — E810 firmware cvl fw 1.7.8.x or later, cold power cycle. One flash, two CVEs.","references":["https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01171.html","https://nvd.nist.gov/vuln/detail/CVE-2025-24851"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-02-10"},{"id":"CVE-2025-27243","cve":"CVE-2025-27243","aliases":["INTEL-SA-01171"],"title":"Intel Ethernet Controller E810 firmware: Out-of-bounds write inside E810 firmware, reachable from a privileged Ring-0","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Ethernet Controller E810 firmware","year":"2025","cvss_score":6,"severity":"medium","kev":false,"impact":"Out-of-bounds write inside E810 firmware, reachable from a privileged Ring-0 software adversary on the host, causing denial of service. The important framing for an operator is that this is a **write** primitive into NIC firmware from the host: on a bare-metal rental where the tenant has kernel privilege, the boundary between 'a tenant had root on the node' and 'the NIC's firmware state was modified' is exactly what this class of bug erodes. The published impact is DoS, but the primitive is the concern.","attack_vector":"Privileged local software on the host (Ring 0 / bare-metal OS). Any tenant with root on a rented bare-metal node qualifies.","remediation":"Flash E810 firmware to cvl fw 1.7.8.x or later; cold power cycle. Beyond the patch: if you rent bare metal, reflash NIC firmware from a known-good image at tenant handoff and verify the version afterwards, because a patched-but-unverified NIC is not a clean NIC.","references":["https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01171.html","https://nvd.nist.gov/vuln/detail/CVE-2025-27243"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2026-02-10"},{"id":"CVE-2025-37149","cve":"CVE-2025-37149","aliases":["HPESBHF04952"],"title":"HPE ProLiant RL300 Gen11 (UEFI firmware, out-of-bounds read): Out-of-bounds reads in the UEFI firmware of the ProLiant","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HPE ProLiant RL300 Gen11 (UEFI firmware, out-of-bounds read)","year":"2025","cvss_score":6,"severity":"medium","kev":false,"impact":"Out-of-bounds reads in the UEFI firmware of the ProLiant RL300 Gen11 - HPE's Arm-based (Ampere) ProLiant. The CVSS vector marks scope as changed with high confidentiality impact, meaning the leak crosses a trust boundary out of the firmware context. Firmware-level memory disclosure is how an attacker recovers the addresses and secrets that make a subsequent firmware write reliable, so treat it as an enabler for a persistence attack rather than as a standalone data-loss event. Relevant to operators running mixed-architecture racks where Arm nodes handle inference or control-plane duty alongside x86 GPU boxes.","attack_vector":"Local to the host with high privileges - a root/administrator account on the operating system. Not remotely reachable and not exposed on the management VLAN.","remediation":"UEFI firmware update on the affected RL300 Gen11 nodes. As with any system firmware, it applies on the next reboot, so it costs a maintenance window per node rather than a live out-of-band flash. Small affected footprint means the campaign should be quick to scope - identify RL300 Gen11 nodes specifically, since the rest of the ProLiant Gen11 line is not in scope. No config-only mitigation.","references":["https://support.hpe.com/hpesc/public/docDisplay?docId=hpesbhf04952en_us&docLocale=en_US","https://nvd.nist.gov/vuln/detail/CVE-2025-37149"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2025-10-14"},{"cwe":["CWE-833"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-38436","cve":"CVE-2025-38436","aliases":[],"title":"Linux kernel (drivers/gpu/drm/scheduler): When one tenant's scheduler entity is killed, its scheduled fences are not","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/scheduler)","year":"2025","cvss_score":6,"severity":"medium","kev":false,"impact":"When one tenant's scheduler entity is killed, its scheduled fences are not signalled, so any other tenant whose job depends on those fences waits forever. One workload dying - or one attacker deliberately killing processes in a loop - permanently hangs unrelated jobs belonging to different tenants on the same GPU, which is a cross-tenant availability break, not just a local crash.","attack_vector":"An unprivileged process in a container with /dev/dri/renderD* creates cross-process fence dependencies (shared dma-buf / sync_file / syncobj, which is how compositors, media pipelines and multi-process GPU workloads normally work) and then exits or is killed. Lives in the shared drm/scheduler layer, so every scheduler-based driver on the node is affected.","remediation":"Update to a kernel containing the fix commits below. Interim: avoid sharing fences/syncobjs across tenant boundaries, and be prepared to reset the GPU (node drain plus device reset) to clear hung dependent jobs.","references":["https://git.kernel.org/stable/c/8342127a8a65b0673863b106ce32b79c91ae3270","https://git.kernel.org/stable/c/c5734f9bab6f0d40577ad0633af4090a5fda2407","https://nvd.nist.gov/vuln/detail/CVE-2025-38436"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-476","CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-40086","cve":"CVE-2025-40086","aliases":[],"title":"Linux kernel (drivers/gpu/drm/xe): A batched array of VM_BIND operations could evict other buffer objects belonging to","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/xe)","year":"2025","cvss_score":6,"severity":"medium","kev":false,"impact":"A batched array of VM_BIND operations could evict other buffer objects belonging to the same VM while that VM's bind pipeline is still walking them, leaving the pipeline dereferencing objects with no backing resource. A tenant can crash the kernel from an ordinary GPU address-space bind, and the crash lands inside the shared xe bind machinery on the node.","attack_vector":"A tenant container holding /dev/dri/renderD* on an Intel xe GPU submits an array of VM_BIND operations sized to force eviction inside its own VM - no privilege beyond the render node, no display path, no host root. The fallout is a kernel oops on a node other tenants are sharing.","remediation":"Boot a kernel carrying the xe_vm eviction-policy fix below. Interim: cap per-tenant VRAM so binds do not push the VM into eviction, and keep panic_on_oops off so a single tenant's oops does not take the whole node with it.","references":["https://git.kernel.org/stable/c/5aa0ab0ba7d94549cfe17d6ef7a4f33ba1de8384","https://git.kernel.org/stable/c/7ac74613e5f2ef3450f44fd2127198662c2563a9","https://nvd.nist.gov/vuln/detail/CVE-2025-40086"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-833","CWE-667"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-40329","cve":"CVE-2025-40329","aliases":[],"title":"Linux kernel (drivers/gpu/drm/scheduler): Tearing down a GPU scheduler entity takes locks from a fence-signalling","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/scheduler)","year":"2025","cvss_score":6,"severity":"medium","kev":false,"impact":"Tearing down a GPU scheduler entity takes locks from a fence-signalling callback that runs in interrupt context, so a tenant whose jobs carry cross-fence dependencies can wedge the CPU that is signalling and the shared GPU scheduler behind it. This is a whole-node outage vector: the scheduler is shared by every tenant on the device, and a deadlock there stalls all of their queues, not just the attacker's.","attack_vector":"An unprivileged process in a container holding /dev/dri/renderD* reaches this by submitting jobs with dependencies on other fences and then dying or being killed, which is exactly what a crashing or OOM-killed workload does. The code is in the shared drm/scheduler layer, so amdgpu, xe, nouveau, panfrost and every other scheduler user is affected, not one vendor.","remediation":"Update to a kernel with the fix commits below, which moves the dependency re-arming out of the fence callback into a work item. No interim control short of removing GPU access; the trigger is normal process teardown.","references":["https://git.kernel.org/stable/c/70150b9443dddf02157d821c68abf438f55a2e8e","https://git.kernel.org/stable/c/0d63031ee4a57be0252cb9a4e09ae921c75cece9","https://nvd.nist.gov/vuln/detail/CVE-2025-40329"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-1386","cve":"CVE-2026-1386","aliases":[],"title":"Firecracker: Symlink following in the jailer lets a local host user with write access to pre-created jailer","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Firecracker","year":"2026","cvss_score":6,"severity":"medium","kev":false,"impact":"Symlink following in the jailer lets a local host user with write access to pre-created jailer directories overwrite arbitrary files","attack_vector":"A local host user","remediation":"Upgrade Firecracker; tighten jailer directory ownership","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-1386"],"status":"curated","published":"2026-01-23"},{"id":"CVE-2026-15792","cve":"CVE-2026-15792","aliases":[],"title":"BuildKit: Malicious client or frontend panics the BuildKit daemon","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"BuildKit","year":"2026","cvss_score":6,"severity":"medium","kev":false,"impact":"Malicious client or frontend panics the BuildKit daemon","attack_vector":"Anyone with build API access","remediation":"Upgrade BuildKit","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-15792"],"status":"curated","published":"2026-07-21"},{"id":"CVE-2026-41184","cve":"CVE-2026-41184","aliases":[],"title":"Calico: install-cni logs the rendered CNI config including the substituted service-account token","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Calico","year":"2026","cvss_score":6,"severity":"medium","kev":false,"impact":"install-cni logs the rendered CNI config including the substituted service-account token","attack_vector":"Anyone with pod-log read access","remediation":"Rolling Calico upgrade; rotate the CNI service-account token; scrub logs","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-41184"],"status":"curated","published":"2026-05-28"},{"id":"CVE-2026-41185","cve":"CVE-2026-41185","aliases":[],"title":"Calico: Azure IPAM helper logs the mutated CNI config including credentials","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Calico","year":"2026","cvss_score":6,"severity":"medium","kev":false,"impact":"Azure IPAM helper logs the mutated CNI config including credentials","attack_vector":"Anyone with pod-log read access","remediation":"Rolling Calico upgrade; rotate leaked credentials","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-41185"],"status":"curated","published":"2026-05-28"},{"id":"CVE-2026-41186","cve":"CVE-2026-41186","aliases":[],"title":"Calico: kube-controllers and Goldmane bind an unauthenticated pprof listener to 0.0.0.0","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Calico","year":"2026","cvss_score":6,"severity":"medium","kev":false,"impact":"kube-controllers and Goldmane bind an unauthenticated pprof listener to 0.0.0.0","attack_vector":"Any pod on the cluster network","remediation":"Rolling Calico upgrade; keep the debug server disabled (it is off by default)","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-41186"],"status":"curated","published":"2026-07-30"},{"id":"CVE-2017-15361","cve":"CVE-2017-15361","aliases":["ROCA","Return of Coppersmith's Attack"],"title":"Infineon TPM firmware (RSA key generation): RSA keys generated inside affected Infineon TPMs are factorable","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Infineon TPM firmware (RSA key generation)","year":"2017","cvss_score":5.9,"severity":"medium","kev":false,"impact":"RSA keys generated inside affected Infineon TPMs are factorable from the public key alone - no access to the machine required. Every key the TPM ever produced is retroactively compromised: attestation identity keys, sealed storage keys, machine certificates, and any SSH or code-signing key an operator generated in the TPM believing it was hardware-protected. For a fleet, that means the attestation evidence you have been collecting is forgeable by anyone who saw a public key.","attack_vector":"No access to the hardware at all. The attacker needs only a public key that the TPM generated - which by definition has been published to whatever service consumed it.","remediation":"Two-part and expensive. First a TPM firmware update from the platform OEM, usually shipped inside a BIOS package, so it is a per-node flash plus reboot. Then - and this is the part that gets skipped - every key generated by the vulnerable TPM must be regenerated and re-enrolled, and the old ones revoked. Sealed data must be unsealed before the update or it becomes unrecoverable. Inventory which nodes carry Infineon TPMs before planning; on old hardware the OEM may never have shipped the fix, in which case move the trust anchor off the TPM.","references":["https://nvd.nist.gov/vuln/detail/CVE-2017-15361","https://www.infineon.com/cms/en/product/promopages/tpm-update/"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2017-10-16"},{"id":"CVE-2018-12130","cve":"CVE-2018-12130","aliases":[],"title":"Intel CPU (MDS / ZombieLoad): Microarchitectural Fill Buffer Data Sampling","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel CPU (MDS / ZombieLoad)","year":"2018","cvss_score":5.9,"severity":"medium","kev":false,"impact":"Microarchitectural Fill Buffer Data Sampling - cross-domain leak from fill buffers, including across SMT siblings","attack_vector":"Any tenant process in a container; tenant VM guest","remediation":"Microcode + kernel VERW mitigation + reboot; SMT disable for full protection, with the corresponding throughput loss","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-12130"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"published":"2019-05-30"},{"id":"CVE-2019-16863","cve":"CVE-2019-16863","aliases":["TPM-FAIL"],"title":"STMicroelectronics ST33 TPM (ECDSA timing): Discrete TPM leaks ECDSA nonce data through timing, allowing private key","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"STMicroelectronics ST33 TPM (ECDSA timing)","year":"2019","cvss_score":5.9,"severity":"medium","kev":false,"impact":"Discrete TPM leaks ECDSA nonce data through timing, allowing private key recovery from observed signatures. The discrete-chip half of TPM-FAIL - which matters because the usual answer to the fTPM flaw is 'use a real TPM', and this shows the real TPM had the same class of problem.","attack_vector":"Local attacker able to request signatures from the TPM, with accurate timing measurement.","remediation":"TPM firmware update from ST, distributed through the platform OEM's BIOS package - per-node flash plus reboot, and vendor availability was patchy. Rotate any long-lived key the TPM produced. Where no update exists, treat that TPM's keys as software-grade rather than hardware-protected.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-16863","https://tpm.fail/"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2019-11-14"},{"id":"CVE-2020-16843","cve":"CVE-2020-16843","aliases":[],"title":"Firecracker: Network stack freezes under heavy ingress","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Firecracker","year":"2020","cvss_score":5.9,"severity":"medium","kev":false,"impact":"Network stack freezes under heavy ingress; microVM DoS","attack_vector":"Unauthenticated network","remediation":"Upgrade Firecracker","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-16843"],"status":"curated","published":"2020-08-04"},{"id":"CVE-2020-26569","cve":"CVE-2020-26569","aliases":[],"title":"Arista EOS (EVPN VXLAN MAC/IP binding): Malformed packets create incorrect MAC-to-IP bindings in an EVPN VXLAN fabric","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (EVPN VXLAN MAC/IP binding)","year":"2020","cvss_score":5.9,"severity":"medium","kev":false,"impact":"Malformed packets create incorrect MAC-to-IP bindings in an EVPN VXLAN fabric, and packets get forwarded across VLAN boundaries as a result. An attacker who can source crafted frames from inside one tenant's overlay can poison the fabric's bindings and cause traffic to cross into the wrong VLAN. The advisory notes traffic is discarded on the receiving VLAN, which limits it to a leak-and-drop rather than a clean interception — but it is still a control-plane-driven breach of the segmentation model.","attack_vector":"A host inside an EVPN VXLAN tenant network able to emit specific malformed packets.","remediation":"EOS upgrade plus reload across the VTEP layer. No live workaround. If you run EVPN multi-tenancy, also audit the MAC/IP binding table for entries that do not correspond to a real workload.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-26569"],"status":"curated","tags":["tenant-isolation"],"published":"2020-12-28"},{"id":"CVE-2020-8553","cve":"CVE-2020-8553","aliases":[],"title":"ingress-nginx: A tenant can overwrite another ingress's basic-auth password file","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2020","cvss_score":5.9,"severity":"medium","kev":false,"impact":"A tenant can overwrite another ingress's basic-auth password file","attack_vector":"Cluster user able to create namespaces and Ingress objects","remediation":"Rolling controller upgrade, no GPU drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-8553"],"status":"curated","published":"2020-07-29"},{"id":"CVE-2021-20199","cve":"CVE-2021-20199","aliases":[],"title":"Podman: Rootless containers see all traffic as coming from 127.0.0.1, defeating localhost-trust checks","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Podman","year":"2021","cvss_score":5.9,"severity":"medium","kev":false,"impact":"Rootless containers see all traffic as coming from 127.0.0.1, defeating localhost-trust checks","attack_vector":"Unauthenticated network","remediation":"Upgrade Podman; never trust source 127.0.0.1 in containerised apps","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-20199"],"status":"curated","published":"2021-02-02"},{"id":"CVE-2022-24769","cve":"CVE-2022-24769","aliases":[],"title":"Docker / moby: Containers started with non-empty inheritable capabilities","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2022","cvss_score":5.9,"severity":"medium","kev":false,"impact":"Containers started with non-empty inheritable capabilities; unexpected privilege retention on setuid binaries","attack_vector":"Any tenant workload","remediation":"Upgrade Docker Engine / moby; restart containers","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-24769"],"status":"curated","published":"2022-03-24"},{"id":"CVE-2022-29162","cve":"CVE-2022-29162","aliases":[],"title":"runc: `runc exec --cap` created processes with non-empty inheritable capabilities","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"runc","year":"2022","cvss_score":5.9,"severity":"medium","kev":false,"impact":"`runc exec --cap` created processes with non-empty inheritable capabilities; unexpected privilege retention","attack_vector":"Any tenant workload","remediation":"Replace runc binary; restart affected containers","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-29162"],"status":"curated","published":"2022-05-17"},{"cwe":["CWE-833","CWE-667"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-49434","cve":"CVE-2022-49434","aliases":[],"title":"Linux kernel (drivers/pci): Pci_dev_lock() and the sysfs SR-IOV path took the device lock and the config-space access","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/pci)","year":"2022","cvss_score":5.9,"severity":"medium","kev":false,"impact":"Pci_dev_lock() and the sysfs SR-IOV path took the device lock and the config-space access lock in opposite orders, so a device reset racing an SR-IOV VF enable/disable deadlocks permanently. Both threads hang in D state holding the PCI device lock, which wedges every subsequent operation on that device - VF provisioning, reset, unbind - and the only way out is a reboot of the node all the tenants are sharing.","attack_vector":"Both halves are reachable in a normal passthrough cluster. A tenant holding /dev/vfio/* can call VFIO_DEVICE_RESET, which goes through pci_dev_lock() - that is side A and needs no host privilege. Side B is the operator's own VF lifecycle: writing to sysfs sriov_numvfs to hand out or reclaim VFs, which reaches pci_cfg_access_lock() through vfio_pci_core_sriov_configure() and pci_disable_sriov(). A tenant that resets its device in a loop while the control plane is reprovisioning VFs on the same PF can win the interleaving. Requires SR-IOV to be in use on the node.","remediation":"Update to a kernel carrying the fix (no fixed_in published by the CNA; the stable commits below are widely backported). Interim: serialise VF provisioning against tenant activity - stop the tenant workload and revoke the VFIO device node before changing sriov_numvfs on a PF, rather than reconfiguring VFs while tenants are live on the same PF.","references":["https://git.kernel.org/stable/c/c3c6dc1853b8bf3c718f96fd8480a6eb09ba4831","https://git.kernel.org/stable/c/aed6d4d519210c28817948f34c53b6e058e0456c","https://nvd.nist.gov/vuln/detail/CVE-2022-49434"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2023-20902","cve":"CVE-2023-20902","aliases":[],"title":"Harbor: Timing condition allows creating and stopping jobs and retrieving job info","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Harbor","year":"2023","cvss_score":5.9,"severity":"medium","kev":false,"impact":"Timing condition allows creating and stopping jobs and retrieving job info","attack_vector":"Unauthenticated network access to Harbor","remediation":"Upgrade Harbor","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-20902"],"status":"curated","published":"2023-11-09"},{"id":"CVE-2023-48795","cve":"CVE-2023-48795","aliases":[],"title":"OpenSSH (transport): Terrapin: prefix-truncation attack on the SSH Binary Packet Protocol","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"OpenSSH (transport)","year":"2023","cvss_score":5.9,"severity":"medium","kev":false,"impact":"Terrapin: prefix-truncation attack on the SSH Binary Packet Protocol - silently downgrades channel integrity","attack_vector":"Unauthenticated network (MITM position)","remediation":"Package update on both ends + sshd restart; no reboot","references":["https://access.redhat.com/security/cve/CVE-2023-48795"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2023-12-18"},{"cwe":["CWE-416","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-54235","cve":"CVE-2023-54235","aliases":[],"title":"Linux kernel (drivers/pci): The DOE state machine signals the caller's completion before destroying the work_struct","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/pci)","year":"2023","cvss_score":5.9,"severity":"medium","kev":false,"impact":"The DOE state machine signals the caller's completion before destroying the work_struct that lives on the caller's stack, so the workqueue can still be touching a stack frame the caller has already left. DOE is the mailbox that carries CMA/SPDM device attestation and IDE link-encryption negotiation, which makes this a race in exactly the machinery a confidential-GPU deployment relies on to decide whether a device is trustworthy.","attack_vector":"Host-side only, on nodes where a device exposes a DOE mailbox and something drives it - CXL, device attestation (CMA/SPDM), or IDE key exchange. There is no tenant-facing entry point: the race is between the DOE workqueue and the kernel thread that submitted the task, and its timing is influenced by how slowly the device answers, so a slow or deliberately laggy device shifts the window. Nodes with no DOE-capable device, or with attestation/IDE unused, never execute the path.","remediation":"Update to 6.1.53 / 6.3 or later, where the work struct is destroyed before the completion is signalled. Interim: on affected kernels avoid driving DOE in production - leave CMA/SPDM attestation and IDE negotiation disabled until patched rather than running them against untrusted devices.","references":["https://git.kernel.org/stable/c/d96799ee3b78962c80e4b6653734f488f999ca09","https://git.kernel.org/stable/c/c4f9c0a3a6df143f2e1092823b7fa9e07d6ab57f","https://nvd.nist.gov/vuln/detail/CVE-2023-54235"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2023-5502","cve":"CVE-2023-5502","aliases":[],"title":"Arista EOS (802.1X on access/trunk ports): With 802.1X configured on access or trunk ports and routing enabled on the","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (802.1X on access/trunk ports)","year":"2023","cvss_score":5.9,"severity":"medium","kev":false,"impact":"With 802.1X configured on access or trunk ports and routing enabled on the access VLAN, a malicious supplicant can skip 802.1X authentication entirely. Port-based admission control is what stops an unauthorized machine being plugged into a rack and joining the fabric — this makes it optional. Companion issue CVE-2024-6858 does the same thing in multi-auth mode via a device in the fallback VLAN.","attack_vector":"A device physically connected to a switch port that has 802.1X configured. Colocation, shared cages, and contractor rack-and-stack are the realistic scenarios.","remediation":"EOS upgrade plus reload. Do not rely on 802.1X alone as the tenant admission boundary; combine it with per-port VLAN pinning and MAC allowlisting (live config) so a bypassed supplicant still lands nowhere useful.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-5502","https://nvd.nist.gov/vuln/detail/CVE-2024-6858"],"status":"curated","tags":["tenant-isolation"],"published":"2026-06-04"},{"id":"CVE-2024-29018","cve":"CVE-2024-29018","aliases":[],"title":"Docker / moby: DNS requests from an internal network can be forwarded to external resolvers, leaking data out","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2024","cvss_score":5.9,"severity":"medium","kev":false,"impact":"DNS requests from an internal network can be forwarded to external resolvers, leaking data out of an isolated network","attack_vector":"Any tenant workload on an \"internal\" Docker network","remediation":"Upgrade moby","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-29018"],"status":"curated","published":"2024-03-20"},{"id":"CVE-2024-9042","cve":"CVE-2024-9042","aliases":[],"title":"Kubernetes (kubelet): Command injection on Windows nodes via the nodes/*/logs/query API","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kubelet)","year":"2024","cvss_score":5.9,"severity":"medium","kev":false,"impact":"Command injection on Windows nodes via the nodes/*/logs/query API","attack_vector":"Cluster user with node log-query rights","remediation":"Rolling kubelet upgrade; Windows node drain; restrict nodes/log RBAC","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated","fleet":{"ubiquity":"Niche in GPU clouds - Windows GPU nodes are rare, but the `nodes/*/logs/query` path is a reminder that kubelet log endpoints reach the host shell","remediation_pain":"`daemon-restart` - kubelet upgrade to v1.32.1 / v1.31.5 / v1.30.9 / v1.29.13, node-by-node","pain_class":"daemon-restart","why_fleet_wide":"Anyone with `nodes/*/logs` read rights injects into PowerShell and executes as SYSTEM on the node; low ubiquity in GPU fleets keeps this off the emergency list"},"published":"2025-03-13"},{"id":"CVE-2025-23333","cve":"CVE-2025-23333","aliases":[],"title":"NVIDIA Triton Inference Server: Manipulating the Python backend's shared memory region produces an out-of-bounds read","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2025","cvss_score":5.9,"severity":"medium","kev":false,"impact":"Manipulating the Python backend's shared memory region produces an out-of-bounds read and leaks data. Where one Triton instance serves several models or tenants, that shared memory holds other requests' tensors. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5687. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23333","https://github.com/NVIDIA/product-security/tree/main/2025/5687"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-125"],"tags":["tenant-isolation"],"published":"2025-08-06"},{"id":"CVE-2025-23334","cve":"CVE-2025-23334","aliases":[],"title":"NVIDIA Triton Inference Server: A crafted request causes an out-of-bounds read in the Python backend, disclosing memory","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2025","cvss_score":5.9,"severity":"medium","kev":false,"impact":"A crafted request causes an out-of-bounds read in the Python backend, disclosing memory that can include co-resident request data. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5687. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23334","https://github.com/NVIDIA/product-security/tree/main/2025/5687"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-125"],"tags":["tenant-isolation"],"published":"2025-08-06"},{"id":"CVE-2025-26466","cve":"CVE-2025-26466","aliases":[],"title":"OpenSSH (sshd): Pre-auth memory/CPU amplification - denial of service against sshd","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"OpenSSH (sshd)","year":"2025","cvss_score":5.9,"severity":"medium","kev":false,"impact":"Pre-auth memory/CPU amplification - denial of service against sshd","attack_vector":"Unauthenticated network","remediation":"Package update + sshd restart; rate-limit with PerSourcePenalties","references":["https://access.redhat.com/security/cve/CVE-2025-26466"],"status":"curated","published":"2025-02-28"},{"id":"CVE-2025-29948","cve":"CVE-2025-29948","aliases":[],"title":"AMD SEV firmware - RMP protection bypass: An access-control failure in SEV firmware lets a malicious hypervisor bypass","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV firmware - RMP protection bypass","year":"2025","cvss_score":5.9,"severity":"medium","kev":false,"impact":"An access-control failure in SEV firmware lets a malicious hypervisor bypass reverse-map table protections and break SEV-SNP guest memory integrity. The RMP is the single structure standing between a hostile host and a confidential guest's pages; a firmware-level bypass of it means the isolation you are selling is not enforced.","attack_vector":"Malicious hypervisor - host-privileged.","remediation":"Fixed in AMD SEV firmware / AGESA and reaches you as an OEM SBIOS package - AMD hands AGESA to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before shipping BIOS. **Budget one to six months of OEM lag**, longer on older platforms and sometimes never on end-of-support SKUs. Applying it means draining the host and doing a full power cycle. Because the fix moves the platform's reported SEV-SNP TCB version, you must also pull fresh VCEK certificates from AMD's Key Distribution Service and update any attestation policy your tenants pin - otherwise guests will start failing launch validation the moment the BIOS lands. Some SEV firmware can alternatively be staged from linux-firmware (amd/amd_sev_*.sbin) and committed via the ccp driver at boot, which is faster than waiting on BIOS - check whether your platform supports firmware hot-load before assuming the OEM is the only route.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-29948","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2026-02-10"},{"id":"CVE-2025-29952","cve":"CVE-2025-29952","aliases":[],"title":"AMD SEV firmware - improper initialization corrupting RMP-covered memory: An initialization defect in SEV firmware lets","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV firmware - improper initialization corrupting RMP-covered memory","year":"2025","cvss_score":5.9,"severity":"medium","kev":false,"impact":"An initialization defect in SEV firmware lets an admin-privileged attacker corrupt memory covered by the RMP, costing confidential guest integrity. Same family as the other RMP issues in this batch and shipped in the same AMD advisory wave - the recurring theme is that SNP's protections are only as good as the firmware that sets them up.","attack_vector":"Local, admin-privileged.","remediation":"Fixed in AMD reference firmware (AGESA / PSP / SEV firmware) and delivered only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo and the ODMs each rebuild and requalify AMD's AGESA drop before shipping. **Expect one to six months of OEM lag**, and on end-of-support platforms expect nothing. Applying it is a drain plus full power cycle, not a driver reload. Verify by reading back the PSP/SMU firmware version afterwards rather than trusting the BIOS version string. This sits inside the SEV-SNP trust boundary, so the update moves the platform's reported TCB version: refresh VCEK certificates from AMD's KDS and update any attestation policy your tenants pin, or confidential guest launches will start failing right after the BIOS lands.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-29952","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2026-02-10"},{"id":"CVE-2025-33242","cve":"CVE-2025-33242","aliases":[],"title":"HGX / DGX B300: Authentication bypass in the hardware management interface","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"HGX / DGX B300","year":"2025","cvss_score":5.9,"severity":"medium","kev":false,"impact":"Authentication bypass in the hardware management interface","attack_vector":"Network-adjacent attacker on the mgmt path","remediation":"Flash HGX/DGX management firmware out-of-band","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33242","https://github.com/NVIDIA/product-security/tree/main/2026/5768"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:H/UI:N/S:U/C:N/I:H/A:H","cwe":["CWE-1234"],"published":"2026-03-24"},{"cwe":["CWE-415"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-38069","cve":"CVE-2025-38069","aliases":[],"title":"Linux kernel (drivers/pci/endpoint/functions): When BAR allocation fails, the endpoint test function frees the backing","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/pci/endpoint/functions)","year":"2025","cvss_score":5.9,"severity":"medium","kev":false,"impact":"When BAR allocation fails, the endpoint test function frees the backing memory but leaves the stale pointer in place, so the next allocation round frees the same buffer again. A double free in the host kernel oopses the node, and double frees are a classic route to heap control rather than a mere crash.","attack_vector":"Endpoint mode required, and the trigger comes from the OTHER side of the link: the connected host rebooting deasserts PERST#, which restarts BAR allocation on the endpoint. If inbound windows are exhausted so pci_epc_set_bar() fails, each host reboot re-runs the failing path and hits the double free. That makes it remotely drivable, pre-authentication, by whoever controls the host the endpoint is plugged into - a real consideration for a card handed to a tenant chassis, and inert on a conventional GPU server.","remediation":"Update to a kernel carrying the fix (no fixed_in published; stable commits below). Interim: configure endpoint functions with no more BARs than the controller has free inbound windows, so the allocation failure that starts this sequence never occurs.","references":["https://git.kernel.org/stable/c/fe2329eff5bee461ebcafadb6ca1df0cbf5945fd","https://git.kernel.org/stable/c/8b83893d1f6c6061a7d58169ecdf9d5ee9f306ee","https://nvd.nist.gov/vuln/detail/CVE-2025-38069"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-54510","cve":"CVE-2025-54510","aliases":[],"title":"AMD Secure Processor firmware - MMIO routing lock (Zen 5): A missing lock check in ASP firmware on some Zen 5 parts","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor firmware - MMIO routing lock (Zen 5)","year":"2025","cvss_score":5.9,"severity":"medium","kev":false,"impact":"A missing lock check in ASP firmware on some Zen 5 parts lets a locally authenticated administrator alter MMIO routing after the point where it should have been frozen. Redirecting MMIO means pointing a device's window somewhere it should not go - which is how a host administrator reaches into a confidential guest's memory despite SEV-SNP.","attack_vector":"Local, administrative privilege on the host. Affects Zen 5 (Turin / EPYC 9005) generation parts.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. Because this sits inside the SEV-SNP trust boundary, the update also moves the platform's reported TCB version: after patching you must refresh VCEK certificates from AMD's KDS and update whatever attestation policy your tenants (or your own confidential-VM control plane) pin against, or every guest launch will start failing validation.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-54510","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2026-04-16"},{"id":"CVE-2025-61971","cve":"CVE-2025-61971","aliases":[],"title":"AMD NBIO register lock bits - MMIO routing configuration: The sibling of the SMN issue: unprotected NBIO lock bits let","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD NBIO register lock bits - MMIO routing configuration","year":"2025","cvss_score":5.9,"severity":"medium","kev":false,"impact":"The sibling of the SMN issue: unprotected NBIO lock bits let a local administrator rewrite MMIO routing configuration and break SEV-SNP guest integrity. The host operator can point address windows at confidential guest memory that the RMP was supposed to fence off.","attack_vector":"Local, host administrator privilege.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. Because this sits inside the SEV-SNP trust boundary, the update also moves the platform's reported TCB version: after patching you must refresh VCEK certificates from AMD's KDS and update whatever attestation policy your tenants (or your own confidential-VM control plane) pin against, or every guest launch will start failing validation. Same OEM BIOS package as the SMN lock-bit issue; do not patch one and leave the other.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-61971","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2026-05-13"},{"id":"CVE-2026-24231","cve":"CVE-2026-24231","aliases":[],"title":"NemoClaw: SSRF","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NemoClaw","year":"2026","cvss_score":5.9,"severity":"medium","kev":false,"impact":"SSRF -> internal service access from a tenant-facing service","attack_vector":"Tenant supplying a crafted URL","remediation":"Upgrade the service; add egress network policy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24231","https://github.com/NVIDIA/product-security/tree/main/2026/5837"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:C/C:H/I:N/A:N","cwe":["CWE-918"],"published":"2026-04-28"},{"id":"CVE-2026-24266","cve":"CVE-2026-24266","aliases":[],"title":"Triton Inference Server: Use-after-free in request processing","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2026","cvss_score":5.9,"severity":"medium","kev":false,"impact":"Use-after-free in request processing","attack_vector":"Any inference client","remediation":"Upgrade Triton; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24266","https://github.com/NVIDIA/product-security/tree/main/2026/5848"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-416"],"published":"2026-07-01"},{"id":"CVE-2026-27482","cve":"CVE-2026-27482","aliases":[],"title":"Ray (dashboard DELETE endpoints): Browser-origin protection covers POST/PUT but not DELETE","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ray (dashboard DELETE endpoints)","year":"2026","cvss_score":5.9,"severity":"medium","kev":false,"impact":"Browser-origin protection covers POST/PUT but not DELETE; key DELETE endpoints unauthenticated","attack_vector":"Malicious page visited by an operator with dashboard access","remediation":"Upgrade past 2.53.0","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-27482"],"status":"curated","published":"2026-02-21"},{"cwe":["CWE-459","CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-53164","cve":"CVE-2026-53164","aliases":[],"title":"Linux kernel (drivers/iommu): An unaligned DMA mapping with no aligned middle section calls into the mapper with length","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu)","year":"2026","cvss_score":5.9,"severity":"medium","kev":false,"impact":"An unaligned DMA mapping with no aligned middle section calls into the mapper with length zero, which the newer page-table backend rejects; the error unwind then starts from the wrong offset and unlinks the wrong range, leaving the mapping corrupted. Parts of an IOVA range stay mapped after the mapping is supposed to be torn down, so a device retains a working DMA window into pages the kernel has already released.","attack_vector":"Driven by unaligned I/O buffers on a device forced through SWIOTLB bouncing. Upstream sees it from NVMe passthrough commands (smartctl) on drives behind forced-bounce paths - and tenants in a GPU cluster do hold /dev/nvme* and can issue passthrough commands with arbitrary buffer alignment. Conditional on forced SWIOTLB being in play (untrusted device, sub-page IOVA granule, or a bounce-forcing config); a node with no bouncing never takes this path.","remediation":"The record lists no fixed release; boot a kernel carrying the stable fix commits below. Interim controls: do not expose raw NVMe passthrough (/dev/nvme* admin/IO passthrough ioctls) to tenants, and avoid configurations that force SWIOTLB bouncing for tenant-facing devices.","references":["https://git.kernel.org/stable/c/ab61c990a87d084f5565ee70340543e3a5394697","https://git.kernel.org/stable/c/b16f8d40bac9ced838d24c9842707af9ecae92e2","https://nvd.nist.gov/vuln/detail/CVE-2026-53164"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-56742","cve":"CVE-2026-56742","aliases":[],"title":"Cilium: A namespaced HTTPRoute can mirror another tenant's HTTP traffic","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2026","cvss_score":5.9,"severity":"medium","kev":false,"impact":"A namespaced HTTPRoute can mirror another tenant's HTTP traffic; cross-tenant data exfiltration","attack_vector":"Cluster user with namespace access and Gateway API rights","remediation":"Rolling Cilium upgrade; restrict HTTPRoute mirroring","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-56742"],"status":"curated","published":"2026-07-15"},{"cwe":["CWE-362","CWE-1233"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-64472","cve":"CVE-2026-64472","aliases":[],"title":"Linux kernel (drivers/vfio/pci/mlx5): Migration and dirty-tracking state flags for an mlx5 VF were packed into shared","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/vfio/pci/mlx5)","year":"2026","cvss_score":5.9,"severity":"medium","kev":false,"impact":"Migration and dirty-tracking state flags for an mlx5 VF were packed into shared bitfields updated non-atomically from concurrent paths, including a reset on one device reaching across to another. A lost update means the driver's view of deferred_reset, dirty-logging active, or error state diverges from reality - a passthrough NIC can be treated as reset when it was not, or dirty-page logging can be believed active when it is off, which silently corrupts a migrated tenant's memory.","attack_vector":"Driven by tenant-visible actions on an mlx5 SR-IOV VF bound to mlx5-vfio-pci: a tenant issuing device resets or migration state changes through /dev/vfio/* while the operator's dirty-tracking or VF event handling runs concurrently. Conditional on Mellanox/NVIDIA ConnectX VFs with the mlx5 vfio variant driver - which is the standard configuration for SR-IOV NIC passthrough in GPU clouds.","remediation":"Update to a stable kernel carrying commits 1dd99b8f / f1db80a6. Interim: bind VFs to plain vfio-pci where live migration and dirty tracking are not needed, and serialize tenant-initiated resets against migration operations in the VMM.","references":["https://git.kernel.org/stable/c/1dd99b8f4e143592e12e5a77e7b538bc698116cb","https://git.kernel.org/stable/c/f1db80a67da928a92ba460ede1be52d8941f46be","https://nvd.nist.gov/vuln/detail/CVE-2026-64472"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-295"],"fleet":{"pain_class":"unpatchable / mitigate-only"},"id":"NCVD-2020-007-etcd-gateway-discovery-srv-secur","cve":null,"aliases":["GHSA-j86v-2vjr-fg8f"],"title":"etcd gateway (--discovery-srv secure endpoint validation): TLS VALIDATION THAT VALIDATES NOTHING: the etcd gateway's","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"etcd gateway (--discovery-srv secure endpoint validation)","year":"2020","cvss_score":5.9,"severity":"medium","kev":false,"impact":"TLS VALIDATION THAT VALIDATES NOTHING: the etcd gateway's secure endpoint validation, run when --discovery-srv is enabled, only confirms that the endpoint is reachable over TCP. It never confirms the endpoint actually accepts TLS, so a plaintext listener sitting behind an HTTPS URL passes the check and the gateway proceeds to connect. An operator reading the configuration sees TLS-validated endpoints; what they have is TCP-reachable ones. Because the gateway fronts access to the cluster datastore, that gap means control-plane traffic can be routed to an endpoint offering no transport protection — and an attacker who can influence the SRV records or stand up a listener at a discovered address inherits that traffic. Surfaced by the etcd security audit; the maintainers accepted documentation plus deprecation of the misleading validation as the resolution rather than a code fix.","attack_vector":"Network: requires the gateway to be started with --discovery-srv, plus the ability to influence DNS SRV discovery results or to occupy a discovered endpoint address with a non-TLS listener.","remediation":"Do not rely on the gateway's endpoint validation as a TLS guarantee — the maintainers' resolution was documentation and deprecation, not a code change. Configure etcd client endpoints explicitly rather than via SRV discovery, verify TLS termination at each endpoint out of band, and lock down the DNS zone serving the SRV records. Review the etcd gateway documentation for the current supported posture before depending on this path.","references":["https://github.com/etcd-io/etcd/security/advisories/GHSA-j86v-2vjr-fg8f"],"status":"curated"},{"id":"CVE-2020-15115","cve":"CVE-2020-15115","aliases":[],"title":"etcd: No password length validation permits one-character etcd passwords","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"etcd","year":"2020","cvss_score":5.8,"severity":"medium","kev":false,"impact":"No password length validation permits one-character etcd passwords","attack_vector":"Unauthenticated network brute force","remediation":"Rolling etcd upgrade; move to certificate auth","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-15115"],"status":"curated","published":"2020-08-06"},{"id":"CVE-2021-25736","cve":"CVE-2021-25736","aliases":[],"title":"Kubernetes (kube-proxy): Windows kube-proxy forwards LoadBalancer traffic to local processes on the same port","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-proxy)","year":"2021","cvss_score":5.8,"severity":"medium","kev":false,"impact":"Windows kube-proxy forwards LoadBalancer traffic to local processes on the same port","attack_vector":"Unauthenticated network via the LB","remediation":"Rolling kube-proxy upgrade on Windows nodes","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-25736"],"status":"curated","published":"2023-10-30"},{"id":"CVE-2021-28511","cve":"CVE-2021-28511","aliases":[],"title":"Arista EOS (security ACL vs NAT rule interaction): A security ACL drop rule is bypassed when a NAT ACL permit rule","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (security ACL vs NAT rule interaction)","year":"2021","cvss_score":5.8,"severity":"medium","kev":false,"impact":"A security ACL drop rule is bypassed when a NAT ACL permit rule matches the same packet. Traffic you explicitly denied is forwarded. Same class of problem as the VXLAN ACL bug — the enforcement does not match the config, so your segmentation audit passes while the boundary is open.","attack_vector":"Any source whose traffic matches both a NAT permit and a security deny. Requires NAT to be configured on the device, which is common on the cluster's egress or storage-gateway leaves.","remediation":"EOS upgrade plus reload. Interim: avoid overlapping NAT and security ACL match spaces on the same device, and verify enforcement with actual traffic tests rather than reading the config. Live config change.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-28511"],"status":"curated","tags":["tenant-isolation"],"published":"2022-08-05"},{"id":"CVE-2021-4228","cve":"CVE-2021-4228","aliases":["AMI-SA-2022001","Nozomi Labs BMC firmware research"],"title":"AMI MegaRAC SPx 12 (BMC default TLS certificate): The BMC ships with a hard-coded default TLS certificate, so HTTPS","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx 12 (BMC default TLS certificate)","year":"2021","cvss_score":5.8,"severity":"medium","kev":false,"impact":"The BMC ships with a hard-coded default TLS certificate, so HTTPS to the management interface can be transparently intercepted by anyone holding the extracted key - which is everyone, since it is identical across all devices built from that firmware. The operator loses exactly what they thought they were buying with HTTPS: BMC admin passwords, Redfish tokens and KVM traffic are readable to an attacker who can sit in the path. Because the same certificate is on every node, a single extraction compromises the whole management plane.","attack_vector":"Requires a man-in-the-middle position on the management network - a compromised jump host, a rogue device on the management VLAN, or control of a switch or DHCP server on that segment. No credentials needed. Disclosed by Nozomi Labs against a Lanner IAC-AST2500A platform; AMI's own advisory confirms the affected code is part of MegaRAC SPx, so the exposure is not limited to that one vendor's box.","remediation":"Firmware flash to SPx_12-update-3.00 or later (AMI states SPx_13 is not affected), but flashing alone does not fix a node whose certificate is already installed. The operative fix is config-only and can be done today across the fleet without a reboot: generate a unique certificate per BMC from your own internal CA and push it over Redfish, then have your management tooling actually pin or verify it rather than skipping certificate validation - which is the default in most homegrown BMC scrapers and is the reason this bug stays exploitable.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2022001.pdf","https://www.nozominetworks.com/labs/vulnerability-advisories/cve-2021-4228/","https://nvd.nist.gov/vuln/detail/CVE-2021-4228"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-10-24"},{"id":"CVE-2021-46279","cve":"CVE-2021-46279","aliases":["AMI-SA-2022001","Nozomi Labs BMC firmware research"],"title":"AMI MegaRAC SPx 12 / SPx 13 (BMC web session management): Session fixation combined with sessions that never properly","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx 12 / SPx 13 (BMC web session management)","year":"2021","cvss_score":5.8,"severity":"medium","kev":false,"impact":"Session fixation combined with sessions that never properly expire. An attacker can plant a session identifier, wait for an administrator to authenticate with it, and inherit a live admin session on the BMC - power control, console, virtual media, firmware update. The never-expiring half means stolen sessions stay valid long after the admin walked away, so a token lifted from a browser, a proxy log or a shared jump host keeps working for as long as the attacker wants it.","attack_vector":"Network access to the BMC web interface plus getting an administrator to interact with an attacker-supplied session - a link, a shared workstation, or a proxy on the management path. High complexity, no credentials required.","remediation":"Firmware flash to SPx_12-update-7.00 / SPx_13-update-5.00 or later, out-of-band per node, ODM-gated. Config-only mitigations that help immediately: restrict BMC web access to a bastion, do not let admins browse anything else from that host, and force logout rather than closing the tab - the sessions this bug leaves behind are the ones that never got explicitly ended.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2022001.pdf","https://www.nozominetworks.com/labs/vulnerability-advisories/cve-2021-46279/","https://nvd.nist.gov/vuln/detail/CVE-2021-46279"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-10-24"},{"id":"CVE-2023-2538","cve":"CVE-2023-2538","aliases":[],"title":"Tyan S5552 BMC web interface, firmware version 3.00: An unauthenticated attacker downloads the BMC's TLS private key","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Tyan S5552 BMC web interface, firmware version 3.00","year":"2023","cvss_score":5.8,"severity":"medium","kev":false,"impact":"An unauthenticated attacker downloads the BMC's TLS private key. With it they can decrypt captured management traffic and impersonate the BMC to your own tooling - meaning your provisioning system, monitoring collector and operators can be fed a controller that is not the controller, and will hand over BMC credentials to it. Because vendors frequently ship the same certificate across a production run, a key pulled from one Tyan S5552 may authenticate an impersonated BMC across every node of that model in the fleet. That turns a medium-scored file disclosure into a fleet-wide management-plane credential harvest. The private key for the TLS certificate the BMC presents is retrievable by forced browsing, with no authentication.","attack_vector":"Unauthenticated HTTP access to the BMC web interface - anything routable to the out-of-band management VLAN. No credential and no host foothold required.","remediation":"Firmware update from Tyan, and this is where an operator hits a wall: Tyan has been folded into MiTAC Computing, www.tyan.com no longer presents a valid TLS certificate for its own hostname, and mitaccomputing.com returns 403 to automated clients - so there is no reachable vendor PSIRT to obtain a fixed image from. The advisory that exists is third-party, from Nozomi Networks. Regardless of firmware state, replace the BMC's TLS certificate with one you generated and control, and rotate any BMC credentials that were transmitted to that BMC over a session an attacker could have decrypted. Certificate replacement is a config-only change and should be done on every BMC in the fleet as standard practice, not just Tyan ones.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-2538","https://www.nozominetworks.com/labs/vulnerability-advisories-cve-2023-2538/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-38490","cve":"CVE-2024-38490","aliases":[],"title":"Dell iDRAC Service Module (out-of-bounds write): Out-of-bounds write allowing a privileged local attacker to execute","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC Service Module (out-of-bounds write)","year":"2024","cvss_score":5.8,"severity":"medium","kev":false,"impact":"Out-of-bounds write allowing a privileged local attacker to execute code, typically ending in denial of service of the host.","attack_vector":"Privileged local user with interaction on the host.","remediation":"Upgrade iSM past 5.3.0.0. Package update only.","references":["https://www.dell.com/support/kbdoc/en-us/000227444/dsa-2024-086-security-update-for-dell-idrac-service-module-for-memory-corruption-vulnerabilities"],"status":"curated"},{"id":"CVE-2024-52529","cve":"CVE-2024-52529","aliases":[],"title":"Cilium: L3 port-range plus L7 allow combination results in over-permissive policy","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2024","cvss_score":5.8,"severity":"medium","kev":false,"impact":"L3 port-range plus L7 allow combination results in over-permissive policy","attack_vector":"Any tenant workload","remediation":"Rolling Cilium upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-52529"],"status":"curated","published":"2024-11-25"},{"id":"CVE-2024-53197","cve":"CVE-2024-53197","aliases":[],"title":"Linux kernel (ALSA usb-audio): Out-of-bounds access for Extigy/Mbox devices","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (ALSA usb-audio)","year":"2024","cvss_score":5.8,"severity":"medium","kev":true,"impact":"Out-of-bounds access for Extigy/Mbox devices; part of a real-world Android forensic-unlock chain [KEV]","attack_vector":"Local user with USB device access","remediation":"Livepatchable; otherwise drain + reboot. Physical-access class - matters mainly for colo cages with weak physical controls","references":["https://access.redhat.com/security/cve/CVE-2024-53197"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-12-27"},{"id":"CVE-2024-6095","cve":"CVE-2024-6095","aliases":[],"title":"LocalAI (`/models/apply`): SSRF and partial local file inclusion","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LocalAI (`/models/apply`)","year":"2024","cvss_score":5.8,"severity":"medium","kev":false,"impact":"SSRF and partial local file inclusion","attack_vector":"Unauthenticated network","remediation":"Upgrade past 2.15.0","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-6095"],"status":"curated","published":"2024-07-06"},{"id":"CVE-2024-6437","cve":"CVE-2024-6437","aliases":[],"title":"Arista EOS (PBR / BGP Flowspec / interface traffic policy): IPv4 packets carrying IP options can bypass policy-based","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (PBR / BGP Flowspec / interface traffic policy)","year":"2024","cvss_score":5.8,"severity":"medium","kev":false,"impact":"IPv4 packets carrying IP options can bypass policy-based routing, BGP Flowspec and interface traffic policy redirection. If you use PBR or Flowspec to steer a tenant's traffic through an inspection or scrubbing path, an attacker sets an IP option and goes around it. Flowspec-based DDoS mitigation on the cluster edge fails the same way.","attack_vector":"Any sender able to emit IPv4 packets with IP options toward an interface with the affected redirection configured.","remediation":"EOS upgrade plus reload. Interim: drop IPv4 packets with IP options at the edge with an ACL — a live config change, and reasonable policy in a datacenter fabric where IP options have no legitimate use.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-6437"],"status":"curated","tags":["tenant-isolation"],"published":"2025-01-10"},{"id":"CVE-2025-13281","cve":"CVE-2025-13281","aliases":[],"title":"Kubernetes (kube-controller-manager): Half-blind SSRF via the Portworx in-tree volume plugin in kube-controller-manager","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-controller-manager)","year":"2025","cvss_score":5.8,"severity":"medium","kev":false,"impact":"Half-blind SSRF via the Portworx in-tree volume plugin in kube-controller-manager","attack_vector":"Cluster user able to create a Portworx volume","remediation":"Rolling control-plane upgrade; remove the in-tree Portworx plugin","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated","published":"2025-12-14"},{"id":"CVE-2025-33043","cve":"CVE-2025-33043","aliases":["AMI-SA-2025005","Binarly rediscovery"],"title":"AMI AptioV UEFI BIOS: Improper input validation in the BIOS with an integrity impact and a changed scope","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI AptioV UEFI BIOS","year":"2025","cvss_score":5.8,"severity":"medium","kev":false,"impact":"Improper input validation in the BIOS with an integrity impact and a changed scope. The score understates why this one matters: AMI states in the advisory that the issue was identified, fixed and disclosed under NDA back in 2018, and that Binarly rediscovered it still unmitigated in multiple systems in the field seven years later. For a fleet operator that is a direct statement about the supply chain - a fix existing at AMI is not the same as a fix reaching your hardware, and the gap can be measured in years. Assume the same is true of every other AptioV CVE in this cluster on any SKU you have not explicitly verified.","attack_vector":"Local access with high privileges and user interaction, at high attack complexity. Requires an attacker who already has administrative control of the host and can induce the right operation - so it is a persistence and privilege-depth bug rather than an entry point.","remediation":"BIOS update to AptioV_5.011 or later - firmware flash plus reboot per node. The specific action this CVE demands is different from the others: do not trust version numbers, verify. Pull the actual firmware image off a representative node per SKU and confirm the fix is present, because the whole point of this advisory is that vendors shipped systems that never got the 2018 fix. Prioritise older and white-box SKUs, and any hardware acquired second-hand or through a broker.","references":["https://go.ami.com/hubfs/Security%20Advisories/2025/AMI-SA-2025005.pdf","https://nvd.nist.gov/vuln/detail/CVE-2025-33043"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2025-05-29"},{"cwe":["CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-40219","cve":"CVE-2025-40219","aliases":[],"title":"Linux kernel (drivers/pci): Enabling or disabling SR-IOV virtual functions was not serialised against PCI hotplug, so","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/pci)","year":"2025","cvss_score":5.8,"severity":"medium","kev":false,"impact":"Enabling or disabling SR-IOV virtual functions was not serialised against PCI hotplug, so VF add/remove can run concurrently with the same devices being torn down by a hotplug event. Two paths mutate the same device list at once - the outcome is corrupted PCI device state and a crash on the node whose VFs are being handed out or reclaimed. The earlier attempt to fix this deadlocked on PF removal and had to be reverted, so systems have been carrying the race for a while.","attack_vector":"The VF half is the operator's own tenant-provisioning path: writing sysfs sriov_numvfs to create or destroy the VFs assigned to tenant VMs. The hotplug half is device- or firmware-driven and does not need a human - a link event, a surprise removal, or a DPC/AER-driven re-enumeration on the same hierarchy supplies it. Requires SR-IOV in use, which is the norm on any node handing out VFs of a ConnectX-class NIC or a virtualised GPU. Not directly tenant-triggerable, but a tenant that can induce link or error events on its own passthrough device improves the odds.","remediation":"Update to a kernel carrying the fix (no fixed_in published; stable commits below). Interim: serialise VF provisioning in your control plane so sriov_numvfs writes never overlap hotplug activity on the same root port, and drain the node before reprovisioning VFs.","references":["https://git.kernel.org/stable/c/3cddde484471c602bea04e6f384819d336a1ff84","https://git.kernel.org/stable/c/d7673ac466eca37ec3e6b7cc9ccdb06de3304e9b","https://nvd.nist.gov/vuln/detail/CVE-2025-40219"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-24201","cve":"CVE-2026-24201","aliases":[],"title":"vGPU Manager: Host impact via GPU command-buffer overflow","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU Manager","year":"2026","cvss_score":5.8,"severity":"medium","kev":false,"impact":"Host impact via GPU command-buffer overflow","attack_vector":"Tenant VM guest","remediation":"vGPU Manager upgrade; evacuate guest VMs","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24201","https://github.com/NVIDIA/product-security/tree/main/2026/5821"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:L/I:L/A:H","cwe":["CWE-787"],"published":"2026-05-26"},{"id":"CVE-2026-44210","cve":"CVE-2026-44210","aliases":[],"title":"Kata Containers: Default configuration allows pod creators more than intended","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kata Containers","year":"2026","cvss_score":5.8,"severity":"medium","kev":false,"impact":"Default configuration allows pod creators more than intended","attack_vector":"Cluster user with namespace access","remediation":"Upgrade Kata to 3.31.0+ and harden the default config","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-44210"],"status":"curated","published":"2026-07-23"},{"id":"CVE-2026-7473","cve":"CVE-2026-7473","aliases":[],"title":"Arista EOS (tunnel decapsulation): With VXLAN, decap-groups or GRE configured, the switch incorrectly decapsulates and","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (tunnel decapsulation)","year":"2026","cvss_score":5.8,"severity":"medium","kev":true,"impact":"With VXLAN, decap-groups or GRE configured, the switch incorrectly decapsulates and forwards unexpected tunnelled packets — a tenant-isolation break on a VXLAN-segmented GPU fabric, exploited in the wild","attack_vector":"Network, fabric-local","remediation":"EOS upgrade with fabric failover; on a multi-tenant VXLAN fabric this is a segmentation failure, so it also demands a review of whether cross-tenant traffic actually occurred","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-7473"],"status":"curated","published":"2026-06-05"},{"id":"CVE-2020-15707","cve":"CVE-2020-15707","aliases":["BootHole family"],"title":"GRUB2 (initrd size handling): Integer overflows in the initrd command's size arithmetic corrupt GRUB's heap","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (initrd size handling)","year":"2020","cvss_score":5.7,"severity":"medium","kev":false,"impact":"Integer overflows in the initrd command's size arithmetic corrupt GRUB's heap. Same end state as the rest of the family - unsigned code running pre-kernel with Secure Boot still claiming to be enforcing.","attack_vector":"Requires the attacker to control the initrd list, i.e. write access to boot configuration on the node.","remediation":"grub2 package update + reboot. No config-only mitigation.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-15707","https://ubuntu.com/security/CVE-2020-15707"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2020-07-29"},{"id":"CVE-2021-3696","cve":"CVE-2021-3696","aliases":[],"title":"GRUB2 (PNG grayscale reader): Out-of-bounds write on the grayscale PNG path","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (PNG grayscale reader)","year":"2021","cvss_score":5.7,"severity":"medium","kev":false,"impact":"Out-of-bounds write on the grayscale PNG path. Same shape as the colour path but a separate code path that the first patch did not cover - relevant if you patched early and stopped.","attack_vector":"Attacker-supplied boot splash image.","remediation":"grub2 package update + reboot. Confirm your package version covers this one specifically, not just the headline PNG CVE.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-3696","https://access.redhat.com/security/cve/CVE-2021-3696"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2022-07-06"},{"id":"CVE-2022-23471","cve":"CVE-2022-23471","aliases":[],"title":"containerd: Goroutine leak in the CRI stream server terminal-resize path exhausts host memory","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2022","cvss_score":5.7,"severity":"medium","kev":false,"impact":"Goroutine leak in the CRI stream server terminal-resize path exhausts host memory; node DoS","attack_vector":"Any tenant workload that opens exec/attach sessions","remediation":"Rolling containerd upgrade with node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-23471"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2022-12-07"},{"id":"CVE-2023-25534","cve":"CVE-2023-25534","aliases":[],"title":"DGX H100 BMC (IPMI): Info disclosure","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX H100 BMC (IPMI)","year":"2023","cvss_score":5.7,"severity":"medium","kev":false,"impact":"Info disclosure","attack_vector":"Network-adjacent IPMI client","remediation":"Flash BMC 23.08.18 out-of-band","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5473/5473.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:H/UI:N/S:U/C:L/I:L/A:H","cwe":["CWE-20"],"published":"2023-09-20"},{"id":"CVE-2023-32690","cve":"CVE-2023-32690","aliases":["libspdm CTExponent"],"title":"DMTF libspdm - SPDM Requester timeout handling: A libspdm Requester stores the Responder's CTExponent","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"DMTF libspdm - SPDM Requester timeout handling","year":"2023","cvss_score":5.7,"severity":"medium","kev":false,"impact":"A libspdm Requester stores the Responder's CTExponent without validating it, so a malicious or faulty responder can force an enormous computed timeout and hang the requester. In an attestation flow this is a denial of service against the thing that decides whether a device is trustworthy - and a hung attestation is often failed open by the surrounding orchestration.","attack_vector":"Adjacent, unauthenticated with user interaction. A device on the link that answers CAPABILITIES dishonestly.","remediation":"Update to libspdm 2.3.3 / 3.0 or later, again through your device vendor's firmware. Separately, check what your orchestration does when attestation times out rather than fails - failing open on timeout is the more damaging half of this.","references":["https://github.com/DMTF/libspdm/security/advisories/GHSA-56h8-4gv5-jf2c","https://nvd.nist.gov/vuln/detail/CVE-2023-32690"],"status":"curated","published":"2023-06-01"},{"id":"CVE-2023-34472","cve":"CVE-2023-34472","aliases":["AMI-SA-2023006","Nozomi Labs BMC audit"],"title":"AMI MegaRAC SPx (BMC web interface, HTTP header handling): CRLF sequences are not neutralised in HTTP headers, so","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx (BMC web interface, HTTP header handling)","year":"2023","cvss_score":5.7,"severity":"medium","kev":false,"impact":"CRLF sequences are not neutralised in HTTP headers, so an attacker can split responses and inject headers of their choosing. Against a BMC web UI the payoff is session and cache manipulation against an administrator's browser - poisoning what the admin sees, planting cookies, or setting up a follow-on credential capture. It is an integrity bug that is useful as a stepping stone toward hijacking an admin's BMC session rather than a direct takeover.","attack_vector":"Adjacent network with a low-privilege BMC account. Requires an administrator to subsequently interact with the BMC web interface for the payoff, so it depends on your ops team actually using the web UI - which most do for KVM and console access.","remediation":"Firmware flash to SPx_12.5 / SPx_13.3 or later, out-of-band per node, ODM-gated. Low priority relative to the rest of this cluster, so fold it into the same flash campaign rather than scheduling separately. Config-only reduction: reach BMC web UIs only from a hardened jump host with a dedicated browser profile, so an admin session cannot be crossed with anything else.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2023006.pdf","https://nvd.nist.gov/vuln/detail/CVE-2023-34472"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-07-05"},{"id":"CVE-2024-21981","cve":"CVE-2024-21981","aliases":[],"title":"AMD Secure Processor - cryptographic key usage control: Once an attacker has arbitrary code execution inside the ASP","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor - cryptographic key usage control","year":"2024","cvss_score":5.7,"severity":"medium","kev":false,"impact":"Once an attacker has arbitrary code execution inside the ASP, weak key-usage controls let them extract the ASP's cryptographic keys rather than merely using them. Extracted platform keys are portable: they can be used off-box to forge attestation material or decrypt data long after you have re-imaged the node, so this turns a contained firmware compromise into a lasting one.","attack_vector":"Local, and requires having already achieved code execution in the ASP - it is a privilege-amplifier chained behind one of the other ASP bugs, not a standalone entry point.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. Because this sits inside the SEV-SNP trust boundary, the update also moves the platform's reported TCB version: after patching you must refresh VCEK certificates from AMD's KDS and update whatever attestation policy your tenants (or your own confidential-VM control plane) pin against, or every guest launch will start failing validation. If you have reason to believe a node's ASP was compromised, patching does not undo key extraction - the platform keys must be considered burned and the node's attestation identity retired.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21981","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2024-08-13"},{"id":"CVE-2024-25944","cve":"CVE-2024-25944","aliases":[],"title":"Dell OpenManage Enterprise (path traversal): An unauthenticated remote attacker reads files from the OME server","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Dell OpenManage Enterprise (path traversal)","year":"2024","cvss_score":5.7,"severity":"medium","kev":false,"impact":"An unauthenticated remote attacker reads files from the OME server filesystem with the web application's privileges.","attack_vector":"Unauthenticated network access to the OME web interface (v4.0 and prior).","remediation":"Apply the DSA-2024-100 update. Application upgrade. OME should never be internet-reachable; verify that while patching.","references":["https://www.dell.com/support/kbdoc/en-us/000223623/dsa-2024-100-security-update-for-dell-openmanage-enterprise-path-traversal-sensitive-data-disclosure-vulnerability"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-362","CWE-1108"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2024-47827","cve":"CVE-2024-47827","aliases":["GHSA-ghjw-32xw-ffwr"],"title":"Argo Workflows (controller, daemon workflow SPDY client race): A data race in a global variable in the Kubernetes SPDY","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo Workflows (controller, daemon workflow SPDY client race)","year":"2024","cvss_score":5.7,"severity":"medium","kev":false,"impact":"A data race in a global variable in the Kubernetes SPDY client path lets any user who can run a workflow panic the workflow controller on demand. One tenant issuing back-to-back daemon workflows halts scheduling for every other tenant sharing that controller.","attack_vector":"Any principal with permission to execute a workflow in a namespace the controller watches. No elevated rights needed.","remediation":"Upgrade the controller off 3.6.0-rc1 to 3.6.0-rc2 or any later release and restart. If you are running a release candidate in production, this is the signal to move to a GA build.","references":["https://github.com/argoproj/argo-workflows/security/advisories/GHSA-ghjw-32xw-ffwr","https://nvd.nist.gov/vuln/detail/CVE-2024-47827"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-23272","cve":"CVE-2025-23272","aliases":[],"title":"CUDA Toolkit: Info disclosure / DoS (buffer over-read)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2025","cvss_score":5.7,"severity":"medium","kev":false,"impact":"Info disclosure / DoS (buffer over-read)","attack_vector":"Malicious artifact","remediation":"Bump CUDA Toolkit; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23272","https://github.com/NVIDIA/product-security/tree/main/2025/5661"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:N/UI:N/S:U/C:L/I:N/A:H","cwe":["CWE-125"],"published":"2025-09-24"},{"id":"CVE-2025-33191","cve":"CVE-2025-33191","aliases":[],"title":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware: An invalid memory read in OSROOT firmware crashes","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware","year":"2025","cvss_score":5.7,"severity":"medium","kev":false,"impact":"An invalid memory read in OSROOT firmware crashes the platform. These sit in the GB10 root-of-trust chain, so a successful exploit undermines the platform's own attestation and secure-boot story rather than just the OS above it.","attack_vector":"Local access to the DGX Spark. Several of the set need no privileges at all; the rest need host root. This is a desk-side developer box, so physical and local access assumptions are much weaker than for a racked DGX.","remediation":"Apply the DGX Spark firmware update from bulletin 5720. Cost: flash plus reboot, low drain cost given the form factor, but not live-patchable and root-of-trust firmware cannot be rolled back once applied.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33191","https://github.com/NVIDIA/product-security/tree/main/2025/5720"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:C/C:L/I:N/A:L","cwe":["CWE-20"],"fleet":{"pain_class":"firmware-flash"},"published":"2025-11-25"},{"id":"CVE-2025-33192","cve":"CVE-2025-33192","aliases":[],"title":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware: An arbitrary memory read in SROOT firmware gives","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware","year":"2025","cvss_score":5.7,"severity":"medium","kev":false,"impact":"An arbitrary memory read in SROOT firmware gives a denial of service and exposes firmware memory layout. These sit in the GB10 root-of-trust chain, so a successful exploit undermines the platform's own attestation and secure-boot story rather than just the OS above it.","attack_vector":"Local access to the DGX Spark. Several of the set need no privileges at all; the rest need host root. This is a desk-side developer box, so physical and local access assumptions are much weaker than for a racked DGX.","remediation":"Apply the DGX Spark firmware update from bulletin 5720. Cost: flash plus reboot, low drain cost given the form factor, but not live-patchable and root-of-trust firmware cannot be rolled back once applied.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33192","https://github.com/NVIDIA/product-security/tree/main/2025/5720"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:C/C:L/I:N/A:L","cwe":["CWE-690"],"fleet":{"pain_class":"firmware-flash"},"published":"2025-11-25"},{"id":"CVE-2025-33193","cve":"CVE-2025-33193","aliases":[],"title":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware: SROOT firmware validates integrity improperly, leaking","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware","year":"2025","cvss_score":5.7,"severity":"medium","kev":false,"impact":"SROOT firmware validates integrity improperly, leaking information that the root of trust was supposed to protect. These sit in the GB10 root-of-trust chain, so a successful exploit undermines the platform's own attestation and secure-boot story rather than just the OS above it.","attack_vector":"Local access to the DGX Spark. Several of the set need no privileges at all; the rest need host root. This is a desk-side developer box, so physical and local access assumptions are much weaker than for a racked DGX.","remediation":"Apply the DGX Spark firmware update from bulletin 5720. Cost: flash plus reboot, low drain cost given the form factor, but not live-patchable and root-of-trust firmware cannot be rolled back once applied.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33193","https://github.com/NVIDIA/product-security/tree/main/2025/5720"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:C/C:L/I:N/A:L","cwe":["CWE-354"],"fleet":{"pain_class":"firmware-flash"},"published":"2025-11-25"},{"id":"CVE-2025-33194","cve":"CVE-2025-33194","aliases":[],"title":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware: Improper input processing in SROOT firmware yields","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware","year":"2025","cvss_score":5.7,"severity":"medium","kev":false,"impact":"Improper input processing in SROOT firmware yields information disclosure or a crash. These sit in the GB10 root-of-trust chain, so a successful exploit undermines the platform's own attestation and secure-boot story rather than just the OS above it.","attack_vector":"Local access to the DGX Spark. Several of the set need no privileges at all; the rest need host root. This is a desk-side developer box, so physical and local access assumptions are much weaker than for a racked DGX.","remediation":"Apply the DGX Spark firmware update from bulletin 5720. Cost: flash plus reboot, low drain cost given the form factor, but not live-patchable and root-of-trust firmware cannot be rolled back once applied.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33194","https://github.com/NVIDIA/product-security/tree/main/2025/5720"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:C/C:L/I:N/A:L","cwe":["CWE-180"],"fleet":{"pain_class":"firmware-flash"},"published":"2025-11-25"},{"id":"CVE-2026-24215","cve":"CVE-2026-24215","aliases":[],"title":"Triton Inference Server: DoS via connection flooding","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2026","cvss_score":5.7,"severity":"medium","kev":false,"impact":"DoS via connection flooding","attack_vector":"Network-adjacent unauthenticated","remediation":"Upgrade Triton; add rate limiting","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24215","https://github.com/NVIDIA/product-security/tree/main/2026/5828"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:R/S:U/C:N/I:N/A:H","cwe":["CWE-400"],"published":"2026-05-20"},{"id":"CVE-2017-5715","cve":"CVE-2017-5715","aliases":["Spectre v2","Branch Target Injection"],"title":"Intel processors (indirect branch prediction): Spectre v2: an attacker trains the indirect branch predictor so that a","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (indirect branch prediction)","year":"2017","cvss_score":5.6,"severity":"medium","kev":false,"impact":"Spectre v2: an attacker trains the indirect branch predictor so that a victim context - another process, another VM, or the kernel - speculatively executes an attacker-chosen gadget and leaks its memory through a cache side channel. On a shared GPU host this is the canonical cross-VM and container-to-host read primitive, and it is still the reason retpoline, IBPB and eIBRS exist in every kernel you run.","attack_vector":"Local code execution anywhere on the host - any container, any VM. No privilege needed.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect. On nodes that host untrusted co-tenants, also disable SMT or enforce core scheduling; that costs real throughput and is a capacity-planning decision, not a free toggle.","references":["https://nvd.nist.gov/vuln/detail/CVE-2017-5715","https://security-center.intel.com/advisory.aspx?intelid=INTEL-SA-00088&languageid=en-fr","https://spectreattack.com/"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"tags":["tenant-isolation"],"published":"2018-01-04"},{"id":"CVE-2017-5753","cve":"CVE-2017-5753","aliases":["Spectre v1","Bounds Check Bypass"],"title":"Intel processors (bounds check bypass): Spectre v1: speculative execution past a bounds check lets an attacker read","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (bounds check bypass)","year":"2017","cvss_score":5.6,"severity":"medium","kev":false,"impact":"Spectre v1: speculative execution past a bounds check lets an attacker read memory the check was supposed to protect, within the same address space. The practical exposure on an AI node is inside anything that JITs or interprets untrusted input - eBPF, a Python runtime, a model-serving framework's custom-op path.","attack_vector":"Local code execution, including code inside a sandbox or interpreter that is meant to be confined.","remediation":"Software mitigation in the kernel and in individual programs (array index masking, speculation barriers) rather than microcode. Take kernel updates, keep runtimes current, and assume any interpreter you expose to untrusted input needs its own hardening. Reboot for the kernel component.","references":["https://nvd.nist.gov/vuln/detail/CVE-2017-5753","https://spectreattack.com/","https://www.intel.com/content/www/us/en/developer/topic-technology/software-security-guidance/advisory-guidance/bounds-check-bypass.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"published":"2018-01-04"},{"id":"CVE-2017-5754","cve":"CVE-2017-5754","aliases":["Meltdown","Rogue Data Cache Load"],"title":"Intel processors (rogue data cache load): Meltdown: unprivileged code reads kernel memory - and on affected parts","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (rogue data cache load)","year":"2017","cvss_score":5.6,"severity":"medium","kev":false,"impact":"Meltdown: unprivileged code reads kernel memory - and on affected parts, memory belonging to other tenants mapped into the kernel's direct map - by exploiting out-of-order execution past a permission check. The mitigation, kernel page-table isolation, is the reason every syscall on affected hardware got measurably slower.","attack_vector":"Any local unprivileged code on an affected processor.","remediation":"Kernel page-table isolation (KPTI/PTI), shipped in the kernel; take the kernel update and reboot. Newer silicon fixes it in hardware. No microcode or BIOS component for the mitigation itself. Expect a syscall-heavy throughput regression on affected parts - that is the fix working.","references":["https://nvd.nist.gov/vuln/detail/CVE-2017-5754","https://meltdownattack.com/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2018-01-04"},{"id":"CVE-2018-12126","cve":"CVE-2018-12126","aliases":["MSBDS","Fallout"],"title":"Intel processors (microarchitectural data sampling): One of the MDS family: store buffers retain data from other","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (microarchitectural data sampling)","year":"2018","cvss_score":5.6,"severity":"medium","kev":false,"impact":"One of the MDS family: store buffers retain data from other contexts and can be sampled speculatively. The operator-relevant consequence is that data crosses between hyperthread siblings, between VMs, and out of enclaves without any architectural access - so on a node with SMT enabled and untrusted co-tenants, tenant isolation is not holding.","attack_vector":"Local code on the same physical core - with SMT enabled that includes a co-tenant on the sibling thread, which is the configuration most density-optimised fleets run.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect. The software half is buffer clearing on context switch (VERW), already in current kernels. On nodes that host untrusted co-tenants, also disable SMT or enforce core scheduling; that costs real throughput and is a capacity-planning decision, not a free toggle.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-12126","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00233.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"tags":["tenant-isolation"],"published":"2019-05-30"},{"id":"CVE-2018-12127","cve":"CVE-2018-12127","aliases":["MLPDS","RIDL"],"title":"Intel processors (microarchitectural data sampling): One of the MDS family: load ports retain data from other contexts","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (microarchitectural data sampling)","year":"2018","cvss_score":5.6,"severity":"medium","kev":false,"impact":"One of the MDS family: load ports retain data from other contexts and can be sampled speculatively. The operator-relevant consequence is that data crosses between hyperthread siblings, between VMs, and out of enclaves without any architectural access - so on a node with SMT enabled and untrusted co-tenants, tenant isolation is not holding.","attack_vector":"Local code on the same physical core - with SMT enabled that includes a co-tenant on the sibling thread, which is the configuration most density-optimised fleets run.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect. The software half is buffer clearing on context switch (VERW), already in current kernels. On nodes that host untrusted co-tenants, also disable SMT or enforce core scheduling; that costs real throughput and is a capacity-planning decision, not a free toggle.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-12127","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00233.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"tags":["tenant-isolation"],"published":"2019-05-30"},{"id":"CVE-2018-3620","cve":"CVE-2018-3620","aliases":["Foreshadow-OS","L1TF"],"title":"Intel processors (L1 terminal fault, OS/SMM): The OS-level variant of L1 terminal fault: a local user can speculatively","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (L1 terminal fault, OS/SMM)","year":"2018","cvss_score":5.6,"severity":"medium","kev":false,"impact":"The OS-level variant of L1 terminal fault: a local user can speculatively read data present in the L1 cache belonging to the kernel or to system-management memory, by faulting on a carefully crafted page-table entry. Sits alongside the hypervisor-level L1TF variant as the reason page-table entry inversion exists in every modern kernel.","attack_vector":"Any local unprivileged code on an affected processor.","remediation":"Microcode update plus the kernel's PTE-inversion mitigation, then reboot. Microcode is late-loadable at boot without an OEM BIOS release. Verify with the kernel's l1tf sysfs vulnerability file after reboot rather than assuming.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-3620","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00161.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"published":"2018-08-14"},{"id":"CVE-2018-3640","cve":"CVE-2018-3640","aliases":["Spectre v3a","RSRE","Rogue System Register Read"],"title":"Intel processors (rogue system register read): Spectre v3a: speculative reads of system registers leak system","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (rogue system register read)","year":"2018","cvss_score":5.6,"severity":"medium","kev":false,"impact":"Spectre v3a: speculative reads of system registers leak system parameters - MSR contents and similar - to unprivileged local code. On its own it exposes configuration rather than data, but that configuration is what an attacker needs to aim the more serious attacks.","attack_vector":"Local unprivileged code.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-3640","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00115.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"published":"2018-05-22"},{"id":"CVE-2018-3646","cve":"CVE-2018-3646","aliases":[],"title":"Intel CPU (L1TF / Foreshadow-NG): L1 Terminal Fault: a guest reads any data present in the L1 data cache, including","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel CPU (L1TF / Foreshadow-NG)","year":"2018","cvss_score":5.6,"severity":"medium","kev":false,"impact":"L1 Terminal Fault: a guest reads any data present in the L1 data cache, including other guests' and the hypervisor's memory","attack_vector":"Tenant VM guest","remediation":"Microcode + hypervisor L1D-flush mitigation + reboot; full mitigation requires core scheduling or disabling SMT, which costs roughly half the CPU throughput on a hyperthreaded host","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-3646"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"published":"2018-08-14"},{"id":"CVE-2018-3665","cve":"CVE-2018-3665","aliases":["LazyFP"],"title":"Intel processors (lazy FP state restore): LazyFP: when the OS restores FPU/vector state lazily, one process can","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (lazy FP state restore)","year":"2018","cvss_score":5.6,"severity":"medium","kev":false,"impact":"LazyFP: when the OS restores FPU/vector state lazily, one process can speculatively read another process's floating-point and vector register contents. On an AI node the vector registers hold tensor data and, in crypto-adjacent code, key material.","attack_vector":"Local code on a host whose OS uses lazy FPU state restore.","remediation":"Fixed in the OS by switching to eager FPU restore - take the kernel update and reboot. No microcode or BIOS component. Modern kernels already default to eager restore; this matters mostly on frozen legacy images.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-3665","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00145.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2018-06-21"},{"id":"CVE-2018-3693","cve":"CVE-2018-3693","aliases":["Spectre 1.1","Bounds Check Bypass Store"],"title":"Intel processors (bounds check bypass store): Spectre 1.1: speculative stores can overflow a bounds-checked buffer","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (bounds check bypass store)","year":"2018","cvss_score":5.6,"severity":"medium","kev":false,"impact":"Spectre 1.1: speculative stores can overflow a bounds-checked buffer, letting an attacker write speculatively into structures that steer later speculation. Extends Spectre v1 from a read primitive to a write-and-redirect one inside a sandbox.","attack_vector":"Local code execution, particularly inside interpreters and JITs handling untrusted input.","remediation":"Software mitigation in compilers, kernels and runtimes rather than microcode. Take kernel and runtime updates; reboot for the kernel component.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-3693","https://www.intel.com/content/www/us/en/developer/topic-technology/software-security-guidance/advisory-guidance/bounds-check-bypass-store.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"published":"2018-07-10"},{"id":"CVE-2019-11091","cve":"CVE-2019-11091","aliases":["MDSUM"],"title":"Intel processors (microarchitectural data sampling): One of the MDS family: uncacheable-memory accesses leave","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (microarchitectural data sampling)","year":"2019","cvss_score":5.6,"severity":"medium","kev":false,"impact":"One of the MDS family: uncacheable-memory accesses leave sampleable residue in microarchitectural buffers. The operator-relevant consequence is that data crosses between hyperthread siblings, between VMs, and out of enclaves without any architectural access - so on a node with SMT enabled and untrusted co-tenants, tenant isolation is not holding.","attack_vector":"Local code on the same physical core - with SMT enabled that includes a co-tenant on the sibling thread, which is the configuration most density-optimised fleets run.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect. The software half is buffer clearing on context switch (VERW), already in current kernels. On nodes that host untrusted co-tenants, also disable SMT or enforce core scheduling; that costs real throughput and is a capacity-planning decision, not a free toggle.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-11091","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00233.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"tags":["tenant-isolation"],"published":"2019-05-30"},{"id":"CVE-2019-1125","cve":"CVE-2019-1125","aliases":["SWAPGS","Spectre v1 SWAPGS variant","Spectre-SWAPGS"],"title":"Intel x86-64 CPUs (Ivy Bridge onward); Windows and Linux kernel entry paths: The kernel's syscall/interrupt entry path","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel x86-64 CPUs (Ivy Bridge onward); Windows and Linux kernel entry paths","year":"2019","cvss_score":5.6,"severity":"medium","kev":false,"impact":"The kernel's syscall/interrupt entry path uses SWAPGS to switch the GS base between user and kernel values; the CPU speculates past the conditional that decides whether to swap, so a user process can make the kernel speculatively dereference a user-controlled GS base and leak kernel memory through a cache channel. It bypasses the standard Spectre v1 mitigations because it lives in the entry code that runs before them. Operator exposure is a local process reading kernel memory - which on a container host means other tenants' data in the page cache and the kernel's credential structures.","attack_vector":"Unprivileged local code on the host or inside a guest, exercised through ordinary syscalls and interrupts. No SMT or core-sharing requirement.","remediation":"Kernel/hypervisor patch only - no microcode, no BIOS flash. The fix adds LFENCE serialization in the SWAPGS entry paths. Reboot into the patched kernel (drain the node), and that is the whole job. Runtime cost is a serializing instruction on kernel entry - measurable on syscall-microbenchmarks, effectively invisible on GPU workloads. Patched in Linux since 5.2 and backported everywhere, so on a current fleet this is closed; the audit item is confirming no host runs a pre-August-2019 kernel and that 'mitigations=off' is not set anywhere.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-1125","https://docs.kernel.org/admin-guide/hw-vuln/spectre.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2019-09-03"},{"id":"CVE-2020-0550","cve":"CVE-2020-0550","aliases":["Snoop-assisted L1D Sampling","SnoopAssist"],"title":"Intel processors (snoop-assisted L1D sampling): Data can be leaked out of L1D during snoop transactions, crossing","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (snoop-assisted L1D sampling)","year":"2020","cvss_score":5.6,"severity":"medium","kev":false,"impact":"Data can be leaked out of L1D during snoop transactions, crossing privilege boundaries. Lower practical yield than the fill-buffer attacks but it targets the same shared L1 that makes SMT co-tenancy risky.","attack_vector":"Local code on the same physical core as the victim.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-0550","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00330.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"published":"2020-03-12"},{"id":"CVE-2020-0551","cve":"CVE-2020-0551","aliases":["LVI","Load Value Injection"],"title":"Intel processors / SGX (load value injection): The inverse of Meltdown: instead of leaking data out of the enclave, the","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors / SGX (load value injection)","year":"2020","cvss_score":5.6,"severity":"medium","kev":false,"impact":"The inverse of Meltdown: instead of leaking data out of the enclave, the attacker injects a value into a faulting load inside the victim enclave and steers its transient execution into attacker-chosen gadgets. That gives enclave-secret extraction from outside the enclave, again breaking the SGX guarantee against a privileged host.","attack_vector":"Local privileged code on the host targeting a victim enclave on the same machine.","remediation":"SGX SDK/PSW update that inserts LFENCE serialisation in enclave code, plus microcode. The software mitigation requires recompiling enclaves with the patched SDK - a code change for whoever ships the enclave, not something the operator can apply unilaterally - and it carries a heavy performance cost. Microcode is late-loadable at boot; enclave recompilation is not. Re-attestation required.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-0551","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00334.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"tags":["tenant-isolation"],"published":"2020-03-12"},{"id":"CVE-2021-26401","cve":"CVE-2021-26401","aliases":["LFENCE/JMP insufficiency"],"title":"AMD processors - LFENCE/JMP mitigation for Spectre v2 (CVE-2017-5715): The LFENCE/JMP sequence AMD originally","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD processors - LFENCE/JMP mitigation for Spectre v2 (CVE-2017-5715)","year":"2021","cvss_score":5.6,"severity":"medium","kev":false,"impact":"The LFENCE/JMP sequence AMD originally recommended as the cheap Spectre-v2 mitigation (mitigation V2-2) turns out not to sufficiently block branch target injection on some AMD CPUs. Anyone who took AMD's early guidance and chose LFENCE/JMP over retpoline because it was faster has been running with a Spectre-v2 mitigation that does not actually hold - a cross-VM and cross-process speculative disclosure channel that people believe is already closed. The dangerous part is the false sense of coverage, not the novelty of the attack.","attack_vector":"Local, cross-privilege and cross-guest speculative execution. Reachable from any tenant workload on affected hardware.","remediation":"Switch from LFENCE/JMP to retpoline or hardware IBRS/IBPB. On Linux, verify what is actually active by reading /sys/devices/system/cpu/vulnerabilities/spectre_v2 on your fleet - do not assume, check, because the string tells you exactly which mitigation the kernel selected. Changing it needs a kernel update and/or boot parameter change plus a reboot; on some platforms full IBRS also needs microcode from an SBIOS update. Retpoline and IBRS both cost performance relative to LFENCE/JMP, which is why the weaker option got chosen in the first place.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26401","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2022-03-11"},{"id":"CVE-2021-46778","cve":"CVE-2021-46778","aliases":["SQUIP"],"title":"AMD Zen 1 / Zen 2 / Zen 3 - execution unit scheduler queue contention (SMT): AMD's split scheduler design gives each","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD Zen 1 / Zen 2 / Zen 3 - execution unit scheduler queue contention (SMT)","year":"2021","cvss_score":5.6,"severity":"medium","kev":false,"impact":"AMD's split scheduler design gives each execution unit its own queue, and contention on those queues is observable from the sibling SMT thread. An attacker running on one hardware thread measures scheduler pressure and reconstructs what the co-resident thread is computing - the SQUIP researchers recovered a full RSA-4096 private key from a victim on the sibling thread. This is a genuine cross-tenant confidentiality break with no memory access involved at all: nothing in your container, VM or cgroup boundary sees it happening, because no boundary is crossed in software.","attack_vector":"Local, unprivileged, requires SMT enabled and the attacker scheduled on the sibling thread of the victim's physical core. On a bin-packed cluster that co-residency happens by default - your scheduler arranges it for you. Affects Zen 1, Zen 2 and Zen 3.","remediation":"AMD's guidance is that software should use constant-time / secret-independent control flow rather than a microcode fix, so **treat this as effectively unpatchable in hardware**. The operator-side controls are the real answer: disable SMT on nodes that mix tenants, or enforce whole-core (not thread) allocation so a physical core is never shared across trust boundaries. Kubernetes operators can get this with the CPU manager's full-pcpus-only policy. Disabling SMT needs a reboot; core-pinning policy needs a kubelet restart and a drain. The zero-cost mitigation available today is scheduling policy rather than patching: these attacks need the attacker and victim co-resident on sibling SMT threads, so either disable SMT (costing roughly 10-25% throughput on most inference and training workloads) or enforce core isolation so no two tenants ever share a physical core. On a GPU fleet the CPU is rarely the bottleneck, which makes disabling SMT a cheaper trade than it looks on paper.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-46778","https://stefangast.eu/papers/squip.pdf","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"tags":["tenant-isolation"],"published":"2022-08-10"},{"id":"CVE-2022-23825","cve":"CVE-2022-23825","aliases":[],"title":"AMD CPU (Branch Type Confusion): Non-Retbleed branch type confusion - speculative cross-domain leak","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"AMD CPU (Branch Type Confusion)","year":"2022","cvss_score":5.6,"severity":"medium","kev":false,"impact":"Non-Retbleed branch type confusion - speculative cross-domain leak","attack_vector":"Any tenant process in a container; tenant VM guest","remediation":"Microcode + kernel mitigation + reboot; standing perf cost","references":["https://access.redhat.com/security/cve/CVE-2022-23825"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"published":"2022-07-14"},{"id":"CVE-2022-23960","cve":"CVE-2022-23960","aliases":["Spectre-BHB","Branch History Injection","BHI","TFV-9","CVE-2022-25368 (Ampere variant)"],"title":"Arm Cortex-A and Neoverse cores (Neoverse N1/N2/V1 among them); Trusted Firmware-A","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arm Cortex-A and Neoverse cores (Neoverse N1/N2/V1 among them); Trusted Firmware-A; also tracked by Ampere as…","year":"2022","cvss_score":5.6,"severity":"medium","kev":false,"impact":"Branch history is shared across privilege boundaries, so an attacker in one context can poison the Branch History Buffer and make a victim context mispredict into a gadget that leaves the secret in the cache. Practically: a guest reads hypervisor memory, or one container reads another's, at side-channel bandwidth. Slow, but it does not crash anything and leaves no trace, which is exactly the profile that matters when the secret is a customer's model weights or an API key sitting in host memory on a shared Arm node.","attack_vector":"Unprivileged code in any guest or container on an affected Arm core that shares a physical core with the victim. Worse when SMT or aggressive core-sharing is enabled; a dedicated-core scheduling policy raises the bar considerably.","remediation":"Layered and none of it is optional. EL3 firmware exposes the mitigation via SMCCC (updated TF-A from the OEM, flash + reboot); the guest and host kernels must run the arm64 Spectre-BHB workarounds (loop or clearbhb sequence per core type). Newer cores get FEAT_CLEARBHB in hardware; older Neoverse N1 parts pay for a software loop on every exception entry, which is a real syscall-path cost. Ampere published AMP-SB-0001 for Altra/Altra Max/AmpereOne. The durable operator-level control is to stop co-tenanting untrusted workloads on one physical core.","references":["https://developer.arm.com/Arm%20Security%20Center/Speculative%20Processor%20Vulnerability","https://trustedfirmware-a.readthedocs.io/en/latest/security_advisories/security-advisory-tfv-9.html","https://amperecomputing.com/products/security-bulletins/impact-of-spectre-bhb-on-ampere.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2022-03-13"},{"id":"CVE-2022-29900","cve":"CVE-2022-29900","aliases":[],"title":"AMD CPU (Retbleed): Retbleed: arbitrary speculative code execution via return instructions","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"AMD CPU (Retbleed)","year":"2022","cvss_score":5.6,"severity":"medium","kev":false,"impact":"Retbleed: arbitrary speculative code execution via return instructions - cross-domain secret disclosure","attack_vector":"Any tenant process in a container; tenant VM guest","remediation":"Kernel mitigation (retbleed=) + reboot. Standing performance cost, historically severe on Zen 1/2 - measure before enabling fleet-wide","references":["https://access.redhat.com/security/cve/CVE-2022-29900"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2022-07-12"},{"id":"CVE-2022-29901","cve":"CVE-2022-29901","aliases":[],"title":"Intel CPU (Retbleed): Retbleed on Intel - speculative execution of return instructions leaks across privilege boundaries","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel CPU (Retbleed)","year":"2022","cvss_score":5.6,"severity":"medium","kev":false,"impact":"Retbleed on Intel - speculative execution of return instructions leaks across privilege boundaries","attack_vector":"Any tenant process in a container; tenant VM guest","remediation":"Kernel mitigation + reboot; standing perf cost (IBRS on Skylake-era parts is expensive)","references":["https://access.redhat.com/security/cve/CVE-2022-29901"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2022-07-12"},{"id":"CVE-2023-20569","cve":"CVE-2023-20569","aliases":[],"title":"AMD CPU (Inception / SRSO): Inception: Speculative Return Stack Overflow","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"AMD CPU (Inception / SRSO)","year":"2023","cvss_score":5.6,"severity":"medium","kev":false,"impact":"Inception: Speculative Return Stack Overflow - attacker-controlled speculative disclosure across privilege domains on Zen 1-4","attack_vector":"Any tenant process in a container; tenant VM guest","remediation":"Microcode + kernel mitigation (`spec_rstack_overflow=`) + reboot. Standing perf cost - the safe-RET mitigation is measurable on syscall-heavy workloads","references":["https://access.redhat.com/security/cve/CVE-2023-20569"],"status":"curated","fleet":{"ubiquity":"very common - spans Zen 1 through Zen 4, i.e. essentially every AMD host CPU generation in service","remediation_pain":"microcode+reboot **and** a kernel update (SRSO safe-return / IBPB-on-entry); on Zen 1/2 the mitigation additionally requires **disabling SMT**, which permanently cuts logical core count on those nodes","pain_class":"microcode + reboot","why_fleet_wide":"Kernel-memory disclosure via speculative return redirection across the whole AMD lineup; the SMT-disable mitigation is a capacity hit the operator has to absorb fleet-wide, not a one-time reboot."},"published":"2023-08-08"},{"id":"CVE-2023-25775","cve":"CVE-2023-25775","aliases":[],"title":"Intel irdma driver (Ethernet Controller RDMA for Linux): Improper access control in the Intel RDMA driver lets an","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel irdma driver (Ethernet Controller RDMA for Linux)","year":"2023","cvss_score":5.6,"severity":"medium","kev":false,"impact":"Improper access control in the Intel RDMA driver lets an unauthenticated user escalate privilege. RDMA is the transport for collective operations across an AI cluster, and the driver maps queue pairs directly into user processes - so an access-control failure here is one workload reaching another's RDMA resources on the same host.","attack_vector":"Unauthenticated, which on an RDMA fabric means anything that can present traffic to the verbs interface. Treat the RDMA fabric as an authentication boundary that is not actually authenticating.","remediation":"Update the Intel irdma driver to 1.9.30 or later. Module reload drops RDMA connections and kills in-flight collectives - drain the node. No firmware component for this one.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25775","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00794.html"],"status":"curated","fleet":{"pain_class":"node-drain"},"tags":["tenant-isolation"],"published":"2023-08-11"},{"id":"CVE-2023-3301","cve":"CVE-2023-3301","aliases":[],"title":"QEMU (net): Triggerable assertion via a race on NIC hot-unplug - guest can abort the host QEMU process","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"QEMU (net)","year":"2023","cvss_score":5.6,"severity":"medium","kev":false,"impact":"Triggerable assertion via a race on NIC hot-unplug - guest can abort the host QEMU process","attack_vector":"Tenant VM guest","remediation":"QEMU update + VM restart/live-migration","references":["https://access.redhat.com/security/cve/CVE-2023-3301"],"status":"curated","published":"2023-09-13"},{"id":"CVE-2024-28956","cve":"CVE-2024-28956","aliases":["XSA-469"],"title":"Xen / Intel CPU (ITS): Indirect Target Selection - speculative execution leak across privilege domains on Intel parts","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen / Intel CPU (ITS)","year":"2024","cvss_score":5.6,"severity":"medium","kev":false,"impact":"Indirect Target Selection - speculative execution leak across privilege domains on Intel parts","attack_vector":"Tenant VM guest; any tenant process in a container","remediation":"Microcode + hypervisor/kernel mitigation + reboot; standing perf cost","references":["https://xenbits.xen.org/xsa/advisory-469.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"published":"2025-05-13"},{"id":"CVE-2024-33607","cve":"CVE-2024-33607","aliases":[],"title":"Intel TDX module: An out-of-bounds read in the TDX module reachable by an authenticated user, leaking information","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel TDX module","year":"2024","cvss_score":5.6,"severity":"medium","kev":false,"impact":"An out-of-bounds read in the TDX module reachable by an authenticated user, leaking information across the TDX boundary. Read primitives inside the module are the ones to watch: they are the shape that leaks other trust domains' state.","attack_vector":"An authenticated user on the host.","remediation":"Update the Intel TDX module. The TDX module is loaded by the SEAM loader at boot, so the practical rollout is: stage the new module, drain every trust domain off the node, and reboot. It is not a live-patchable component and running TDs cannot be migrated through it. After the update, every TD must re-attest because the TDX module SVN is part of the attestation report - so anything that pinned the old measurement will fail until you update your attestation policy too. No OEM BIOS release needed for the module itself, which makes this materially faster than a platform firmware update.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-33607","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01192.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2025-08-12"},{"id":"CVE-2024-36350","cve":"CVE-2024-36350","aliases":[],"title":"AMD CPU (TSA-L1): Transient Scheduler Attack - store-to-load forwarding leak across contexts on Zen 3/4","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"AMD CPU (TSA-L1)","year":"2024","cvss_score":5.6,"severity":"medium","kev":false,"impact":"Transient Scheduler Attack - store-to-load forwarding leak across contexts on Zen 3/4","attack_vector":"Any tenant process in a container; tenant VM guest","remediation":"Microcode + kernel mitigation (VERW on transitions) + reboot; standing perf cost. Also tracked as Xen XSA-471","references":["https://access.redhat.com/security/cve/CVE-2024-36350"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"published":"2025-07-08"},{"id":"CVE-2024-36357","cve":"CVE-2024-36357","aliases":[],"title":"AMD CPU (TSA-SQ): Transient Scheduler Attack via the store queue - cross-context information leak","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"AMD CPU (TSA-SQ)","year":"2024","cvss_score":5.6,"severity":"medium","kev":false,"impact":"Transient Scheduler Attack via the store queue - cross-context information leak","attack_vector":"Any tenant process in a container; tenant VM guest","remediation":"Microcode + kernel mitigation + reboot; standing perf cost; consider disabling SMT for isolation-sensitive tenants","references":["https://access.redhat.com/security/cve/CVE-2024-36357"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"published":"2025-07-08"},{"id":"CVE-2024-43420","cve":"CVE-2024-43420","aliases":[],"title":"Intel Atom processors (shared predictor transient execution): Shared microarchitectural predictor state influences","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Atom processors (shared predictor transient execution)","year":"2024","cvss_score":5.6,"severity":"medium","kev":false,"impact":"Shared microarchitectural predictor state influences transient execution on Atom parts, allowing local information disclosure. Same advisory wave as the branch privilege injection work; relevant to Atom-based appliances inside the datacenter footprint.","attack_vector":"Local authenticated code on an affected Atom platform.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-43420","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01247.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"published":"2025-05-13"},{"id":"CVE-2024-45332","cve":"CVE-2024-45332","aliases":["Branch Privilege Injection","BPI","Branch Predictor Race Conditions"],"title":"Intel processors (indirect branch predictor race): Branch Privilege Injection: a race in how the indirect branch","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (indirect branch predictor race)","year":"2024","cvss_score":5.6,"severity":"medium","kev":false,"impact":"Branch Privilege Injection: a race in how the indirect branch predictor associates predictions with privilege level lets unprivileged code get its predictions applied in kernel context, reading kernel memory even on parts with hardware Spectre-v2 mitigations. The researchers demonstrated reading /etc/shadow on a fully patched machine - it defeats the mitigations operators had been told were sufficient.","attack_vector":"Local unprivileged code on the node - any container or VM.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect. Confirm the specific microcode revision Intel names for your stepping; this one is not fully closed by kernel changes alone.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-45332","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01247.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"tags":["tenant-isolation"],"published":"2025-05-13"},{"id":"CVE-2025-20623","cve":"CVE-2025-20623","aliases":[],"title":"Intel Core processors, 10th generation (shared predictor state): Shared predictor state influencing transient execution","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Core processors, 10th generation (shared predictor state)","year":"2025","cvss_score":5.6,"severity":"medium","kev":false,"impact":"Shared predictor state influencing transient execution on 10th-generation Core parts, leaking information to a local authenticated user. Matters where Core-class hardware runs management, build or edge-inference roles alongside the Xeon fleet.","attack_vector":"Local authenticated code.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-20623","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01247.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"published":"2025-05-13"},{"id":"CVE-2025-24495","cve":"CVE-2025-24495","aliases":["Training Solo"],"title":"Intel Core Ultra processors (branch prediction unit initialisation): Part of the Training Solo family: incorrect","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Core Ultra processors (branch prediction unit initialisation)","year":"2025","cvss_score":5.6,"severity":"medium","kev":false,"impact":"Part of the Training Solo family: incorrect initialisation of the branch prediction unit lets an attacker self-train a predictor within the victim's own domain, so the leak works without the cross-domain training that existing mitigations assume. That is the significance - it sidesteps domain-isolation mitigations rather than defeating them head-on, and it reopens guest-to-host and user-to-kernel leakage on parts believed fixed.","attack_vector":"Local unprivileged code on an affected processor.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect. Training Solo also needs kernel-side changes for the eBPF and indirect-branch paths; take both.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-24495","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01322.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"tags":["tenant-isolation"],"published":"2025-05-13"},{"cwe":["CWE-125"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-39823","cve":"CVE-2025-39823","aliases":[],"title":"Linux kernel (arch/x86/kvm): Guest-supplied array indices in the host's local-APIC emulation (an IPI destination id and","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (arch/x86/kvm)","year":"2025","cvss_score":5.6,"severity":"medium","kev":false,"impact":"Guest-supplied array indices in the host's local-APIC emulation (an IPI destination id and a lowest-priority search index) were bounds-checked but never clamped, leaving a Spectre-v1 gadget that the guest controls. A tenant can train the bounds check and have the host speculatively read memory past those arrays, then recover it through a cache side channel - host kernel data disclosed into the VM.","attack_vector":"Guest-driven and unprivileged inside the VM: the tenant issues IPIs and APIC/MSR writes with out-of-range destination values from any vCPU. No host privilege, no device node, no VMM cooperation, and it works on both Intel and AMD hosts with in-kernel APIC emulation (the default).","remediation":"Update to a stable kernel with the linked fix (no fixed release enumerated; take the branch carrying commit d51e381beed5). No practical interim control - moving APIC emulation to userspace is not a realistic production option.","references":["https://git.kernel.org/stable/c/d51e381beed5e2f50f85f49f6c90e023754efa12","https://git.kernel.org/stable/c/72777fc31aa7ab2ce00f44bfa3929c6eabbeaf48","https://nvd.nist.gov/vuln/detail/CVE-2025-39823"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-15788","cve":"CVE-2026-15788","aliases":[],"title":"BuildKit: NTFS junctions inside the cache root escape the cache mount on Windows container workers","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"BuildKit","year":"2026","cvss_score":5.6,"severity":"medium","kev":false,"impact":"NTFS junctions inside the cache root escape the cache mount on Windows container workers","attack_vector":"Untrusted build author on a WCOW builder","remediation":"Upgrade BuildKit; avoid shared WCOW builders","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-15788"],"status":"curated","published":"2026-07-20"},{"id":"CVE-2026-24198","cve":"CVE-2026-24198","aliases":[],"title":"GPU Display Driver: Improper access control in driver operations","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2026","cvss_score":5.6,"severity":"medium","kev":false,"impact":"Improper access control in driver operations","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; rolling reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24198","https://github.com/NVIDIA/product-security/tree/main/2026/5821"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:L/I:L/A:H","cwe":["CWE-200"],"fleet":{"pain_class":"node-reboot"},"published":"2026-05-26"},{"id":"CVE-2026-50195","cve":"CVE-2026-50195","aliases":[],"title":"containerd: CRI checkpoint import does not validate image references in checkpoint metadata","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2026","cvss_score":5.6,"severity":"medium","kev":false,"impact":"CRI checkpoint import does not validate image references in checkpoint metadata","attack_vector":"Malicious checkpoint image","remediation":"Rolling containerd upgrade with node drain; disable checkpoint import","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-50195"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2026-07-01"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-909","CWE-908"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2011-1044","cve":"CVE-2011-1044","aliases":[],"title":"Linux kernel InfiniBand uverbs drivers/infiniband/core/uverbs_cmd.c - ib_uverbs_poll_cq: Kernel memory disclosure","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel InfiniBand uverbs drivers/infiniband/core/uverbs_cmd.c - ib_uverbs_poll_cq","year":"2011","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Kernel memory disclosure straight into a tenant's address space. The POLL_CQ response buffer is not zeroed before being copied out, so when fewer completions are returned than the buffer holds, the remaining bytes are whatever was previously in that kernel allocation - other tenants' RDMA metadata, pointers useful for defeating KASLR, and residual heap contents. A tenant can poll in a loop and harvest a continuous stream of kernel memory with no crash, no log entry, and no anomalous behaviour to detect. Low severity on paper, high value in practice as the reconnaissance step that makes the heap-corruption bugs on this same device node reliable.","attack_vector":"Local, unprivileged - /dev/infiniband/uverbsN access. Fully silent; nothing in the host's telemetry distinguishes it from normal verbs traffic.","remediation":"Kernel upgrade past 2.6.37 or a vendor backport that memsets the response structure; rolling reboot. There is no runtime mitigation short of revoking uverbs access, because the leak happens in the normal, expected code path rather than an error path - you cannot rate-limit or alert your way out of it.","references":["https://access.redhat.com/security/cve/CVE-2011-1044","https://bugzilla.redhat.com/show_bug.cgi?id=677260","https://nvd.nist.gov/vuln/detail/CVE-2011-1044"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.0/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2016-6327","cve":"CVE-2016-6327","aliases":[],"title":"Linux kernel SRP target drivers/infiniband/ulp/srpt/ib_srpt.c: An SRP initiator that issues an ABORT_TASK against an","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel SRP target drivers/infiniband/ulp/srpt/ib_srpt.c","year":"2016","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An SRP initiator that issues an ABORT_TASK against an in-flight write dereferences NULL in the target and panics the box. In an InfiniBand cluster the SRP target is the shared block-storage head that many compute nodes mount, so a single tenant sending one malformed SCSI task-management command takes storage away from every job on the fabric simultaneously - the highest-fan-out DoS in an IB cluster. Worth cataloguing because the SRP protocol has no initiator authentication beyond fabric membership: if you are on the fabric, you are an authorized initiator.","attack_vector":"An SRP initiator on the fabric - i.e. any node or tenant permitted to mount SRP LUNs. Authorization is fabric membership, not a credential.","remediation":"Kernel upgrade to 4.5.1+ on the storage target nodes, or a vendor backport; requires draining and rebooting the storage head, which for a single-headed SRP target means a storage outage unless the target is clustered. Longer-term control: restrict SRP initiator access with IB partition keys and target-side ACLs so that fabric membership alone does not grant initiator rights.","references":["https://access.redhat.com/security/cve/CVE-2016-6327","https://bugzilla.redhat.com/show_bug.cgi?id=1367293","https://nvd.nist.gov/vuln/detail/CVE-2016-6327"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.0/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-399","CWE-400"],"fleet":{"pain_class":"node-drain"},"id":"CVE-2016-8826","cve":"CVE-2016-8826","aliases":[],"title":"NVIDIA GPU Display Driver kernel mode layer (nvidia.ko on Linux, nvlddmkm.sys on Windows): A tenant can drive the GPU","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver kernel mode layer (nvidia.ko on Linux, nvlddmkm.sys on Windows)","year":"2016","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A tenant can drive the GPU into an interrupt storm and wedge the node. All driver versions were affected. The reason this belongs in an AI-datacenter catalogue despite being 'only' a DoS is the shape of the failure: one tenant's workload saturates the host's interrupt handling and every other job on that machine - including jobs on the other seven GPUs - stalls or dies. There is no per-tenant quota on GPU interrupt generation, so this is a noisy-neighbour attack with node-wide blast radius and no scheduler-level control that limits it.","attack_vector":"Local, unprivileged - any workload with a GPU handle on the node. Works from inside a container with the GPU mapped in.","remediation":"Update the NVIDIA driver to a fixed branch. Requires draining the node so nvidia.ko can be unloaded. Because the practical exposure is availability rather than confidentiality, the compensating control worth building is detection: watch per-node interrupt rates and GPU health counters and attribute spikes back to the owning tenant, so the same tenant cannot repeat it unnoticed across the fleet.","references":["https://www.cve.org/CVERecord?id=CVE-2016-8826","https://nvd.nist.gov/vuln/detail/CVE-2016-8826","https://www.nvidia.com/en-us/security/"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2017-7262","cve":"CVE-2017-7262","aliases":[],"title":"AMD Ryzen with AGESA microcode - FMA3 instruction sequence hang: A long series of FMA3 instructions hangs the system","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Ryzen with AGESA microcode - FMA3 instruction sequence hang","year":"2017","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A long series of FMA3 instructions hangs the system. FMA3 is the fused multiply-add path that every dense linear-algebra kernel hammers, so on an ML host this is not an exotic instruction sequence - it is the normal workload. An unprivileged tenant running a numerical benchmark can hang the node, deliberately or by accident.","attack_vector":"Local, unprivileged. Reachable from any container running numerical code.","remediation":"Fixed by an AGESA microcode update from 2017 onward, delivered as an OEM BIOS package. Client Ryzen silicon rather than EPYC, so on a server fleet this is mostly a non-issue - verify your SKUs before spending a window on it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2017-7262"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2017-03-25"},{"cvss_vector":"CVSS:3.0/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-200"],"fleet":{"pain_class":"node-drain"},"id":"CVE-2018-1723","cve":"CVE-2018-1723","aliases":[],"title":"IBM Spectrum Scale / GPFS node file access path: An unprivileged but authenticated user on a GPFS node reads arbitrary","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"IBM Spectrum Scale / GPFS node file access path","year":"2018","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An unprivileged but authenticated user on a GPFS node reads arbitrary files available to that node, which on a shared cluster includes other tenants' data and any credential material sitting on the node.","attack_vector":"An account on any GPFS node - the exact situation on a multi-user training cluster where researchers get shells on the same login or compute nodes.","remediation":"Upgrade to the fixed Spectrum Scale level in IBM's bulletin (4.1.1.21 / 4.2.3.11 / 5.0.1.3 or later). Also re-check whether shared login nodes should mount the whole filesystem namespace at all, or only per-tenant filesets.","references":["https://www.ibm.com/support/docview.wss?uid=ibm10732713","https://nvd.nist.gov/vuln/detail/CVE-2018-1723"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.0/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","fleet":{"pain_class":"node-drain"},"id":"CVE-2018-1783","cve":"CVE-2018-1783","aliases":[],"title":"IBM GPFS command line utility: Any unprivileged user with a shell on a GPFS node can force GPFS down on that node","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"IBM GPFS command line utility","year":"2018","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Any unprivileged user with a shell on a GPFS node can force GPFS down on that node, cutting every job on it off from the filesystem. On a training cluster that means in-flight runs lose their checkpoint target and die.","attack_vector":"Local account on a GPFS node. No admin rights required - the command line utility itself is the lever.","remediation":"Upgrade to the fixed Spectrum Scale level. As a compensating control, restrict execution of the GPFS admin utilities to an admin group rather than leaving them generally executable on shared nodes.","references":["https://www.ibm.com/support/docview.wss?uid=ibm10732717","https://nvd.nist.gov/vuln/detail/CVE-2018-1783"],"status":"curated"},{"id":"CVE-2018-3639","cve":"CVE-2018-3639","aliases":["Spectre v4","SSB","Speculative Store Bypass"],"title":"Intel processors (speculative store bypass): Spectre v4: a load speculatively executes before an older store","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (speculative store bypass)","year":"2018","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Spectre v4: a load speculatively executes before an older store to the same address is resolved, so it reads stale data. The exposure that matters is inside language runtimes and sandboxes where the attacker supplies the code being JITed - which describes most of a model-serving stack.","attack_vector":"Local code execution, including inside sandboxes and JIT runtimes.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect. The mitigation (SSBD) is opt-in per process on Linux because it costs measurable performance; decide deliberately whether your JIT-hosting workloads get it via prctl or whether you enable it system-wide.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-3639","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00115.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"published":"2018-05-22"},{"id":"CVE-2018-3689","cve":"CVE-2018-3689","aliases":[],"title":"Intel SGX Platform Software for Linux (AESM daemon): A local attacker can disable the AESM daemon","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SGX Platform Software for Linux (AESM daemon)","year":"2018","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A local attacker can disable the AESM daemon, which is the component that services remote attestation. Kill it and enclaves on that node can no longer prove what they are - so attestation-gated workloads stop being admitted.","attack_vector":"Local attacker on the node, no special privilege required.","remediation":"Upgrade SGX PSW for Linux to 2.1.102 or later, and monitor AESM liveness as a first-class signal on confidential-compute nodes. Userspace daemon update and restart.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-3689","https://access.redhat.com/security/cve/CVE-2018-3689"],"status":"curated","published":"2018-04-03"},{"id":"CVE-2018-6252","cve":"CVE-2018-6252","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): The escape handler exposes functionality that should never","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2018","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The escape handler exposes functionality that should never have shipped to production callers; an unprivileged process can reach it and crash the display driver.","attack_vector":"Any local user on the host.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-6252"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2018-04-02"},{"id":"CVE-2018-6253","cve":"CVE-2018-6253","aliases":[],"title":"NVIDIA GPU Display Driver, DirectX and OpenGL user-mode drivers: A crafted pixel shader drives the user-mode driver","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver, DirectX and OpenGL user-mode drivers","year":"2018","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A crafted pixel shader drives the user-mode driver into infinite recursion and kills the rendering process or the session. On a shared render farm or VDI host, one tenant's shader takes out their own session and can pin a core doing it.","attack_vector":"Anyone who can submit shaders - remote session users, tenant VMs, or local processes.","remediation":"Install the fixed GPU Display Driver branch on both Windows and Linux nodes. The kernel component (nvlddmkm.sys / nvidia.ko) cannot be hot-swapped under load, so this is a node drain and reboot per host; restart the container runtime afterwards so mounted driver libraries match the kernel module. No VBIOS or BMC flash.","references":["https://usn.ubuntu.com/3662-1/","https://www.talosintelligence.com/vulnerability_reports/TALOS-2018-0522","https://nvd.nist.gov/vuln/detail/CVE-2018-6253"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2018-04-02"},{"id":"CVE-2018-6260","cve":"CVE-2018-6260","aliases":[],"title":"NVIDIA GPU Display Driver, GPU hardware performance counters: GPU performance counters are readable by any local user","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver, GPU hardware performance counters","year":"2018","cvss_score":5.5,"severity":"medium","kev":false,"impact":"GPU performance counters are readable by any local user and leak enough about another process's GPU activity to reconstruct data it is processing - published work recovered neural-network structure and input images this way. On a shared GPU node, one tenant's job can profile a co-tenant's inference or training work through the counters. NVIDIA's answer was not a code fix but an access-control change, and the restriction was not on by default for a long time in several driver branches, so a large installed base ran exposed for years. Treat any node where profiling is unrestricted as offering no side-channel isolation between GPU contexts.","attack_vector":"Any local user or container that can open the GPU device and run a profiling-capable API (CUPTI, nvprof, Nsight) against the same physical GPU as the victim.","remediation":"Update to a driver branch that ships the profiling restriction, then verify it is actually enforced - this is the step operators skip. On Linux check the nvidia module parameter NVreg_RestrictProfilingToAdminUsers and set it to 1 in /etc/modprobe.d, then reload the module (node drain) or reboot; on Windows apply the driver update and confirm the developer-mode registry setting is not re-enabling profiling. On a multi-tenant node also stop mapping profiling-capable device nodes into tenant containers. Note this is a mitigation, not a fix: MPS and time-sliced sharing still leave the counters shared at the hardware level, so the only hard isolation is one tenant per physical GPU or MIG.","references":["http://support.lenovo.com/us/en/solutions/LEN-26250","https://usn.ubuntu.com/3904-1/","https://nvd.nist.gov/vuln/detail/CVE-2018-6260"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2018-11-13"},{"id":"CVE-2019-0157","cve":"CVE-2019-0157","aliases":[],"title":"Intel SGX driver for Linux: Insufficient input validation in the out-of-tree SGX Linux driver lets a local","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SGX driver for Linux","year":"2019","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Insufficient input validation in the out-of-tree SGX Linux driver lets a local authenticated user deny service. Relevant on hosts still running the legacy Intel SGX DKMS driver rather than the in-kernel driver.","attack_vector":"Local authenticated user with access to the SGX device node.","remediation":"Move to the in-kernel SGX driver where the kernel supports it, otherwise update the Intel SGX DKMS driver. Driver reload or reboot; no firmware or microcode.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-0157","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00235.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2019-06-13"},{"id":"CVE-2019-11090","cve":"CVE-2019-11090","aliases":["TPM-FAIL"],"title":"Intel PTT / fTPM (ECDSA and ECSchnorr timing): The firmware TPM's signing operation leaks nonce information through","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel PTT / fTPM (ECDSA and ECSchnorr timing)","year":"2019","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The firmware TPM's signing operation leaks nonce information through timing, letting an attacker recover the private key after observing a few hundred signatures. Intel PTT is the TPM on a large share of server boards where nobody fitted a discrete chip - so this is the default configuration, not an edge case. Recovered attestation keys let an attacker sign forged quotes and make a compromised node present as measured-clean.","attack_vector":"Local unprivileged user who can request signatures, or a network attacker where the TPM key backs a network-facing service such as a VPN or TLS client certificate. No physical access needed, which is what separated this from earlier TPM side channels.","remediation":"Intel CSME/PTT firmware update, delivered as a BIOS/ME package from the server OEM - per-node flash and reboot. Regenerate and re-enroll every key the fTPM signed with, because the firmware fix does not un-leak an already-extracted key. If attestation is load-bearing for your product, this is the argument for a discrete TPM over the chipset one.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-11090","https://tpm.fail/"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2019-12-18"},{"id":"CVE-2019-19077","cve":"CVE-2019-19077","aliases":[],"title":"Linux bnxt_re RoCE driver (bnxt_re_create_srq memory leak): A tenant can exhaust host memory by repeatedly triggering","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux bnxt_re RoCE driver (bnxt_re_create_srq memory leak)","year":"2019","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A tenant can exhaust host memory by repeatedly triggering shared-receive-queue creation failures through the RDMA verbs interface. This is a straightforward noisy-neighbour denial of service available to any container or VM granted RDMA access, and it needs no privilege beyond opening verbs.","attack_vector":"Any local process with access to the RDMA verbs device — in practice, any tenant container given RDMA.","remediation":"Kernel upgrade plus host reboot. Independently: apply memory cgroup limits to RDMA-capable workloads and restrict verbs device access to workloads that actually need it — both container-runtime config changes, and both good practice regardless of this CVE.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-19077"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2019-11-18"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-732"],"fleet":{"pain_class":"hot-patch"},"id":"CVE-2019-19727","cve":"CVE-2019-19727","aliases":[],"title":"Slurm (slurmdbd.conf file permissions): slurmdbd.conf is installed world-readable, which leaks the accounting","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Slurm (slurmdbd.conf file permissions)","year":"2019","cvss_score":5.5,"severity":"medium","kev":false,"impact":"slurmdbd.conf is installed world-readable, which leaks the accounting database's credentials to every local account on the host. Whoever reads it connects to MySQL as slurmdbd and owns the accounting data directly, bypassing Slurm entirely.","attack_vector":"Any local user on the host running slurmdbd - which on smaller clusters is the same box as slurmctld or even a login node.","remediation":"Upgrade to Slurm 18.08.9 or 19.05.5, then chmod 600 slurmdbd.conf and chown it to the SlurmUser. Rotate the database password too - assume it was readable for the whole time the file sat at the default mode.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-19727","https://bugzilla.suse.com/show_bug.cgi?id=1155784"],"status":"curated"},{"id":"CVE-2019-5671","cve":"CVE-2019-5671","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): Resource leak in the escape handler","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2019","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Resource leak in the escape handler - a local process can exhaust kernel resources and take the display driver, and with it the node's GPUs, out of service.","attack_vector":"Any local user with GPU device access.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["http://support.lenovo.com/us/en/solutions/LEN-26250","https://nvd.nist.gov/vuln/detail/CVE-2019-5671"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2019-02-27"},{"id":"CVE-2019-5677","cve":"CVE-2019-5677","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): Out-of-bounds read in the DeviceIoControl handler","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2019","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Out-of-bounds read in the DeviceIoControl handler - reads past the target buffer and crashes the driver.","attack_vector":"Any local user with GPU device access.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-5677"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2019-05-10"},{"id":"CVE-2019-5686","cve":"CVE-2019-5686","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): The escape handler relies on invariants that are not guaranteed, so","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2019","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The escape handler relies on invariants that are not guaranteed, so a local caller can drive the kernel driver into an unsupported state and crash the node.","attack_vector":"Any local user with GPU device access.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://support.lenovo.com/us/en/product_security/LEN-28096","https://nvd.nist.gov/vuln/detail/CVE-2019-5686"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2019-08-06"},{"id":"CVE-2019-5693","cve":"CVE-2019-5693","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): Uninitialised pointer use in nvlddmkm.sys. Local denial of service","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2019","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Uninitialised pointer use in nvlddmkm.sys. Local denial of service; uninitialised-memory bugs of this shape are often better than their rating suggests.","attack_vector":"Any local user with GPU device access.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-5693"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2019-11-09"},{"id":"CVE-2019-5696","cve":"CVE-2019-5696","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): A guest VM hands the vGPU Manager a wrongly sized buffer and drives the host into an","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2019","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A guest VM hands the vGPU Manager a wrongly sized buffer and drives the host into an out-of-bounds GPU access. One tenant VM can take down the vGPU host, which means every other tenant sharing that physical GPU goes with it.","attack_vector":"Any unprivileged user inside any guest VM assigned a vGPU on the host.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-5696"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2019-11-09"},{"id":"CVE-2020-0543","cve":"CVE-2020-0543","aliases":[],"title":"Intel CPU (SRBDS / CrossTalk): Special Register Buffer Data Sampling","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel CPU (SRBDS / CrossTalk)","year":"2020","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Special Register Buffer Data Sampling - leaks RDRAND/RDSEED output across cores, i.e. across tenants on the same socket","attack_vector":"Any tenant process in a container; tenant VM guest","remediation":"Microcode + reboot; the mitigation serialises RDRAND/RDSEED with a large measured cost on RNG-heavy workloads","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-0543"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"published":"2020-06-15"},{"id":"CVE-2020-0548","cve":"CVE-2020-0548","aliases":["Vector Register Sampling","VRS"],"title":"Intel processors (vector register sampling): Stale values left in vector registers can be sampled by other contexts","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (vector register sampling)","year":"2020","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Stale values left in vector registers can be sampled by other contexts, leaking data across the process and enclave boundary. Lower yield than the fill-buffer attacks but it targets the vector registers - which is where floating-point tensor data lives on an AI node.","attack_vector":"Local code on the same physical core as the victim.","remediation":"Microcode update plus the OS buffer-clearing mitigation, then reboot. Microcode is late-loadable at boot; no OEM BIOS release strictly needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-0548","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00329.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"published":"2020-01-28"},{"id":"CVE-2020-0549","cve":"CVE-2020-0549","aliases":["CacheOut","L1DES","SGAxe"],"title":"Intel processors (L1D eviction sampling) / SGX attestation keys: Stale data can be sampled out of L1D fill buffers","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (L1D eviction sampling) / SGX attestation keys","year":"2020","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Stale data can be sampled out of L1D fill buffers during cache-line eviction, leaking across privilege and enclave boundaries. The consequence operators care about is the SGAxe result built on it: recovery of the machine's SGX attestation keys, which lets an attacker produce quotes that pass Intel's attestation service for a machine whose enclaves are entirely under their control. Once that happens, remote attestation stops proving anything about that platform.","attack_vector":"Local code on the same core as the victim; with SMT enabled, a sibling-thread co-tenant.","remediation":"Microcode update, which is late-loadable at boot without an OEM BIOS release, plus a TCB recovery and re-attestation. Also disable SMT or enforce core scheduling on nodes serving untrusted tenants. Any attestation key material provisioned before the microcode update must be treated as compromised - the fix does not revoke it for you.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-0549","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00329.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"tags":["tenant-isolation"],"published":"2020-01-28"},{"id":"CVE-2020-12966","cve":"CVE-2020-12966","aliases":[],"title":"AMD EPYC SEV-ES / SEV-SNP - information disclosure: An information-disclosure flaw in SEV-ES and SEV-SNP on EPYC lets a","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD EPYC SEV-ES / SEV-SNP - information disclosure","year":"2020","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An information-disclosure flaw in SEV-ES and SEV-SNP on EPYC lets a locally authenticated attacker recover data that the encrypted-state protections were meant to keep opaque. For an operator this is a confidentiality gap in the feature you are charging for when you sell confidential VMs on EPYC.","attack_vector":"Local, authenticated attacker on the host.","remediation":"Fixed in AMD SEV firmware / AGESA and reaches you as an OEM SBIOS package - AMD hands AGESA to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before shipping BIOS. **Budget one to six months of OEM lag**, longer on older platforms and sometimes never on end-of-support SKUs. Applying it means draining the host and doing a full power cycle. Because the fix moves the platform's reported SEV-SNP TCB version, you must also pull fresh VCEK certificates from AMD's Key Distribution Service and update any attestation policy your tenants pin - otherwise guests will start failing launch validation the moment the BIOS lands. Some SEV firmware can alternatively be staged from linux-firmware (amd/amd_sev_*.sbin) and committed via the ccp driver at boot, which is faster than waiting on BIOS - check whether your platform supports firmware hot-load before assuming the OEM is the only route.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-12966","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2022-02-04"},{"id":"CVE-2020-24502","cve":"CVE-2020-24502","aliases":[],"title":"Intel E810 adapter driver for Linux (< 1.0.4): Early E810 Linux driver flaw (improper input validation) reachable","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel E810 adapter driver for Linux (< 1.0.4)","year":"2020","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Early E810 Linux driver flaw (improper input validation) reachable by an authenticated local user. Present on nodes still running the original E810 driver, which happens when the out-of-tree driver was pinned at deployment and never revisited.","attack_vector":"Authenticated local user on the node.","remediation":"Fixed in the Intel out-of-tree ice driver (or the equivalent in-kernel version). Updating the driver requires unloading and reloading the module, which drops every link on that NIC - on a node whose RDMA fabric carries collective traffic, that is a job-killing event, so drain first. If you take it via a distro kernel update instead, it is a reboot. No firmware flash for the driver-side fixes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-24502","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00462.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-02-17"},{"id":"CVE-2020-24503","cve":"CVE-2020-24503","aliases":[],"title":"Intel E810 adapter driver for Linux (< 1.0.4): Early E810 Linux driver flaw (insufficient access control leading","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel E810 adapter driver for Linux (< 1.0.4)","year":"2020","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Early E810 Linux driver flaw (insufficient access control leading to information disclosure) reachable by an authenticated local user. Present on nodes still running the original E810 driver, which happens when the out-of-tree driver was pinned at deployment and never revisited.","attack_vector":"Authenticated local user on the node.","remediation":"Fixed in the Intel out-of-tree ice driver (or the equivalent in-kernel version). Updating the driver requires unloading and reloading the module, which drops every link on that NIC - on a node whose RDMA fabric carries collective traffic, that is a job-killing event, so drain first. If you take it via a distro kernel update instead, it is a reboot. No firmware flash for the driver-side fixes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-24503","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00462.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-02-17"},{"id":"CVE-2020-24504","cve":"CVE-2020-24504","aliases":[],"title":"Intel E810 adapter driver for Linux (< 1.0.4): Early E810 Linux driver flaw (uncontrolled resource consumption)","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel E810 adapter driver for Linux (< 1.0.4)","year":"2020","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Early E810 Linux driver flaw (uncontrolled resource consumption) reachable by an authenticated local user. Present on nodes still running the original E810 driver, which happens when the out-of-tree driver was pinned at deployment and never revisited.","attack_vector":"Authenticated local user on the node.","remediation":"Fixed in the Intel out-of-tree ice driver (or the equivalent in-kernel version). Updating the driver requires unloading and reloading the module, which drops every link on that NIC - on a node whose RDMA fabric carries collective traffic, that is a job-killing event, so drain first. If you take it via a distro kernel update instead, it is a reboot. No firmware flash for the driver-side fixes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-24504","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00462.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-02-17"},{"id":"CVE-2020-25596","cve":"CVE-2020-25596","aliases":[],"title":"Xen - x86 PV guest denial of service via SYSENTER: SYSENTER leaves state sanitisation to software, and on AMD hardware","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen - x86 PV guest denial of service via SYSENTER","year":"2020","cvss_score":5.5,"severity":"medium","kev":false,"impact":"SYSENTER leaves state sanitisation to software, and on AMD hardware Xen's PV guest handling got it wrong, letting a PV guest kernel deny service to itself and destabilise the host path. Legacy PV territory, included for completeness of the Xen-on-AMD picture.","attack_vector":"From inside an x86 PV guest.","remediation":"Fixed in Xen (XSA-339). Hypervisor update plus host reboot. The durable answer is to stop running PV guests - HVM/PVH is the supported path and carries less of this legacy surface.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-25596","https://xenbits.xen.org/xsa/advisory-339.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2020-09-23"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","fleet":{"pain_class":"daemon-restart"},"id":"CVE-2020-4491","cve":"CVE-2020-4491","aliases":[],"title":"IBM Spectrum Scale mmfsd daemon (RPC request handling): A local attacker floods mmfsd with RPC requests and crashes it","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"IBM Spectrum Scale mmfsd daemon (RPC request handling)","year":"2020","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A local attacker floods mmfsd with RPC requests and crashes it, taking the filesystem away from every workload on that node. Repeatable, so it is a persistent denial of storage rather than a one-shot crash.","attack_vector":"Local account on a node running Spectrum Scale 4.2.x up to 4.2.3.22 or 5.0.x up to 5.0.5. No privileges beyond being able to issue RPCs to the local daemon.","remediation":"Upgrade to 4.2.3.23 / 5.0.5.1 or later. There is no configuration workaround that keeps the daemon reachable to legitimate clients while blocking this, so the patch is the fix.","references":["https://www.ibm.com/support/pages/node/6349465","https://nvd.nist.gov/vuln/detail/CVE-2020-4491"],"status":"curated"},{"id":"CVE-2020-5959","cve":"CVE-2020-5959","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): Unvalidated guest-supplied index in the vGPU plugin. A tenant VM crashes the host","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2020","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Unvalidated guest-supplied index in the vGPU plugin. A tenant VM crashes the host plugin and takes out the GPU for every co-tenant.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5959"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2020-03-12"},{"id":"CVE-2020-5960","cve":"CVE-2020-5960","aliases":[],"title":"NVIDIA vGPU Manager kernel module (nvidia.ko, host): NULL dereference in the host-side nvidia.ko under the vGPU","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager kernel module (nvidia.ko, host)","year":"2020","cvss_score":5.5,"severity":"medium","kev":false,"impact":"NULL dereference in the host-side nvidia.ko under the vGPU Manager. A guest can panic the hypervisor host's kernel module, killing every VM sharing that GPU.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5960"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2020-03-12"},{"id":"CVE-2020-5961","cve":"CVE-2020-5961","aliases":[],"title":"NVIDIA vGPU guest graphics driver: Bad cleanup on a failure path in the guest driver crashes the tenant's own VM","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU guest graphics driver","year":"2020","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Bad cleanup on a failure path in the guest driver crashes the tenant's own VM. Self-inflicted blast radius only, but it is a reliable way for a tenant to lose their own workload and blame the platform.","attack_vector":"A user inside the guest VM (or any workload that triggers the failure path).","remediation":"Upgrade the vGPU Manager on the host and the vGPU guest driver inside each tenant VM to the fixed release. Host side is a node drain plus reboot; guest side is a per-VM driver install and reboot. Because the guest driver is inside tenant-controlled VMs, in a multi-tenant estate you cannot fully remediate the guest half yourself - the host-side upgrade is the control you own. No VBIOS flash.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5961"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2020-03-12"},{"id":"CVE-2020-5965","cve":"CVE-2020-5965","aliases":[],"title":"NVIDIA Windows GPU Display Driver, DirectX 11 user-mode driver (nvwgf2um.dll): A crafted shader causes an out-of-bounds","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver, DirectX 11 user-mode driver (nvwgf2um.dll)","year":"2020","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A crafted shader causes an out-of-bounds access in the DX11 user-mode driver and kills the rendering process. Anywhere tenants supply shaders - VDI, cloud gaming, remote rendering - a tenant can knock over their own and neighbouring sessions.","attack_vector":"Anyone who can submit a shader to the host, including from inside a guest VM or a remote session.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5965"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2020-06-25"},{"id":"CVE-2020-5986","cve":"CVE-2020-5986","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): Guest-supplied size not validated in the vGPU plugin","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2020","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Guest-supplied size not validated in the vGPU plugin; a tenant tampers with host state or crashes the shared GPU. vGPU 8.x before 8.5, 10.x before 10.4, and 11.0.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5986"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2020-10-02"},{"id":"CVE-2020-5989","cve":"CVE-2020-5989","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): NULL dereference in the vGPU plugin reachable from a guest - one tenant crashes the","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2020","cvss_score":5.5,"severity":"medium","kev":false,"impact":"NULL dereference in the vGPU plugin reachable from a guest - one tenant crashes the plugin and the shared GPU. vGPU 8.x before 8.5, 10.x before 10.4, and 11.0.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5989"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2020-10-02"},{"id":"CVE-2020-8557","cve":"CVE-2020-8557","aliases":[],"title":"Kubernetes (kubelet): Pod writes to its own /etc/hosts unaccounted for in eviction","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kubelet)","year":"2020","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Pod writes to its own /etc/hosts unaccounted for in eviction; fills node disk","attack_vector":"Any tenant workload","remediation":"Rolling kubelet upgrade with node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-8557"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2020-07-23"},{"id":"CVE-2020-8698","cve":"CVE-2020-8698","aliases":["Fast Store Forwarding Predictor","FSFP"],"title":"Intel processors (fast store forwarding predictor): Improper isolation of a shared microarchitectural resource lets","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (fast store forwarding predictor)","year":"2020","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Improper isolation of a shared microarchitectural resource lets a local authenticated user infer data from another context. Part of the steady drip of shared-predictor leaks that each cost another microcode revision and another small performance tax.","attack_vector":"Local authenticated code on the host.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-8698","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00381"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"published":"2020-11-12"},{"id":"CVE-2021-0145","cve":"CVE-2021-0145","aliases":[],"title":"Intel processors (fast store forwarding predictor initialisation): Improper initialisation of a shared predictor","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (fast store forwarding predictor initialisation)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Improper initialisation of a shared predictor resource allows a local authenticated user to infer data across contexts. Same family as the earlier fast-store-forwarding issue.","attack_vector":"Local authenticated code.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-0145","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00561.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"published":"2022-02-09"},{"id":"CVE-2021-1053","cve":"CVE-2021-1053","aliases":[],"title":"NVIDIA GPU Display Driver (Windows nvlddmkm.sys + Linux nvidia.ko): Improper validation of a user-supplied pointer","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver (Windows nvlddmkm.sys + Linux nvidia.ko)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Improper validation of a user-supplied pointer in the kernel-mode layer. An unprivileged local caller or GPU container crashes the driver and takes every GPU job on the node with it.","attack_vector":"Any local user or GPU container with access to the NVIDIA device nodes.","remediation":"Install the fixed GPU Display Driver branch on both Windows and Linux nodes. The kernel component (nvlddmkm.sys / nvidia.ko) cannot be hot-swapped under load, so this is a node drain and reboot per host; restart the container runtime afterwards so mounted driver libraries match the kernel module. No VBIOS or BMC flash.","references":["https://security.gentoo.org/glsa/202310-02","https://nvd.nist.gov/vuln/detail/CVE-2021-1053"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-01-08"},{"id":"CVE-2021-1054","cve":"CVE-2021-1054","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): Missing authorization check in the escape handler","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing authorization check in the escape handler - an unprivileged caller performs an action the driver should have refused, resulting in denial of service.","attack_vector":"Any local user with GPU device access.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1054"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-01-08"},{"id":"CVE-2021-1066","cve":"CVE-2021-1066","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): Unvalidated guest input causes unbounded resource consumption on the host, so one","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Unvalidated guest input causes unbounded resource consumption on the host, so one tenant starves the vGPU host and denies service to co-tenants. vGPU 8.x before 8.6, 11.0 before 11.3.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1066"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2021-01-08"},{"id":"CVE-2021-1078","cve":"CVE-2021-1078","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): NULL dereference in nvlddmkm.sys leading to a system crash","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"NULL dereference in nvlddmkm.sys leading to a system crash - unprivileged local caller kills the whole node and every job on it.","attack_vector":"Any local user with GPU device access.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1078"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-04-21"},{"id":"CVE-2021-1087","cve":"CVE-2021-1087","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): Information retrievable by a guest that defeats ASLR on the host side. On its own it","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Information retrievable by a guest that defeats ASLR on the host side. On its own it does nothing; combined with any of the memory-corruption bugs in the same bulletin it is what makes a guest-to-host escape reliable instead of a coin flip. vGPU 12.x before 12.2, 11.x before 11.4, 8.x before 8.7.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1087"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2021-04-29"},{"id":"CVE-2021-1095","cve":"CVE-2021-1095","aliases":[],"title":"NVIDIA GPU Display Driver (Windows nvlddmkm.sys + Linux nvidia.ko): All control calls with embedded parameters","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver (Windows nvlddmkm.sys + Linux nvidia.ko)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"All control calls with embedded parameters dereference an untrusted pointer - not one handler but the whole family of them. Unprivileged local caller crashes the driver on Windows or Linux.","attack_vector":"Any local user or GPU container with access to the NVIDIA device nodes.","remediation":"Install the fixed GPU Display Driver branch on both Windows and Linux nodes. The kernel component (nvlddmkm.sys / nvidia.ko) cannot be hot-swapped under load, so this is a node drain and reboot per host; restart the container runtime afterwards so mounted driver libraries match the kernel module. No VBIOS or BMC flash.","references":["https://lists.debian.org/debian-lts-announce/2022/01/msg00013.html","https://security.gentoo.org/glsa/202310-02","https://nvd.nist.gov/vuln/detail/CVE-2021-1095"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-07-22"},{"id":"CVE-2021-1096","cve":"CVE-2021-1096","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): NULL dereference in the escape handler causing a system crash","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"NULL dereference in the escape handler causing a system crash. Any local process with GPU access can reboot the node.","attack_vector":"Any local user with GPU device access on a Windows host.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1096"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-07-22"},{"id":"CVE-2021-1101","cve":"CVE-2021-1101","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): NULL dereference in the vGPU plugin, guest-reachable","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"NULL dereference in the vGPU plugin, guest-reachable; one tenant crashes the shared GPU. vGPU 12.x before 12.3, 11.x before 11.5, 8.x before 8.8.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1101"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2021-07-21"},{"id":"CVE-2021-1102","cve":"CVE-2021-1102","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): Guest input drives the vGPU plugin into a floating-point exception and takes the","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Guest input drives the vGPU plugin into a floating-point exception and takes the shared GPU down. vGPU 12.x before 12.3, 11.x before 11.5, 8.x before 8.8.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1102"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2021-07-21"},{"id":"CVE-2021-1116","cve":"CVE-2021-1116","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): A NULL pointer created in user-mode code is dereferenced","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer created in user-mode code is dereferenced in the kernel, crashing the system. Unprivileged local user takes the node down.","attack_vector":"Any local user with GPU device access.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1116"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-10-27"},{"id":"CVE-2021-1121","cve":"CVE-2021-1121","aliases":[],"title":"NVIDIA vGPU Manager kernel module (nvidia.ko, host): One vGPU can starve the other vGPUs hosted on the same physical","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager kernel module (nvidia.ko, host)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"One vGPU can starve the other vGPUs hosted on the same physical GPU of resources. This is the noisy-neighbour problem as a security bug: a tenant deliberately degrades every co-tenant on the card, and nothing in the vGPU scheduler stops them.","attack_vector":"Any user inside a guest VM sharing a physical GPU with other tenants.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1121"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2021-10-29"},{"id":"CVE-2021-1122","cve":"CVE-2021-1122","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): Guest-reachable NULL dereference in the vGPU plugin causing denial of service across","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Guest-reachable NULL dereference in the vGPU plugin causing denial of service across the shared GPU.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1122"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2021-10-29"},{"id":"CVE-2021-1123","cve":"CVE-2021-1123","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): A guest can deadlock the vGPU plugin. A hung plugin does not crash-and-restart","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A guest can deadlock the vGPU plugin. A hung plugin does not crash-and-restart cleanly - it wedges the GPU for every tenant on it until the host is reset.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1123"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2021-10-29"},{"id":"CVE-2021-26320","cve":"CVE-2021-26320","aliases":[],"title":"AMD SEV firmware - ASK validation in SEND_START: Insufficient validation of the AMD SEV Signing Key in the SEND_START","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV firmware - ASK validation in SEND_START","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Insufficient validation of the AMD SEV Signing Key in the SEND_START command lets a locally authenticated attacker wedge the PSP. SEND_START is part of the guest migration/export flow, so a tenant-triggered migration path can take the secure processor out - and with the PSP down, every confidential guest on the node loses its attestation and key services.","attack_vector":"Local, authenticated. Exercised through the SEV guest export path.","remediation":"Fixed in AMD reference firmware (AGESA / PSP / SEV firmware) and delivered only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo and the ODMs each rebuild and requalify AMD's AGESA drop before shipping. **Expect one to six months of OEM lag**, and on end-of-support platforms expect nothing. Applying it is a drain plus full power cycle, not a driver reload. Verify by reading back the PSP/SMU firmware version afterwards rather than trusting the BIOS version string. This sits inside the SEV-SNP trust boundary, so the update moves the platform's reported TCB version: refresh VCEK certificates from AMD's KDS and update any attestation policy your tenants pin, or confidential guest launches will start failing right after the BIOS lands.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26320","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2021-11-16"},{"id":"CVE-2021-26333","cve":"CVE-2021-26333","aliases":[],"title":"AMD PSP chipset driver - permissive device DACL: The PSP chipset driver's discretionary access control list lets","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD PSP chipset driver - permissive device DACL","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The PSP chipset driver's discretionary access control list lets low-privileged users open a handle to the PSP device. Once open, an unprivileged process can query the driver and read back information it should not have - and, more importantly, it has a handle on the secure processor interface that every other PSP bug then becomes reachable through. Weak device permissions are what turn 'requires privilege' into 'requires a shell'.","attack_vector":"Local, unprivileged - which is the notable part. Worth explicitly checking on any node where you hand semi-trusted workloads local execution.","remediation":"Fixed by updating the AMD PSP chipset driver package and reloading it or rebooting. Driver-speed rather than BIOS-speed, so this is one you can actually close quickly. Audit the permissions on your PSP/ccp device node as part of the same pass - a device that unprivileged users can open is a standing invitation.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26333","https://www.amd.com/en/resources/product-security.html"],"status":"curated","published":"2021-09-21"},{"id":"CVE-2021-26339","cve":"CVE-2021-26339","aliases":[],"title":"AMD CPU core logic - core hang triggered from an unprivileged VM: Specific code executed from an unprivileged VM can","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD CPU core logic - core hang triggered from an unprivileged VM","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Specific code executed from an unprivileged VM can hang an AMD CPU core outright. A guest wedges a physical core on the host, denying it to every other workload scheduled there - and on a GPU node whose CPU cores feed data to accelerators, losing cores starves the GPUs. Availability attack from a tenant against the host, requiring no privilege.","attack_vector":"From inside an unprivileged guest VM. No escalation needed - the guest simply executes a particular sequence.","remediation":"Mitigated by AMD microcode plus, on most of these, a kernel-side change - and the durable delivery vehicle is the OEM SBIOS/AGESA package, which carries **one to six months of OEM lag** and needs a drained node and a full power cycle. The linux-firmware amd-ucode blobs get you the microcode sooner via initramfs early-load and a reboot, but AMD does not support late-loading microcode on a running EPYC host, so either way this is reboot-required, not a live patch.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26339","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2022-05-11"},{"id":"CVE-2021-26349","cve":"CVE-2021-26349","aliases":[],"title":"AMD SEV-SNP migration agent (report ID assignment): An imported SEV-SNP guest is not assigned a fresh report ID, so the","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV-SNP migration agent (report ID assignment)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An imported SEV-SNP guest is not assigned a fresh report ID, so the guest can be tricked into trusting a dishonest Migration Agent. Live migration is the moment a confidential VM is most exposed - it has to hand its state to something - and this lets a malicious host present an MA the guest will accept. The guest then migrates its secrets into an attacker-controlled destination believing it is talking to a legitimate peer.","attack_vector":"Requires a malicious or compromised hypervisor driving guest migration. Only exercised if you actually use SEV-SNP live migration.","remediation":"Fixed in AMD SEV firmware / AGESA and reaches you as an OEM SBIOS package - AMD hands AGESA to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before shipping BIOS. **Budget one to six months of OEM lag**, longer on older platforms and sometimes never on end-of-support SKUs. Applying it means draining the host and doing a full power cycle. Because the fix moves the platform's reported SEV-SNP TCB version, you must also pull fresh VCEK certificates from AMD's Key Distribution Service and update any attestation policy your tenants pin - otherwise guests will start failing launch validation the moment the BIOS lands. Some SEV firmware can alternatively be staged from linux-firmware (amd/amd_sev_*.sbin) and committed via the ccp driver at boot, which is faster than waiting on BIOS - check whether your platform supports firmware hot-load before assuming the OEM is the only route. If you do not offer live migration for confidential VMs, the path is not exercised and this can sit in the normal patch queue. If you do, it is a top-of-queue item, and worth reviewing whether migration should be disabled for confidential tenants until the fleet is patched.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26349","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2022-05-11"},{"id":"CVE-2021-26361","cve":"CVE-2021-26361","aliases":[],"title":"AMD AGESA Boot Loader (ABL) / ASP stage-2 bootloader: A malicious or compromised User Application or AGESA Boot Loader","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD AGESA Boot Loader (ABL) / ASP stage-2 bootloader","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A malicious or compromised User Application or AGESA Boot Loader can exfiltrate arbitrary memory from the ASP stage-2 bootloader. What leaks is firmware memory at the deepest pre-boot stage - the material an attacker needs to build a reliable secure-processor exploit, and potentially key state handled during early boot.","attack_vector":"Local, requires control of a UApp or the ABL itself, so firmware-level access rather than OS-level.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26361","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2022-05-12"},{"id":"CVE-2021-26393","cve":"CVE-2021-26393","aliases":[],"title":"AMD Secure Processor TEE - memory cleanup between trusted applications: The ASP's trusted execution environment fails","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor TEE - memory cleanup between trusted applications","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The ASP's trusted execution environment fails to scrub memory between uses, so an attacker able to get a validly signed trusted application loaded can read or poison residue left by a previous TA. On a platform where the ASP handles fTPM state and SEV key material, that residue is exactly the material you least want leaking sideways.","attack_vector":"Local and privileged: the attacker must be able to produce and load a validly signed trusted application, which normally means a signing-key or supply-chain compromise rather than plain root.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26393","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2022-11-09"},{"id":"CVE-2021-27562","cve":"CVE-2021-27562","aliases":[],"title":"Arm Trusted Firmware-M: Non-secure world can halt the system, overwrite secure data, or leak secure data via the NSPE","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arm Trusted Firmware-M","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Non-secure world can halt the system, overwrite secure data, or leak secure data via the NSPE handler — relevant to BMC SoCs and DPUs built on Arm TrustZone","attack_vector":"Local","remediation":"Firmware update of the affected Arm-based management controller; on a BMC this is again an ODM-gated rebase","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-27562"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2021-05-25"},{"id":"CVE-2021-28689","cve":"CVE-2021-28689","aliases":[],"title":"Xen on x86 - speculative vulnerabilities with bare 32-bit PV guests: Bare (non-shim) 32-bit PV guests run in ring 1, an","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen on x86 - speculative vulnerabilities with bare 32-bit PV guests","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Bare (non-shim) 32-bit PV guests run in ring 1, an arrangement that leaves them exposed to speculative-execution attacks against the hypervisor and other guests. Listed here because the mitigation story on AMD hardware differs from Intel's and gets overlooked - if you still run 32-bit PV guests anywhere, this is a standing cross-guest speculative exposure.","attack_vector":"From inside a 32-bit PV guest.","remediation":"Fixed in Xen (XSA-370) by running 32-bit PV guests under the PV shim rather than bare. Update Xen, switch affected guests to shim mode, and reboot. The durable answer is to stop running 32-bit PV guests at all - on a modern AI fleet there is no reason to have any.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-28689","https://xenbits.xen.org/xsa/advisory-370.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2021-06-11"},{"id":"CVE-2021-33096","cve":"CVE-2021-33096","aliases":["INTEL-SA-00571","CVE-2021-33061"],"title":"Intel 82599 Ethernet Controllers and Adapters - network-on-chip shared-resource isolation: Improper isolation of shared","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel 82599 Ethernet Controllers and Adapters - network-on-chip shared-resource isolation","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Improper isolation of shared resources inside the controller's network-on-chip lets an authenticated user cause denial of service. This is the multi-tenant failure mode that matters for SR-IOV: one function on the adapter - a virtual function assigned to a tenant VM - can exhaust or wedge a resource shared with the physical function and the other VFs. A tenant who is only supposed to control their own VF takes down networking for every other tenant sharing that adapter, and for the host. Intel documented no firmware fix for the 82599 here, which makes it a design limit of the part rather than a bug you close.","attack_vector":"An authenticated local user on any function of the adapter - concretely, a tenant VM that has been assigned an SR-IOV virtual function on a shared 82599.","remediation":"Intel's guidance for the 82599 is mitigation, not a firmware fix: do not share a single adapter's virtual functions across mutually untrusted tenants. Practically that means either dedicating a physical adapter per tenant, moving multi-tenant workloads off 82599-class parts onto controllers with stronger VF isolation, or accepting the shared-fate risk and rate-limiting at the switch. If your multi-tenant story depends on SR-IOV VF isolation, this CVE is the reason to test that assumption on your actual silicon rather than reading it off a datasheet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-33096","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00571.html","https://security.netapp.com/advisory/ntap-20220210-0010/"],"status":"curated"},{"id":"CVE-2021-33135","cve":"CVE-2021-33135","aliases":[],"title":"Intel SGX Linux kernel driver: Uncontrolled resource consumption in the in-kernel SGX driver lets a local authenticated","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SGX Linux kernel driver","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Uncontrolled resource consumption in the in-kernel SGX driver lets a local authenticated user exhaust EPC or driver resources and deny SGX to everyone else on the node.","attack_vector":"Local authenticated user with SGX device access - on a confidential-compute node, any tenant.","remediation":"Kernel update and reboot. Kernel-only, no firmware.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-33135","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00603.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2022-05-12"},{"id":"CVE-2021-4453","cve":"CVE-2021-4453","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): A memory or reference-count leak in the amdgpu power","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu power management (SMU/powerplay). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amd/pm: fix a potential gpu_metrics_table memory leak","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-4453","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-02-26"},{"id":"CVE-2021-47042","cve":"CVE-2021-47042","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A memory or reference-count leak in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu display core (DC/DM). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amd/display: Free local data after use","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-47042","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-02-28"},{"cwe":["CWE-369"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2021-47080","cve":"CVE-2021-47080","aliases":[],"title":"Linux kernel RDMA core (UVERBS_METHOD_QUERY_GID_TABLE): The GID-table query handler used a user-supplied entry size","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel RDMA core (UVERBS_METHOD_QUERY_GID_TABLE)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The GID-table query handler used a user-supplied entry size directly as a divisor. Passing zero divides by zero in kernel context and takes the node down. It is a small bug with an outsized operational cost on a GPU cluster: one unprivileged tenant issuing one ioctl kills a machine that may be holding several other tenants' multi-day training jobs, and the checkpointing loss is the real damage, not the crash.","attack_vector":"Local ioctl on /dev/infiniband/uverbs* by any RDMA-capable tenant. Unprivileged, single call, no race.","remediation":"Kernel update validating user_entry_size before the division. This is the canonical example of why 'just a DoS' scores differently on shared training infrastructure than on a single-tenant box - prioritise it accordingly when triaging.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=54d87913f147a983589923c7f651f97de9af5be1","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2021/CVE-2021-47080.json"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-476","CWE-665"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2021-47140","cve":"CVE-2021-47140","aliases":[],"title":"Linux kernel (drivers/iommu/amd): On AMD hosts, switching a device's IOMMU group between a DMA domain and an identity","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/amd)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"On AMD hosts, switching a device's IOMMU group between a DMA domain and an identity (passthrough) domain left the stale dma-iommu operations installed on the device, so the DMA layer then calls IOMMU helpers against a domain that has none. The node oopses on the next allocation - and the switch in question is precisely the identity/DMA transition an operator performs when preparing or reclaiming a card for passthrough.","attack_vector":"Needs host root: unbind the driver, write to /sys/bus/pci/devices/<bdf>/iommu_group/type, rebind. That is node-provisioning automation, not a tenant surface. Affects AMD-Vi hosts; the equivalent VT-d path was already fixed.","remediation":"Update to a stable kernel carrying commits f3f2cf46 / d6177a65 - this is old enough that every supported distro kernel has it, so treat it as a floor check rather than an action. Interim: reboot the node after changing iommu_group/type instead of rebinding drivers in place.","references":["https://git.kernel.org/stable/c/f3f2cf46291a693eab21adb94171b0128c2a9ec1","https://git.kernel.org/stable/c/d6177a6556f853785867e2ec6d5b7f4906f0d809","https://nvd.nist.gov/vuln/detail/CVE-2021-47140"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2021-47253","cve":"CVE-2021-47253","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A memory or reference-count leak in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu display core (DC/DM). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amd/display: Fix potential memory leak in DMUB hw_init","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-47253","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-21"},{"id":"CVE-2021-47362","cve":"CVE-2021-47362","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): A NULL pointer dereference in the amdgpu power","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/pm: Update intermediate power state for SI","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-47362","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-21"},{"id":"CVE-2021-47410","cve":"CVE-2021-47410","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A race condition or locking defect in the amdkfd (KFD","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdkfd (KFD compute driver, /dev/kfd). Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdkfd: fix svm_migrate_fini warning","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-47410","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-21"},{"id":"CVE-2021-47420","cve":"CVE-2021-47420","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A memory or reference-count leak in the amdkfd (KFD","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdkfd (KFD compute driver, /dev/kfd). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amdkfd: fix a potential ttm->sg memory leak","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-47420","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-21"},{"id":"CVE-2021-47431","cve":"CVE-2021-47431","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A memory or reference-count leak","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu GEM/VM/command-submission ioctl surface. Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amdgpu: fix gart.bo pin_count leak","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-47431","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-21"},{"id":"CVE-2021-47550","cve":"CVE-2021-47550","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amd/amdgpu): A memory or reference-count leak in the amdgpu kernel driver","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amd/amdgpu)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu kernel driver core. Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amd/amdgpu: fix potential memleak","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-47550","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-24"},{"id":"CVE-2021-47658","cve":"CVE-2021-47658","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): A memory or reference-count leak in the amdgpu power","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu power management (SMU/powerplay). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amd/pm: fix a potential gpu_metrics_table memory leak","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-47658","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-02-26"},{"id":"CVE-2022-0171","cve":"CVE-2022-0171","aliases":[],"title":"Linux KVM SEV API - host kernel crash from unprivileged guest creation: A non-root host user-level application can","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux KVM SEV API - host kernel crash from unprivileged guest creation","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A non-root host user-level application can crash the host kernel simply by creating a confidential guest through the KVM SEV API. Denial of service against the whole node from an unprivileged local process - on a GPU host that means every training job on the box dies because somebody with shell access called an ioctl.","attack_vector":"Local, **unprivileged** - the notable part. Only needs access to /dev/kvm, which on many hosts is more widely granted than people assume.","remediation":"Fixed in the Linux kernel. Take the distro kernel update (RHEL/Rocky, Ubuntu, SLES) and reboot the host - no firmware, VBIOS or AGESA step. On a GPU fleet this is a cordon, drain and rolling reboot; plan it as normal kernel maintenance. Also audit who has /dev/kvm on your GPU hosts; if nothing on the node runs VMs, the device should not be world-accessible.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-0171"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2022-08-26"},{"id":"CVE-2022-0563","cve":"CVE-2022-0563","aliases":[],"title":"util-linux (chfn/chsh): Partial disclosure of arbitrary files via libreadline in setuid chfn/chsh","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"util-linux (chfn/chsh)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Partial disclosure of arbitrary files via libreadline in setuid chfn/chsh","attack_vector":"Local user","remediation":"Package update; no reboot","references":["https://access.redhat.com/security/cve/CVE-2022-0563"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2022-02-21"},{"id":"CVE-2022-20235","cve":"CVE-2022-20235","aliases":["PowerVR information page"],"title":"Imagination PowerVR GPU driver - cache subsystem information page: The driver's cache-subsystem information page","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Imagination PowerVR GPU driver - cache subsystem information page","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The driver's cache-subsystem information page, intended to be writable only by the driver, was mapped writable to userspace before DDK 1.18. A tenant process could alter driver cache state. Included as another instance of the GPU-driver-maps-too-much class that recurs across vendors.","attack_vector":"Unprivileged local application holding the GPU device node.","remediation":"Update to Imagination DDK 1.18 or later. Not a datacenter part; treat as vendor-evaluation intelligence rather than a fleet action.","references":["https://source.android.com/security/bulletin/2022-08-01","https://nvd.nist.gov/vuln/detail/CVE-2022-20235"],"status":"curated","published":"2023-01-26"},{"id":"CVE-2022-21125","cve":"CVE-2022-21125","aliases":["SBDS","MMIO Stale Data"],"title":"Intel processors (shared buffers data sampling): Incomplete cleanup of microarchitectural fill buffers lets a local","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (shared buffers data sampling)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Incomplete cleanup of microarchitectural fill buffers lets a local user sample data left behind by other contexts. Part of the MMIO stale-data cluster, whose distinguishing feature is that a guest can pull data across the VM boundary through device MMIO accesses - relevant on any node passing devices through to tenants, which describes every GPU node.","attack_vector":"Local authenticated code, including inside a guest with a passed-through device.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect. The hypervisor half of the mitigation matters here: make sure your VMM's stale-data mitigations are enabled, not just the host microcode. On nodes that host untrusted co-tenants, also disable SMT or enforce core scheduling; that costs real throughput and is a capacity-planning decision, not a free toggle.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-21125","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00615.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"tags":["tenant-isolation"],"published":"2022-06-15"},{"id":"CVE-2022-2153","cve":"CVE-2022-2153","aliases":[],"title":"KVM: NULL pointer dereference in kvm_irq_delivery_to_apic_fast() - guest crashes the host","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"KVM","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"NULL pointer dereference in kvm_irq_delivery_to_apic_fast() - guest crashes the host","attack_vector":"Tenant VM guest","remediation":"Kernel patch; KVM-module scope usually forces drain + reboot rather than livepatch","references":["https://access.redhat.com/security/cve/CVE-2022-2153"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2022-08-31"},{"id":"CVE-2022-21815","cve":"CVE-2022-21815","aliases":[],"title":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys): A null pointer dereference created from user mode","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A null pointer dereference created from user mode inside the kernel driver bluescreens the node. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5312. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-21815","https://github.com/NVIDIA/product-security/tree/main/2022/5312"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-drain"},"published":"2022-02-07"},{"id":"CVE-2022-21816","cve":"CVE-2022-21816","aliases":[],"title":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko): A user inside a guest VM triggers a GPU","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A user inside a guest VM triggers a GPU interrupt storm that lands on the hypervisor host, wedging the physical GPU. The attacker sits inside a tenant guest VM and the blast radius is the hypervisor host, which is running every other tenant's vGPU on the same physical GPU. This is precisely the boundary a vGPU-based multi-tenant offering is sold on. One tenant can take down every other tenant sharing that GPU with no memory-corruption skill required - it is a pure availability attack that any guest user can run.","attack_vector":"A tenant inside their own guest VM, driving the paravirtualised vGPU control interface. Several of these need only an unprivileged process in the guest; the rest need guest root, which a tenant already has on a VM they rented. No host credentials are involved at any point.","remediation":"Patch the vGPU Manager on the hypervisor host per bulletin 5312. Cost: the highest of any class here. The host driver cannot be reloaded while vGPUs are attached, so every tenant VM on that hypervisor must be live-migrated or powered off - a full host drain. NVIDIA also enforces a supported host/guest driver skew, so budget a matching guest-driver campaign in the same window or tenants lose their vGPU on next boot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-21816","https://github.com/NVIDIA/product-security/tree/main/2022/5312"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-284"],"fleet":{"pain_class":"node-drain"},"tags":["tenant-isolation"],"published":"2022-02-07"},{"id":"CVE-2022-26373","cve":"CVE-2022-26373","aliases":["PBRSB","Post-barrier Return Stack Buffer"],"title":"Intel processors (post-barrier return stack buffer): PBRSB: return predictions made after an IBPB barrier can still use","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (post-barrier return stack buffer)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"PBRSB: return predictions made after an IBPB barrier can still use pre-barrier state, so the barrier that hypervisors rely on to separate guests does not fully separate them. The specific worry is a guest reading host memory on a machine where the operator believed IBPB closed that door.","attack_vector":"Local code in a guest or unprivileged context on an affected host.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect. The mitigation also requires a hypervisor/kernel change that stuffs the RSB after VM exit - patch both, and confirm through the spectre_v2 sysfs entry.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-26373","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00706.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"tags":["tenant-isolation"],"published":"2022-08-18"},{"id":"CVE-2022-28187","cve":"CVE-2022-28187","aliases":[],"title":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys): The kernel mode layer fails to release a resource","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The kernel mode layer fails to release a resource after its lifetime ends; a local user leaks it until the node falls over. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5353. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28187","https://github.com/NVIDIA/product-security/tree/main/2022/5353"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-772"],"fleet":{"pain_class":"node-drain"},"published":"2022-05-17"},{"id":"CVE-2022-28188","cve":"CVE-2022-28188","aliases":[],"title":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys): Improper input validation in the DxgkDdiEscape","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Improper input validation in the DxgkDdiEscape handler ends in a node crash. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5353. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28188","https://github.com/NVIDIA/product-security/tree/main/2022/5353"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-20"],"fleet":{"pain_class":"node-drain"},"published":"2022-05-17"},{"id":"CVE-2022-28189","cve":"CVE-2022-28189","aliases":[],"title":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys): A null-pointer dereference in the DxgkDdiEscape","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A null-pointer dereference in the DxgkDdiEscape handler bluescreens the node. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5353. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28189","https://github.com/NVIDIA/product-security/tree/main/2022/5353"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-drain"},"published":"2022-05-17"},{"id":"CVE-2022-28190","cve":"CVE-2022-28190","aliases":[],"title":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys): Improper input validation in the DxgkDdiEscape","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Improper input validation in the DxgkDdiEscape handler gives a local user a denial of service. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5353. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28190","https://github.com/NVIDIA/product-security/tree/main/2022/5353"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-20"],"fleet":{"pain_class":"node-drain"},"published":"2022-05-17"},{"id":"CVE-2022-28191","cve":"CVE-2022-28191","aliases":[],"title":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko): An unprivileged guest user drives","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An unprivileged guest user drives uncontrolled resource consumption in the host vGPU Manager until the host driver stops serving. The attacker sits inside a tenant guest VM and the blast radius is the hypervisor host, which is running every other tenant's vGPU on the same physical GPU. This is precisely the boundary a vGPU-based multi-tenant offering is sold on. This is the classic vGPU noisy-neighbour-as-attack: no exploit engineering, just resource exhaustion crossing from guest to host.","attack_vector":"A tenant inside their own guest VM, driving the paravirtualised vGPU control interface. Several of these need only an unprivileged process in the guest; the rest need guest root, which a tenant already has on a VM they rented. No host credentials are involved at any point.","remediation":"Patch the vGPU Manager on the hypervisor host per bulletin 5353. Cost: the highest of any class here. The host driver cannot be reloaded while vGPUs are attached, so every tenant VM on that hypervisor must be live-migrated or powered off - a full host drain. NVIDIA also enforces a supported host/guest driver skew, so budget a matching guest-driver campaign in the same window or tenants lose their vGPU on next boot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28191","https://github.com/NVIDIA/product-security/tree/main/2022/5353"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-400"],"fleet":{"pain_class":"node-drain"},"tags":["tenant-isolation"],"published":"2022-05-17"},{"id":"CVE-2022-28224","cve":"CVE-2022-28224","aliases":[],"title":"Calico: Route hijacking via the floating IP feature","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Calico","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Route hijacking via the floating IP feature","attack_vector":"Privileged cluster user","remediation":"Rolling Calico upgrade; disable floating IPs for tenants","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28224"],"status":"curated","published":"2022-06-06"},{"id":"CVE-2022-31030","cve":"CVE-2022-31030","aliases":[],"title":"containerd: Unbounded memory consumption in containerd daemon via repeated ExecSync","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Unbounded memory consumption in containerd daemon via repeated ExecSync; node DoS","attack_vector":"Any tenant workload / anyone with kube API exec rights","remediation":"Rolling containerd upgrade with node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31030"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2022-06-09"},{"id":"CVE-2022-31615","cve":"CVE-2022-31615","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): A basic local user triggers a null-pointer dereference","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A basic local user triggers a null-pointer dereference in the kernel mode layer and panics the node. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5383. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31615","https://github.com/NVIDIA/product-security/tree/main/2022/5383"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-drain"},"published":"2022-11-19"},{"id":"CVE-2022-31618","cve":"CVE-2022-31618","aliases":[],"title":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko): A null-pointer dereference in the vGPU","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A null-pointer dereference in the vGPU plugin lets a guest crash the host GPU stack. The attacker sits inside a tenant guest VM and the blast radius is the hypervisor host, which is running every other tenant's vGPU on the same physical GPU. This is precisely the boundary a vGPU-based multi-tenant offering is sold on.","attack_vector":"A tenant inside their own guest VM, driving the paravirtualised vGPU control interface. Several of these need only an unprivileged process in the guest; the rest need guest root, which a tenant already has on a VM they rented. No host credentials are involved at any point.","remediation":"Patch the vGPU Manager on the hypervisor host per bulletin 5383. Cost: the highest of any class here. The host driver cannot be reloaded while vGPUs are attached, so every tenant VM on that hypervisor must be live-migrated or powered off - a full host drain. NVIDIA also enforces a supported host/guest driver skew, so budget a matching guest-driver campaign in the same window or tenants lose their vGPU on next boot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31618","https://github.com/NVIDIA/product-security/tree/main/2022/5383"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-drain"},"tags":["tenant-isolation"],"published":"2022-08-05"},{"id":"CVE-2022-34675","cve":"CVE-2022-34675","aliases":[],"title":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko): The Virtual GPU Manager ignores a","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The Virtual GPU Manager ignores a return value and dereferences null, giving a guest a host-side denial of service. The attacker sits inside a tenant guest VM and the blast radius is the hypervisor host, which is running every other tenant's vGPU on the same physical GPU. This is precisely the boundary a vGPU-based multi-tenant offering is sold on.","attack_vector":"A tenant inside their own guest VM, driving the paravirtualised vGPU control interface. Several of these need only an unprivileged process in the guest; the rest need guest root, which a tenant already has on a VM they rented. No host credentials are involved at any point.","remediation":"Patch the vGPU Manager on the hypervisor host per bulletin 5415. Cost: the highest of any class here. The host driver cannot be reloaded while vGPUs are attached, so every tenant VM on that hypervisor must be live-migrated or powered off - a full host drain. NVIDIA also enforces a supported host/guest driver skew, so budget a matching guest-driver campaign in the same window or tenants lose their vGPU on next boot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34675","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-drain"},"tags":["tenant-isolation"],"published":"2022-12-30"},{"id":"CVE-2022-34677","cve":"CVE-2022-34677","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): An unprivileged user forces an integer truncation","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An unprivileged user forces an integer truncation in the kernel handler, producing a crash or silent data corruption. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5415. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34677","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-125"],"fleet":{"pain_class":"node-drain"},"published":"2022-12-30"},{"id":"CVE-2022-34679","cve":"CVE-2022-34679","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): An unhandled return value leads to a null-pointer","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An unhandled return value leads to a null-pointer dereference in the kernel handler, crashing the node from an unprivileged account. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5415. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34679","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-drain"},"published":"2022-12-30"},{"id":"CVE-2022-34680","cve":"CVE-2022-34680","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): Integer truncation leads to an out-of-bounds read","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Integer truncation leads to an out-of-bounds read in the kernel handler and a node crash. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5415. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34680","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-197"],"fleet":{"pain_class":"node-drain"},"published":"2022-12-30"},{"id":"CVE-2022-34681","cve":"CVE-2022-34681","aliases":[],"title":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys): Improper validation of a display-related data","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Improper validation of a display-related data structure in the kernel handler crashes the node. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5415. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34681","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-20"],"fleet":{"pain_class":"node-drain"},"published":"2022-12-30"},{"id":"CVE-2022-34682","cve":"CVE-2022-34682","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): An unprivileged user triggers a null-pointer dereference","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An unprivileged user triggers a null-pointer dereference in the kernel mode layer and panics the node. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5415. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34682","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-drain"},"published":"2022-12-30"},{"id":"CVE-2022-34683","cve":"CVE-2022-34683","aliases":[],"title":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys): A null-pointer dereference in the DxgkDdiEscape","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A null-pointer dereference in the DxgkDdiEscape handler crashes the node. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5415. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34683","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-drain"},"published":"2022-12-30"},{"id":"CVE-2022-36056","cve":"CVE-2022-36056","aliases":[],"title":"cosign / sigstore: Multiple verify-blob flaws cause successful verification of unsigned or wrongly-signed artifacts","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"cosign / sigstore","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Multiple verify-blob flaws cause successful verification of unsigned or wrongly-signed artifacts","attack_vector":"Malicious artifact","remediation":"Upgrade cosign; re-verify","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-36056"],"status":"curated","published":"2022-09-14"},{"id":"CVE-2022-42266","cve":"CVE-2022-42266","aliases":[],"title":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys): The DxgkDdiEscape handler exposes information","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The DxgkDdiEscape handler exposes information to a caller that should not have it - limited kernel information disclosure to any local user. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5415. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42266","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-200"],"fleet":{"pain_class":"node-drain"},"published":"2022-12-30"},{"id":"CVE-2022-42703","cve":"CVE-2022-42703","aliases":[],"title":"Linux kernel (mm anon_vma): Use-after-free from leaf anon_vma double reuse - memory corruption / privesc primitive","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (mm anon_vma)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Use-after-free from leaf anon_vma double reuse - memory corruption / privesc primitive","attack_vector":"Local user / tenant process","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2022-42703"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2022-10-09"},{"id":"CVE-2022-43309","cve":"CVE-2022-43309","aliases":[],"title":"Supermicro X11SSL-CF hardware revision 1.01, BMC firmware v1.63: A local low-privilege actor gains write access","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro X11SSL-CF hardware revision 1.01, BMC firmware v1.63","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A local low-privilege actor gains write access to something they should not be able to modify. What makes this worth tracking despite the modest score is where it sits: the VRM advisory track covers voltage regulator firmware, and write access to power-delivery firmware on a server is a physical-damage and availability primitive, not just an integrity one. On dense GPU nodes where the VRMs are already running near their limits, an attacker able to tamper with regulator configuration has a plausible path to hardware damage or node-level denial of service that no software remediation reverses. Insecure permissions, disclosed under Supermicro's VRM (voltage regulator module) advisory track rather than the BMC track.","attack_vector":"Local access to the node with low privilege - a user account on the host, not necessarily root. The permissions problem is on the node itself rather than across the management network.","remediation":"Firmware update per Supermicro's January 2023 VRM advisory. VRM firmware is updated separately from BIOS and BMC on Supermicro platforms, which is the operational trap here: an operator who believes they have a fully patched node because BIOS and BMC are current may still be running vulnerable regulator firmware. Add VRM firmware to whatever inventory you use to track BIOS and BMC versions, because most fleet tooling does not enumerate it by default.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-43309","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2022/43xxx/CVE-2022-43309.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-4543","cve":"CVE-2022-4543","aliases":[],"title":"Linux kernel (KASLR): EntryBleed: prefetch side channel defeats KASLR even with KPTI","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (KASLR)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"EntryBleed: prefetch side channel defeats KASLR even with KPTI - enabling primitive for every other kernel exploit","attack_vector":"Any tenant process in a container","remediation":"Kernel patch, drain + reboot. Not independently exploitable, but it removes the main mitigation everything else relies on - treat as a severity multiplier on the whole kernel-privesc set","references":["https://access.redhat.com/security/cve/CVE-2022-4543"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2023-01-11"},{"id":"CVE-2022-48766","cve":"CVE-2022-48766","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Wrap dcn301_calculate_wm_and_dlg for FPU.","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-48766","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-06-20"},{"id":"CVE-2022-48849","cve":"CVE-2022-48849","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): A race condition or locking defect","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdgpu firmware, ACPI and IP-block initialisation. Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdgpu: bypass tiling flag check in virtual display case (v2)","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-48849","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-07-16"},{"id":"CVE-2022-48853","cve":"CVE-2022-48853","aliases":[],"title":"Linux swiotlb - info leak with DMA_FROM_DEVICE bounce buffers: The software IO TLB leaks information through bounce","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux swiotlb - info leak with DMA_FROM_DEVICE bounce buffers","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The software IO TLB leaks information through bounce buffers on DMA_FROM_DEVICE transfers - stale buffer contents are exposed rather than being overwritten by the device. swiotlb is the bounce-buffer layer that SEV and SEV-SNP guests are forced to use for all DMA, because a confidential guest cannot let a device write directly into encrypted memory. So this leak sits precisely on the path every confidential VM's I/O takes, and what leaks is whatever the previous user of that bounce buffer left behind.","attack_vector":"Local, through DMA operations that use bounce buffers - which is all device I/O in an SEV/SNP guest, and any DMA above the device's addressing limit on a normal host.","remediation":"Fixed in the Linux kernel. Distro kernel update plus reboot; no firmware step. Prioritise on confidential-computing hosts and inside confidential guest images, since SEV guests route all I/O through swiotlb by design.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-48853"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-07-16"},{"id":"CVE-2022-49055","cve":"CVE-2022-49055","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A NULL pointer dereference in the amdkfd (KFD compute","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdkfd (KFD compute driver, /dev/kfd). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdkfd: Check for potential null return of kmalloc_array()","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49055","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-02-26"},{"id":"CVE-2022-49069","cve":"CVE-2022-49069","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Fix by adding FPU protection for dcn30_internal_validate_bw","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49069","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-02-26"},{"id":"CVE-2022-49102","cve":"CVE-2022-49102","aliases":[],"title":"habanalabs kernel driver (MMU shadow teardown): A memory leak on the habanalabs MMU teardown path","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"habanalabs kernel driver (MMU shadow teardown)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory leak on the habanalabs MMU teardown path. Each affected teardown leaks host-resident shadow page-table memory, so a tenant that repeatedly opens and closes the device can grind the node into OOM over a long-running shift. Slow-burn availability problem, not a confidentiality one.","attack_vector":"Local user with the habanalabs device node - repeated device open/close cycles are enough.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49102","https://git.kernel.org/stable/c/12e49aefda2e04b07604f13e03f40027cbeb0dc6"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-02-26"},{"id":"CVE-2022-49133","cve":"CVE-2022-49133","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A race condition or locking defect in the amdkfd (KFD","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdkfd (KFD compute driver, /dev/kfd). Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdkfd: svm range restore work deadlock when process exit","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49133","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-02-26"},{"id":"CVE-2022-49135","cve":"CVE-2022-49135","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A memory or reference-count leak in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu display core (DC/DM). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amd/display: Fix memory leak","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49135","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-02-26"},{"id":"CVE-2022-49137","cve":"CVE-2022-49137","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amd/amdgpu/amdgpu_cs): A memory or reference-count","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amd/amdgpu/amdgpu_cs)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu GEM/VM/command-submission ioctl surface. Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amd/amdgpu/amdgpu_cs: fix refcount leak of a dma_fence obj","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49137","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-02-26"},{"cwe":["CWE-401"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-49219","cve":"CVE-2022-49219","aliases":[],"title":"Linux kernel (drivers/vfio/pci): Whoever holds the VFIO device fd for a passed-through PCI function can make the host","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/vfio/pci)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Whoever holds the VFIO device fd for a passed-through PCI function can make the host kernel allocate and never free a saved-PCI-state buffer, once per iteration of a tight loop. The upstream fix explicitly describes this as a malicious sequence that drives the host into OOM, which kills or stalls every other tenant sharing the node, not just the one doing it.","attack_vector":"A tenant VM's VMM (or any container given /dev/vfio/*) that owns a passthrough device: put the device into D3hot, then issue VFIO_DEVICE_RESET or VFIO_DEVICE_PCI_HOT_RESET, repeat. The reset path silently returns the device to D0 so the driver skips the free, and the previously saved state buffer leaks. Only affects devices whose PMCSR lacks the No_Soft_Reset bit, so the driver takes the software power-state-save path. Requires vfio-pci bound to a device and the device node exposed to the tenant.","remediation":"No fixed release is listed in this record; pick up the fix from the linked stable commits and run a current stable/LTS kernel on every node that does PCI passthrough. Interim: do not hand raw VFIO_DEVICE_RESET / VFIO_DEVICE_PCI_HOT_RESET rights to untrusted tenants, cap the VMM process with a memory cgroup so the leak hits the tenant's own limit rather than the node, and alert on unexplained host slab growth on passthrough nodes.","references":["https://git.kernel.org/stable/c/da426ad86027b849b877d4628b277ffbbd2f5325","https://git.kernel.org/stable/c/4319f17fb8264ba39352b611dfa913a4d8c1d1a0","https://nvd.nist.gov/vuln/detail/CVE-2022-49219"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2022-49232","cve":"CVE-2022-49232","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Fix a NULL pointer dereference in amdgpu_dm_connector_add_common_modes()","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49232","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-02-26"},{"id":"CVE-2022-49233","cve":"CVE-2022-49233","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A memory or reference-count leak in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu display core (DC/DM). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amd/display: Call dc_stream_release for remove link enc assignment","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49233","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-02-26"},{"id":"CVE-2022-49294","cve":"CVE-2022-49294","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A division by zero in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A division by zero in the amdgpu display core (DC/DM), reachable with attacker-influenced parameters. The kernel takes a divide fault and the node goes down, taking every co-resident job with it. Upstream fix: drm/amd/display: Check if modulo is 0 before dividing.","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49294","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-02-26"},{"id":"CVE-2022-49335","cve":"CVE-2022-49335","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu/cs): Missing or insufficient validation of","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu/cs)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu GEM/VM/command-submission ioctl surface. A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdgpu/cs: make commands with 0 chunks illegal behaviour.","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49335","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-02-26"},{"id":"CVE-2022-49365","cve":"CVE-2022-49365","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amdgpu): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amdgpu)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: Off by one in dm_dmub_outbox1_low_irq()","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49365","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-02-26"},{"id":"CVE-2022-49529","cve":"CVE-2022-49529","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu/pm): A NULL pointer dereference in the amdgpu power","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu/pm)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu/pm: fix the null pointer while the smu is disabled","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49529","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-02-26"},{"cwe":["CWE-1303"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-49610","cve":"CVE-2022-49610","aliases":[],"title":"Linux kernel (arch/x86/kvm/vmx): Between the point where KVM loads the guest's SPEC_CTRL value and the actual VM entry","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (arch/x86/kvm/vmx)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Between the point where KVM loads the guest's SPEC_CTRL value and the actual VM entry there were returns that could be resolved from a depleted or guest-influenced RSB, giving speculative execution in host context while host branch protections are already relaxed. The payoff for a tenant is the same as the vmexit case: speculative reads of host kernel memory recovered over a side channel.","attack_vector":"Guest-driven on Intel hosts: the tenant primes predictor state and relies on host NMI activity to drain the RSB in the entry window. High complexity and probabilistic, but requires nothing more than an ordinary vCPU - no host privilege, no passthrough device, no VMM cooperation.","remediation":"Patch and reboot into a kernel carrying this fix alongside the vmexit RSB fill (CVE-2022-49611); the two are halves of one mitigation and should be deployed together. No interim runtime control.","references":["https://git.kernel.org/stable/c/afd743f6dde87296c6f3414706964c491bb85862","https://git.kernel.org/stable/c/07853adc29a058c5fd143c14e5ac528448a72ed9","https://nvd.nist.gov/vuln/detail/CVE-2022-49610"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2022-49773","cve":"CVE-2022-49773","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A race condition or locking defect in the amdgpu display","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdgpu display core (DC/DM). Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amd/display: Fix optc2_configure warning on dcn314","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49773","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-05-01"},{"id":"CVE-2022-49784","cve":"CVE-2022-49784","aliases":[],"title":"Linux perf/x86/amd/uncore - memory leak in the events array: Per-CPU northbridge and last-level-cache uncore contexts","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Linux perf/x86/amd/uncore - memory leak in the events array","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Per-CPU northbridge and last-level-cache uncore contexts are allocated but not freed when a CPU comes online, leaking memory on every hotplug. Uncore counters are what you use to measure memory bandwidth and cache behaviour on AMD - i.e. the telemetry an AI operator actually cares about - so this leaks in proportion to how much you monitor.","attack_vector":"Local, driven by CPU hotplug with uncore perf events in use.","remediation":"Distro kernel update plus reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49784"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-05-01"},{"cwe":["CWE-416","CWE-401"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-49829","cve":"CVE-2022-49829","aliases":[],"title":"Linux kernel (drivers/gpu/drm/scheduler): When a process is killed with GPU work still queued, the scheduler entity","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/scheduler)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"When a process is killed with GPU work still queued, the scheduler entity tears down without releasing the dependency fences it holds and without taking a reference on the fence it last scheduled. A tenant that repeatedly starts and kills jobs grows kernel memory on the shared node and exercises an unreferenced fence in the common scheduler that amdgpu and xe both sit on.","attack_vector":"A tenant container holding /dev/dri/renderD* (or /dev/kfd on an AMD node) submits GPU work with outstanding dependencies and kills the process before it completes, in a loop. No privilege beyond the compute/render device node; the affected code is the shared DRM scheduler, so it applies on the mainstream datacenter GPU drivers rather than a niche one.","remediation":"Boot a kernel carrying the drm/sched entity fence-reference fix below. Interim: enforce per-container kernel-memory limits so repeated kill cycles cannot exhaust the node, and watch for unexplained slab growth on nodes running short-lived GPU jobs.","references":["https://git.kernel.org/stable/c/e5f4b38362df93594cb426b04979d8834122f159","https://git.kernel.org/stable/c/b3af84383e7abdc5e63435817bb73a268e7c3637","https://nvd.nist.gov/vuln/detail/CVE-2022-49829"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2022-49864","cve":"CVE-2022-49864","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A NULL pointer dereference in the amdkfd (KFD compute","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdkfd (KFD compute driver, /dev/kfd). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdkfd: Fix NULL pointer dereference in svm_migrate_to_ram()","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49864","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-05-01"},{"id":"CVE-2022-49965","cve":"CVE-2022-49965","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): A memory or reference-count leak in the amdgpu power","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu power management (SMU/powerplay). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amd/pm: add missing ->fini_xxxx interfaces for some SMU13 asics","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49965","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-06-18"},{"id":"CVE-2022-49966","cve":"CVE-2022-49966","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): A memory or reference-count leak in the amdgpu power","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu power management (SMU/powerplay). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amd/pm: add missing ->fini_microcode interface for Sienna Cichlid","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49966","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-06-18"},{"id":"CVE-2022-49971","cve":"CVE-2022-49971","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): A memory or reference-count leak in the amdgpu power","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu power management (SMU/powerplay). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amd/pm: Fix a potential gpu_metrics_table memory leak","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49971","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-06-18"},{"cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-50004","cve":"CVE-2022-50004","aliases":[],"title":"Linux kernel (net/xfrm): Transmitting a packet carrying a metadata dst (no dst","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Transmitting a packet carrying a metadata dst (no dst->dev) through an xfrm interface dereferences a NULL device pointer in the policy lookup and oopses the kernel in the transmit path. On a node running an encrypted tenant overlay this is a crash triggered by ordinary forwarded traffic.","attack_vector":"Requires an xfrm interface in the path plus a packet whose dst is a metadata dst - produced by collect_md tunnels (VXLAN/Geneve in metadata mode) or BPF redirect, which is exactly how per-tenant overlays are wired. A tenant that can get such a packet onto the overlay reaches the oops; no device node or privilege inside the container is needed beyond sending traffic.","remediation":"Boot a kernel carrying the linked stable commits (5.10.140 and later in the 5.10 series). Interim: avoid combining collect_md metadata tunnels with xfrm interfaces on the same path.","references":["https://git.kernel.org/stable/c/2761612bcde9776dd93ce60ce55ef0b7c7329153","https://git.kernel.org/stable/c/96f2758a6d028d1ac08616de9c3c7ff2a122ecf1","https://nvd.nist.gov/vuln/detail/CVE-2022-50004"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-226","CWE-200"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-50037","cve":"CVE-2022-50037","aliases":[],"title":"Linux kernel (drivers/gpu/drm/i915/gt): Compression (CCS) metadata attached to local memory was not cleared when the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/i915/gt)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Compression (CCS) metadata attached to local memory was not cleared when the memory was handed to a new owner, so a tenant receiving recycled VRAM inherits the previous tenant's compression state. The upstream fix is worded exactly that way - the kernel was leaking the CCS state from the previous user - which is residual-data exposure across workloads sharing a discrete Intel GPU.","attack_vector":"A tenant container holding /dev/dri/renderD* on a discrete Intel GPU allocates local memory that a previous tenant freed and reads it back with the compression state still attached. Purely local, unprivileged, no display or profiling access needed; only affects hosts with discrete Intel (lmem-capable) GPUs shared serially between workloads.","remediation":"Boot a kernel carrying the i915 CCS-state fix below. Interim: do not recycle a GPU between tenants without a full device reset/scrub cycle, and prefer whole-device-per-tenant scheduling on discrete Intel cards until patched.","references":["https://git.kernel.org/stable/c/b431cffb4883b9e90d48f0c408674c50fef428a5","https://git.kernel.org/stable/c/232d150fa15606e96c0e01e5c7a2d4e03f621787","https://nvd.nist.gov/vuln/detail/CVE-2022-50037"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-908","CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-50117","cve":"CVE-2022-50117","aliases":[],"title":"Linux kernel (drivers/vfio): VFIO core advertised migration ioctls for devices whose driver never actually initialised","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/vfio)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"VFIO core advertised migration ioctls for devices whose driver never actually initialised its migration state, so a tenant holding the device fd could drive the host into an oops through uninitialised kernel state. On a shared node that is a host crash triggered from inside one tenant's passthrough VM.","attack_vector":"A tenant VMM holding /dev/vfio/* for an mlx5 (ConnectX SR-IOV VF) or HiSilicon accelerator function issues the migration state get/set operations even though the device does not report migration support. VFIO core called straight into the driver op, which touched a state mutex that was never initialised. Reachable purely from ioctls on the device fd, no host privilege needed; conditional on the mlx5-vfio-pci or hisi_acc_vfio_pci variant driver being bound and the VF being assigned to the tenant.","remediation":"No fixed release is listed in this record; take the fix from the linked stable commits and move passthrough nodes to a current stable/LTS kernel. Interim: unbind the mlx5-vfio-pci / hisi_acc_vfio_pci variant drivers and fall back to plain vfio-pci where live migration of the VF is not a product requirement, or block the migration ioctls at the VMM layer.","references":["https://git.kernel.org/stable/c/bba6b12d73d36e0ddbc2c3ac5668a667b00d4345","https://git.kernel.org/stable/c/6e97eba8ad8748fabb795cffc5d9e1a7dcfd7367","https://nvd.nist.gov/vuln/detail/CVE-2022-50117"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-665","CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-50127","cve":"CVE-2022-50127","aliases":[],"title":"Linux kernel (drivers/infiniband/sw/rxe): Any tenant that can open an RDMA verbs device can oops the node. A queue-pair","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/sw/rxe)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Any tenant that can open an RDMA verbs device can oops the node. A queue-pair creation that fails part-way unwinds into cleanup code that grabs a spinlock which was never initialized, killing the kernel thread in an unrecoverable state and taking every other tenant sharing the box down with it.","attack_vector":"A tenant container holding /dev/infiniband/uverbs* issues an ordinary create_qp with attributes that make initialization fail (oversized capabilities, exhausted resources). No privilege beyond the device node is required, and the failure is tenant-controlled so it is repeatable rather than a rare race. Only applies where the soft-RoCE driver rdma_rxe is loaded; hardware HCAs do not run this path.","remediation":"No fixed release is published in this record - pull the stable fix commits (backported across the 5.x series) or run a current stable kernel. Interim: blacklist/unload rdma_rxe if soft-RoCE is not deliberately in use, and stop mapping /dev/infiniband/* into containers that do not need verbs access.","references":["https://git.kernel.org/stable/c/3c838ca6fbdb173102780d7bdf18f2f7d9e30979","https://git.kernel.org/stable/c/1a63f24e724f677db1ab21251f4d0011ae0bb5b5","https://nvd.nist.gov/vuln/detail/CVE-2022-50127"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-400","CWE-674"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-50445","cve":"CVE-2022-50445","aliases":[],"title":"Linux kernel (net/xfrm): Transport-mode IPsec packets were reinjected in the same execution context instead of being","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Transport-mode IPsec packets were reinjected in the same execution context instead of being handed to a workqueue, so a sustained stream of ESP traffic can keep a CPU inside the crypto path long enough for the watchdog to fire (soft lockup, CPU stuck 22s). On a node whose east-west traffic all rides IPsec, that is a stall every tenant on the box feels.","attack_vector":"Driven by traffic, not by a device node. Any peer on the fabric - or any tenant workload generating heavy traffic through an IPsec transport-mode SA - can push the node into the reinject loop. Conditional on transport-mode (not tunnel-mode) xfrm being configured, which is the common shape for node-to-node encryption on a flat cluster fabric.","remediation":"Boot a kernel carrying the linked stable commits. Interim: prefer tunnel mode over transport mode for node-to-node SAs, or rate-limit per-tenant egress so a single workload cannot saturate the ESP path.","references":["https://git.kernel.org/stable/c/7d98b26684cb2390729525b341ea099f0badbe18","https://git.kernel.org/stable/c/f520075da484306bbb8425afd2c42404ba74816f","https://nvd.nist.gov/vuln/detail/CVE-2022-50445"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2022-50479","cve":"CVE-2022-50479","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amd): A memory or reference-count leak in the amdgpu kernel driver core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amd)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu kernel driver core. Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amd: fix potential memory leak","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-50479","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-10-04"},{"id":"CVE-2022-50515","cve":"CVE-2022-50515","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amdgpu): A memory or reference-count leak in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amdgpu)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu display core (DC/DM). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amdgpu: Fix memory leak in hpd_rx_irq_create_workqueue()","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-50515","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-10-07"},{"id":"CVE-2022-50527","cve":"CVE-2022-50527","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): Missing or insufficient validation of user-supplied parameters in","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu kernel driver core. A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdgpu: Fix size validation for non-exclusive domains (v4)","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-50527","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-10-07"},{"id":"CVE-2022-50535","cve":"CVE-2022-50535","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Fix potential null-deref in dm_resume","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-50535","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-10-07"},{"id":"CVE-2023-0188","cve":"CVE-2023-0188","aliases":[],"title":"GPU Display Driver: DoS (stack overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"DoS (stack overflow)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; rolling node reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0188","https://github.com/NVIDIA/product-security/tree/main/2023/5452"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-119"],"fleet":{"pain_class":"node-reboot"},"published":"2023-04-01"},{"id":"CVE-2023-0190","cve":"CVE-2023-0190","aliases":[],"title":"GPU Display Driver: DoS (null deref)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"DoS (null deref)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; rolling node reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0190","https://github.com/NVIDIA/product-security/tree/main/2023/5452"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"published":"2023-04-22"},{"id":"CVE-2023-0197","cve":"CVE-2023-0197","aliases":[],"title":"vGPU Manager (Cloud Gaming): Guest-triggered host DoS (null deref)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU Manager (Cloud Gaming)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Guest-triggered host DoS (null deref)","attack_vector":"Tenant VM guest","remediation":"Upgrade vGPU Manager on hypervisor; migrate/evict guest VMs, reboot host","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0197","https://github.com/NVIDIA/product-security/tree/main/2023/5452"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"published":"2023-04-01"},{"id":"CVE-2023-1018","cve":"CVE-2023-1018","aliases":[],"title":"TPM 2.0 reference implementation: Out-of-bounds read in the same routine — disclosure of TPM-resident data","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"TPM 2.0 reference implementation","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Out-of-bounds read in the same routine — disclosure of TPM-resident data","attack_vector":"Local, low privilege","remediation":"Same TPM firmware update and the same key-loss problem","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-1018"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-02-28"},{"id":"CVE-2023-31021","cve":"CVE-2023-31021","aliases":[],"title":"vGPU software: Guest-triggered host DoS (null deref)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU software","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Guest-triggered host DoS (null deref)","attack_vector":"Tenant VM guest","remediation":"Upgrade vGPU Manager; migrate VMs, reboot host","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5491/5491.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"published":"2023-11-02"},{"id":"CVE-2023-31022","cve":"CVE-2023-31022","aliases":[],"title":"GPU Display Driver / vGPU: DoS (null deref)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver / vGPU","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"DoS (null deref)","attack_vector":"Tenant container or VM guest","remediation":"Driver + vGPU Manager upgrade; rolling reboot","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5491/5491.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"published":"2023-11-02"},{"id":"CVE-2023-31023","cve":"CVE-2023-31023","aliases":[],"title":"GPU Display Driver (Windows): DoS (untrusted pointer deref)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver (Windows)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"DoS (untrusted pointer deref)","attack_vector":"Local user","remediation":"Upgrade Oct-2023 driver branch","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5491/5491.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-822"],"published":"2023-11-02"},{"id":"CVE-2023-34327","cve":"CVE-2023-34327","aliases":[],"title":"Xen on AMD - debug extensions (DBEXT) exposure to guests: AMD CPUs since roughly 2014 carry extensions to x86 debugging","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen on AMD - debug extensions (DBEXT) exposure to guests","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"AMD CPUs since roughly 2014 carry extensions to x86 debugging that Xen exposed to guests without properly managing the associated MSR state across context switches. A guest can use debug facilities to observe or interfere with state belonging to another context - debug hardware is designed to see everything, which is precisely why leaking it across a VM boundary matters.","attack_vector":"From inside a guest VM on AMD hardware under Xen.","remediation":"Fixed in Xen (XSA-444). Update the hypervisor and reboot the host.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-34327","https://xenbits.xen.org/xsa/advisory-444.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-01-05"},{"id":"CVE-2023-34328","cve":"CVE-2023-34328","aliases":[],"title":"Xen on AMD - debug extensions (DBEXT) exposure to guests: Companion to the other XSA-444 debug-extension issue on AMD.","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen on AMD - debug extensions (DBEXT) exposure to guests","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Companion to the other XSA-444 debug-extension issue on AMD. Guest-accessible debug state that is not correctly isolated across context switches.","attack_vector":"From inside a guest VM on AMD hardware under Xen.","remediation":"Fixed in Xen (XSA-444). Hypervisor update plus host reboot; patch both XSA-444 CVEs together.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-34328","https://xenbits.xen.org/xsa/advisory-444.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-01-05"},{"id":"CVE-2023-38575","cve":"CVE-2023-38575","aliases":[],"title":"Intel processors (return predictor target sharing): Return predictor targets are shared non-transparently","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (return predictor target sharing)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Return predictor targets are shared non-transparently between contexts, giving an authorised local user an information-disclosure channel. Fixed in the same microcode wave as the 2024 return-predictor advisories.","attack_vector":"Local authorised code.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-38575","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00982.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"published":"2024-03-14"},{"id":"CVE-2023-40238","cve":"CVE-2023-40238","aliases":["LogoFAIL"],"title":"Insyde InsydeH2O BmpDecoderDxe: Crafted BMP logo copies data to a chosen address during DXE","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O BmpDecoderDxe","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Crafted BMP logo copies data to a chosen address during DXE — arbitrary write before Secure Boot","attack_vector":"Local, ESP write","remediation":"Insyde kernel update shipped through each OEM; the CVSS understates it because the outcome is a firmware implant","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-40238"],"status":"curated","published":"2023-12-07"},{"id":"CVE-2023-40549","cve":"CVE-2023-40549","aliases":["shim 15.8 batch"],"title":"shim (verify_buffer_authenticode): Out-of-bounds read on a malformed PE file crashes shim and blocks boot","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"shim (verify_buffer_authenticode)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Out-of-bounds read on a malformed PE file crashes shim and blocks boot. Same operational cost as the other shim DoS - a node that will not boot needs hands or BMC per box.","attack_vector":"A malformed EFI binary in the boot path.","remediation":"shim package update + reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-40549","https://www.openwall.com/lists/oss-security/2024/01/26/1"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-01-29"},{"id":"CVE-2023-40550","cve":"CVE-2023-40550","aliases":["shim 15.8 batch"],"title":"shim (verify_buffer_sbat): Out-of-bounds read in SBAT verification discloses adjacent boot-time memory to an attacker","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"shim (verify_buffer_sbat)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Out-of-bounds read in SBAT verification discloses adjacent boot-time memory to an attacker who can already run in the boot path. Reconnaissance value rather than direct compromise.","attack_vector":"Crafted SBAT metadata in a binary shim is asked to verify.","remediation":"shim package update + reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-40550","https://www.openwall.com/lists/oss-security/2024/01/26/1"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-01-29"},{"id":"CVE-2023-4211","cve":"CVE-2023-4211","aliases":[],"title":"Arm Mali GPU kernel driver: Use-after-free via improper GPU memory processing","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Arm Mali GPU kernel driver","year":"2023","cvss_score":5.5,"severity":"medium","kev":true,"impact":"Use-after-free via improper GPU memory processing - local non-privileged user reaches freed memory; exploited in the wild [KEV]","attack_vector":"Local user with GPU device access","remediation":"Driver update + reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-4211"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2023-10-01"},{"id":"CVE-2023-4327","cve":"CVE-2023-4327","aliases":["CVE-2023-4328"],"title":"Broadcom LSI Storage Authority (LSA) - on-disk credential/key storage on Linux and Windows: The keys LSA uses","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Broadcom LSI Storage Authority (LSA) - on-disk credential/key storage on Linux and Windows","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The keys LSA uses to encrypt its stored secrets sit in files any local user can read. On a bare-metal node that means anyone who lands non-root code on the host - a container escape, a compromised monitoring agent, a tenant with shell on a shared management jumpbox - recovers the credentials LSA uses to talk to the RAID controller and, in Gateway deployments, to other nodes. That turns a low-privilege foothold into control of the storage controller on the fleet, which is the layer that decides whether the next tenant sees the previous tenant's blocks.","attack_vector":"Any unprivileged local account on a host running the LSA agent (Linux or Windows). No network exposure required, no controller access required first.","remediation":"Upgrade LSA to the fixed 7.017.011.000 build. If your OEM has not shipped it, tighten the file permissions on the LSA install directory and its key material by hand and re-apply after every OEM tooling update, since OEM installers reset them. Rotate any controller/gateway credentials that were stored under the old keys - patching alone does not invalidate what has already leaked. No reboot, no array impact.","references":["https://www.broadcom.com/support/resources/product-security-center","https://nvd.nist.gov/vuln/detail/CVE-2023-4327"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2023-08-15"},{"id":"CVE-2023-52460","cve":"CVE-2023-52460","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Fix NULL pointer dereference at hibernate","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52460","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-02-23"},{"id":"CVE-2023-52485","cve":"CVE-2023-52485","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A race condition or locking defect in the amdgpu display","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdgpu display core (DC/DM). Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amd/display: Wake DMCUB before sending a command","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52485","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-02-29"},{"id":"CVE-2023-52585","cve":"CVE-2023-52585","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): A NULL pointer dereference in the amdgpu RAS / GPU","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu RAS / GPU reset and recovery path. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: Fix possible NULL dereference in amdgpu_ras_query_error_status_helper()","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52585","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-03-06"},{"id":"CVE-2023-52625","cve":"CVE-2023-52625","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Refactor DMCUB enter/exit idle interface","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52625","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-03-26"},{"id":"CVE-2023-52632","cve":"CVE-2023-52632","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A race condition or locking defect in the amdkfd (KFD","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdkfd (KFD compute driver, /dev/kfd). Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdkfd: Fix lock dependency warning with srcu","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52632","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-04-02"},{"id":"CVE-2023-52634","cve":"CVE-2023-52634","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Fix disable_otg_wa logic","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52634","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-04-02"},{"id":"CVE-2023-52659","cve":"CVE-2023-52659","aliases":[],"title":"Linux x86/mm - pfn_to_kaddr() 64-bit input handling (SNP support code): On 64-bit platforms the pfn_to_kaddr() macro","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux x86/mm - pfn_to_kaddr() 64-bit input handling (SNP support code)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"On 64-bit platforms the pfn_to_kaddr() macro dropped high address bits when handed a narrower type, producing wrong kernel addresses in the SEV-SNP support paths that use it. Wrong addresses in code that manages confidential-guest page state means operating on memory that is not the memory intended - a correctness failure right underneath the mechanism enforcing guest isolation.","attack_vector":"Local, in the host kernel's SNP page-management paths.","remediation":"Fixed in the Linux kernel - KVM/x86 SEV code or the ccp/PSP driver. Take the distro kernel update (RHEL/Rocky, Ubuntu, SLES) and **reboot the host**; SEV/SNP hypervisor paths cannot be live-patched in any meaningful way, and SNP platform init/shutdown is not safe to cycle under running guests. Drain confidential-VM tenants, reboot, then re-admit. No firmware, VBIOS or AGESA step needed, which makes this one of the cheaper classes of SEV fix to roll out.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52659"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-17"},{"id":"CVE-2023-52671","cve":"CVE-2023-52671","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Fix hang/underflow when transitioning to ODM4:1","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52671","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-17"},{"id":"CVE-2023-52673","cve":"CVE-2023-52673","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Fix a debugfs null pointer error","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52673","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-17"},{"id":"CVE-2023-52695","cve":"CVE-2023-52695","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Check writeback connectors in create_validate_stream_for_sink","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52695","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-17"},{"id":"CVE-2023-52753","cve":"CVE-2023-52753","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Avoid NULL dereference of timing generator","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52753","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-21"},{"cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-52767","cve":"CVE-2023-52767","aliases":[],"title":"Linux kernel (net/tls): Sendfile() on a kTLS socket whose plaintext and ciphertext buffers are both empty drives the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/tls)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Sendfile() on a kTLS socket whose plaintext and ciphertext buffers are both empty drives the splice EOF path into the BPF-only 'split record' branch and then into record merging, which assumes a populated buffer - a NULL dereference and kernel oops in the transmit path.","attack_vector":"Local and unprivileged: any process with a kTLS socket calling sendfile()/splice() after a send that bailed out and trimmed both buffers. Every tenant container reaches this with ordinary syscalls - no device node, no capability, no cooperating peer. On kernels with panic_on_oops this is a tenant-triggerable node kill.","remediation":"Boot a kernel carrying the linked stable commits. Interim: none at the tenant boundary; review panic_on_oops policy, since it converts a task oops into a full-node outage here.","references":["https://git.kernel.org/stable/c/944900fe2736c07288efe2d9394db4d3ca23f2c9","https://git.kernel.org/stable/c/2214e2bb5489145aba944874d0ee1652a0a63dc8","https://nvd.nist.gov/vuln/detail/CVE-2023-52767"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2023-52773","cve":"CVE-2023-52773","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: fix a NULL pointer dereference in amdgpu_dm_i2c_xfer()","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52773","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-21"},{"id":"CVE-2023-52814","cve":"CVE-2023-52814","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): A NULL pointer dereference in the amdgpu RAS / GPU","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu RAS / GPU reset and recovery path. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: Fix potential null pointer derefernce","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52814","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-21"},{"id":"CVE-2023-52815","cve":"CVE-2023-52815","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu/vkms): A NULL pointer dereference in the amdgpu kernel driver core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu/vkms)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu kernel driver core. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu/vkms: fix a possible null pointer dereference","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52815","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-21"},{"id":"CVE-2023-52817","cve":"CVE-2023-52817","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu): A NULL pointer dereference in the amdgpu power","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: Fix a null pointer access when the smc_rreg pointer is NULL","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52817","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-21"},{"id":"CVE-2023-52908","cve":"CVE-2023-52908","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): A NULL pointer dereference in the amdgpu kernel driver core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu kernel driver core. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: Fix potential NULL dereference","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52908","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-08-21"},{"id":"CVE-2023-52912","cve":"CVE-2023-52912","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A race condition or locking defect","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdgpu GEM/VM/command-submission ioctl surface. Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdgpu: Fixed bug on error when unloading amdgpu","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52912","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-08-21"},{"id":"CVE-2023-53036","cve":"CVE-2023-53036","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A race condition or locking defect","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdgpu GEM/VM/command-submission ioctl surface. Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdgpu: Fix call trace warning and hang when removing amdgpu device","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53036","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-05-02"},{"id":"CVE-2023-53042","cve":"CVE-2023-53042","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Do not set DRR on pipe Commit","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53042","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-05-02"},{"id":"CVE-2023-53073","cve":"CVE-2023-53073","aliases":[],"title":"Linux perf/x86/amd/core - overflow status not cleared for unhandled indices: Unhandled overflow bits are left set","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux perf/x86/amd/core - overflow status not cleared for unhandled indices","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Unhandled overflow bits are left set in the PMU status register, so stale overflow state persists and subsequent interrupts are misattributed. The visible effect is corrupted performance data and spurious NMIs - which on a GPU cluster means your capacity and efficiency measurements are quietly wrong.","attack_vector":"Local, through perf counter overflow handling.","remediation":"Distro kernel update plus reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53073"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-05-02"},{"id":"CVE-2023-53074","cve":"CVE-2023-53074","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): A race condition or locking defect in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdgpu RAS / GPU reset and recovery path. Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdgpu: fix ttm_bo calltrace warning in psp_hw_fini","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53074","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-05-02"},{"id":"CVE-2023-53152","cve":"CVE-2023-53152","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu): A race condition or locking defect in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdgpu power management (SMU/powerplay). Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdgpu: fix calltrace warning in amddrm_buddy_fini","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53152","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-09-15"},{"id":"CVE-2023-53193","cve":"CVE-2023-53193","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A race condition or locking defect","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdgpu GEM/VM/command-submission ioctl surface. Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdgpu: fix amdgpu_irq_put call trace in gmc_v10_0_hw_fini","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53193","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-09-15"},{"id":"CVE-2023-53228","cve":"CVE-2023-53228","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A NULL pointer dereference in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu GEM/VM/command-submission ioctl surface. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: drop redundant sched job cleanup when cs is aborted","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53228","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-09-15"},{"id":"CVE-2023-53237","cve":"CVE-2023-53237","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A race condition or locking defect","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdgpu GEM/VM/command-submission ioctl surface. Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdgpu: fix amdgpu_irq_put call trace in gmc_v11_0_hw_fini","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53237","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-09-15"},{"id":"CVE-2023-53248","cve":"CVE-2023-53248","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): A NULL pointer dereference in the amdgpu kernel driver core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu kernel driver core. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: install stub fence into potential unused fence pointers","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53248","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-09-15"},{"id":"CVE-2023-53258","cve":"CVE-2023-53258","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Fix possible underflow for displays with large vblank","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53258","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-09-15"},{"id":"CVE-2023-53351","cve":"CVE-2023-53351","aliases":[],"title":"Linux kernel DRM scheduler / TTM / dma-buf shared layer used by amdgpu (drm/sched): Memory is handed to a consumer","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel DRM scheduler / TTM / dma-buf shared layer used by amdgpu (drm/sched)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Memory is handed to a consumer without being initialised or cleared in the DRM scheduler / TTM / dma-buf shared layer used by amdgpu. Whatever the previous owner left behind is readable - and on a GPU node the previous owner is very often a different tenant's job. This is the classic residual-data leak between workloads sharing a card: model weights, activations, keys or tokens from the prior tenant can surface in a fresh allocation. Upstream fix: drm/sched: Check scheduler work queue before calling timeout handling","attack_vector":"Local. Reachable by any process that can submit GPU work or import/export a dma-buf - i.e. any ROCm or graphics tenant on the node. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53351","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-09-17"},{"id":"CVE-2023-53352","cve":"CVE-2023-53352","aliases":[],"title":"Linux kernel DRM scheduler / TTM / dma-buf shared layer used by amdgpu (drm/ttm): Missing or insufficient validation of","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel DRM scheduler / TTM / dma-buf shared layer used by amdgpu (drm/ttm)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the DRM scheduler / TTM / dma-buf shared layer used by amdgpu. A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/ttm: check null pointer before accessing when swapping","attack_vector":"Local. Reachable by any process that can submit GPU work or import/export a dma-buf - i.e. any ROCm or graphics tenant on the node. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53352","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-09-17"},{"id":"CVE-2023-53370","cve":"CVE-2023-53370","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): A memory or reference-count leak","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu firmware, ACPI and IP-block initialisation. Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amdgpu: fix memory leak in mes self test","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53370","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-09-18"},{"id":"CVE-2023-53438","cve":"CVE-2023-53438","aliases":[],"title":"Linux x86/MCE - CS register not saved on AMD Zen Instruction Fetch Poison errors: On AMD Zen systems, the Instruction","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux x86/MCE - CS register not saved on AMD Zen Instruction Fetch Poison errors","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"On AMD Zen systems, the Instruction Fetch unit does not report the address of a poisoned (uncorrectable-error) instruction fetch, and the kernel was not saving the CS register either - so when memory corruption hits an instruction fetch, you lose the information needed to attribute it. On a GPU training fleet where uncorrectable memory errors are a routine operational event, losing attribution means you cannot tell which tenant's job hit the bad memory or which DIMM to replace.","attack_vector":"Not attacker-driven. This is an observability failure on the machine-check path.","remediation":"Fixed in the Linux kernel. Distro kernel update plus reboot. Worth taking on any fleet where you do RAS-driven node retirement - without it your machine-check records are missing the field you need to act on.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53438"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-09-18"},{"id":"CVE-2023-53498","cve":"CVE-2023-53498","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Fix potential null dereference","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53498","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-10-01"},{"cwe":["CWE-476","CWE-665"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-53528","cve":"CVE-2023-53528","aliases":[],"title":"Linux kernel (drivers/infiniband/sw/rxe): Soft-RoCE queue-pair cleanup drains send and receive work queues that a","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/sw/rxe)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Soft-RoCE queue-pair cleanup drains send and receive work queues that a partially failed create-QP never allocated, faulting in the kernel. Same shape as the other rxe unwind bug: a tenant that can make create-QP fail gets a reliable node crash, so one container can reboot a box full of other tenants' jobs.","attack_vector":"Local and unprivileged: reached from /dev/infiniband/uverbs* on an rdma_rxe (soft-RoCE) device by issuing a create-QP that fails after partial setup. No hardware adapter, no fabric peer. Conditional on rdma_rxe being loaded.","remediation":"No fixed version is recorded in this entry; boot a stable kernel carrying the drain-queue guard (commits da572f6313ae / d366642b3099). Interim: blacklist rdma_rxe where soft-RoCE is not needed, and remove /dev/infiniband from tenant containers.","references":["https://git.kernel.org/stable/c/da572f6313aeead1f79e0810666bd8d8ffc794d4","https://git.kernel.org/stable/c/d366642b3099bd322375f5b71ba84ab1d586cd6d","https://nvd.nist.gov/vuln/detail/CVE-2023-53528"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2023-53547","cve":"CVE-2023-53547","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): A race condition or locking defect","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdgpu firmware, ACPI and IP-block initialisation. Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdgpu: Fix sdma v4 sw fini error","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53547","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-10-04"},{"id":"CVE-2023-53563","cve":"CVE-2023-53563","aliases":[],"title":"Linux cpufreq/amd-pstate-ut - kernel panic when loading the unit-test driver: Loading the amd-pstate unit-test module","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux cpufreq/amd-pstate-ut - kernel panic when loading the unit-test driver","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Loading the amd-pstate unit-test module panics the kernel. A test module shipped in production kernels that crashes the host when loaded - relevant because automated hardware-validation tooling on a new GPU fleet is exactly the thing that loads every available module to see what happens.","attack_vector":"Local, requires the ability to load the amd_pstate_ut module - so root, or automated burn-in tooling running as root.","remediation":"Distro kernel update plus reboot. Meanwhile, keep amd_pstate_ut out of any module-loading sweep in your node acceptance testing.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53563"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-10-04"},{"id":"CVE-2023-53625","cve":"CVE-2023-53625","aliases":[],"title":"Linux i915 GVT-g mediated GPU virtualisation: Unsafe cleanup of per-vGPU debugfs state when a mediated vGPU is","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux i915 GVT-g mediated GPU virtualisation","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Unsafe cleanup of per-vGPU debugfs state when a mediated vGPU is destroyed. GVT-g is the mechanism that carves one physical Intel GPU into vGPUs handed to different VMs, so bugs in its lifecycle paths sit directly on the tenant boundary. Practical effect here is a host kernel crash triggered by a vGPU teardown.","attack_vector":"Reachable by whoever can cause a vGPU to be created and destroyed - the virtualisation control plane, or a tenant that can start and stop VMs.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53625","https://git.kernel.org/stable/c/44c0e07e3972e3f2609d69ad873d4f342f8a68ec"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-10-07"},{"id":"CVE-2023-53628","cve":"CVE-2023-53628","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A race condition or locking defect","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdgpu GEM/VM/command-submission ioctl surface. Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdgpu: drop gfx_v11_0_cp_ecc_error_irq_funcs","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53628","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-10-07"},{"cwe":["CWE-200","CWE-909"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-53684","cve":"CVE-2023-53684","aliases":[],"title":"Linux kernel (net/xfrm): Structure padding in the xfrm algorithm and encapsulation templates was copied to userspace","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Structure padding in the xfrm algorithm and encapsulation templates was copied to userspace without being zeroed, so every SA dump hands out whatever was in those slab bytes. Small but reliable kernel-memory disclosure from inside a namespace.","attack_vector":"XFRM_MSG_GETSA / policy dumps over xfrm netlink, requiring CAP_NET_ADMIN in the network namespace - which a container granted NET_ADMIN with its own netns has. No fabric access needed; this is a container-to-host information leak.","remediation":"Boot a kernel carrying the linked stable commits. Interim: drop CAP_NET_ADMIN from tenant containers.","references":["https://git.kernel.org/stable/c/0725daaa9a879388ed312110f62dbd5ea2d75f8f","https://git.kernel.org/stable/c/5218af4ad5d8948faac19f71583bcd786c3852df","https://nvd.nist.gov/vuln/detail/CVE-2023-53684"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-53789","cve":"CVE-2023-53789","aliases":[],"title":"Linux kernel (drivers/iommu/amd): The AMD-Vi interrupt thread dereferences a NULL domain while reporting an IOMMU page","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/amd)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The AMD-Vi interrupt thread dereferences a NULL domain while reporting an IOMMU page fault for a device group whose domain was never set up, oopsing the host from inside the fault handler. The faults themselves are generated by devices doing DMA to addresses they are not permitted to touch - exactly what a misbehaving or tenant-programmed device produces.","attack_vector":"An IOMMU page fault raised by a device whose group has no domain configured. A tenant holding a passthrough device can generate IOMMU faults freely by programming bad DMA addresses, but the 'no domain configured' precondition is host-side, so this is not a clean tenant-only path. AMD-Vi (EPYC) hosts; the upstream trace involves an amdgpu device.","remediation":"The record lists no fixed release; boot a kernel carrying the stable fix commits below. Interim control: verify every passthrough group has a domain attached before exposing the device to a tenant.","references":["https://git.kernel.org/stable/c/be8301e2d5a8b95c04ae8e35d7bfee7b0f03f83a","https://git.kernel.org/stable/c/446080b353f048b1fddaec1434cb3d27b5de7efe","https://nvd.nist.gov/vuln/detail/CVE-2023-53789"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-665","CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-54028","cve":"CVE-2023-54028","aliases":[],"title":"Linux kernel (drivers/infiniband/sw/rxe): If soft-RoCE queue-pair creation fails partway, the unwind path runs cleanup","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/sw/rxe)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"If soft-RoCE queue-pair creation fails partway, the unwind path runs cleanup on task state that was never initialized and oopses on an uninitialized spinlock. A tenant that submits a create-QP that is guaranteed to fail can panic the node at will - a repeatable noisy-neighbour outage for everyone sharing it.","attack_vector":"Local and unprivileged: a tenant holding /dev/infiniband/uverbs* on an rdma_rxe (soft-RoCE) device issues a create-QP with parameters that fail after partial initialization. No hardware adapter and no fabric peer needed - rxe runs over any Ethernet NIC. Conditional on the rdma_rxe module being loaded and an rxe device present.","remediation":"No fixed version is recorded in this entry; boot a stable kernel carrying the rxe cleanup-task fix (commits 3236221bb8e4 / c8473cd5b301). Interim: blacklist rdma_rxe on nodes that do not need soft-RoCE, and remove /dev/infiniband from tenant containers.","references":["https://git.kernel.org/stable/c/3236221bb8e4de8e3d0c8385f634064fb26b8e38","https://git.kernel.org/stable/c/c8473cd5b301279a41dc75e5afb26b3d5223b6c7","https://nvd.nist.gov/vuln/detail/CVE-2023-54028"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-476","CWE-457"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-54174","cve":"CVE-2023-54174","aliases":[],"title":"Linux kernel (drivers/vfio): An uninitialized pointer in the VFIO group structure is dereferenced from a group ioctl","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/vfio)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An uninitialized pointer in the VFIO group structure is dereferenced from a group ioctl, oopsing the kernel from inside a tenant's own device-setup path. On a node configured with panic_on_oops that is a reboot for every tenant sharing it; without it the faulting task dies mid-operation holding group state.","attack_vector":"A tenant holding /dev/vfio/<group> calling the group ioctls out of the expected order - unbinding or querying before an iommufd bind has actually succeeded, so group->iommufd was never set. Plain ioctl sequence, no race to win, no host root.","remediation":"The record lists no fixed release; boot a kernel carrying the stable fix commits below. Interim control: drop the /dev/vfio group node from containers that do not require passthrough, and mediate group binding through the host VMM.","references":["https://git.kernel.org/stable/c/8f24eef598ce7cce0bbefe0ec642bcc031d0f528","https://git.kernel.org/stable/c/d649c34cb916b015fdcb487e51409fcc5caeca8d","https://nvd.nist.gov/vuln/detail/CVE-2023-54174"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2023-6548","cve":"CVE-2023-6548","aliases":[],"title":"Citrix NetScaler ADC/Gateway: Code injection on the management interface","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Citrix NetScaler ADC/Gateway","year":"2023","cvss_score":5.5,"severity":"medium","kev":true,"impact":"Code injection on the management interface -> authenticated RCE via NSIP/CLIP/SNIP","attack_vector":"Adjacent network","remediation":"Control-plane: patch; management IPs must never be internet-reachable","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-6548"],"status":"curated","published":"2024-01-17"},{"id":"CVE-2024-0086","cve":"CVE-2024-0086","aliases":[],"title":"vGPU Manager: Host DoS (null deref)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU Manager","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Host DoS (null deref)","attack_vector":"Tenant VM guest","remediation":"vGPU Manager upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0086","https://github.com/NVIDIA/product-security/tree/main/2024/5551"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"published":"2024-06-13"},{"id":"CVE-2024-0088","cve":"CVE-2024-0088","aliases":[],"title":"Triton Inference Server: DoS (stack buffer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"DoS (stack buffer overflow)","attack_vector":"Client of the inference endpoint","remediation":"Upgrade Triton; redeploy serving images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0088","https://github.com/NVIDIA/product-security/tree/main/2024/5535"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:H/UI:N/S:U/C:L/I:L/A:H","cwe":["CWE-119"],"published":"2024-05-14"},{"id":"CVE-2024-0092","cve":"CVE-2024-0092","aliases":[],"title":"GPU Display Driver: DoS","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"DoS","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; rolling reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0092","https://github.com/NVIDIA/product-security/tree/main/2024/5551"],"status":"curated","fleet":{"ubiquity":"Universal - same driver","remediation_pain":"`node-reboot`","pain_class":"node-reboot","why_fleet_wide":"Improper exception handling gives a local tenant a reliable node-level DoS: one customer can knock an 8-GPU box offline for everyone on it"},"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-703"],"published":"2024-06-13"},{"id":"CVE-2024-0094","cve":"CVE-2024-0094","aliases":[],"title":"vGPU Manager: Host resource exhaustion / DoS","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU Manager","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Host resource exhaustion / DoS","attack_vector":"Tenant VM guest","remediation":"vGPU Manager upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0094","https://github.com/NVIDIA/product-security/tree/main/2024/5551"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-799"],"published":"2024-06-13"},{"id":"CVE-2024-0098","cve":"CVE-2024-0098","aliases":[],"title":"ChatRTX: Unencrypted credential storage","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"ChatRTX","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Unencrypted credential storage","attack_vector":"Local user","remediation":"Consumer app; no DC action","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0098","https://github.com/NVIDIA/product-security/tree/main/2024/5533"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-319"],"published":"2024-05-14"},{"id":"CVE-2024-0137","cve":"CVE-2024-0137","aliases":[],"title":"Container Toolkit / GPU Operator: Host DoS / info disclosure","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Container Toolkit / GPU Operator","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Host DoS / info disclosure","attack_vector":"Any tenant with a container","remediation":"Bump toolkit + restart runtime","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0137","https://github.com/NVIDIA/product-security/tree/main/2025/5599"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:L/UI:R/S:C/C:L/I:L/A:L","cwe":["CWE-653"],"published":"2025-01-28"},{"id":"CVE-2024-0147","cve":"CVE-2024-0147","aliases":[],"title":"GPU Display Driver: DoS (use-after-free)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"DoS (use-after-free)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; rolling reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0147","https://github.com/NVIDIA/product-security/tree/main/2025/5614"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"},"published":"2025-01-28"},{"id":"CVE-2024-1441","cve":"CVE-2024-1441","aliases":[],"title":"libvirt: Off-by-one in udevListInterfacesByStatus() - libvirtd crash / info leak","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"libvirt","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Off-by-one in udevListInterfacesByStatus() - libvirtd crash / info leak","attack_vector":"Local user or API client with libvirt access","remediation":"Package update + libvirtd restart; running domains survive","references":["https://access.redhat.com/security/cve/CVE-2024-1441"],"status":"curated","published":"2024-03-11"},{"id":"CVE-2024-26629","cve":"CVE-2024-26629","aliases":[],"title":"Linux nfsd (NFS server): Broken RELEASE_LOCKOWNER handling in nfsd causing state corruption","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux nfsd (NFS server)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Broken RELEASE_LOCKOWNER handling in nfsd causing state corruption","attack_vector":"Local","remediation":"Data-plane: kernel update batched into the next reboot window","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26629"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-03-13"},{"id":"CVE-2024-26647","cve":"CVE-2024-26647","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Fix late derefrence 'dsc' check in 'link_set_dsc_pps_packet()'","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26647","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-03-26"},{"id":"CVE-2024-26648","cve":"CVE-2024-26648","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Fix variable deferencing before NULL check in edp_setup_replay()","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26648","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-03-26"},{"id":"CVE-2024-26649","cve":"CVE-2024-26649","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): A NULL pointer dereference in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu firmware, ACPI and IP-block initialisation. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: Fix the null pointer when load rlc firmware","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26649","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-03-26"},{"id":"CVE-2024-26657","cve":"CVE-2024-26657","aliases":[],"title":"Linux kernel DRM scheduler / TTM / dma-buf shared layer used by amdgpu (drm/sched): A NULL pointer dereference","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel DRM scheduler / TTM / dma-buf shared layer used by amdgpu (drm/sched)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the DRM scheduler / TTM / dma-buf shared layer used by amdgpu. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/sched: fix null-ptr-deref in init entity","attack_vector":"Local. Reachable by any process that can submit GPU work or import/export a dma-buf - i.e. any ROCm or graphics tenant on the node. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26657","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-04-02"},{"id":"CVE-2024-26661","cve":"CVE-2024-26661","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Add NULL test for 'timing generator' in 'dcn21_set_pipe()'","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26661","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-04-02"},{"id":"CVE-2024-26662","cve":"CVE-2024-26662","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Fix 'panel_cntl' could be null in 'dcn21_set_backlight_level()'","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26662","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-04-02"},{"id":"CVE-2024-26700","cve":"CVE-2024-26700","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Fix MST Null Ptr for RV","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26700","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-04-03"},{"id":"CVE-2024-26729","cve":"CVE-2024-26729","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Fix potential null pointer dereference in dc_dmub_srv","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26729","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-04-03"},{"id":"CVE-2024-26767","cve":"CVE-2024-26767","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: fixed integer types and null check locations","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26767","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-04-03"},{"id":"CVE-2024-26817","cve":"CVE-2024-26817","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (amdkfd): An out-of-bounds access in the amdkfd (KFD compute driver","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (amdkfd)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An out-of-bounds access in the amdkfd (KFD compute driver, /dev/kfd) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: amdkfd: use calloc instead of kzalloc to avoid integer overflow","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26817","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-04-13"},{"id":"CVE-2024-26833","cve":"CVE-2024-26833","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A memory or reference-count leak in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu display core (DC/DM). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amd/display: Fix memory leak in dm_sw_fini()","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26833","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-04-17"},{"id":"CVE-2024-26915","cve":"CVE-2024-26915","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): An out-of-bounds access in the amdgpu RAS / GPU","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An out-of-bounds access in the amdgpu RAS / GPU reset and recovery path - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: Reset IH OVERFLOW_CLEAR bit","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26915","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-04-17"},{"id":"CVE-2024-26948","cve":"CVE-2024-26948","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add a dc_state NULL check in dc_state_release","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26948","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-01"},{"id":"CVE-2024-26949","cve":"CVE-2024-26949","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu/pm): A NULL pointer dereference in the amdgpu power","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu/pm)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu/pm: Fix NULL pointer dereference when get power limit","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26949","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-01"},{"cwe":["CWE-362","CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-26984","cve":"CVE-2024-26984","aliases":[],"title":"Linux kernel (drivers/gpu/drm/nouveau/nvkm/subdev/instmem): Concurrent GPU work races the instance-memory pointer","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/nouveau/nvkm/subdev/instmem)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Concurrent GPU work races the instance-memory pointer stores used while walking GPU page tables, and the kernel dereferences a torn pointer. One tenant's ordinary parallel workload oopses the kernel and takes the whole node down with every other tenant on it.","attack_vector":"Unprivileged tenant holding /dev/dri/renderD* on a nouveau device, running many concurrent submissions - a parallel Vulkan conformance-style load reproduced it within hours on stock hardware. No special capabilities, no device misconfiguration required. nouveau only.","remediation":"Update to a stable kernel carrying the fix (no fixed_in published in the record; use the stable commits below). Interim: keep nouveau off shared GPU nodes, or drop /dev/dri from tenant containers on nouveau-driven hosts.","references":["https://git.kernel.org/stable/c/bba8ec5e9b16649d85bc9e9086bf7ae5b5716ff9","https://git.kernel.org/stable/c/1bc4825d4c3ec6abe43cf06c3c39d664d044cbf7","https://nvd.nist.gov/vuln/detail/CVE-2024-26984"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-26986","cve":"CVE-2024-26986","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A memory or reference-count leak in the amdkfd (KFD","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdkfd (KFD compute driver, /dev/kfd). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amdkfd: Fix memory leak in create_process failure","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26986","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-01"},{"id":"CVE-2024-27041","cve":"CVE-2024-27041","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: fix NULL checks for adev->dm.dc in amdgpu_dm_fini()","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-27041","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-01"},{"id":"CVE-2024-27044","cve":"CVE-2024-27044","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Fix potential NULL pointer dereferences in 'dcn10_set_output_transfer_func()'","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-27044","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-01"},{"cwe":["CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-27437","cve":"CVE-2024-27437","aliases":[],"title":"Linux kernel (drivers/vfio/pci): For passthrough devices whose INTx has to be masked at the irqchip, the IRQ is enabled","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/vfio/pci)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"For passthrough devices whose INTx has to be masked at the irqchip, the IRQ is enabled before vfio disables it, so an interrupt arriving in that window double-increments the disable depth. The line stays disabled with no way for the tenant to recover it through vfio, and where that IRQ line is shared with other devices on the host, they stop receiving interrupts too.","attack_vector":"A tenant holding a vfio-pci device fd enabling INTx, with the device asserting its interrupt inside the request_irq window. No host root, and the tenant does not have to win a tight race - it controls when the device asserts. Conditional on a passthrough function that lacks DisINTx support, which is what makes vfio use the exclusive masked-INTx path.","remediation":"The record lists no fixed release; boot a kernel carrying the stable fix commits below. Interim controls: pass through only MSI/MSI-X-capable functions, and avoid assigning devices whose INTx line is shared with host-owned devices.","references":["https://git.kernel.org/stable/c/26389925d6c2126fb777821a0a983adca7ee6351","https://git.kernel.org/stable/c/561d5e1998d58b54ce2bbbb3e843b669aa0b3db5","https://nvd.nist.gov/vuln/detail/CVE-2024-27437"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-31584","cve":"CVE-2024-31584","aliases":[],"title":"PyTorch (flatbuffer loader): Out-of-bounds read parsing flatbuffer model","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"PyTorch (flatbuffer loader)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Out-of-bounds read parsing flatbuffer model","attack_vector":"Customer-supplied model file","remediation":"Ship torch >= 2.2.0","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-31584"],"status":"curated","published":"2024-04-19"},{"cwe":["CWE-667","CWE-833"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-35786","cve":"CVE-2024-35786","aliases":[],"title":"Linux kernel (drivers/gpu/drm/nouveau): Calling the legacy pushbuf submission ioctl on a client that has VM_BIND","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/nouveau)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Calling the legacy pushbuf submission ioctl on a client that has VM_BIND enabled returns an error with the client mutex still held. The tenant's next ioctl and its file close both deadlock, leaving an unkillable task that pins GPU contexts and memory which never return to the pool for other tenants.","attack_vector":"Tenant container holding /dev/dri/renderD* on nouveau: enable VM_BIND on the client, then call DRM_IOCTL_NOUVEAU_GEM_PUSHBUF. One ioctl, deterministic, unprivileged - no race window to win.","remediation":"Update to a kernel carrying the fix (stable commits below; no fixed_in published). Interim: none - both ioctls are on the normal submission interface; drain and reboot nodes with stuck nouveau clients.","references":["https://git.kernel.org/stable/c/c288a61a48ddb77ec097e11ab81b81027cd4e197","https://git.kernel.org/stable/c/b466416bdd6ecbde15ce987226ea633a0268fbb1","https://nvd.nist.gov/vuln/detail/CVE-2024-35786"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-35795","cve":"CVE-2024-35795","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A race condition or locking defect","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdgpu GEM/VM/command-submission ioctl surface. Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdgpu: fix deadlock while reading mqd from debugfs","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-35795","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-17"},{"id":"CVE-2024-35799","cve":"CVE-2024-35799","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Prevent crash when disable stream","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-35799","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-17"},{"cwe":["CWE-787","CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-35841","cve":"CVE-2024-35841","aliases":[],"title":"Linux kernel (net/tls): Splice with MSG_SPLICE_PAGES and MSG_MORE could push more pages into the plaintext scatterlist","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/tls)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Splice with MSG_SPLICE_PAGES and MSG_MORE could push more pages into the plaintext scatterlist than MAX_MSG_FRAGS allows. The code failed to mark the record full, fell through to the continue path and kept trying to add data to an already-full scatterlist - the bounds check catches it with a warning, but this is the zerocopy splice path overrunning its own frag budget on attacker-chosen input sizes.","attack_vector":"Local and unprivileged: any process with a kTLS socket calling splice()/sendfile() with MSG_MORE and more pages than fit in one record. Every tenant container can do this with plain socket and file syscalls - no device node, no capability. On a kernel booted with panic_on_warn this becomes a node-wide outage.","remediation":"Boot a kernel carrying the linked stable commits. Interim: if the fleet runs panic_on_warn, that setting turns this into a tenant-triggerable node kill - weigh disabling it until patched.","references":["https://git.kernel.org/stable/c/02e368eb1444a4af649b73cbe2edd51780511d86","https://git.kernel.org/stable/c/294e7ea85f34748f04e5f3f9dba6f6b911d31aa8","https://nvd.nist.gov/vuln/detail/CVE-2024-35841"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-36026","cve":"CVE-2024-36026","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): A correctness defect in the amdgpu power management","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu power management (SMU/powerplay) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/pm: fixes a random hang in S4 for SMU v13.0.4/11","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36026","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-30"},{"id":"CVE-2024-36284","cve":"CVE-2024-36284","aliases":[],"title":"Intel Neural Compressor: Input-validation failure reachable by an authenticated user, ending in privilege escalation","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel Neural Compressor","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Input-validation failure reachable by an authenticated user, ending in privilege escalation inside the Neural Compressor service.","attack_vector":"Authenticated user of the service.","remediation":"Upgrade to v3.0 or later.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36284","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01219.html"],"status":"curated","published":"2024-11-13"},{"id":"CVE-2024-36316","cve":"CVE-2024-36316","aliases":[],"title":"AMD Graphics Driver - integer overflow bypassing size checks: An integer overflow in the AMD graphics driver lets","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD Graphics Driver - integer overflow bypassing size checks","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An integer overflow in the AMD graphics driver lets an attacker wrap a size calculation and slip past the driver's own bounds checks. The advisory scopes the outcome to denial of service, but overflow-defeats-size-check is the standard front half of a heap corruption chain, so treat the ceiling as higher than the score.","attack_vector":"Local, via the graphics driver interface.","remediation":"Update the AMD graphics driver and reload or reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36316","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-02-11"},{"cwe":["CWE-476","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-36489","cve":"CVE-2024-36489","aliases":[],"title":"Linux kernel (net/tls): Tls_init published the new sk_prot before the TLS context was fully initialized, so a","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/tls)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Tls_init published the new sk_prot before the TLS context was fully initialized, so a concurrent setsockopt/getsockopt on the same socket can see ctx->sk_proto as NULL and dereference it. An unprivileged tenant gets a kernel NULL dereference out of the kTLS crypto-info setsockopt path.","attack_vector":"Local and unprivileged: one thread enables the TLS ULP while another calls setsockopt/getsockopt(SOL_TLS) on the same socket. No device node, no capability. The store-store reordering the race needs is real on weakly-ordered CPUs (arm64 - which is what Grace/GH200-class head nodes are), and much harder to hit on x86.","remediation":"Boot a kernel carrying the linked stable commits. Interim: none at the tenant boundary; on arm64 nodes consider blacklisting the tls module where kTLS is not required.","references":["https://git.kernel.org/stable/c/d72e126e9a36d3d33889829df8fc90100bb0e071","https://git.kernel.org/stable/c/2c260a24cf1c4d30ea3646124f766ee46169280b","https://nvd.nist.gov/vuln/detail/CVE-2024-36489"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-36897","cve":"CVE-2024-36897","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Atom Integrated System Info v2_2 for DCN35","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36897","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-30"},{"id":"CVE-2024-36951","cve":"CVE-2024-36951","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): Missing or insufficient validation of user-supplied","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdkfd (KFD compute driver, /dev/kfd). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdkfd: range check cp bad op exception interrupts","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36951","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-05-30"},{"id":"CVE-2024-36969","cve":"CVE-2024-36969","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A division by zero in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A division by zero in the amdgpu display core (DC/DM), reachable with attacker-influenced parameters. The kernel takes a divide fault and the node goes down, taking every co-resident job with it. Upstream fix: drm/amd/display: Fix division by zero in setup_dsc_config","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36969","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-06-08"},{"cwe":["CWE-833","CWE-667"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-37026","cve":"CVE-2024-37026","aliases":[],"title":"Linux kernel (drivers/gpu/drm/xe): A tenant's jobs can occupy the same copy engines the driver needs to service GPU","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/xe)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A tenant's jobs can occupy the same copy engines the driver needs to service GPU page faults, and the GuC scheduling queue is only two deep. The migration path ends up waiting behind a fault it is itself required to resolve, the GPU deadlocks, and every tenant sharing that device stops making progress until the node is reset.","attack_vector":"Tenant container holding /dev/dri/renderD* on an Intel xe device with recoverable GPU page faults enabled (USM/SVM platforms): submit sustained blitter (BCS) work so user jobs and the migration queue contend for the same engine instances. No capabilities required.","remediation":"Update to a kernel carrying the fix (stable commits below; no fixed_in published). Interim: disable recoverable page faults / USM on shared xe devices, or give each tenant its own device rather than sharing engines.","references":["https://git.kernel.org/stable/c/92deed4a9bfd9ef187764225bba530116c49e15c","https://git.kernel.org/stable/c/c8ea2c31f5ea437199b239d76ad5db27343edb0c","https://nvd.nist.gov/vuln/detail/CVE-2024-37026"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-617","CWE-20"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-39497","cve":"CVE-2024-39497","aliases":[],"title":"Linux kernel (drivers/gpu/drm): Three lines of userspace - mmap a GEM object with PROT_WRITE and MAP_PRIVATE, then","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Three lines of userspace - mmap a GEM object with PROT_WRITE and MAP_PRIVATE, then write to it - hit a BUG_ON and panic the kernel. Any tenant can reboot the node out from under every other tenant on it, with no race and no setup. The upstream fix states this affects every DRM driver using the default shmem helpers.","attack_vector":"Unprivileged process holding /dev/dri/renderD* on any shmem-helper driver. In a GPU fleet the datacenter-relevant instance is virtio-gpu inside a tenant VM, and any node running vkms/panfrost-class drivers. Reproducer is published in the commit message.","remediation":"Update to a stable kernel carrying the fix (commits below; no fixed_in published). Interim: remove /dev/dri from guests and containers that do not need GPU access - there is no way to block the mmap flags from userspace policy.","references":["https://git.kernel.org/stable/c/a508a102edf8735adc9bb73d37dd13c38d1a1b10","https://git.kernel.org/stable/c/3ae63a8c1685e16958560ec08d30defdc5b9cca0","https://nvd.nist.gov/vuln/detail/CVE-2024-39497"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-40987","cve":"CVE-2024-40987","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu): A correctness defect in the amdgpu power management","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu power management (SMU/powerplay) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu: fix UBSAN warning in kv_dpm.c","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-40987","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-07-12"},{"id":"CVE-2024-40997","cve":"CVE-2024-40997","aliases":[],"title":"Linux cpufreq/amd-pstate - memory leak on CPU EPP exit: The amd-pstate driver leaks its per-CPU allocation when a CPU's","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux cpufreq/amd-pstate - memory leak on CPU EPP exit","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The amd-pstate driver leaks its per-CPU allocation when a CPU's energy-performance-preference path exits. Each CPU hotplug or governor transition drops memory on the floor - slow, but on a long-lived host that cycles power states under variable AI load it accumulates into unreclaimable kernel memory and eventually pressures every workload on the node.","attack_vector":"Local, driven by CPU hotplug and power-governor transitions rather than by an attacker directly - though a tenant that can influence CPU frequency governors can accelerate it.","remediation":"Fixed in the Linux kernel. Distro kernel update plus reboot; no firmware step. Group it with the other amd-pstate fixes rather than scheduling a window for it alone.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-40997"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-07-12"},{"cwe":["CWE-457","CWE-908"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-41052","cve":"CVE-2024-41052","aliases":[],"title":"Linux kernel (drivers/vfio/pci): An uninitialized stack variable is used as the device count when a tenant asks vfio","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/vfio/pci)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An uninitialized stack variable is used as the device count when a tenant asks vfio which devices are in its PCI hot-reset group. The count returned is wrong and the caller crashes - and that same group enumeration is the data vfio works from when reasoning about which devices a bus reset would take with it.","attack_vector":"A tenant holding a vfio-pci device fd calling VFIO_DEVICE_GET_PCI_HOT_RESET_INFO. One plain ioctl, no race to win, no host root. Upstream's stated observable is a wrong device count and a userspace crash, not a demonstrated kernel write.","remediation":"Update to 6.6.41 or 6.9.10 or later. Interim control: drop /dev/vfio device nodes from containers that do not need passthrough.","references":["https://git.kernel.org/stable/c/f476dffc52ea70745dcabf63288e770e50ac9ab3","https://git.kernel.org/stable/c/f44136b9652291ac1fc39ca67c053ac624d0d11b","https://nvd.nist.gov/vuln/detail/CVE-2024-41052"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-41093","cve":"CVE-2024-41093","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): A correctness defect in the amdgpu kernel driver core reachable","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu kernel driver core reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu: avoid using null object of framebuffer","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-41093","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-07-29"},{"id":"CVE-2024-42122","cve":"CVE-2024-42122","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add NULL pointer check for kzalloc","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-42122","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-07-30"},{"id":"CVE-2024-43827","cve":"CVE-2024-43827","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add null check before access structs","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-43827","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-08-17"},{"id":"CVE-2024-43886","cve":"CVE-2024-43886","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add null check in resource_log_pipe_topology_update","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-43886","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-08-26"},{"id":"CVE-2024-43899","cve":"CVE-2024-43899","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Fix null pointer deref in dcn20_resource.c","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-43899","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-08-26"},{"id":"CVE-2024-43901","cve":"CVE-2024-43901","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Fix NULL pointer dereference for DTN log in DCN401","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-43901","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-08-26"},{"id":"CVE-2024-43902","cve":"CVE-2024-43902","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add null checker before passing variables","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-43902","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-08-26"},{"id":"CVE-2024-43904","cve":"CVE-2024-43904","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add null checks for 'stream' and 'plane' before dereferencing","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-43904","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-08-26"},{"id":"CVE-2024-43905","cve":"CVE-2024-43905","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): A NULL pointer dereference in the amdgpu power","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/pm: Fix the null pointer dereference for vega10_hwmgr","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-43905","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-08-26"},{"id":"CVE-2024-43907","cve":"CVE-2024-43907","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu/pm): A NULL pointer dereference in the amdgpu power","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu/pm)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu/pm: Fix the null pointer dereference in apply_state_adjust_rules","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-43907","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-08-26"},{"id":"CVE-2024-43908","cve":"CVE-2024-43908","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): A NULL pointer dereference in the amdgpu RAS / GPU","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu RAS / GPU reset and recovery path. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: Fix the null pointer dereference to ras_manager","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-43908","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-08-26"},{"id":"CVE-2024-43909","cve":"CVE-2024-43909","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu/pm): A NULL pointer dereference in the amdgpu power","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu/pm)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu/pm: Fix the null pointer dereference for smu7","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-43909","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-08-26"},{"id":"CVE-2024-44961","cve":"CVE-2024-44961","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): A correctness defect in the amdgpu RAS / GPU reset","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu: Forward soft recovery errors to userspace","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-44961","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-04"},{"id":"CVE-2024-44980","cve":"CVE-2024-44980","aliases":[],"title":"Linux drm/xe GPU kernel driver (display opregion): A resource leak in the xe driver's display opregion handling","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux drm/xe GPU kernel driver (display opregion)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A resource leak in the xe driver's display opregion handling. xe is the newer Intel GPU driver that Data Center GPU Max and later parts move to, so this is worth tracking as the xe driver takes over from i915 in production images.","attack_vector":"Local, on driver load/unload cycles.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-44980","https://git.kernel.org/stable/c/f4b2a0ae1a31fd3d1b5ca18ee08319b479cf9b5f"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-04"},{"id":"CVE-2024-46694","cve":"CVE-2024-46694","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: avoid using null object of framebuffer","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46694","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-13"},{"id":"CVE-2024-46714","cve":"CVE-2024-46714","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Skip wbscl_set_scaler_filter if filter is null","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46714","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-18"},{"id":"CVE-2024-46720","cve":"CVE-2024-46720","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): A NULL pointer dereference in the amdgpu kernel driver core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu kernel driver core. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: fix dereference after null check","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46720","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-18"},{"id":"CVE-2024-46726","cve":"CVE-2024-46726","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Ensure index calculation will not overflow","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46726","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-18"},{"id":"CVE-2024-46727","cve":"CVE-2024-46727","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add otg_master NULL check within resource_log_pipe_topology_update","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46727","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-18"},{"id":"CVE-2024-46728","cve":"CVE-2024-46728","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Check index for aux_rd_interval before using","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46728","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-18"},{"id":"CVE-2024-46732","cve":"CVE-2024-46732","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A division by zero in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A division by zero in the amdgpu display core (DC/DM), reachable with attacker-influenced parameters. The kernel takes a divide fault and the node goes down, taking every co-resident job with it. Upstream fix: drm/amd/display: Assign linear_pitch_alignment even for VM","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46732","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-18"},{"cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-46761","cve":"CVE-2024-46761","aliases":[],"title":"Linux kernel (drivers/pci/hotplug): The hotplug driver disables MSI/MSI-X during slot unregistration after the MSI data","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/pci/hotplug)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The hotplug driver disables MSI/MSI-X during slot unregistration after the MSI data structure has already been freed and NULLed on the disable path. The NULL dereference crashes the host kernel, so hot-unplugging a PCIe switch or bridge takes the whole node down rather than removing one device.","attack_vector":"Triggered by disabling or hot-unplugging a PCIe switch/bridge from the host bridge - an operator sysfs action or a physical removal, neither requiring tenant privilege. PowerNV only (pnv_php on OpenPOWER); inert on x86 and ARM GPU nodes. Relevant only if IBM POWER hosts are in the fleet.","remediation":"Update to a kernel carrying the fix (fixed_in lists only very old branches; use the stable commits below against your running branch). Interim on PowerNV: drain the node before removing bridges or switches.","references":["https://git.kernel.org/stable/c/4eb4085c1346d19d4a05c55246eb93e74e671048","https://git.kernel.org/stable/c/c4c681999d385e28f84808bbf3a85ea8e982da55","https://nvd.nist.gov/vuln/detail/CVE-2024-46761"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-46772","cve":"CVE-2024-46772","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Check denominator crb_pipes before used","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46772","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-18"},{"id":"CVE-2024-46773","cve":"CVE-2024-46773","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Check denominator pbn_div before used","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46773","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-18"},{"id":"CVE-2024-46775","cve":"CVE-2024-46775","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Validate function returns","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46775","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-18"},{"id":"CVE-2024-46776","cve":"CVE-2024-46776","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Run DC_LOG_DC after checking link->link_enc","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46776","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-18"},{"id":"CVE-2024-46778","cve":"CVE-2024-46778","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Check UnboundedRequestEnabled's value","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46778","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-18"},{"id":"CVE-2024-46802","cve":"CVE-2024-46802","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: added NULL check at start of dc_validate_stream","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46802","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-27"},{"id":"CVE-2024-46805","cve":"CVE-2024-46805","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): A NULL pointer dereference in the amdgpu kernel driver core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu kernel driver core. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: fix the waring dereferencing hive","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46805","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-27"},{"id":"CVE-2024-46806","cve":"CVE-2024-46806","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): A correctness defect in the amdgpu kernel driver core reachable","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu kernel driver core reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu: Fix the warning division or modulo by zero","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46806","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-27"},{"id":"CVE-2024-46807","cve":"CVE-2024-46807","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amd/amdgpu): Missing or insufficient validation of user-supplied parameters","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amd/amdgpu)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu kernel driver core. A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/amdgpu: Check tbo resource pointer","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46807","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-09-27"},{"id":"CVE-2024-46808","cve":"CVE-2024-46808","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add missing NULL pointer check within dpcd_extend_address_range","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46808","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-27"},{"id":"CVE-2024-46809","cve":"CVE-2024-46809","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Check BIOS images before it is used","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46809","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-27"},{"id":"CVE-2024-46819","cve":"CVE-2024-46819","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): A correctness defect in the amdgpu RAS / GPU reset","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu: the warning dereferencing obj for nbio_v7_4","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46819","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-27"},{"id":"CVE-2024-46835","cve":"CVE-2024-46835","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): A correctness defect in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu: Fix smatch static checker warning","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46835","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-27"},{"cwe":["CWE-833","CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-46867","cve":"CVE-2024-46867","aliases":[],"title":"Linux kernel (drivers/gpu/drm/xe): Same per-client accounting path, different failure - if the fdinfo read drops the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/xe)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Same per-client accounting path, different failure - if the fdinfo read drops the last reference to a buffer object, destruction wants sleeping locks while the caller holds a spinlock. The result is a hard deadlock (and sleep-in-atomic) with driver locks held, so the xe device stops serving every tenant on it, not just the one that triggered it.","attack_vector":"Tenant holding /dev/dri/renderD* on Intel xe arranges for an fdinfo read to drop the final BO reference - achievable by freeing BOs from one thread while reading /proc/<pid>/fdinfo/<drmfd> from another. Reachable by the tenant itself and by any monitoring agent reading tenant fdinfo.","remediation":"Update to a kernel with the fix (stable commits below; no fixed_in published). Interim: disable DRM fdinfo scraping in your GPU telemetry stack on xe nodes.","references":["https://git.kernel.org/stable/c/9d3de463e23bfb1ff1567a32b099b1b3e5286a48","https://git.kernel.org/stable/c/9bd7ff293fc84792514aeafa06c5a17f05cb5f4b","https://nvd.nist.gov/vuln/detail/CVE-2024-46867"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-46896","cve":"CVE-2024-46896","aliases":[],"title":"Linux kernel DRM scheduler / TTM / dma-buf shared layer used by amdgpu (drm/amdgpu): A NULL pointer dereference","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel DRM scheduler / TTM / dma-buf shared layer used by amdgpu (drm/amdgpu)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the DRM scheduler / TTM / dma-buf shared layer used by amdgpu. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: don't access invalid sched","attack_vector":"Local. Reachable by any process that can submit GPU work or import/export a dma-buf - i.e. any ROCm or graphics tenant on the node. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46896","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-01-11"},{"id":"CVE-2024-47661","cve":"CVE-2024-47661","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Avoid overflow from uint32_t to uint8_t","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-47661","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-09"},{"id":"CVE-2024-47662","cve":"CVE-2024-47662","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Remove register from DCN35 DMCUB diagnostic collection","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-47662","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-09"},{"id":"CVE-2024-47683","cve":"CVE-2024-47683","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Skip Recompute DSC Params if no Stream on Link","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-47683","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"id":"CVE-2024-47704","cve":"CVE-2024-47704","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Check link_res->hpo_dp_link_enc before using it","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-47704","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"id":"CVE-2024-47720","cve":"CVE-2024-47720","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add null check for set_output_gamma in dcn30_set_output_transfer_func","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-47720","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"cwe":["CWE-833","CWE-667"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-47729","cve":"CVE-2024-47729","aliases":[],"title":"Linux kernel (drivers/gpu/drm/xe): User VM_BIND work is scheduled onto engines that can themselves take page faults","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/xe)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"User VM_BIND work is scheduled onto engines that can themselves take page faults, and resolving those faults depends on the bind completing. A tenant issuing binds on a fault-capable device deadlocks the GPU, and the device stops serving every other tenant on it until the node is reset.","attack_vector":"Tenant container holding /dev/dri/renderD* on an Intel xe device with recoverable page faults enabled: issue VM_BIND operations against a faulting VM. No capabilities required. Not reachable on xe platforms without recoverable faults.","remediation":"Update to a kernel carrying the fix (stable commits below; no fixed_in published). Interim: disable recoverable page faults / SVM on shared xe devices.","references":["https://git.kernel.org/stable/c/439fc1e569c57669dbb842d0a77c7ba0a82a9f5d","https://git.kernel.org/stable/c/852856e3b6f679c694dd5ec41e5a3c11aa46640b","https://nvd.nist.gov/vuln/detail/CVE-2024-47729"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-49890","cve":"CVE-2024-49890","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): A NULL pointer dereference in the amdgpu power","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/pm: ensure the fw_info is not null before using it","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49890","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"id":"CVE-2024-49892","cve":"CVE-2024-49892","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A division by zero in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A division by zero in the amdgpu display core (DC/DM), reachable with attacker-influenced parameters. The kernel takes a divide fault and the node goes down, taking every co-resident job with it. Upstream fix: drm/amd/display: Initialize get_bytes_per_element's default to 1","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49892","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"id":"CVE-2024-49893","cve":"CVE-2024-49893","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Check stream_status before it is used","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49893","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"id":"CVE-2024-49896","cve":"CVE-2024-49896","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Check stream before comparing them","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49896","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"id":"CVE-2024-49897","cve":"CVE-2024-49897","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Check phantom_stream before it is used","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49897","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"id":"CVE-2024-49898","cve":"CVE-2024-49898","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Check null-initialized variables","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49898","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"id":"CVE-2024-49899","cve":"CVE-2024-49899","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A division by zero in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A division by zero in the amdgpu display core (DC/DM), reachable with attacker-influenced parameters. The kernel takes a divide fault and the node goes down, taking every co-resident job with it. Upstream fix: drm/amd/display: Initialize denominators' default to 1","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49899","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"id":"CVE-2024-49904","cve":"CVE-2024-49904","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): A NULL pointer dereference in the amdgpu kernel driver core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu kernel driver core. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: add list empty check to avoid null pointer issue","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49904","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"id":"CVE-2024-49905","cve":"CVE-2024-49905","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add null check for 'afb' in amdgpu_dm_plane_handle_cursor_update (v2)","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49905","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"id":"CVE-2024-49906","cve":"CVE-2024-49906","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Check null pointer before try to access it","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49906","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"id":"CVE-2024-49907","cve":"CVE-2024-49907","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Check null pointers before using dc->clk_mgr","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49907","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"id":"CVE-2024-49908","cve":"CVE-2024-49908","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add null check for 'afb' in amdgpu_dm_update_cursor (v2)","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49908","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"id":"CVE-2024-49909","cve":"CVE-2024-49909","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add NULL check for function pointer in dcn32_set_output_transfer_func","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49909","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"id":"CVE-2024-49910","cve":"CVE-2024-49910","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add NULL check for function pointer in dcn401_set_output_transfer_func","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49910","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"id":"CVE-2024-49911","cve":"CVE-2024-49911","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add NULL check for function pointer in dcn20_set_output_transfer_func","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49911","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"id":"CVE-2024-49912","cve":"CVE-2024-49912","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Handle null 'stream_status' in 'planes_changed_for_existing_stream'","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49912","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"id":"CVE-2024-49913","cve":"CVE-2024-49913","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add null check for top_pipe_to_program in commit_planes_for_stream","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49913","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"id":"CVE-2024-49914","cve":"CVE-2024-49914","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Add null check for pipe_ctx->plane_state in dcn20_program_pipe","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49914","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"id":"CVE-2024-49915","cve":"CVE-2024-49915","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add NULL check for clk_mgr in dcn32_init_hw","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49915","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"id":"CVE-2024-49916","cve":"CVE-2024-49916","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Add NULL check for clk_mgr and clk_mgr->funcs in dcn401_init_hw","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49916","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"id":"CVE-2024-49917","cve":"CVE-2024-49917","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Add NULL check for clk_mgr and clk_mgr->funcs in dcn30_init_hw","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49917","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"id":"CVE-2024-49918","cve":"CVE-2024-49918","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add null check for head_pipe in dcn32_acquire_idle_pipe_for_head_pipe_in_layer","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49918","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"id":"CVE-2024-49919","cve":"CVE-2024-49919","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add null check for head_pipe in dcn201_acquire_free_pipe_for_layer","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49919","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"id":"CVE-2024-49920","cve":"CVE-2024-49920","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Check null pointers before multiple uses","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49920","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"id":"CVE-2024-49921","cve":"CVE-2024-49921","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Check null pointers before used","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49921","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"id":"CVE-2024-49922","cve":"CVE-2024-49922","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Check null pointers before using them","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49922","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"id":"CVE-2024-49923","cve":"CVE-2024-49923","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Pass non-null to dcn20_validate_apply_pipe_split_flags","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49923","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"id":"CVE-2024-49970","cve":"CVE-2024-49970","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Implement bounds check for stream encoder creation in DCN401","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49970","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"id":"CVE-2024-49971","cve":"CVE-2024-49971","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Increase array size of dummy_boolean","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49971","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"id":"CVE-2024-49972","cve":"CVE-2024-49972","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Memory is handed to a consumer without being initialised","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Memory is handed to a consumer without being initialised or cleared in the amdgpu display core (DC/DM). Whatever the previous owner left behind is readable - and on a GPU node the previous owner is very often a different tenant's job. This is the classic residual-data leak between workloads sharing a card: model weights, activations, keys or tokens from the prior tenant can surface in a fresh allocation. Upstream fix: drm/amd/display: Deallocate DML memory if allocation fails","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49972","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"id":"CVE-2024-50003","cve":"CVE-2024-50003","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Fix system hang while resume with TBT monitor","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-50003","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"id":"CVE-2024-50004","cve":"CVE-2024-50004","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: update DML2 policy EnhancedPrefetchScheduleAccelerationFinal DCN35","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-50004","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"id":"CVE-2024-50009","cve":"CVE-2024-50009","aliases":[],"title":"Linux cpufreq/amd-pstate - unchecked cpufreq_cpu_get() return value: cpufreq_cpu_get() can return NULL and amd-pstate","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux cpufreq/amd-pstate - unchecked cpufreq_cpu_get() return value","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"cpufreq_cpu_get() can return NULL and amd-pstate did not check it, giving a kernel NULL dereference and a host panic. Availability failure in always-running platform code.","attack_vector":"Local, on the amd-pstate path.","remediation":"Distro kernel update plus reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-50009"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"id":"CVE-2024-50049","cve":"CVE-2024-50049","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Check null pointer before dereferencing se","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-50049","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-21"},{"id":"CVE-2024-50108","cve":"CVE-2024-50108","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A race condition or locking defect in the amdgpu display","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdgpu display core (DC/DM). Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amd/display: Disable PSR-SU on Parade 08-01 TCON too","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-50108","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-11-05"},{"cwe":["CWE-908","CWE-200"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-50110","cve":"CVE-2024-50110","aliases":[],"title":"Linux kernel (net/xfrm): Dumping SAs over xfrm netlink copies algorithm structures that were never fully initialized","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Dumping SAs over xfrm netlink copies algorithm structures that were never fully initialized, so ~50 bytes of uninitialized kernel heap per algorithm are handed to userspace. That is a kernel-memory read primitive useful for defeating KASLR or recovering adjacent slab contents from inside a container.","attack_vector":"XFRM_MSG_GETSA dump over netlink, which needs CAP_NET_ADMIN in the network namespace - satisfied by any container granted NET_ADMIN with its own netns, and by the node's IKE daemon. The caller adds an SA (attach_auth allocates the buffer) and then dumps it back to read the uninitialized tail.","remediation":"Boot a kernel carrying the linked stable commits. Interim: drop CAP_NET_ADMIN from tenant containers so the xfrm netlink dump surface is not exposed to them.","references":["https://git.kernel.org/stable/c/610d4cea9b442b22b4820695fc3335e64849725e","https://git.kernel.org/stable/c/dc2ad8e8818e4bf1a93db78d81745b4877b32972","https://nvd.nist.gov/vuln/detail/CVE-2024-50110"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-50172","cve":"CVE-2024-50172","aliases":[],"title":"Linux bnxt_re RoCE driver (chip context memory leak): Memory leak in the Broadcom RoCE driver when doorbell BAR mapping","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux bnxt_re RoCE driver (chip context memory leak)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Memory leak in the Broadcom RoCE driver when doorbell BAR mapping fails. Low severity on its own, but doorbell-BAR mapping is the mechanism that gives a userspace RDMA process direct hardware access, and leaks in its error path are worth tracking on a fabric where that mapping is the tenant boundary.","attack_vector":"Local, via repeated RDMA device setup failures.","remediation":"Kernel/driver upgrade plus host reboot; bundle with the other bnxt_re fixes rather than scheduling separately.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-50172"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-11-07"},{"id":"CVE-2024-50177","cve":"CVE-2024-50177","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: fix a UBSAN warning in DML2.1","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-50177","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-11-08"},{"id":"CVE-2024-53060","cve":"CVE-2024-53060","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): A NULL pointer dereference in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu firmware, ACPI and IP-block initialisation. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: prevent NULL pointer dereference if ATIF is not supported","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53060","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-11-19"},{"id":"CVE-2024-53072","cve":"CVE-2024-53072","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (platform/x86/amd/pmc): A correctness defect in the amdgpu power","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (platform/x86/amd/pmc)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu power management (SMU/powerplay) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: platform/x86/amd/pmc: Detect when STB is not available","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53072","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-11-19"},{"cwe":["CWE-667","CWE-833"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-53086","cve":"CVE-2024-53086","aliases":[],"title":"Linux kernel (drivers/gpu/drm/xe): Passing a sync object that fails fence lookup makes the exec ioctl return to","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/xe)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Passing a sync object that fails fence lookup makes the exec ioctl return to userspace while still holding the VM's dma-resv lock. That VM's buffers can then never be evicted, migrated or freed, so TTM eviction stalls and other tenants on the device are starved of VRAM until the node is rebooted.","attack_vector":"Tenant container holding /dev/dri/renderD* on Intel xe: call the exec ioctl with a syncobj handle whose in-fence lookup fails. One ioctl, deterministic, unprivileged.","remediation":"Update to a kernel carrying the fix (stable commits below; no fixed_in published). Interim: none that is real - the lock is taken on the normal submit path; evacuate and reboot affected nodes.","references":["https://git.kernel.org/stable/c/96397b1e25dda8389dea63ec914038a170bf953d","https://git.kernel.org/stable/c/64a2b6ed4bfd890a0e91955dd8ef8422a3944ed9","https://nvd.nist.gov/vuln/detail/CVE-2024-53086"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-772","CWE-401"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-53087","cve":"CVE-2024-53087","aliases":[],"title":"Linux kernel (drivers/gpu/drm/xe): Every exec ioctl that bails on an input-validation error leaves an exec-queue","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/xe)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Every exec ioctl that bails on an input-validation error leaves an exec-queue reference permanently pinned. A tenant loops the ioctl with deliberately bad arguments and exhausts the finite GuC context id space and the GPU memory backing those queues, after which co-tenants can no longer create queues on that device.","attack_vector":"Tenant container holding /dev/dri/renderD* on Intel xe calls DRM_IOCTL_XE_EXEC in a loop with arguments chosen to fail validation after the exec-queue lookup. Deterministic, unprivileged, no race required.","remediation":"Update to a kernel carrying the fix (stable commits below; no fixed_in published). Interim: cap per-tenant GPU context counts if your runtime supports it; otherwise a node reboot is the only way to reclaim the leaked queues.","references":["https://git.kernel.org/stable/c/2f92b77a8ce043fbda2664d9be4b66bdc57f67b7","https://git.kernel.org/stable/c/af797b831d8975cb4610f396dcb7f03f4b9908e7","https://nvd.nist.gov/vuln/detail/CVE-2024-53087"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-53200","cve":"CVE-2024-53200","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Fix null check for pipe_ctx->plane_state in hwss_setup_dpp","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53200","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-12-27"},{"id":"CVE-2024-53201","cve":"CVE-2024-53201","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Fix null check for pipe_ctx->plane_state in dcn20_program_pipe","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53201","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-12-27"},{"id":"CVE-2024-53869","cve":"CVE-2024-53869","aliases":[],"title":"GPU Display Driver: Info disclosure (missing initialization)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Info disclosure (missing initialization)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; rolling reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53869","https://github.com/NVIDIA/product-security/tree/main/2025/5614"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-459"],"fleet":{"pain_class":"node-reboot"},"published":"2025-01-28"},{"id":"CVE-2024-53881","cve":"CVE-2024-53881","aliases":[],"title":"vGPU Manager: Host DoS (missing error handling)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU Manager","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Host DoS (missing error handling)","attack_vector":"Tenant VM guest","remediation":"vGPU Manager upgrade; evacuate VMs","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53881","https://github.com/NVIDIA/product-security/tree/main/2025/5614"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-459"],"published":"2025-01-28"},{"id":"CVE-2024-55881","cve":"CVE-2024-55881","aliases":[],"title":"Linux KVM x86 - hypercall completion for protected guests: KVM used the wrong helper to decide whether a hypercall","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux KVM x86 - hypercall completion for protected guests","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"KVM used the wrong helper to decide whether a hypercall was 64-bit when completing it, which misbehaves for guests with protected state such as SEV-ES and SEV-SNP where the host cannot see guest registers. The result is host-side state confusion driven by a confidential guest - a crash or incorrect emulation on the hypervisor path that every other guest on the node shares.","attack_vector":"From inside a protected (SEV-ES/SNP) guest issuing hypercalls.","remediation":"Fixed in the Linux kernel - KVM/x86 SEV code or the ccp/PSP driver. Take the distro kernel update (RHEL/Rocky, Ubuntu, SLES) and **reboot the host**; SEV/SNP hypervisor paths cannot be live-patched in any meaningful way, and SNP platform init/shutdown is not safe to cycle under running guests. Drain confidential-VM tenants, reboot, then re-admit. No firmware, VBIOS or AGESA step needed, which makes this one of the cheaper classes of SEV fix to roll out.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-55881"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-01-11"},{"id":"CVE-2024-56542","cve":"CVE-2024-56542","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A memory or reference-count leak in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu display core (DC/DM). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amd/display: fix a memleak issue when driver is removed","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-56542","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-12-27"},{"id":"CVE-2024-56594","cve":"CVE-2024-56594","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdgpu): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdgpu)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdkfd (KFD compute driver, /dev/kfd). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdgpu: set the right AMDGPU sg segment limitation","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-56594","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-12-27"},{"id":"CVE-2024-56666","cve":"CVE-2024-56666","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A correctness defect in the amdkfd (KFD compute","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdkfd (KFD compute driver, /dev/kfd) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdkfd: Dereference null return value","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-56666","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-12-27"},{"id":"CVE-2024-56697","cve":"CVE-2024-56697","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): An out-of-bounds access in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An out-of-bounds access in the amdgpu firmware, ACPI and IP-block initialisation - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: Fix the memory allocation issue in amdgpu_discovery_get_nps_info()","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-56697","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-12-28"},{"id":"CVE-2024-56753","cve":"CVE-2024-56753","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/gfx9): A memory or reference-count leak","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/gfx9)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu firmware, ACPI and IP-block initialisation. Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amdgpu/gfx9: Add Cleaner Shader Deinitialization in gfx_v9_0 Module","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-56753","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-12-29"},{"id":"CVE-2024-57897","cve":"CVE-2024-57897","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A race condition or locking defect in the amdkfd (KFD","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdkfd (KFD compute driver, /dev/kfd). Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdkfd: Correct the migration DMA map direction","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-57897","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-01-15"},{"id":"CVE-2024-57919","cve":"CVE-2024-57919","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A division by zero in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A division by zero in the amdgpu display core (DC/DM), reachable with attacker-influenced parameters. The kernel takes a divide fault and the node goes down, taking every co-resident job with it. Upstream fix: drm/amd/display: fix divide error in DM plane scale calcs","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-57919","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-01-19"},{"id":"CVE-2024-57922","cve":"CVE-2024-57922","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A division by zero in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A division by zero in the amdgpu display core (DC/DM), reachable with attacker-influenced parameters. The kernel takes a divide fault and the node goes down, taking every co-resident job with it. Upstream fix: drm/amd/display: Add check for granularity in dml ceil/floor helpers","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-57922","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-01-19"},{"id":"CVE-2024-57950","cve":"CVE-2024-57950","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A division by zero in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A division by zero in the amdgpu display core (DC/DM), reachable with attacker-influenced parameters. The kernel takes a divide fault and the node goes down, taking every co-resident job with it. Upstream fix: drm/amd/display: Initialize denominator defaults to 1","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-57950","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-02-10"},{"cwe":["CWE-787","CWE-131"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-58018","cve":"CVE-2024-58018","aliases":[],"title":"Linux kernel (drivers/gpu/drm/nouveau/nvkm/subdev/gsp): The driver miscalculates free space in the GSP firmware command","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel (drivers/gpu/drm/nouveau/nvkm/subdev/gsp)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The driver miscalculates free space in the GSP firmware command ring when the write pointer wraps, and overwrites the RPC the GPU firmware is still reading. The firmware wedges on a corrupted request and the GPU is dead for every tenant on it until the node is rebooted.","attack_vector":"Indirect but tenant-driven: any tenant holding /dev/dri/renderD* on a GSP-firmware NVIDIA GPU under nouveau generates the RPC traffic (allocations, mappings, large requests) that wraps the ring, and a tenant issuing large or high-volume allocations makes the wrap case likely. Only applies where nouveau with GSP-RM is the driver, not the NVIDIA proprietary/open kernel module.","remediation":"Update to a kernel carrying the fix (stable commits below; no fixed_in published). Interim: run the vendor kernel module rather than nouveau on shared GSP-class GPUs.","references":["https://git.kernel.org/stable/c/56e6c7f6d2a6b4e0aae0528c502e56825bb40598","https://git.kernel.org/stable/c/6b6b75728c86f60c1fc596f0d4542427d0e6065b","https://nvd.nist.gov/vuln/detail/CVE-2024-58018"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-58052","cve":"CVE-2024-58052","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu): A NULL pointer dereference in the amdgpu power","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: Fix potential NULL pointer dereference in atomctrl_get_smc_sclk_range_table","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-58052","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-03-06"},{"id":"CVE-2025-21645","cve":"CVE-2025-21645","aliases":[],"title":"Linux platform/x86/amd/pmc - IRQ1 wakeup disabled unconditionally: The AMD PMC driver disabled IRQ1 wakeup in cases","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux platform/x86/amd/pmc - IRQ1 wakeup disabled unconditionally","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The AMD PMC driver disabled IRQ1 wakeup in cases where i8042 had never enabled it, corrupting interrupt wakeup configuration. Low-severity platform-driver correctness bug; included because the amd/pmc driver autoloads on AMD hosts whether or not the platform needs it, and unnecessary loaded drivers are unnecessary attack surface.","attack_vector":"Local, through platform power-management paths.","remediation":"Distro kernel update plus reboot, or blacklist the module on server images where power management is handled by the BMC and BIOS anyway.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21645"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-01-19"},{"cwe":["CWE-667","CWE-404"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-21662","cve":"CVE-2025-21662","aliases":[],"title":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core): When the driver runs out of firmware command slots, the work","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"When the driver runs out of firmware command slots, the work handler bails out without signalling the waiting task, so that task blocks forever. The mlx5 command interface is the single control path for the NIC, so one wedged caller cascades into hung kernel workers and a NIC whose queues, RDMA objects and interrupt moderation can no longer be reconfigured - a fabric stall for every tenant on the node, cleared only by reboot.","attack_vector":"Requires exhausting the mlx5 firmware command index pool, which a tenant can drive from inside a container holding /dev/infiniband/uverbs* (each verbs/DEVX object creation issues a firmware command) or from a VF assigned into a tenant VM. No host privilege and no fabric access needed; the only precondition is enough concurrent command pressure to make cmd_alloc_index() fail.","remediation":"Update to 6.1.125 / 6.6.72 or later on those stable branches, or to 6.9/6.10 and later mainline. Interim controls: rate-limit or cap per-tenant RDMA resource creation, and drop /dev/infiniband/* from containers that do not need verbs.","references":["https://git.kernel.org/stable/c/229cc10284373fbe754e623b7033dca7e7470ec8","https://git.kernel.org/stable/c/36124081f6ffd9dfaad48830bdf106bb82a9457d","https://nvd.nist.gov/vuln/detail/CVE-2025-21662"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-21784","cve":"CVE-2025-21784","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): A correctness defect in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu: bail out when failed to load fw in psp_init_cap_microcode()","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21784","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-02-27"},{"cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-21833","cve":"CVE-2025-21833","aliases":[],"title":"Linux kernel (drivers/iommu/intel): On the VT-d PASID detach path, if the PASID being removed is not found the code","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/intel)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"On the VT-d PASID detach path, if the PASID being removed is not found the code warned and then used the NULL result anyway, panicking the host. This is the teardown path that revokes a device's access to a tenant address space, so it runs every time a tenant's SVA context or assigned-device PASID goes away - and a panic there is a whole-node outage for every co-tenant.","attack_vector":"Local, on Intel VT-d scalable mode with PASID in use - SVA-capable accelerators or PASID-based device assignment through iommufd. Reached on PASID detach when the PASID is already absent from the domain. Upstream treats this as a should-not-happen state guarded by WARN_ON_ONCE, so a tenant needs to drive the PASID attach/detach path into an inconsistent state first; no host root is required to exercise attach/detach itself.","remediation":"No fixed release is listed in this record; apply the linked stable commits or run a current stable/LTS kernel on Intel passthrough nodes. Interim: limit which tenants can create and tear down PASID contexts, and do not expose SVA-capable device nodes into untrusted containers.","references":["https://git.kernel.org/stable/c/68ec78beb4a3fb0877cbaaf49758c85410c05977","https://git.kernel.org/stable/c/df96876be3b064aefc493f760e0639765d13ed0d","https://nvd.nist.gov/vuln/detail/CVE-2025-21833"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-21841","cve":"CVE-2025-21841","aliases":[],"title":"Linux cpufreq/amd-pstate - cpufreq_policy reference counting: amd_pstate_update_limits() takes a cpufreq_policy","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux cpufreq/amd-pstate - cpufreq_policy reference counting","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"amd_pstate_update_limits() takes a cpufreq_policy reference and never drops it. A leaked reference means the policy object can never be freed, so CPU hotplug and driver unbind hang or leak indefinitely - the sort of defect that makes a node impossible to cleanly reconfigure without a reboot.","attack_vector":"Local, through the amd-pstate limits-update path.","remediation":"Distro kernel update plus reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21841"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-03-07"},{"cwe":["CWE-911","CWE-833"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-21886","cve":"CVE-2025-21886","aliases":[],"title":"Linux kernel (drivers/infiniband/hw/mlx5): Memory-region deregistration hangs forever on the flagship AI-cluster NIC. A","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/hw/mlx5)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Memory-region deregistration hangs forever on the flagship AI-cluster NIC. A reference taken on the parent region is never dropped, so the deregister call blocks in uninterruptible sleep - the task cannot be killed, it holds device references that prevent teardown, and the tenant's GPU job wedges without releasing its RDMA resources back to the node.","attack_vector":"A tenant container holding /dev/infiniband/uverbs* on mlx5 hardware with on-demand paging (implicit ODP) in use - deregistering a parent MR that still has implicit children is enough, which is normal application shutdown behaviour and therefore trivially repeatable. Only applies where ODP is enabled; non-ODP registrations do not reach this path.","remediation":"Update to 6.12.18 / 6.13.6 or later. Interim: disable on-demand paging for tenant workloads (do not advertise ODP capability) so implicit MRs are never created, and drain jobs off nodes that have accumulated stuck deregistration tasks.","references":["https://git.kernel.org/stable/c/cb96ae783e7249e8e5a50c22952c0bb2983133df","https://git.kernel.org/stable/c/a095ede2daca49d15e74d66d014883f2fa8bb924","https://nvd.nist.gov/vuln/detail/CVE-2025-21886"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-833"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-21892","cve":"CVE-2025-21892","aliases":[],"title":"Linux kernel (drivers/infiniband/hw/mlx5): The memory-registration engine on the primary AI-cluster NIC wedges","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/hw/mlx5)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The memory-registration engine on the primary AI-cluster NIC wedges permanently. Recovery resets the internal UMR queue-pair without waiting for outstanding work, the firmware discards those completions, and every waiter blocks forever - registration and deregistration stall for all tenants on the node, not just the one that hit the error.","attack_vector":"A tenant container holding /dev/infiniband/uverbs* on mlx5 hardware doing normal memory-region register/deregister work; once anything pushes the shared UMR queue-pair into error and recovery runs, the node-wide stall follows. The UMR queue-pair is per-device and shared, so a single tenant's error becomes everyone's outage.","remediation":"The fixed-version metadata in this record is unreliable - apply the listed stable fix commits or move to a current stable kernel. Interim: drain workloads off nodes showing hung MR-registration tasks and reboot; there is no runtime way to reset the UMR queue-pair.","references":["https://git.kernel.org/stable/c/3e3bf255992cc02404e9d209b127c1c9944239cf","https://git.kernel.org/stable/c/1d2b84d8d054313deed2b2fcafe1168bbcb9e99f","https://nvd.nist.gov/vuln/detail/CVE-2025-21892"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-21940","cve":"CVE-2025-21940","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A NULL pointer dereference in the amdkfd (KFD compute","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdkfd (KFD compute driver, /dev/kfd). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdkfd: Fix NULL Pointer Dereference in KFD queue","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21940","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-04-01"},{"id":"CVE-2025-21941","cve":"CVE-2025-21941","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Fix null check for pipe_ctx->plane_state in resource_build_scaling_params","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21941","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-04-01"},{"id":"CVE-2025-21956","cve":"CVE-2025-21956","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Assign normalized_pix_clk when color depth = 14","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21956","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-04-01"},{"id":"CVE-2025-21987","cve":"CVE-2025-21987","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): Memory is handed to a consumer without being initialised or","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Memory is handed to a consumer without being initialised or cleared in the amdgpu kernel driver core. Whatever the previous owner left behind is readable - and on a GPU node the previous owner is very often a different tenant's job. This is the classic residual-data leak between workloads sharing a card: model weights, activations, keys or tokens from the prior tenant can surface in a fresh allocation. Upstream fix: drm/amdgpu: init return value in amdgpu_ttm_clear_buffer","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21987","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-04-02"},{"id":"CVE-2025-21989","cve":"CVE-2025-21989","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: fix missing .is_two_pixels_per_container","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21989","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-04-02"},{"id":"CVE-2025-21990","cve":"CVE-2025-21990","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): A correctness defect in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu: NULL-check BO's backing store when determining GFX12 PTE flags","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21990","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-04-02"},{"cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-22092","cve":"CVE-2025-22092","aliases":[],"title":"Linux kernel (drivers/pci): When setting up an SR-IOV virtual function fails partway through, the half-initialised VF","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/pci)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"When setting up an SR-IOV virtual function fails partway through, the half-initialised VF is left registered and is dereferenced later during removal, panicking the host. A single VF that fails to come up takes down the whole node and every tenant sitting on it, in the middle of the routine operation that hands VFs out to tenants.","attack_vector":"The trigger path is the operator's own VF provisioning - writing sriov_numvfs, which lands in sriov_enable() via the driver's sriov_configure callback (the upstream report is mlx5_core, i.e. the ConnectX fabric NIC these clusters run on). It needs host root to initiate plus a pci_setup_device() failure on one VF, so it is not directly tenant-reachable; the realistic scenario is a NIC or GPU left in a bad state by the previous tenant, or a flaky device, causing VF setup to fail during recycle and panicking the node instead of erroring out. Not applicable if SR-IOV is unused on the node.","remediation":"Update to a kernel carrying the fix (no fixed_in published; take the stable commits below into your 6.1.y / 6.6.y / 6.12.y branch). Interim: reset the PF and confirm all VFs enumerate cleanly before rescheduling tenants onto a node, and drain a node before changing sriov_numvfs so a panic during reprovisioning does not take live tenants with it.","references":["https://git.kernel.org/stable/c/ef421b4d206f0d3681804b8f94f06a8458a53aaf","https://git.kernel.org/stable/c/c67a233834b778b8c78f8b62c072ccf87a9eb6d0","https://nvd.nist.gov/vuln/detail/CVE-2025-22092"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-22093","cve":"CVE-2025-22093","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: avoid NPD when ASIC does not support DMUB","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-22093","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-04-16"},{"id":"CVE-2025-23137","cve":"CVE-2025-23137","aliases":[],"title":"Linux cpufreq/amd-pstate - missing NULL check in amd_pstate_update: amd_pstate_update() dereferences the cpufreq policy","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux cpufreq/amd-pstate - missing NULL check in amd_pstate_update","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"amd_pstate_update() dereferences the cpufreq policy without checking it for NULL, panicking the host. The CPU frequency driver runs constantly on every AMD node, so a NULL dereference here takes the machine down and every GPU job on it with no warning and no attacker involvement.","attack_vector":"Local, on the amd-pstate update path. Reachable through normal frequency-governor activity.","remediation":"Distro kernel update plus reboot; no firmware step.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23137"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-04-16"},{"id":"CVE-2025-23245","cve":"CVE-2025-23245","aliases":[],"title":"vGPU Manager: Local privesc via improper file permissions","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU Manager","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Local privesc via improper file permissions","attack_vector":"Local operator on the hypervisor","remediation":"vGPU Manager upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23245","https://github.com/NVIDIA/product-security/tree/main/2025/5630"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-732"],"published":"2025-05-01"},{"id":"CVE-2025-23246","cve":"CVE-2025-23246","aliases":[],"title":"vGPU Manager: Host DoS via resource exhaustion","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU Manager","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Host DoS via resource exhaustion","attack_vector":"Tenant VM guest","remediation":"vGPU Manager upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23246","https://github.com/NVIDIA/product-security/tree/main/2025/5630"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-400"],"published":"2025-05-01"},{"id":"CVE-2025-23261","cve":"CVE-2025-23261","aliases":[],"title":"Cumulus Linux: Sensitive data exposure on the switch","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Cumulus Linux","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Sensitive data exposure on the switch","attack_vector":"Authenticated switch user","remediation":"Upgrade Cumulus Linux; rolling switch upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23261","https://github.com/NVIDIA/product-security/tree/main/2025/5655"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-532"],"published":"2025-09-04"},{"id":"CVE-2025-23285","cve":"CVE-2025-23285","aliases":[],"title":"vGPU Manager: Local privesc via file permissions","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU Manager","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Local privesc via file permissions","attack_vector":"Local operator on the hypervisor","remediation":"vGPU Manager upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23285","https://github.com/NVIDIA/product-security/tree/main/2025/5670"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-732"],"published":"2025-08-02"},{"id":"CVE-2025-23300","cve":"CVE-2025-23300","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): Allocating a specific GPU memory resource drives","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Allocating a specific GPU memory resource drives the kernel driver into a null-pointer dereference, crashing the node. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5703. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23300","https://github.com/NVIDIA/product-security/tree/main/2025/5703"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-drain"},"published":"2025-10-23"},{"id":"CVE-2025-23330","cve":"CVE-2025-23330","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): A null-pointer dereference in the Linux display driver","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A null-pointer dereference in the Linux display driver crashes the node from an unprivileged local account. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5703. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23330","https://github.com/NVIDIA/product-security/tree/main/2025/5703"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-drain"},"published":"2025-10-23"},{"id":"CVE-2025-27249","cve":"CVE-2025-27249","aliases":[],"title":"Intel Gaudi software suite (SynapseAI stack): An authenticated local user can drive the Gaudi software stack","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Intel Gaudi software suite (SynapseAI stack)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An authenticated local user can drive the Gaudi software stack into unbounded resource consumption and deny the accelerator to everything else on the node. On a bin-packed training cluster that is one tenant stalling a whole 8-card box.","attack_vector":"A local authenticated user on the node, which on most Gaudi deployments means anyone with a container that has the Gaudi devices mapped in.","remediation":"Upgrade the Gaudi software suite to 1.21.0 or later. Userspace and driver package update; plan a node drain because the habanalabs driver has to be reloaded, but no BIOS or accelerator firmware flash is required.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-27249","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01374.html"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2025-11-11"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-400"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-33177","cve":"CVE-2025-33177","aliases":[],"title":"NVIDIA Jetson Linux / IGX OS (NvMap): NvMap does not track memory allocations correctly, so one unprivileged process","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Jetson Linux / IGX OS (NvMap)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"NvMap does not track memory allocations correctly, so one unprivileged process can drive the allocator into overallocation and starve everything else sharing the integrated GPU. On an IGX or Jetson node running several containerized workloads against one GPU, any single container can take GPU memory away from all its neighbours and hold the node down. The break is availability only - no data crosses the boundary - but it is a genuine noisy-neighbour escape hatch on a device class that is increasingly used to run more than one workload.","attack_vector":"A local, unprivileged process on the device - which in practice means any container or workload you scheduled onto the node. No privilege escalation and no user interaction needed first; the allocator path is reachable from ordinary GPU API use.","remediation":"Move Jetson Linux to 35.6.3 or later on the L4T 35 branch, 36.4.6 or later on L4T 36, or 38.2.2 or later on L4T 38; on IGX OS take Kernel SRU 1035 or newer. The fix is in the kernel-side allocator, so the update does not take effect until the device reboots. Until then, the only real mitigation is capping GPU memory per workload if your runtime supports it.","references":["https://github.com/NVIDIA/product-security/tree/main/2025/5716","https://nvd.nist.gov/vuln/detail/CVE-2025-33177"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-33237","cve":"CVE-2025-33237","aliases":[],"title":"GPU Display Driver: DoS (null deref)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"DoS (null deref)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; rolling reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33237","https://github.com/NVIDIA/product-security/tree/main/2026/5747"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"published":"2026-01-28"},{"id":"CVE-2025-37754","cve":"CVE-2025-37754","aliases":[],"title":"Linux i915 GPU kernel driver (HuC firmware load): The HuC delayed-loading fence is not released when probe fails early","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux i915 GPU kernel driver (HuC firmware load)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The HuC delayed-loading fence is not released when probe fails early, leaving the driver in a broken state. Matters on nodes where the HuC firmware blob is missing or mismatched, which is a common outcome of an incomplete linux-firmware package in a slim container host image.","attack_vector":"Triggered on driver probe - so on node boot, not by a tenant.","remediation":"Kernel update plus reboot. Also confirm the linux-firmware package on the node actually contains the matching GuC/HuC blobs for the installed silicon; a missing blob is what exposes this path in the first place.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-37754","https://git.kernel.org/stable/c/4bd4bf79bcfe101f0385ab81dbabb6e3f7d96c00"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-05-01"},{"id":"CVE-2025-37766","cve":"CVE-2025-37766","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): A division by zero in the amdgpu power management","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A division by zero in the amdgpu power management (SMU/powerplay), reachable with attacker-influenced parameters. The kernel takes a divide fault and the node goes down, taking every co-resident job with it. Upstream fix: drm/amd/pm: Prevent division by zero","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-37766","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-05-01"},{"id":"CVE-2025-37767","cve":"CVE-2025-37767","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): A division by zero in the amdgpu power management","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A division by zero in the amdgpu power management (SMU/powerplay), reachable with attacker-influenced parameters. The kernel takes a divide fault and the node goes down, taking every co-resident job with it. Upstream fix: drm/amd/pm: Prevent division by zero","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-37767","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-05-01"},{"id":"CVE-2025-37768","cve":"CVE-2025-37768","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): A division by zero in the amdgpu power management","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A division by zero in the amdgpu power management (SMU/powerplay), reachable with attacker-influenced parameters. The kernel takes a divide fault and the node goes down, taking every co-resident job with it. Upstream fix: drm/amd/pm: Prevent division by zero","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-37768","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-05-01"},{"id":"CVE-2025-37769","cve":"CVE-2025-37769","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm/smu11): A division by zero in the amdgpu power","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm/smu11)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A division by zero in the amdgpu power management (SMU/powerplay), reachable with attacker-influenced parameters. The kernel takes a divide fault and the node goes down, taking every co-resident job with it. Upstream fix: drm/amd/pm/smu11: Prevent division by zero","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-37769","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-05-01"},{"id":"CVE-2025-37770","cve":"CVE-2025-37770","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): A division by zero in the amdgpu power management","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A division by zero in the amdgpu power management (SMU/powerplay), reachable with attacker-influenced parameters. The kernel takes a divide fault and the node goes down, taking every co-resident job with it. Upstream fix: drm/amd/pm: Prevent division by zero","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-37770","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-05-01"},{"id":"CVE-2025-37771","cve":"CVE-2025-37771","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): A division by zero in the amdgpu power management","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A division by zero in the amdgpu power management (SMU/powerplay), reachable with attacker-influenced parameters. The kernel takes a divide fault and the node goes down, taking every co-resident job with it. Upstream fix: drm/amd/pm: Prevent division by zero","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-37771","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-05-01"},{"cwe":["CWE-833"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-37843","cve":"CVE-2025-37843","aliases":[],"title":"Linux kernel (drivers/pci/hotplug): Removing nested PCIe hotplug ports can deadlock - a parent hotplug port holds the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/pci/hotplug)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Removing nested PCIe hotplug ports can deadlock - a parent hotplug port holds the global PCI rescan/remove lock while waiting for the child port's driver to unbind, and the child then blocks trying to take the same lock. Nothing breaks the cycle: PCI enumeration and removal are wedged host-wide from that point on, so no device can be added, removed or reprovisioned until the node reboots.","attack_vector":"Device-driven. The lock cycle lives in the nested hotplug removal path, which is the topology of any GPU or NVMe chassis where hotplug-capable downstream ports sit behind another hotplug port. The upstream commit only reduces the frequency - it removes a device-replacement check that fired on plain removal too, and notes a proper fix is still being worked on. The reported reproducer is device removal during system sleep with Thunderbolt, which does not apply to a server that never suspends; the underlying long-standing race does, and it needs no tenant privilege, just a removal event in a nested hotplug hierarchy.","remediation":"Update to a kernel carrying the fix (no fixed_in published; stable commits below) understanding it narrows rather than closes the race. Interim: drain the node before any removal under a nested hotplug hierarchy, and avoid concurrent removals of parent and child hotplug ports.","references":["https://git.kernel.org/stable/c/e4a1d7defbc2d806540720a5adebe24ec3488683","https://git.kernel.org/stable/c/0d0bbd01f7c0ac7d1be9f85aaf2cd0baec34655f","https://nvd.nist.gov/vuln/detail/CVE-2025-37843"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-37852","cve":"CVE-2025-37852","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu): A NULL pointer dereference in the amdgpu power","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: handle amdgpu_cgs_create_device() errors in amd_powerplay_create()","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-37852","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-05-09"},{"id":"CVE-2025-37853","cve":"CVE-2025-37853","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A NULL pointer dereference in the amdkfd (KFD compute","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdkfd (KFD compute driver, /dev/kfd). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdkfd: debugfs hang_hws skip GPU with MES","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-37853","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-05-09"},{"id":"CVE-2025-37855","cve":"CVE-2025-37855","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Guard Possible Null Pointer Dereference","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-37855","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-05-09"},{"id":"CVE-2025-37870","cve":"CVE-2025-37870","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: prevent hang on link training fail","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-37870","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-05-09"},{"cwe":["CWE-404","CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-37877","cve":"CVE-2025-37877","aliases":[],"title":"Linux kernel (drivers/iommu): When IOMMU registration fails, the core tore down groups and default domains but left","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"When IOMMU registration fails, the core tore down groups and default domains but left devices still flagged as using iommu-dma. Those devices then keep routing DMA setup through an iommu-dma layer whose domain is gone - a stale binding to a torn-down translation context, and a crash inside iommu-dma. On a node where the IOMMU driver failed to come up, this is the difference between failing closed and devices doing DMA with no functioning translation layer behind them.","attack_vector":"Not tenant-reachable. Triggered on the host during boot or IOMMU driver probe when iommu_device_register() fails - firmware/ACPI-table problems, a broken IOMMU unit, or a driver bug. Matters to an operator because a node in this state should be treated as having no working IOMMU at all, which means no safe device passthrough, rather than as a node that merely logged an error.","remediation":"No fixed release is listed in this record; apply the linked stable commits or run a current stable/LTS kernel. Interim: fail nodes closed - if the IOMMU driver did not register cleanly, do not schedule passthrough or device-assigned workloads on that node; check dmesg for IOMMU registration failure as part of node admission.","references":["https://git.kernel.org/stable/c/b14d98641312d972bb3f38e82eddf92898522389","https://git.kernel.org/stable/c/104a84276821aed0ed241ce0d82d6c3267e3fcb8","https://nvd.nist.gov/vuln/detail/CVE-2025-37877"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-37965","cve":"CVE-2025-37965","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Fix invalid context error in dml helper","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-37965","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-05-20"},{"id":"CVE-2025-38011","cve":"CVE-2025-38011","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): A memory or reference-count leak in the amdgpu kernel driver core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu kernel driver core. Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amdgpu: csa unmap use uninterruptible lock","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38011","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-06-18"},{"id":"CVE-2025-38021","cve":"CVE-2025-38021","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Fix null check of pipe_ctx->plane_state for update_dchubp_dpp","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38021","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-06-18"},{"id":"CVE-2025-38145","cve":"CVE-2025-38145","aliases":[],"title":"ASPEED LPC snoop driver (drivers/soc/aspeed/aspeed-lpc-snoop.c): Under memory pressure an allocation in the LPC snoop","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ASPEED LPC snoop driver (drivers/soc/aspeed/aspeed-lpc-snoop.c)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Under memory pressure an allocation in the LPC snoop setup path returns NULL and is dereferenced, oopsing the BMC kernel. LPC snoop is the channel that captures host POST codes and BIOS progress, so the practical loss is the BMC panicking during host boot - exactly the window where an operator is watching POST codes to diagnose a node that will not come up. Availability only, but on a large fleet the BMC dropping out mid-boot means losing remote power control on a node you now have to touch physically.","attack_vector":"Local on the BMC, and needs the BMC to be under memory pressure when the snoop channel is enabled. Practically this surfaces as a reliability bug rather than something an external attacker drives, though anything that inflates BMC memory usage (a pre-auth bmcweb allocation bug, for example) raises the odds.","remediation":"Kernel one-liner, backported to stable. Reaches nodes only via a BMC firmware image update: per-node, out-of-band, ODM-rebase dependent. No config workaround worth the tradeoff - disabling LPC snoop costs you host POST-code visibility, which is one of the main reasons the BMC is there. Low enough severity that batching it into your next scheduled firmware refresh is the right call.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38145","https://git.kernel.org/stable/c/c550999f939b529d28a914d5034cc4290066aea6"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2025-07-03"},{"cwe":["CWE-682"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-38158","cve":"CVE-2025-38158","aliases":[],"title":"Linux kernel (drivers/vfio/pci/hisilicon): The VFIO migration driver reassembled the device's event-queue DMA addresses","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/vfio/pci/hisilicon)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The VFIO migration driver reassembled the device's event-queue DMA addresses from hardware registers in the wrong byte order, so after a live migration the accelerator was programmed to DMA at addresses that were never the ones the guest set up. The device ends up writing to the wrong place inside the assigned domain and the guest's crypto services fail; the same class of mistake in an address-composition path is what turns a migration into a stray-DMA event, and the fix has to carry a magic-number check because guests migrated from old kernels otherwise land on bad addresses silently.","attack_vector":"Not attacker-initiated: triggered by live-migrating a guest that has a HiSilicon accelerator VF assigned under hisi_acc_vfio_pci, including migrations from an older kernel to a newer one. The resulting DMA is still bounded by the guest's IOMMU domain, so this is corruption and device malfunction inside the tenant rather than a host escape. Conditional on hisi_acc_vfio_pci being bound and live migration being in use - not reachable on an NVIDIA/AMD GPU fleet that never loads this driver.","remediation":"No fixed release is listed in this record; apply the linked stable commits or run a current stable/LTS kernel on nodes using hisi_acc_vfio_pci, and patch source and destination hosts together so the migration magic-number handling matches. Interim: disable live migration for HiSilicon accelerator VFs.","references":["https://git.kernel.org/stable/c/809a9c10274e1bcf6d05f1c0341459a425a4f05f","https://git.kernel.org/stable/c/f0423873e7aeb69cb68f4e8fa3827832e7b037ba","https://nvd.nist.gov/vuln/detail/CVE-2025-38158"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-38205","cve":"CVE-2025-38205","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A division by zero in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A division by zero in the amdgpu display core (DC/DM), reachable with attacker-influenced parameters. The kernel takes a divide fault and the node goes down, taking every co-resident job with it. Upstream fix: drm/amd/display: Avoid divide by zero by initializing dummy pitch to 1","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38205","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-07-04"},{"id":"CVE-2025-38254","cve":"CVE-2025-38254","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Add sanity checks for drm_edid_raw()","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38254","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-07-09"},{"id":"CVE-2025-38319","cve":"CVE-2025-38319","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pp): A NULL pointer dereference in the amdgpu power","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pp)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/pp: Fix potential NULL pointer dereference in atomctrl_initialize_mc_reg_table","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38319","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-07-10"},{"id":"CVE-2025-38360","cve":"CVE-2025-38360","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Add more checks for DSC / HUBP ONO guarantees","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38360","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-07-25"},{"id":"CVE-2025-38362","cve":"CVE-2025-38362","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add null pointer check for get_first_active_display()","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38362","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-07-25"},{"cwe":["CWE-833","CWE-667"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-38373","cve":"CVE-2025-38373","aliases":[],"title":"Linux kernel (drivers/infiniband/hw/mlx5): Memory-region deregistration self-deadlocks under memory pressure. An","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/hw/mlx5)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Memory-region deregistration self-deadlocks under memory pressure. An allocation made while holding the page-mapping lock can drive reclaim, which calls back into the same driver's invalidation handler and blocks on the lock it already holds - the tenant's task hangs unkillably and the mlx5 registration path is blocked for everything else on the node.","attack_vector":"A tenant container holding /dev/infiniband/uverbs* on mlx5 hardware deregistering an on-demand-paging memory region while the node is under memory pressure. Both halves are tenant-influenceable: the tenant chooses when to deregister and can raise memory pressure itself, which is exactly the condition a busy shared GPU node is already in.","remediation":"Update to 6.12.37 / 6.14 or later. Interim: disable on-demand paging for tenant workloads so ODP regions are never created, and keep node memory headroom high enough that reclaim is not entered during RDMA teardown.","references":["https://git.kernel.org/stable/c/beb89ada5715e7bd1518c58863eedce89ec051a7","https://git.kernel.org/stable/c/727eb1be65a370572edf307558ec3396b8573156","https://nvd.nist.gov/vuln/detail/CVE-2025-38373"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-665","CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-38387","cve":"CVE-2025-38387","aliases":[],"title":"Linux kernel (drivers/infiniband/hw/mlx5): An event subscription is published to the lookup table before its list head","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/hw/mlx5)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An event subscription is published to the lookup table before its list head is initialized, so a device event that arrives in the same instant follows a poison pointer and faults in kernel context. One tenant's event subscription can panic a shared mlx5 node.","attack_vector":"A container holding /dev/infiniband/uverbs* that uses the mlx5 DEVX interface to subscribe to device events, with hardware events arriving concurrently - the tenant controls both the subscription rate and much of the event traffic. DEVX normally requires CAP_NET_RAW, so this needs a privileged container or a tenant explicitly granted raw-network capability; plain verbs users cannot reach it.","remediation":"No fixed release is published in this record - apply the listed stable fix commits or run a current stable kernel. Interim: drop CAP_NET_RAW from tenant containers so the DEVX interface is unavailable, which closes this path entirely.","references":["https://git.kernel.org/stable/c/716b555fc0580c2aa4c2c32ae4401c7e3ad9873e","https://git.kernel.org/stable/c/972e968aac0dce8fe8faad54f6106de576695d8e","https://nvd.nist.gov/vuln/detail/CVE-2025-38387"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-38426","cve":"CVE-2025-38426","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): A correctness defect in the amdgpu RAS / GPU reset","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu: Add basic validation for RAS header","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38426","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-07-25"},{"id":"CVE-2025-38487","cve":"CVE-2025-38487","aliases":[],"title":"ASPEED LPC snoop driver channel teardown (drivers/soc/aspeed/aspeed-lpc-snoop.c): Unbinding the LPC snoop driver tears","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ASPEED LPC snoop driver channel teardown (drivers/soc/aspeed/aspeed-lpc-snoop.c)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Unbinding the LPC snoop driver tears down channels that were never brought up, dereferencing NULL and panicking the BMC kernel. The reproducer is a single write to the driver's sysfs unbind file. Anyone with root on the BMC can hard-crash the management processor on demand; more usefully for an operator, it fires during ordinary driver reload and platform-teardown sequences, so it shows up as BMC instability on ASPEED platforms that only wire up a subset of the snoop channels.","attack_vector":"Root on the BMC (write access to the platform driver's sysfs bind/unbind), or any BMC-side maintenance flow that unbinds the driver. Not reachable from the host or the network on its own.","remediation":"Kernel patch, backported to stable. Delivered only in a new BMC firmware image - per-node, out-of-band flash, gated on the ODM. Low priority as a standalone item; treat it as one more reason not to run BMC firmware images that are years behind upstream, and roll it in with the other lpc-snoop and video-engine fixes in a single flash rather than a dedicated campaign.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38487","https://git.kernel.org/stable/c/9e1d2b97f5e2a36a2fd30a8bd30ead9dac5e3a51"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2025-07-28"},{"id":"CVE-2025-38506","cve":"CVE-2025-38506","aliases":[],"title":"Linux KVM - CPU soft lockup setting per-page memory attributes on large SNP guests: Running an SEV-SNP guest","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux KVM - CPU soft lockup setting per-page memory attributes on large SNP guests","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Running an SEV-SNP guest with a large memory footprint - 1TB and up, which is exactly the shape of an AI training VM - drives the host into CPU soft lockups while KVM walks per-page memory attributes without rescheduling. The host stalls, the watchdog fires, and co-resident workloads suffer. A tenant does not need to attack anything: simply asking for a big confidential VM is enough.","attack_vector":"Triggered by a guest with a very large memory allocation. Tenant-reachable through normal VM sizing, which makes it as much a capacity-planning hazard as a security one.","remediation":"Fixed in the Linux kernel - KVM/x86 SEV code or the ccp/PSP driver. Take the distro kernel update (RHEL/Rocky, Ubuntu, SLES) and **reboot the host**; SEV/SNP hypervisor paths cannot be live-patched in any meaningful way, and SNP platform init/shutdown is not safe to cycle under running guests. Drain confidential-VM tenants, reboot, then re-admit. No firmware, VBIOS or AGESA step needed, which makes this one of the cheaper classes of SEV fix to roll out. Directly relevant to AI hosts, where terabyte-class confidential VMs are the norm rather than the exception - do not dismiss this as a corner case on a GPU fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38506"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-08-16"},{"id":"CVE-2025-38518","cve":"CVE-2025-38518","aliases":[],"title":"Linux x86/CPU/AMD - INVLPGB on Zen 2 (Cyan Skillfish): Using broadcast TLB invalidation (INVLPGB) on affected Zen 2","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux x86/CPU/AMD - INVLPGB on Zen 2 (Cyan Skillfish)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Using broadcast TLB invalidation (INVLPGB) on affected Zen 2 parts oopses the system. TLB invalidation is core memory-management machinery, so a defect here is both a stability problem and, in principle, a correctness problem for the mappings that separate address spaces.","attack_vector":"Local, triggered by normal kernel memory management on affected silicon rather than by an attacker.","remediation":"Fixed in the Linux kernel by disabling INVLPGB on affected parts. Distro kernel update plus reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38518"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-08-16"},{"id":"CVE-2025-38520","cve":"CVE-2025-38520","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): Missing or insufficient validation of user-supplied","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdkfd (KFD compute driver, /dev/kfd). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdkfd: Don't call mmput from MMU notifier callback","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38520","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-08-16"},{"cwe":["CWE-754"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-38623","cve":"CVE-2025-38623","aliases":[],"title":"Linux kernel (drivers/pci/hotplug): A surprise device removal freezes the PCI host bridge's partitionable endpoint and","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/pci/hotplug)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A surprise device removal freezes the PCI host bridge's partitionable endpoint and the driver never clears the freeze. MSI interrupts from the hotplug logic stop arriving, so every slot behind that PE goes deaf - no further plug or unplug events are processed on any of them, and new devices are not detected until the machine is rebooted. One tenant's device dropping off silently disables device management for the whole PE.","attack_vector":"Device-driven and needs no credentials: a card that drops off the link, an unplanned removal, or a device error that causes the upstream bridge to freeze the PE is enough. Applies only to PowerNV - OpenPOWER POWER9/POWER10 hosts running the pnv_php hotplug driver; it is inert on x86 and ARM GPU nodes. Worth tracking only if IBM POWER machines are part of the fleet.","remediation":"Update to a kernel carrying the fix (no fixed_in published; stable commits below). Interim on PowerNV: monitor for PE freezes and treat a frozen PE as a drain-and-reboot condition rather than expecting hotplug to recover on its own.","references":["https://git.kernel.org/stable/c/6e7b5f922901585b8f11e0d6cda12bda5c59fc8a","https://git.kernel.org/stable/c/2ec8ec57bb8ebde3e2a015eff80e5d66e6634fe3","https://nvd.nist.gov/vuln/detail/CVE-2025-38623"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-401","CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-38624","cve":"CVE-2025-38624","aliases":[],"title":"Linux kernel (drivers/pci/hotplug): Unplugging the root of a nested PCIe bridge hierarchy leaks the IRQ resources the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/pci/hotplug)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Unplugging the root of a nested PCIe bridge hierarchy leaks the IRQ resources the child bridges were using for hotplug notification, and the stale MSI state then trips a warning and a panic in the device teardown path. The node goes down on a device removal, taking every tenant on it.","attack_vector":"Reached by removing a bridge that has child bridges below it - either an administrative slot power-off through sysfs or a physical removal, so both an operator action and a device-driven event qualify. PowerNV only (the pnv_php driver on OpenPOWER POWER9/POWER10); inert on x86 and ARM GPU nodes. No tenant privilege involved on the affected platforms.","remediation":"Update to a kernel carrying the fix (no fixed_in published; stable commits below). Interim on PowerNV: drain the node before removing any bridge that has child bridges behind it.","references":["https://git.kernel.org/stable/c/8c1ad4af160691e157d688ad9619ced2df556aac","https://git.kernel.org/stable/c/912e200240b6f9758f0b126e64a61c9227f4ad37","https://nvd.nist.gov/vuln/detail/CVE-2025-38624"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-476","CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-38674","cve":"CVE-2025-38674","aliases":[],"title":"Linux kernel (drivers/gpu/drm): The dma_buf pointer cached on a GEM object goes stale the moment userspace drops the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The dma_buf pointer cached on a GEM object goes stale the moment userspace drops the last handle on that buffer, so a tenant that closes a handle at the right time while a PRIME export or import path is still walking the object gets the kernel to follow a dead pointer. The result is a kernel oops in shared DRM core code - code every driver on the node runs, not just one vendor's.","attack_vector":"A tenant container holding /dev/dri/renderD* drives it entirely with buffer-sharing ioctls: export a GEM object to a dma-buf fd, then release the last GEM handle while another thread is still operating on the shared buffer. No display access, no privileged capability, and the affected file is drm_prime.c in DRM core, so amdgpu, i915, xe, nouveau and virtio-gpu nodes are all in scope.","remediation":"Update to a kernel containing the revert commits below. There is no useful interim control other than denying /dev/dri access to untrusted workloads, since PRIME buffer sharing is used by essentially every GPU client.","references":["https://git.kernel.org/stable/c/5f05d83ce689a8930a70dfa73f879604aef8cc03","https://git.kernel.org/stable/c/fb4ef4a52b79a22ad382bfe77332642d02aef773","https://nvd.nist.gov/vuln/detail/CVE-2025-38674"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-674"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-38690","cve":"CVE-2025-38690","aliases":[],"title":"Linux kernel (drivers/gpu/drm/xe): The migration copy path falls back to a stack bounce buffer when the tenant's buffer","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/gpu/drm/xe)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The migration copy path falls back to a stack bounce buffer when the tenant's buffer is not cacheline-aligned, but that bounce buffer has no alignment guarantee either, so the function recurses into itself until the kernel stack is exhausted. The result is a kernel panic triggered on demand by one tenant, taking the node - and every other tenant's job on it - down.","attack_vector":"A tenant container holding /dev/dri/renderD* on an Intel xe GPU passes a misaligned buffer/offset into the VRAM access path; upstream hit it through the GPU debug interface, and any caller that reaches xe_migrate's non-aligned copy fallback does the same. Unprivileged, render-node only.","remediation":"Boot a kernel carrying the xe_migrate bounce-buffer fix below. Interim: deny the GPU debug/eudebug interface to tenant containers, and drop /dev/dri/renderD* from workloads that do not need direct VRAM access.","references":["https://git.kernel.org/stable/c/89f511c024879c5812cc0c010a6663b5e49950f3","https://git.kernel.org/stable/c/9d7a1cbebbb691891671def57407ba2f8ee914e8","https://nvd.nist.gov/vuln/detail/CVE-2025-38690"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-38705","cve":"CVE-2025-38705","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): A NULL pointer dereference in the amdgpu power","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/pm: fix null pointer access","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38705","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-09-04"},{"id":"CVE-2025-39675","cve":"CVE-2025-39675","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add null pointer check in mod_hdcp_hdcp1_create_session()","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-39675","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-09-05"},{"id":"CVE-2025-39693","cve":"CVE-2025-39693","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Avoid a NULL pointer dereference","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-39693","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-09-05"},{"id":"CVE-2025-39705","cve":"CVE-2025-39705","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: fix a Null pointer dereference vulnerability","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-39705","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-09-05"},{"id":"CVE-2025-39706","cve":"CVE-2025-39706","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A NULL pointer dereference in the amdkfd (KFD compute","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdkfd (KFD compute driver, /dev/kfd). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdkfd: Destroy KFD debugfs after destroy KFD wq","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-39706","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-09-05"},{"id":"CVE-2025-39707","cve":"CVE-2025-39707","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amdgpu): A NULL pointer dereference in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amdgpu)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: check if hubbub is NULL in debugfs/amdgpu_dm_capabilities","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-39707","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-09-05"},{"id":"CVE-2025-39762","cve":"CVE-2025-39762","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: add null check","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-39762","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-09-11"},{"id":"CVE-2025-39936","cve":"CVE-2025-39936","aliases":[],"title":"Linux crypto/ccp - SEV platform shutdown error handling: The ccp driver's SEV/SNP platform shutdown path could","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux crypto/ccp - SEV platform shutdown error handling","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The ccp driver's SEV/SNP platform shutdown path could be called without a valid error pointer, dereferencing it and panicking the host. Since this is on the SEV platform teardown path, the crash lands during operations like driver unload or SNP re-initialisation - the exact moments you are already doing maintenance on a confidential-computing host.","attack_vector":"Local, in the host's SEV platform management path; reachable by whatever drives SEV init/shutdown, i.e. host administration rather than tenants.","remediation":"Fixed in the Linux kernel - KVM/x86 SEV code or the ccp/PSP driver. Take the distro kernel update (RHEL/Rocky, Ubuntu, SLES) and **reboot the host**; SEV/SNP hypervisor paths cannot be live-patched in any meaningful way, and SNP platform init/shutdown is not safe to cycle under running guests. Drain confidential-VM tenants, reboot, then re-admit. No firmware, VBIOS or AGESA step needed, which makes this one of the cheaper classes of SEV fix to roll out.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-39936"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-10-04"},{"cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-40032","cve":"CVE-2025-40032","aliases":[],"title":"Linux kernel (drivers/pci/endpoint/functions): The endpoint test function releases DMA channels it may never have","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/pci/endpoint/functions)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The endpoint test function releases DMA channels it may never have acquired, dereferencing NULL inside the DMA core. The dereference happens in an interrupt thread on the endpoint machine and panics it, so the card or DPU drops off the link entirely.","attack_vector":"Endpoint mode required, and the trigger comes from across the link: the upstream panic runs from tegra_pcie_ep_pex_rst_irq, the interrupt raised when the connected HOST asserts PERST#. Any host reset or reboot therefore reaches this path with no authentication of any kind. Inert on a conventional GPU server, live on hardware running Linux as a PCIe endpoint whose host side is not under your control.","remediation":"Update to a kernel carrying the fix (no fixed_in published; stable commits below). Interim: unbind the pci-epf-test function driver - it is a test function and should not be bound in production endpoint configurations at all.","references":["https://git.kernel.org/stable/c/6411f840a9b5c47c00ca8e004733de232553870d","https://git.kernel.org/stable/c/0c5ce6b6ccc22d486cc7239ed908cb0ae5363a7b","https://nvd.nist.gov/vuln/detail/CVE-2025-40032"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-40034","cve":"CVE-2025-40034","aliases":[],"title":"Linux kernel (drivers/pci/pcie): AER's rate limiter dereferences per-device error state without checking it exists.","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/pci/pcie)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"AER's rate limiter dereferences per-device error state without checking it exists. When firmware reports an error against a device that has no AER capability, that state is NULL and the node panics inside the error-handling worker - the machine dies while processing an error report rather than logging it.","attack_vector":"Reached through the ACPI APEI/GHES path: platform firmware reports a hardware error and names a source device that does not advertise an AER capability. The upstream crash names an Intel Sky Lake-E DMI device - i.e. a root-complex-internal device on a mainstream server chipset, so the affected hardware is ordinary datacenter silicon, not embedded. The error events that make this path run come from real PCIe traffic, which on a passthrough node includes a tenant driving its own GPU or NIC into generating errors; the tenant does not choose which device firmware blames, so treat this as device/firmware-driven rather than precisely targetable.","remediation":"Update to a kernel carrying the fix (no fixed_in published; stable commits below). Interim: on affected Intel server platforms, check whether firmware-first error handling (GHES) is enabled in BIOS and consider native AER handling instead, and monitor GHES-reported corrected-error rates.","references":["https://git.kernel.org/stable/c/41683624cbff0a26bb7e0627f4a7e1b51a8779a8","https://git.kernel.org/stable/c/deb2f228388ff3a9d0623e3b59a053e9235c341d","https://nvd.nist.gov/vuln/detail/CVE-2025-40034"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-369","CWE-190"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-40293","cve":"CVE-2025-40293","aliases":[],"title":"Linux kernel (drivers/iommu/iommufd): A user-supplied page shift of 63 overflows the divisor in the iommufd","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/iommufd)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A user-supplied page shift of 63 overflows the divisor in the iommufd dirty-tracking bitmap math to zero, giving a divide-by-zero in kernel context. A tenant holding /dev/iommu takes down its own host thread and, on a panic_on_oops fleet, the whole node - a noisy-neighbour outage for every other tenant on the box.","attack_vector":"Any process with /dev/iommu open requests a dirty-tracking bitmap read with an absurd page size, so pgshift reaches 63 and BITS_PER_TYPE * pgsize wraps to zero. Pure ioctl input validation on an fd a passthrough tenant already holds; no device or host root required. Conditional on iommufd being in use and dirty tracking being reachable by the caller.","remediation":"No fixed release is listed in this record; apply the linked stable commits or run a current stable/LTS kernel. Interim: keep /dev/iommu out of containers that do not perform passthrough, and validate page-size arguments in the VMM layer that proxies dirty tracking.","references":["https://git.kernel.org/stable/c/07105e61882ff4a7d58db63cc5f9e90c6c60506c","https://git.kernel.org/stable/c/4c8a4f1d34eced168cc0b3a3dfe7b6dcc2090f69","https://nvd.nist.gov/vuln/detail/CVE-2025-40293"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-209"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2025-40760","cve":"CVE-2025-40760","aliases":["SSA-514895"],"title":"Altair Grid Engine (error message handling): Grid Engine leaks password hash material in error messages. Hashes lifted","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Altair Grid Engine (error message handling)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Grid Engine leaks password hash material in error messages. Hashes lifted from a scheduler are directly crackable offline, and the accounts they belong to are the ones with rights over job placement.","attack_vector":"A local user able to provoke the error condition on a Grid Engine host. Affects all versions before V2026.0.0.","remediation":"Upgrade Altair Grid Engine to V2026.0.0 and restart the daemons. Rotate the credentials for any account whose hash could have been exposed. Note that Altair advisories now ship through Siemens ProductCERT, not Altair's own site.","references":["https://cert-portal.siemens.com/productcert/html/ssa-514895.html","https://nvd.nist.gov/vuln/detail/CVE-2025-40760"],"status":"curated"},{"id":"CVE-2025-57275","cve":"CVE-2025-57275","aliases":["SPDK NVMe-oF target buffer overflow","SPDK lib/nvmf"],"title":"SPDK (Storage Performance Development Kit) 25.05 - NVMe-oF target, lib/nvmf: A buffer overflow in the NVMe-oF target","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"SPDK (Storage Performance Development Kit) 25.05 - NVMe-oF target, lib/nvmf","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A buffer overflow in the NVMe-oF target component of SPDK 25.05. SPDK is the userspace, poll-mode NVMe-oF target most commonly deployed by neoclouds and storage vendors precisely because it outperforms the kernel target, so it fronts tenant namespaces in exactly the environments this database is aimed at. NVD scores it as requiring high privileges with limited integrity impact plus availability loss, so the realistic outcome is a target crash - taking every attached tenant's I/O with it - rather than a clean takeover. Worth tracking because NeVerMore separately verified seven NVMe-oF protocol attacks against SPDK, so the target's overall exposure is broader than this single defect.","attack_vector":"Reached through the NVMe-oF target path in lib/nvmf on SPDK 25.05. NVD's vector puts it at network-adjacent reachability with high privileges required, which in practice means an authenticated or otherwise privileged initiator context rather than an anonymous peer. Public detail is thin - the advisory text is a one-line description with no reproducer.","remediation":"Upgrade SPDK past 25.05 and restart the target process - no kernel change, no host reboot, but the restart drops all NVMe-oF connections, so run it behind multipath initiators or during a maintenance window per storage node. Combine with the NVMe-oF hardening in the discovery-controller entry: in-band DH-HMAC-CHAP, per-subsystem host allow-lists, and discovery on a management-only interface. Since SPDK is a library embedded in vendor and in-house appliances, check with your storage vendor which SPDK release their firmware ships rather than assuming the host package version is what is running.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-57275","https://arxiv.org/abs/2202.08080"],"status":"curated","tags":["tenant-isolation"],"published":"2025-10-01"},{"cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-68309","cve":"CVE-2025-68309","aliases":[],"title":"Linux kernel (drivers/pci/pcie): The AER subsystem allocates its per-device error-tracking structure without checking","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/pci/pcie)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The AER subsystem allocates its per-device error-tracking structure without checking for failure, then dereferences it unconditionally. If that allocation fails, every subsequent AER access is a NULL dereference and the node panics - so a burst of PCIe errors arriving while the machine is under memory pressure turns into a full node outage for every tenant.","attack_vector":"Precondition is an allocation failure, which is the honest limiter here - but both halves are things a tenant supplies on a busy GPU node. Memory pressure is the steady state on a node packed with tenants, and the error events are generated by the devices themselves: a tenant with a passthrough GPU, NIC or NVMe behind /dev/vfio/* can drive its own device into producing correctable and uncorrectable PCIe errors at will, which is what makes the AER path run in the first place. No host credentials required for the error-generating half.","remediation":"Update to a kernel carrying the fix (no fixed_in published; stable commits below). Interim: keep real headroom on nodes so the allocation does not fail, and alert on AER error rates per device so a tenant hammering its passthrough device into an error storm is visible before it matters.","references":["https://git.kernel.org/stable/c/6618243bcc3f60825f761a41ed65fef9fe97eb25","https://git.kernel.org/stable/c/0a27bdb14b028fed30a10cec2f945c38cb5ca4fa","https://nvd.nist.gov/vuln/detail/CVE-2025-68309"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-908","CWE-457"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-71096","cve":"CVE-2025-71096","aliases":[],"title":"Linux kernel RDMA core address resolution (RDMA_NL_LS_OP_IP_RESOLVE netlink handler): The netlink handler for","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel RDMA core address resolution (RDMA_NL_LS_OP_IP_RESOLVE netlink handler)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The netlink handler for kernel-initiated RDMA address resolution assumed userspace would always supply the destination GID attribute and never verified it. Omitting it leaves the DGID buffer uninitialized and the kernel then formats and consumes leftover stack contents as a fabric address - a KMSAN-confirmed uninitialized read whose value userspace can influence by shaping the stack beforehand. The practical worry on a shared node is not just the info leak but that a bogus DGID is what the kernel then tries to route to.","attack_vector":"Local. A userspace RDMA netlink service (the rdma-ndd / ibacm style daemon slot) replies to a kernel LS_IP_RESOLVE query without the DGID attribute. Requires the ability to bind the RDMA netlink LS service in the namespace.","remediation":"Kernel update that parses the attributes with nla_parse_deprecated() and fails cleanly when DGID is absent. Verify no tenant workload can register itself as the RDMA netlink LS responder - that slot should belong to a host daemon only.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=0b948afc1ded88b3562c893114387f34389eeb94","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2025/CVE-2025-71096.json"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-476","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-71233","cve":"CVE-2025-71233","aliases":[],"title":"Linux kernel (drivers/pci/endpoint): Endpoint function sub-groups were created asynchronously by a delayed work item","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/pci/endpoint)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Endpoint function sub-groups were created asynchronously by a delayed work item, so removing the directory before the work ran left the worker dereferencing a freed parent - a NULL/dangling dereference inside a kernel workqueue that panics the machine.","attack_vector":"Configfs-driven and trivially reproducible: a loop of mkdir/rmdir under /sys/kernel/config/pci_ep/functions/<driver>/ crashes the kernel within about twenty iterations. That is host root on a machine with the PCI endpoint framework and configfs mounted - so the realistic exposure is an endpoint device whose function configuration is scripted or exposed to a management agent, not a tenant. Inert on a conventional GPU server with no endpoint controller.","remediation":"Update to a kernel carrying the fix (no fixed_in published; stable commits below). Interim: do not expose /sys/kernel/config/pci_ep to any automation that creates and destroys function directories rapidly, and keep configfs unmounted on endpoint machines whose function set is static.","references":["https://git.kernel.org/stable/c/fa9fb38f5fe9c80094c2138354d45cdc8d094d69","https://git.kernel.org/stable/c/5f609b3bffd4207cf9f2c9b41e1978457a5a1ea9","https://nvd.nist.gov/vuln/detail/CVE-2025-71233"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-71293","cve":"CVE-2025-71293","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/ras): A NULL pointer dereference in the amdgpu RAS /","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/ras)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu RAS / GPU reset and recovery path. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu/ras: Move ras data alloc before bad page check","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-71293","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-05-06"},{"id":"CVE-2025-71294","cve":"CVE-2025-71294","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): A NULL pointer dereference in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu firmware, ACPI and IP-block initialisation. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: fix NULL pointer issue buffer funcs","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-71294","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-05-06"},{"id":"CVE-2025-8404","cve":"CVE-2025-8404","aliases":[],"title":"A shared library inside Supermicro BMC firmware that parses request headers: An authenticated attacker overflows","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"A shared library inside Supermicro BMC firmware that parses request headers","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An authenticated attacker overflows a stack buffer during header parsing and executes code in the BMC firmware operating system. The shared-library location is what makes this worth flagging separately: fixing one web endpoint does not fix it, and the same primitive is likely reachable from whichever BMC service an operator has left enabled. Outcome is the usual BMC-root outcome - out-of-band power, console, virtual media and firmware persistence under the host. Because it is shared, the same overflow is reachable from more than one front-end service on the controller rather than from a single CGI endpoint.","attack_vector":"Any authenticated session that reaches a BMC service using this library over the network. Because it is shared code, restricting one interface does not close it.","remediation":"Firmware flash from Supermicro's November 2025 BMC/IPMI batch. Disabling individual BMC services is a weaker mitigation than usual here, since the bug lives in shared parsing code rather than one handler - so treat network isolation of the management VLAN plus per-node unique BMC credentials as the interim control, and prioritise the flash. Expect the fixed image to be board-specific.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-8404","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2025/8xxx/CVE-2025-8404.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2026-12480","cve":"CVE-2026-12480","aliases":[],"title":"Keras (HDF5 ExternalLink, incomplete fix): Arbitrary HDF5 file read","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Keras (HDF5 ExternalLink, incomplete fix)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Arbitrary HDF5 file read; incomplete fix for CVE-2026-1669","attack_vector":"Customer-supplied model file","remediation":"Upgrade past 3.13.2","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-12480"],"status":"curated","published":"2026-07-01"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-532"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2026-19483","cve":"CVE-2026-19483","aliases":[],"title":"IBM Storage Scale management GUI (deploy and upgrade logging): The Storage Scale admin password is written in the clear","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"IBM Storage Scale management GUI (deploy and upgrade logging)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The Storage Scale admin password is written in the clear into GUI logs, so anyone who can read those logs - including support bundle recipients - gets the credential that controls the storage management plane.","attack_vector":"Local read access to GUI log files on a Storage Scale 5.2.3.0-5.2.3.8 or 6.0.0.0-6.0.1.0 management node, or possession of any support snap collected from one.","remediation":"Apply the fix from IBM's bulletin, then rotate the admin password, purge the affected log files, and recall or destroy any support bundles already sent off site.","references":["https://www.ibm.com/support/pages/node/7283308","https://nvd.nist.gov/vuln/detail/CVE-2026-19483"],"status":"curated"},{"id":"CVE-2026-22703","cve":"CVE-2026-22703","aliases":[],"title":"cosign / sigstore: A crafted bundle verifies successfully even though the embedded Rekor entry does not reference","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"cosign / sigstore","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A crafted bundle verifies successfully even though the embedded Rekor entry does not reference the artifact; signature policy bypass","attack_vector":"Malicious image","remediation":"Upgrade cosign to 2.6.2/3.0.4+; re-verify admitted images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-22703"],"status":"curated","published":"2026-01-10"},{"id":"CVE-2026-23163","cve":"CVE-2026-23163","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A NULL pointer dereference in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu GEM/VM/command-submission ioctl surface. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: fix NULL pointer dereference in amdgpu_gmc_filter_faults_remove","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-23163","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-02-14"},{"id":"CVE-2026-23213","cve":"CVE-2026-23213","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): A correctness defect in the amdgpu power management","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu power management (SMU/powerplay) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/pm: Disable MMIO access during SMU Mode 1 reset","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-23213","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-02-18"},{"id":"CVE-2026-23338","cve":"CVE-2026-23338","aliases":[],"title":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu/userq): A race condition or locking defect","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu/userq)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdgpu user-mode queues (doorbell submission path). Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdgpu/userq: Do not allow userspace to trivially triger kernel warnings","attack_vector":"Local. Reachable by any process with a render node open that can create user-mode queues - the normal ROCm submission path, reachable from an unprivileged container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-23338","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-03-25"},{"id":"CVE-2026-23358","cve":"CVE-2026-23358","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): Memory is handed to a consumer without being","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Memory is handed to a consumer without being initialised or cleared in the amdgpu RAS / GPU reset and recovery path. Whatever the previous owner left behind is readable - and on a GPU node the previous owner is very often a different tenant's job. This is the classic residual-data leak between workloads sharing a card: model weights, activations, keys or tokens from the prior tenant can surface in a fresh allocation. Upstream fix: drm/amdgpu: Fix error handling in slot reset","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-23358","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-03-25"},{"id":"CVE-2026-23435","cve":"CVE-2026-23435","aliases":[],"title":"Linux perf/x86 - event pointer setup ordering in x86_pmu_enable(): A NULL pointer dereference in the x86 PMU enable","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux perf/x86 - event pointer setup ordering in x86_pmu_enable()","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the x86 PMU enable path, reported from a production AMD EPYC system. Performance counters are enabled by every profiling and observability agent on a fleet, so this crashes hosts through the monitoring stack rather than through anything a tenant did - and it takes co-resident GPU jobs with it.","attack_vector":"Local, through perf event enablement. Reachable by whatever has perf access, which on many clusters includes node-level observability agents and, if perf_event_paranoid is relaxed, tenants.","remediation":"Fixed in the Linux kernel. Take the distro kernel update (RHEL/Rocky, Ubuntu, SLES) and reboot the host - no firmware, VBIOS or AGESA step. On a GPU fleet this is a cordon, drain and rolling reboot; plan it as normal kernel maintenance. Check perf_event_paranoid on GPU nodes: if you have loosened it so tenants can profile their own kernels - which is a reasonable thing to want on an AI cluster - you have also widened who can reach this.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-23435"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-04-03"},{"id":"CVE-2026-23468","cve":"CVE-2026-23468","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): An out-of-bounds access in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An out-of-bounds access in the amdgpu firmware, ACPI and IP-block initialisation - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: Limit BO list entry count to prevent resource exhaustion","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-23468","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-04-03"},{"id":"CVE-2026-24160","cve":"CVE-2026-24160","aliases":[],"title":"TensorRT-LLM: DoS (null deref in tensor ops)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT-LLM","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"DoS (null deref in tensor ops)","attack_vector":"Malicious inference input","remediation":"Bump TensorRT-LLM; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24160","https://github.com/NVIDIA/product-security/tree/main/2026/5805"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:H","cwe":["CWE-690"],"published":"2026-05-20"},{"id":"CVE-2026-31460","cve":"CVE-2026-31460","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: check if ext_caps is valid in BL setup","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-31460","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-04-22"},{"id":"CVE-2026-31461","cve":"CVE-2026-31461","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A memory or reference-count leak in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu display core (DC/DM). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amd/display: Fix drm_edid leak in amdgpu_dm","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-31461","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-04-22"},{"id":"CVE-2026-31462","cve":"CVE-2026-31462","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): A correctness defect in the amdgpu kernel driver core reachable","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu kernel driver core reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu: prevent immediate PASID reuse case","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-31462","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-04-22"},{"cwe":["CWE-665"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-31492","cve":"CVE-2026-31492","aliases":[],"title":"Linux kernel (drivers/infiniband/hw/irdma): If copying the queue-pair response back to userspace fails, irdma's","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/hw/irdma)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"If copying the queue-pair response back to userspace fails, irdma's create-QP error path calls destroy, which waits on a completion structure that was never initialized. A tenant that supplies a bad response buffer walks the kernel into a wait on uninitialized memory - an indefinite hang or corruption on a node shared with other tenants.","attack_vector":"Local and unprivileged: a tenant holding /dev/infiniband/uverbs* on an Intel irdma adapter issues create-QP with a udata buffer that cannot be written (unmapped or partially mapped), forcing ib_copy_to_udata to fail. No fabric peer or root needed. Conditional on irdma hardware (E810 / X722 RDMA) being present.","remediation":"No fixed version is recorded in this entry; boot a stable kernel carrying the completion-init fix (commits ac1da7bd224d / af310407f79d). Interim: remove /dev/infiniband device nodes from untrusted tenant containers on irdma nodes.","references":["https://git.kernel.org/stable/c/ac1da7bd224d406b6f1b84414f0f652ab43b6bd8","https://git.kernel.org/stable/c/af310407f79d5816fc0ab3638e1588b6193316dd","https://nvd.nist.gov/vuln/detail/CVE-2026-31492"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-31540","cve":"CVE-2026-31540","aliases":[],"title":"Linux i915 GPU kernel driver (submission backend setup): i915 dereferences the submission backend before checking","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux i915 GPU kernel driver (submission backend setup)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"i915 dereferences the submission backend before checking it is set, which happens when the GuC/HuC firmware binaries are absent. Node panics on GPU init instead of degrading. Common on minimal host images that omit linux-firmware.","attack_vector":"No attacker; triggered by a node image missing GPU firmware blobs.","remediation":"Kernel update plus reboot, and fix the node image so linux-firmware carries the GuC/HuC blobs for the installed GPUs.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-31540","https://git.kernel.org/stable/c/0162ab3220bac870e43e229e6e3024d1a21c3f26"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-04-24"},{"id":"CVE-2026-31591","cve":"CVE-2026-31591","aliases":[],"title":"Linux KVM/SEV - vCPU locking when synchronizing VMSAs for SNP launch finish: KVM did not lock all vCPUs","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux KVM/SEV - vCPU locking when synchronizing VMSAs for SNP launch finish","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"KVM did not lock all vCPUs while synchronising and encrypting VMSAs at SNP launch finish, so vCPU state could change underneath the encryption step. The VMSA is what the launch measurement covers; if it can move while being measured, the attestation you hand the tenant does not necessarily describe the VM that actually ran.","attack_vector":"Through the KVM SNP launch path, from the VMM process.","remediation":"Fixed in the Linux kernel. Take the distro kernel update (RHEL/Rocky, Ubuntu, SLES) and reboot the host - no firmware, VBIOS or AGESA step. On a GPU fleet this is a cordon, drain and rolling reboot; plan it as normal kernel maintenance.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-31591"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-04-24"},{"id":"CVE-2026-31593","cve":"CVE-2026-31593","aliases":[],"title":"Linux KVM - VMSA sync on an already-launched SEV vCPU: KVM allowed synchronising vCPU state into the VMSA","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux KVM - VMSA sync on an already-launched SEV vCPU","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"KVM allowed synchronising vCPU state into the VMSA after the VMSA had already been encrypted and the guest launched. Writing to an encrypted VMSA behind the guest's back corrupts the confidential vCPU state that attestation covered - the guest is no longer the thing that was measured, and the failure is silent rather than loud.","attack_vector":"Via the KVM ioctl surface, from the VMM process managing the guest.","remediation":"Fixed in the Linux kernel - KVM/x86 SEV code or the ccp/PSP driver. Take the distro kernel update (RHEL/Rocky, Ubuntu, SLES) and **reboot the host**; SEV/SNP hypervisor paths cannot be live-patched in any meaningful way, and SNP platform init/shutdown is not safe to cycle under running guests. Drain confidential-VM tenants, reboot, then re-admit. No firmware, VBIOS or AGESA step needed, which makes this one of the cheaper classes of SEV fix to roll out.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-31593"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-04-24"},{"id":"CVE-2026-31689","cve":"CVE-2026-31689","aliases":[],"title":"Linux EDAC/mc - error path ordering in edac_mc_alloc(): When a private-data allocation fails in edac_mc_alloc()","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux EDAC/mc - error path ordering in edac_mc_alloc()","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"When a private-data allocation fails in edac_mc_alloc(), the error path unwinds in the wrong order and touches memory it has already released. EDAC is the memory error-detection and correction subsystem - the thing telling you which DIMM is going bad on a node full of expensive HBM-adjacent DRAM - so a defect in its allocation path costs you both stability and the RAS visibility you were relying on.","attack_vector":"Local, on the EDAC allocation error path - hit under memory pressure rather than by an attacker.","remediation":"Distro kernel update plus reboot; no firmware step. Worth taking on any fleet where you drive node retirement off EDAC data.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-31689"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-04-27"},{"id":"CVE-2026-31765","cve":"CVE-2026-31765","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdgpu): A NULL pointer dereference in the amdkfd (KFD compute","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdgpu)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdkfd (KFD compute driver, /dev/kfd). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: Change AMDGPU_VA_RESERVED_TRAP_SIZE to 64KB","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-31765","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-05-01"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:N/A:N","cwe":["CWE-59"],"fleet":{"pain_class":"hot-patch"},"id":"CVE-2026-40610","cve":"CVE-2026-40610","aliases":["GHSA-mcfx-4vc6-qgxv"],"title":"BentoML (bentoml build, symlink dereferencing in the build context): bentoml build follows symlinks inside the build","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"BentoML (bentoml build, symlink dereferencing in the build context)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"bentoml build follows symlinks inside the build context and copies the target's contents into the bento. A symlink planted in an untrusted repository pulls whatever the builder can read - config files, tokens, key material - into the packaged artifact without the builder noticing.","attack_vector":"A victim who runs bentoml build against an attacker-supplied repository or build context. Local to the build machine.","remediation":"Upgrade BentoML to 1.4.39 or later. Until then, build only in a container whose filesystem contains nothing you would not ship, and inspect the file list of bentos built from third-party sources before distributing them.","references":["https://github.com/bentoml/BentoML/security/advisories/GHSA-mcfx-4vc6-qgxv","https://nvd.nist.gov/vuln/detail/CVE-2026-40610"],"status":"curated"},{"id":"CVE-2026-43034","cve":"CVE-2026-43034","aliases":[],"title":"Linux bnxt_en driver (backing store type from firmware response): A second firmware-controlled-index bug in the same","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux bnxt_en driver (backing store type from firmware response)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A second firmware-controlled-index bug in the same driver: the backing-store type returned in a firmware response is stored and later used to index a fixed array. Same class as the DBG_BUF_PRODUCER issue and listed separately because it is fixed by a different commit — an operator matching only one CVE will still be running the other.","attack_vector":"The NIC firmware's response content.","remediation":"Kernel/driver upgrade plus host reboot; same patch cycle as the other bnxt_en firmware-input issues.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43034"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-05-01"},{"id":"CVE-2026-43130","cve":"CVE-2026-43130","aliases":[],"title":"Linux iommu/vt-d (dev-IOTLB flush in scalable mode): The scalable-mode half of the device-IOTLB invalidation problem —","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux iommu/vt-d (dev-IOTLB flush in scalable mode)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The scalable-mode half of the device-IOTLB invalidation problem — ATS invalidation is issued or skipped based on device accessibility, and the earlier fix in this area left a gap. Scalable mode is what modern Intel platforms use for PASID-based device assignment, so this is the current-generation path for handing NIC and accelerator functions to tenants. Same underlying concern: a device retaining stale translations after the host revoked them.","attack_vector":"A tenant with a scalable-mode-assigned, ATS-capable PCIe function.","remediation":"Kernel upgrade plus host reboot, rolling across passthrough-capable nodes. Verify after patching that IOMMU is in enforcing (not passthrough/`iommu=pt`) mode for tenant-assigned devices — a surprising number of performance-tuned GPU hosts run with IOMMU translation effectively disabled, which makes this class of bug moot only because the isolation was never there.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43130"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-05-06"},{"id":"CVE-2026-43131","cve":"CVE-2026-43131","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): A NULL pointer dereference in the amdgpu power","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/pm: Fix null pointer dereference issue","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43131","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-05-06"},{"id":"CVE-2026-43161","cve":"CVE-2026-43161","aliases":[],"title":"Linux iommu/vt-d (dev-IOTLB flush for passed-through PCIe devices): The Intel IOMMU driver skips device-IOTLB","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux iommu/vt-d (dev-IOTLB flush for passed-through PCIe devices)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The Intel IOMMU driver skips device-IOTLB invalidation for PCIe endpoints that are ATS-enabled and passed through to userspace — the exact configuration used for SR-IOV NIC VFs handed to a VM, and for DPDK and RDMA userspace drivers on a GPU node. The device-IOTLB is the device's own cached copy of the address translations; if it is not flushed when a mapping is revoked, the device can keep reaching memory the kernel believes it has taken away. That is the core mechanism DMA isolation depends on, and passthrough NICs are precisely the devices you hand to untrusted tenants. Companion issue CVE-2026-43130 covers the scalable-mode variant.","attack_vector":"A tenant holding a passed-through, ATS-enabled PCIe device — an SR-IOV VF assigned to their VM, or a DPDK/RDMA userspace-bound NIC — able to exercise the device after a mapping has been torn down.","remediation":"Kernel upgrade plus host reboot — rolling across every node that does device passthrough, which on a GPU cloud is all of them. Nothing to flash. If you cannot patch immediately, the meaningful mitigation is disabling ATS on passed-through endpoints (a BIOS/kernel-parameter change, at a measurable performance cost) or not passing devices through to untrusted tenants at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43161","https://nvd.nist.gov/vuln/detail/CVE-2026-43130"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-05-06"},{"id":"CVE-2026-43191","cve":"CVE-2026-43191","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Adjust PHY FSM transition to TX_EN-to-PLL_ON for TMDS on DCN35","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43191","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-05-06"},{"id":"CVE-2026-43195","cve":"CVE-2026-43195","aliases":[],"title":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu): Missing or insufficient validation of","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu user-mode queues (doorbell submission path). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdgpu: validate user queue size constraints","attack_vector":"Local. Reachable by any process with a render node open that can create user-mode queues - the normal ROCm submission path, reachable from an unprivileged container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43195","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-05-06"},{"id":"CVE-2026-43243","cve":"CVE-2026-43243","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Add signal type check for dcn401 get_phyd32clk_src","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43243","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-05-06"},{"id":"CVE-2026-43298","cve":"CVE-2026-43298","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amdgpu): A race condition or locking defect in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amdgpu)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdgpu display core (DC/DM). Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdgpu: Skip vcn poison irq release on VF","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43298","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-05-08"},{"id":"CVE-2026-43305","cve":"CVE-2026-43305","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A race condition or locking defect in the amdgpu display","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdgpu display core (DC/DM). Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amd/display: Fix mismatched unlock for DMUB HW lock in HWSS fast path","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43305","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-05-08"},{"id":"CVE-2026-43318","cve":"CVE-2026-43318","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amdgpu): A correctness defect in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amdgpu)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu: fix sync handling in amdgpu_dma_buf_move_notify","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43318","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-05-08"},{"id":"CVE-2026-43320","cve":"CVE-2026-43320","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Fix dsc eDP issue","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43320","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-05-08"},{"id":"CVE-2026-43337","cve":"CVE-2026-43337","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Fix NULL pointer dereference in dcn401_init_hw()","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43337","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-05-08"},{"id":"CVE-2026-43367","cve":"CVE-2026-43367","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amd): A NULL pointer dereference in the amdgpu kernel driver core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amd)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu kernel driver core. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd: Fix a few more NULL pointer dereference in device cleanup","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43367","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-05-08"},{"id":"CVE-2026-43369","cve":"CVE-2026-43369","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amd): A NULL pointer dereference in the amdgpu kernel driver core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amd)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu kernel driver core. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd: Fix NULL pointer dereference in device cleanup","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43369","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-05-08"},{"id":"CVE-2026-43398","cve":"CVE-2026-43398","aliases":[],"title":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu): Missing or insufficient validation of","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu user-mode queues (doorbell submission path). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdgpu: add upper bound check on user inputs in wait ioctl","attack_vector":"Local. Reachable by any process with a render node open that can create user-mode queues - the normal ROCm submission path, reachable from an unprivileged container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43398","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-05-08"},{"id":"CVE-2026-43399","cve":"CVE-2026-43399","aliases":[],"title":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu/userq): A memory or reference-count leak","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu/userq)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu user-mode queues (doorbell submission path). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amdgpu/userq: Fix reference leak in amdgpu_userq_wait_ioctl","attack_vector":"Local. Reachable by any process with a render node open that can create user-mode queues - the normal ROCm submission path, reachable from an unprivileged container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43399","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-05-08"},{"id":"CVE-2026-43400","cve":"CVE-2026-43400","aliases":[],"title":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu): Missing or insufficient validation of","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu user-mode queues (doorbell submission path). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdgpu: add upper bound check on user inputs in signal ioctl","attack_vector":"Local. Reachable by any process with a render node open that can create user-mode queues - the normal ROCm submission path, reachable from an unprivileged container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43400","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-05-08"},{"id":"CVE-2026-43444","cve":"CVE-2026-43444","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A correctness defect in the amdkfd (KFD compute","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdkfd (KFD compute driver, /dev/kfd) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdkfd: Unreserve bo if queue update failed","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43444","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-05-08"},{"cwe":["CWE-401","CWE-667"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-45880","cve":"CVE-2026-45880","aliases":[],"title":"Linux kernel (drivers/pci): A failed mmap of peer-to-peer DMA memory leaks the pgmap reference it took, and the leak is","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/pci)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A failed mmap of peer-to-peer DMA memory leaks the pgmap reference it took, and the leak is permanent - memunmap_pages() then blocks forever when the PCI device is removed. The node keeps a P2PDMA-capable device it can never release, and the removal thread is stuck unkillably, so device reclaim and driver replacement on that node stop working until it is rebooted.","attack_vector":"Reached from userspace by mmap()ing a p2pmem region, so it needs a process with access to the P2PDMA-exposing device - relevant on GPU/NVMe nodes because peer-to-peer DMA is what GPUDirect-style paths use. The leak occurs on the vm_insert_page() failure path, so an attacker needs to make that insert fail, which in practice means driving the node under memory pressure while looping the mapping. Conditional on CONFIG_PCI_P2PDMA and a driver that publishes p2pmem; nodes without it are unaffected.","remediation":"Boot a kernel with the percpu_ref_put() added to the p2pmem_alloc_mmap() failure path. Interim: do not expose p2pmem device nodes to tenant containers, and treat a hung PCI remove on a P2PDMA node as needing a reboot rather than waiting it out.","references":["https://git.kernel.org/stable/c/baa42b756d183a59572f3890981a3d32b8d05d40","https://git.kernel.org/stable/c/51b7181cfbedf289ce794b6d97a1c596c309ec38","https://nvd.nist.gov/vuln/detail/CVE-2026-45880"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-45947","cve":"CVE-2026-45947","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): A memory or reference-count leak","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu firmware, ACPI and IP-block initialisation. Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amdgpu: Fix memory leak in amdgpu_acpi_enumerate_xcc()","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-45947","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-05-27"},{"id":"CVE-2026-45976","cve":"CVE-2026-45976","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): A memory or reference-count leak in the amdgpu RAS","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu RAS / GPU reset and recovery path. Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amdgpu: Fix memory leak in amdgpu_ras_init()","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-45976","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-05-27"},{"id":"CVE-2026-45979","cve":"CVE-2026-45979","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A correctness defect in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu GEM/VM/command-submission ioctl surface reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu: clean up the amdgpu_cs_parser_bos","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-45979","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-05-27"},{"id":"CVE-2026-46220","cve":"CVE-2026-46220","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu/sdma4): A correctness defect in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu/sdma4)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu GEM/VM/command-submission ioctl surface reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/sdma4: replace BUG_ON with WARN_ON in fence emission","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-46220","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-05-28"},{"id":"CVE-2026-46229","cve":"CVE-2026-46229","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): Memory is handed to a consumer without being","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Memory is handed to a consumer without being initialised or cleared in the amdkfd (KFD compute driver, /dev/kfd). Whatever the previous owner left behind is readable - and on a GPU node the previous owner is very often a different tenant's job. This is the classic residual-data leak between workloads sharing a card: model weights, activations, keys or tokens from the prior tenant can surface in a fresh allocation. Upstream fix: drm/amdkfd: Clear VRAM on allocation to prevent stale data exposure","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-46229","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-05-28"},{"id":"CVE-2026-46245","cve":"CVE-2026-46245","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Fix dc_link NULL handling in HPD init","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-46245","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-06-03"},{"id":"CVE-2026-46276","cve":"CVE-2026-46276","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): Memory is handed to a consumer without being","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Memory is handed to a consumer without being initialised or cleared in the amdgpu RAS / GPU reset and recovery path. Whatever the previous owner left behind is readable - and on a GPU node the previous owner is very often a different tenant's job. This is the classic residual-data leak between workloads sharing a card: model weights, activations, keys or tokens from the prior tenant can surface in a fresh allocation. Upstream fix: drm/amdgpu: fix zero-size GDS range init on RDNA4","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-46276","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-06-08"},{"id":"CVE-2026-47630","cve":"CVE-2026-47630","aliases":[],"title":"NVIDIA Triton Inference Server: An absolute path traversal reachable from a local low-privileged account reaches code","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An absolute path traversal reachable from a local low-privileged account reaches code execution. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5865. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47630","https://github.com/NVIDIA/product-security/tree/main/2026/5865"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-36"],"published":"2026-08-18"},{"cwe":["CWE-908","CWE-200"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-52995","cve":"CVE-2026-52995","aliases":[],"title":"Linux kernel RDS connection info (uninitialised per-item buffer copied to userspace): The connection-info walkers hand","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel RDS connection info (uninitialised per-item buffer copied to userspace)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The connection-info walkers hand a per-item stack buffer to a visitor and then copy the full declared item length back to userspace regardless of how much the visitor filled in. When a connection is not in the UP state the IB visitors write only a subset - several u32 fields and an alignment hole are left holding whatever was on the kernel stack - and all of it is copied out. Any unprivileged process that can query RDS info harvests kernel stack bytes, which is the standard first step for defeating KASLR before using a corruption bug.","attack_vector":"Local, unprivileged. Query RDS connection info while at least one connection is not in the UP state - trivially arranged by the querying tenant.","remediation":"Kernel update zeroing the item buffer before each visitor call. Blacklisting rds removes the surface immediately.","references":["https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=0797b2e6901827694aa9c34c4c72118c8c97fba1","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-52995.json"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-53121","cve":"CVE-2026-53121","aliases":[],"title":"Linux amd-pstate - memory leak in amd_pstate_epp_cpu_init(): On failure to set the energy-performance preference","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux amd-pstate - memory leak in amd_pstate_epp_cpu_init()","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"On failure to set the energy-performance preference, amd_pstate_epp_cpu_init() returns without freeing what it allocated. Another slow leak in the AMD CPU frequency driver, hit on the error path - which is the path a misconfigured or partially-supported platform takes repeatedly rather than once.","attack_vector":"Local, on the amd-pstate initialisation error path.","remediation":"Distro kernel update plus reboot. Batch with the other amd-pstate fixes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53121"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-06-24"},{"id":"CVE-2026-53135","cve":"CVE-2026-53135","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Fix NULL deref and buffer over-read in SDP debugfs","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53135","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-06-25"},{"id":"CVE-2026-53142","cve":"CVE-2026-53142","aliases":[],"title":"Linux drm/xe GPU kernel driver (suspend/shutdown without display): The xe driver oopses on suspend or shutdown","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux drm/xe GPU kernel driver (suspend/shutdown without display)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The xe driver oopses on suspend or shutdown on configurations with no display probed - which is exactly the headless datacenter GPU configuration. Effect is a dirty shutdown rather than a clean one, which risks filesystem and checkpoint state on nodes that reboot under automation.","attack_vector":"No attacker; fires on normal shutdown of a headless xe node.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53142","https://git.kernel.org/stable/c/0f68ddfaaebfbb5581ee931779757d31f4dc9e24"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-06-25"},{"id":"CVE-2026-53144","cve":"CVE-2026-53144","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): Missing or insufficient validation of user-supplied","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdkfd (KFD compute driver, /dev/kfd). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdkfd: fix NULL dereference in get_queue_ids()","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53144","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-06-25"},{"cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-53280","cve":"CVE-2026-53280","aliases":[],"title":"Linux kernel (drivers/iommu): The reset-completion path re-attaches an IOMMU group's domain without checking that the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The reset-completion path re-attaches an IOMMU group's domain without checking that the group has one, so a group whose default domain never allocated crashes the host on the next device reset. A tenant resetting its own passthrough device takes the node down for everyone on it.","attack_vector":"A tenant holding /dev/vfio/* triggers a device reset (VFIO_DEVICE_RESET or device-fd release), on a device whose IOMMU group has a NULL domain. That NULL state requires a default-domain allocation failure at first probe - memory pressure or a driver error during node bring-up - so it is conditional, but once a node is in that state the crash is one tenant ioctl away.","remediation":"Update to a stable kernel carrying commits 17194cd0 / d769711f. Interim: check dmesg for default-domain allocation failures during boot and refuse to schedule tenants onto a node that logged one.","references":["https://git.kernel.org/stable/c/17194cd0dd236e732d116d50840d795ca50ef196","https://git.kernel.org/stable/c/d769711fcddd005f1e654b3bde547140917fe696","https://nvd.nist.gov/vuln/detail/CVE-2026-53280"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-53283","cve":"CVE-2026-53283","aliases":[],"title":"Linux iommu/amd - devid bounds check in __rlookup_amd_iommu(): The AMD IOMMU driver looked up device IDs without","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux iommu/amd - devid bounds check in __rlookup_amd_iommu()","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The AMD IOMMU driver looked up device IDs without bounds-checking them, so a device ID outside the expected range indexes past the array. Device enumeration walks every device on the PCI bus, and on a dense GPU node that bus is crowded - many accelerators, switches, NICs and bridges. An out-of-bounds read in the IOMMU's device lookup is a kernel memory-safety issue in the component enforcing DMA isolation.","attack_vector":"Local, triggered during IOMMU device registration and lookup. Influenced by what is on the PCI bus, so a malicious or malfunctioning device - or a device presented by a compromised BMC - can reach it.","remediation":"Fixed in the Linux kernel. Distro kernel update plus a node reboot; no firmware step.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53283"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-06-26"},{"id":"CVE-2026-53285","cve":"CVE-2026-53285","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Wrap DCN32 phantom-plane allocation in DC_RUN_WITH_PREEMPTION_ENABLED","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53285","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-06-26"},{"id":"CVE-2026-53293","cve":"CVE-2026-53293","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A race condition or locking defect","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdgpu GEM/VM/command-submission ioctl surface. Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdgpu: fix AMDGPU_INFO_READ_MMR_REG","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53293","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-06-26"},{"id":"CVE-2026-53313","cve":"CVE-2026-53313","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Avoid NULL dereference in dc_dmub_srv error paths","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53313","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-06-26"},{"id":"CVE-2026-53315","cve":"CVE-2026-53315","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amd/ras): A NULL pointer dereference in the amdgpu RAS / GPU","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amd/ras)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu RAS / GPU reset and recovery path. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/ras: Fix NULL deref in ras_core_get_utc_second_timestamp()","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53315","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-06-26"},{"id":"CVE-2026-53316","cve":"CVE-2026-53316","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amd/ras): A NULL pointer dereference in the amdgpu RAS / GPU","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amd/ras)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu RAS / GPU reset and recovery path. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/ras: Fix NULL deref in ras_core_ras_interrupt_detected()","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53316","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-06-26"},{"id":"CVE-2026-53345","cve":"CVE-2026-53345","aliases":[],"title":"Linux KVM - dirty-page tracking without a vCPU on a dying VM: KVM warned (and on panic_on_warn hosts, panicked)","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux KVM - dirty-page tracking without a vCPU on a dying VM","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"KVM warned (and on panic_on_warn hosts, panicked) when a page was marked dirty without an associated vCPU while a VM was being torn down. A guest that can arrange the teardown timing turns a warning into a host crash on any fleet running panic_on_warn - which plenty of hardened kernels do.","attack_vector":"From inside a guest, by timing VM teardown. Tenant-reachable.","remediation":"Fixed in the Linux kernel. Take the distro kernel update (RHEL/Rocky, Ubuntu, SLES) and reboot the host - no firmware, VBIOS or AGESA step. On a GPU fleet this is a cordon, drain and rolling reboot; plan it as normal kernel maintenance. If you run panic_on_warn on GPU hosts for crash-dump fidelity, note that it converts warning-class kernel bugs like this into fleet availability incidents - worth reviewing that setting alongside the patch.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53345"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-07-01"},{"cwe":["CWE-20","CWE-665"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-53372","cve":"CVE-2026-53372","aliases":[],"title":"Linux kernel (drivers/iommu/intel): VT-d accepted a PASID attachment to a nested domain whose parent has dirty tracking","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/intel)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"VT-d accepted a PASID attachment to a nested domain whose parent has dirty tracking configured, even though the kernel cannot track dirty pages in that configuration. Dirty pages are then silently dropped, so a tenant VM that is live-migrated or checkpointed comes back with stale memory - guest data corruption the operator produced, with no error anywhere.","attack_vector":"Reached by the VMM / control plane driving iommufd: attach a PASID to a nested domain while the nesting parent has IOMMU dirty tracking enabled. Conditional on VT-d scalable mode with nested translation and dirty-tracking-based live migration in use. Not a tenant-driven memory-safety break - the exposure is to the operator's own migration path.","remediation":"Update to a stable kernel carrying commits 3ea9ce75 / 9009c1af, which fails the attach up front instead of losing pages. Interim: do not combine PASID attachment to nested domains with IOMMU dirty tracking - disable dirty-tracking-based live migration for guests using vIOMMU nesting.","references":["https://git.kernel.org/stable/c/3ea9ce757bd3de955b56e7bc5672fc479e40b045","https://git.kernel.org/stable/c/9009c1af5458322469fa9a4371081a4449c5947d","https://nvd.nist.gov/vuln/detail/CVE-2026-53372"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-53376","cve":"CVE-2026-53376","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): Missing or insufficient validation of user-supplied","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdkfd (KFD compute driver, /dev/kfd). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdkfd: Add upper bound check for num_of_nodes","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53376","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-07-19"},{"cwe":["CWE-911","CWE-362"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-64476","cve":"CVE-2026-64476","aliases":[],"title":"Linux kernel (drivers/vfio/pci): The disable_idle_d3 power-management flag was a module-wide global that could change","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/vfio/pci)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The disable_idle_d3 power-management flag was a module-wide global that could change out from under devices already bound to vfio-pci, so runtime-PM get/put operations on a passed-through device become unbalanced. An unbalanced PM reference means a device can be dropped into D3hot while a tenant still owns it, or held in D0 forever - device state confusion on the passthrough boundary and a path to refcount underflow.","attack_vector":"Needs host root: loading, unloading, or reloading vfio-pci with a different disable_idle_d3 value, or writing the module parameter through sysfs, while devices are already bound to a vfio-pci variant driver. This is an operator/automation hazard rather than a tenant-reachable one - but it lands directly on devices tenants are actively using.","remediation":"Update to a stable kernel carrying commits 332d785f / 654710ef, which latches the flag per device at init. Interim: set disable_idle_d3 once at boot via modprobe.d and never change it or reload vfio-pci while tenant devices are bound.","references":["https://git.kernel.org/stable/c/332d785f9ae426eeeb92527872adf09d84101ba3","https://git.kernel.org/stable/c/654710ef3135c4546b20a903bc23a51b0c44d6c8","https://nvd.nist.gov/vuln/detail/CVE-2026-64476"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-696","CWE-754"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-64591","cve":"CVE-2026-64591","aliases":[],"title":"Linux kernel (drivers/iommu/intel): SVA bind and unbind are asymmetric on VT-d hardware without PCI/PRI - bind skips","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/intel)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"SVA bind and unbind are asymmetric on VT-d hardware without PCI/PRI - bind skips enabling I/O page faulting, unbind tries to disable it anyway. The unbind path hits a kernel WARNING every time a tenant closes an SVA-using GPU context, which on any node booted with panic_on_warn is an immediate node kill, and it leaves the device's IOPF enablement state out of step with reality.","attack_vector":"A tenant holding /dev/dri/renderD* on an Intel Xe GPU (the upstream report came from xe_vm_close_and_put) or any SVA-capable accelerator simply closes its GPU VM. Conditional on VT-d with SVA enabled and the device lacking PRI support. Repeatable at will from inside the container, with no host privilege.","remediation":"Update to 6.18.39 / 6.20 or later. Interim: do not boot tenant nodes with panic_on_warn, which is what turns this from a log line into an outage, and disable SVA where the workload does not need it.","references":["https://git.kernel.org/stable/c/bb354384f40bb087e7c44f0f0743a23fc945f91d","https://git.kernel.org/stable/c/477f8dec3b5ae54e8ef3c33bdd2257e47896512e","https://nvd.nist.gov/vuln/detail/CVE-2026-64591"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-772"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-68081","cve":"CVE-2026-68081","aliases":[],"title":"Linux kernel (arch/x86/kvm/vmx): When a nested VM-Enter fails on invalid guest state, KVM took an open-coded exit path","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (arch/x86/kvm/vmx)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"When a nested VM-Enter fails on invalid guest state, KVM took an open-coded exit path that never released the vmcs12 pages it had already pinned and mapped. An L1 guest that retries VMLAUNCH in a loop leaks pinned, unswappable host pages on every attempt, letting one tenant grind the node's memory down until the host OOMs - a noisy-neighbour outage for everything else on the box.","attack_vector":"Guest-driven and unprivileged inside the VM: the tenant loops VMLAUNCH/VMRESUME with deliberately invalid guest state in its vmcs12. Requires nested VMX to be exposed - Intel host with kvm_intel nested=1 (the default) and VMX in the guest's CPUID.","remediation":"Update to a stable kernel with the linked fix (no fixed release enumerated; take the branch carrying commit 2f2312c422fd). Interim controls: do not expose nested virtualization to tenants (kvm_intel.nested=0), and cap per-VM host memory with cgroup limits so a leaking guest cannot consume the whole node.","references":["https://git.kernel.org/stable/c/2f2312c422fd2695da772cecb30c69994b795964","https://git.kernel.org/stable/c/7996013b85687034d2e820cef94d6404192e3a3d","https://nvd.nist.gov/vuln/detail/CVE-2026-68081"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-68418","cve":"CVE-2026-68418","aliases":[],"title":"Linux kernel (drivers/infiniband/hw/irdma): A tenant that asks for a user QP while declaring a zero-size work-queue","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/infiniband/hw/irdma)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A tenant that asks for a user QP while declaring a zero-size work-queue buffer leaves a pointer unpopulated that the driver then dereferences unconditionally. The kernel takes a NULL dereference in a syscall path, which oopses the task and, in the wrong context, brings the node down - a noisy-neighbour outage for every other tenant sharing the box.","attack_vector":"One ibv_create_qp call with user_wqe_bufs set to zero from any container holding /dev/infiniband/uverbs* on an Intel irdma node. Trivially reachable, no fabric peer, no privilege beyond the device node.","remediation":"No fixed version is listed in the record - take the stable kernel carrying ec675b4cdfd3 (or 728211c815f6 / b9b0889071569) and reboot. Interim: remove /dev/infiniband/* from untrusted containers on irdma nodes.","references":["https://git.kernel.org/stable/c/ec675b4cdfd378d8c9dd8c93126c024f2469bd79","https://git.kernel.org/stable/c/728211c815f6eef28dd3df2a5b6297483185aa20","https://nvd.nist.gov/vuln/detail/CVE-2026-68418"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-401"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-72006","cve":"CVE-2026-72006","aliases":[],"title":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core/lib): Every memory-key allocation carrying a steering-tag hint","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/net/ethernet/mellanox/mlx5/core/lib)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Every memory-key allocation carrying a steering-tag hint leaks a small kernel structure when the key is released. A tenant that registers and deregisters RDMA memory regions in a loop grows the host's kernel slab without bound until the node runs out of memory - a noisy-neighbour path from an unprivileged RDMA workload to an out-of-memory event for every tenant on the box.","attack_vector":"A tenant container holding /dev/infiniband/uverbs* that churns memory-region registration/deregistration on an mlx5 device drives it directly; this is the normal steady-state path, not an error path. Conditional on the mkeys carrying TPH steering-tag hints, which the RDMA stack sets for hinted registrations on supporting hardware.","remediation":"Update to a kernel carrying the fix on your stream. Interim: cap tenant RDMA registration rates where possible, monitor kernel slab growth on nodes exporting verbs to tenants, and withhold /dev/infiniband/* from workloads that do not need it.","references":["https://git.kernel.org/stable/c/262da8b6ea03d01ee7ed01ad309e4c89941f6b14","https://git.kernel.org/stable/c/6eb4cf2fa8997f62c11e0006dc010a1fd89c5a75","https://nvd.nist.gov/vuln/detail/CVE-2026-72006"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-276"],"fleet":{"pain_class":"node-drain"},"id":"NCVD-2020-005-runc-linux-resources-devices-cgr","cve":null,"aliases":["GHSA-g54h-m393-cpwq"],"title":"runc (linux.resources.devices cgroup list handling): MULTI-TENANT DEVICE ISOLATION: runc implemented the","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"runc (linux.resources.devices cgroup list handling)","year":"2020","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT DEVICE ISOLATION: runc implemented the linux.resources.devices list as a blacklist by default, contrary to the OCI runtime specification, which mandates a whitelist. Any caller that builds its own config.json without prefixing an explicit deny-all rule gets no protection from the devices cgroup at all. On a GPU node this is the control that is supposed to guarantee a container sees only the accelerators assigned to it: with the list defaulting open, a sufficiently privileged container holding CAP_MKNOD can create arbitrary device inodes and operate on any device node its Unix DAC permissions allow — including /dev/nvidia* entries belonging to another tenant's allocation, and other host devices entirely. The bug hid for years precisely because the common runtimes all write their own deny-all rule and the spec examples include one, so it only bites custom or hand-rolled OCI configs — which is exactly what bespoke GPU scheduling and device-plugin integration work tends to produce.","attack_vector":"Local, from inside a container, on a host whose OCI config.json was generated without a leading deny-all devices rule. The container additionally needs CAP_MKNOD to create inodes, and normal Unix DAC permissions on whatever device it then targets.","remediation":"Upgrade runc to a release that treats the devices list as a whitelist per the OCI spec, and drain the node for the runtime change. Independently, audit any custom config.json generation in your stack for a leading deny-all rule ({\"allow\": false, \"permissions\": \"rwm\"}), and drop CAP_MKNOD from tenant containers, which removes the inode-creation half of the primitive regardless of runtime version.","references":["https://github.com/opencontainers/runc/security/advisories/GHSA-g54h-m393-cpwq"],"status":"curated"},{"cwe":["CWE-312"],"fleet":{"pain_class":"unpatchable / mitigate-only"},"id":"NCVD-2020-006-etcd-write-ahead-log-user-authen","cve":null,"aliases":["GHSA-528j-9r78-wffx"],"title":"etcd (write-ahead log, user authentication entries): CONTROL-PLANE CREDENTIALS AT REST IN CLEARTEXT: etcd writes the","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"etcd (write-ahead log, user authentication entries)","year":"2020","cvss_score":5.5,"severity":"medium","kev":false,"impact":"CONTROL-PLANE CREDENTIALS AT REST IN CLEARTEXT: etcd writes the user login and password into a WAL entry on each authentication, and etcd does not encrypt key/value data on disk at all. So the datastore behind the Kubernetes control plane keeps a plaintext record of every credential used against it, sitting in files whose protection is entirely the operator's problem. Anyone who reaches those files — through a node compromise, an unencrypted disk that leaves the rack, a backup or snapshot copied to object storage with looser permissions than the cluster itself, or a debug tarball — harvests working etcd credentials rather than hashes. In a GPU cluster the etcd backups are frequently the least-guarded copy of the most sensitive data in the estate. This is a documented design position rather than a bug the project intends to fix, which is why it carries no CVE and why it never expires.","attack_vector":"Local / at-rest file access: read access to etcd WAL files on a control-plane node, or to any backup, snapshot or disk image containing them. No etcd credentials are needed to start, since the credentials are the payload.","remediation":"There is no patch — the project's stated position is that etcd assumes on-disk files are secure and securing them is the operator's responsibility. Encrypt the volumes holding etcd data and WAL, restrict filesystem access on control-plane nodes, and treat etcd backups and snapshots with the same handling rules as the credentials themselves: encrypted at rest, tightly scoped access, short retention. Prefer certificate-based authentication over username/password so there is no password to write into the WAL in the first place.","references":["https://github.com/etcd-io/etcd/security/advisories/GHSA-528j-9r78-wffx"],"status":"curated"},{"fleet":{"pain_class":"node-drain"},"id":"NCVD-2023-009-containerd-default-mounts-sys-de","cve":null,"aliases":["GHSA-7ww5-4wqc-m92c"],"title":"containerd (default mounts, /sys/devices/virtual/powercap RAPL): Containers get read access to Intel RAPL power","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd (default mounts, /sys/devices/virtual/powercap RAPL)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Containers get read access to Intel RAPL power telemetry by default, which is the PLATYPUS attack surface. Fine-grained energy readings leak information about what the host CPU is doing, and were demonstrated against AES-NI keys (including inside SGX enclaves) and against KASLR. The kernel restricted this to root in 5.10, but that mitigation misses the container case: without user namespaces, root in a container is root for this purpose, and sysfs being mounted read-only does not help because reading is the entire attack. On a bare-metal multi-tenant GPU host — the standard neocloud shape, where a tenant gets a real container on real hardware rather than a VM — one tenant can measure the power behaviour of workloads belonging to another tenant or of the host itself. RAPL is only present on bare metal with the powercap module compiled in, so cloud VMs are out of scope and bare-metal GPU fleets are precisely in scope.","attack_vector":"Local, from inside any container on an affected bare-metal host with RAPL/powercap available and an unpatched or unmitigated CPU. Root inside the container, without user namespaces, is enough; only read access to sysfs is required.","remediation":"Upgrade containerd past 1.6.25 / 1.7.10 so /sys/devices/virtual/powercap is masked in the default mount configuration, and drain the node for the runtime restart. Independently: apply the CPU microcode updates that coarsened RAPL sampling resolution, enable user namespaces for tenant containers, and confirm the mask is actually present in running containers rather than assuming the default applies to workloads started before the upgrade.","references":["https://github.com/containerd/containerd/security/advisories/GHSA-7ww5-4wqc-m92c","https://platypusattack.com/"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-312","CWE-532"],"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2026-058-harbor-audit-log-redaction-ldap","cve":null,"aliases":["GHSA-prh4-vhfh-24mj"],"title":"Harbor (audit log redaction, LDAP password and OIDC client secret): CREDENTIAL DISCLOSURE VIA THE AUDIT TRAIL: Harbor","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Harbor (audit log redaction, LDAP password and OIDC client secret)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"CREDENTIAL DISCLOSURE VIA THE AUDIT TRAIL: Harbor records the LDAP bind password and the OIDC client secret into its audit log without redaction. The audit log is by design a widely-readable artifact — it is shipped to SIEM, retained long, and read by compliance and platform staff who are deliberately not granted registry admin. So the identity-provider credentials for the registry end up in the one store an operator has consciously made broadly accessible. An LDAP bind credential typically reaches far beyond Harbor into the rest of the directory, and an OIDC client secret allows impersonating the registry integration to the identity provider. For a cluster operator this is a lateral-movement seed sitting in a low-suspicion location, and its lifetime is set by log retention rather than by patching.","attack_vector":"Local / log access: anyone able to read Harbor audit logs directly or via downstream aggregation. No registry privileges are required beyond audit-log visibility.","remediation":"Upgrade Harbor to a release carrying the redaction fix and restart the components. Then rotate the LDAP bind password and the OIDC client secret, since patching does not retract what is already written. Purge or expire affected audit-log ranges in your SIEM and archives, and re-check which roles can read Harbor audit data.","references":["https://github.com/goharbor/harbor/security/advisories/GHSA-prh4-vhfh-24mj"],"status":"curated"},{"id":"CVE-2019-14456","cve":"CVE-2019-14456","aliases":[],"title":"Opengear console server (serial port logging): Stored XSS injected from a device *connected to* a serial port","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Opengear console server (serial port logging)","year":"2019","cvss_score":5.4,"severity":"medium","kev":false,"impact":"Stored XSS injected from a device *connected to* a serial port — a compromised switch can attack the operator's console-server UI, inverting the expected trust direction","attack_vector":"Local device to OOB management UI","remediation":"Console-server firmware upgrade to 4.5.0+; notable as an example of the OOB network being attackable from the devices it manages","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-14456"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2019-07-31"},{"id":"CVE-2020-8558","cve":"CVE-2020-8558","aliases":[],"title":"Kubernetes (kubelet/kube-proxy): Node's 127.0.0.1-bound services reachable from adjacent hosts and pods","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kubelet/kube-proxy)","year":"2020","cvss_score":5.4,"severity":"medium","kev":false,"impact":"Node's 127.0.0.1-bound services reachable from adjacent hosts and pods; commonly reaches an unauthenticated kubelet or etcd","attack_vector":"Any pod on the node, or an adjacent host on the node network","remediation":"Rolling kubelet/kube-proxy upgrade with node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-8558"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2020-07-27"},{"id":"CVE-2022-21818","cve":"CVE-2022-21818","aliases":[],"title":"NVIDIA License System - DLS virtual appliance: Installation scripts on the DLS appliance leave other users' credentials","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA License System - DLS virtual appliance","year":"2022","cvss_score":5.4,"severity":"medium","kev":false,"impact":"Installation scripts on the DLS appliance leave other users' credentials readable to any signed-in portal user, giving lateral privilege escalation inside your licensing infrastructure.","attack_vector":"Network, authenticated as any portal user. Anyone you gave a licensing-portal account to.","remediation":"Patch the DLS appliance per bulletin 5319 and rotate every credential the appliance held - the patch does not undo the exposure. Cost: appliance restart; vGPU guests keep running on cached licences through a short DLS outage.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-21818","https://github.com/NVIDIA/product-security/tree/main/2022/5319"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:L/I:L/A:N","cwe":["CWE-312"],"published":"2022-02-15"},{"cvss_vector":"CVSS:3.0/AV:N/AC:L/PR:L/UI:R/S:C/C:L/I:L/A:N","cwe":["CWE-79"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2023-6571","cve":"CVE-2023-6571","aliases":[],"title":"Kubeflow (central dashboard, reflected cross-site scripting): Reflected XSS in the Kubeflow dashboard runs attacker","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Kubeflow (central dashboard, reflected cross-site scripting)","year":"2023","cvss_score":5.4,"severity":"medium","kev":false,"impact":"Reflected XSS in the Kubeflow dashboard runs attacker JavaScript in a logged-in user's browser under the Kubeflow origin. The attacker acts as that user against the Kubeflow API - which for an admin means creating notebooks and pipelines and reading other namespaces' resources.","attack_vector":"An authenticated Kubeflow user who follows an attacker-crafted link, so it needs both a valid session and a click.","remediation":"Upgrade Kubeflow past the fixed dashboard release and redeploy the central dashboard. Add a Content-Security-Policy at the ingress in front of Kubeflow as a standing mitigation for this class.","references":["https://huntr.com/bounties/f02781e7-2a53-4c66-aa32-babb16434632","https://nvd.nist.gov/vuln/detail/CVE-2023-6571"],"status":"curated"},{"id":"CVE-2024-0103","cve":"CVE-2024-0103","aliases":[],"title":"Triton Inference Server: Insufficient access-control granularity","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2024","cvss_score":5.4,"severity":"medium","kev":false,"impact":"Insufficient access-control granularity","attack_vector":"Authenticated inference client","remediation":"Upgrade Triton; redeploy serving images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0103","https://github.com/NVIDIA/product-security/tree/main/2024/5546"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:L/I:L/A:N","cwe":["CWE-1419"],"published":"2024-06-13"},{"id":"CVE-2024-9526","cve":"CVE-2024-9526","aliases":[],"title":"Kubeflow (Pipelines UI): Stored XSS in the pipeline view","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Kubeflow (Pipelines UI)","year":"2024","cvss_score":5.4,"severity":"medium","kev":false,"impact":"Stored XSS in the pipeline view","attack_vector":"Tenant-supplied pipeline definition rendered to another user","remediation":"Upgrade; multi-tenant Kubeflow UIs share an origin","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-9526"],"status":"curated","published":"2024-11-18"},{"id":"CVE-2025-48515","cve":"CVE-2025-48515","aliases":[],"title":"AMD Secure Processor bootloader - SPIROM upgrade path: An attacker who can drive the SPIROM upgrade path can pass","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor bootloader - SPIROM upgrade path","year":"2025","cvss_score":5.4,"severity":"medium","kev":false,"impact":"An attacker who can drive the SPIROM upgrade path can pass unsanitised parameters to the ASP bootloader and overwrite memory, reaching arbitrary code execution in the secure processor. This is the classic firmware-update-as-attack-surface problem: the mechanism you use to patch the platform is itself the way in.","attack_vector":"Local, requires access to the SPI ROM upgrade mechanism - typically root plus flash write, or a compromised BMC that can drive host SPI.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. Worth pairing with BMC hardening: on most server designs the BMC can write host SPI, so a BMC compromise reaches this directly. Restrict who can invoke firmware updates and require signed update packages end to end.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-48515","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2026-02-10"},{"id":"CVE-2025-7623","cve":"CVE-2025-7623","aliases":[],"title":"Supermicro BMC SMASH-CLP shell on MBD-X13SEDW-F: Full control of the instruction pointer inside the BMC's firmware OS","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC SMASH-CLP shell on MBD-X13SEDW-F","year":"2025","cvss_score":5.4,"severity":"medium","kev":false,"impact":"Full control of the instruction pointer inside the BMC's firmware OS from a shell that operators routinely hand to junior staff and monitoring tooling. The attacker converts a limited management login into arbitrary code on the controller, and from there into the standard BMC prize set: power control, console capture, virtual media, and firmware-level persistence. The published CVSS understates this - the vector was scored conservatively, but the described primitive is return-address control. A stack buffer overflow reached by a crafted SMASH command, with control of the saved return address and registers.","attack_vector":"An authenticated low-privilege BMC account with SSH access to the controller. Any operator-tier credential works; no administrator role is needed.","remediation":"Firmware flash from Supermicro's November 2025 BMC/IPMI advisory batch, matched to the board SKU. The cheap and immediate mitigation is config-only: turn off SSH/SMASH on the BMC if your management path is Redfish or IPMI-over-LAN. That single change also covers CVE-2026-3821 and the rest of the SMASH overflow cluster, so it is the highest-leverage action available before a flash window opens.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-7623","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2025/7xxx/CVE-2025-7623.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2026-32677","cve":"CVE-2026-32677","aliases":[],"title":"Intel Gaudi / gaudi-container-runtime: A path-traversal bug in the container runtime shim that wires Gaudi devices into","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Intel Gaudi / gaudi-container-runtime","year":"2026","cvss_score":5.4,"severity":"medium","kev":false,"impact":"A path-traversal bug in the container runtime shim that wires Gaudi devices into containers lets a tenant who controls the container spec cause the runtime to touch host paths outside the container root. On a shared Gaudi box the runtime executes as root on the host, so this is the classic accelerator-runtime escape shape: tenant container -> host filesystem -> every other tenant's job on that node.","attack_vector":"Any tenant who can launch a container on a Gaudi node through the normal scheduler. No host account and no physical access required - the container spec is the attack surface.","remediation":"Upgrade gaudi-container-runtime to 1.24.0 or later across every Gaudi node. This is a host-side userspace package, so no BIOS, firmware or microcode update is involved, but the runtime binary is in the path of every new container start - roll it per node and restart the container engine, which means draining running jobs on that node.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-32677","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01489.html"],"status":"curated","tags":["tenant-isolation"],"published":"2026-08-11"},{"id":"CVE-2026-33726","cve":"CVE-2026-33726","aliases":[],"title":"Cilium: Ingress NetworkPolicies not enforced for pod traffic to L7 services","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2026","cvss_score":5.4,"severity":"medium","kev":false,"impact":"Ingress NetworkPolicies not enforced for pod traffic to L7 services","attack_vector":"Any pod on the cluster network","remediation":"Rolling Cilium upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-33726"],"status":"curated","published":"2026-03-27"},{"id":"CVE-2026-39350","cve":"CVE-2026-39350","aliases":[],"title":"Istio: serviceAccounts and notServiceAccounts in AuthorizationPolicy are evaluated incorrectly","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Istio","year":"2026","cvss_score":5.4,"severity":"medium","kev":false,"impact":"serviceAccounts and notServiceAccounts in AuthorizationPolicy are evaluated incorrectly","attack_vector":"Any pod on the mesh","remediation":"Rolling istiod upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-39350"],"status":"curated","published":"2026-04-15"},{"id":"CVE-2026-56743","cve":"CVE-2026-56743","aliases":[],"title":"Cilium: CIDR ipBlock rules without selectors generate a wildcard, over-permitting traffic","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2026","cvss_score":5.4,"severity":"medium","kev":false,"impact":"CIDR ipBlock rules without selectors generate a wildcard, over-permitting traffic","attack_vector":"Any tenant workload","remediation":"Rolling Cilium upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-56743"],"status":"curated","published":"2026-07-15"},{"cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:N/UI:N/S:U/C:L/I:N/A:L","cwe":["CWE-306"],"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2026-050-cloudnativepg-instance-manager-s","cve":null,"aliases":["GHSA-7qwx-x8ff-3px9"],"title":"CloudNativePG instance manager (status server, TCP/8000 control endpoints): A set of operator-only control endpoints","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"CloudNativePG instance manager (status server, TCP/8000 control endpoints)","year":"2026","cvss_score":5.4,"severity":"medium","kev":false,"impact":"A set of operator-only control endpoints (/pg/mode/backup, /update, /pg/archive/partial, /pg/controldata) sit on the same unauthenticated mux as the kubelet probes on TCP/8000 of every PostgreSQL pod, with network isolation as the only intended boundary — and even with status-port-tls enabled the server never requested or verified a client certificate. Any workload that can reach that port without a NetworkPolicy in the way can drive backup orchestration, trigger WAL and checkpoint activity, and read operational metadata including replication and WAL state, the system identifier, and pg_controldata output. The vendor is careful to bound this: single-flight and self-cleaning backup control means bounded DoS and transient storage pressure, not unbounded WAL exhaustion, secret disclosure or code execution. The lesson for a cluster operator is the assumption, not the severity — a component shipped with 'protected by NetworkPolicy' as its security model, in clusters where NetworkPolicies are frequently absent or permissive.","attack_vector":"Adjacent network / in-cluster, unauthenticated. Any pod able to open TCP/8000 on a PostgreSQL instance pod qualifies, which in a cluster with no restricting NetworkPolicy means every pod including tenant workloads.","remediation":"Apply a NetworkPolicy restricting TCP/8000 on instance pods to the operator — this is the mitigation for all currently supported releases, since the hardening was not backported. Upgrading to 1.30.0 adds in-process operator authentication as defence in depth: the operator generates an in-memory ECDSA P-256 client certificate at startup, publishes the SHA-256 fingerprint in .status.operatorCertificateFingerprint, and the instance manager rejects calls to the sensitive endpoints without it. Probe endpoints and /pg/status stay unauthenticated by design.","references":["https://github.com/cloudnative-pg/cloudnative-pg/security/advisories/GHSA-7qwx-x8ff-3px9","https://github.com/cloudnative-pg/cloudnative-pg/pull/10579"],"status":"curated"},{"id":"CVE-2018-10892","cve":"CVE-2018-10892","aliases":[],"title":"Docker / moby: Default OCI spec does not mask /proc/acpi, so a container can change host hardware state","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2018","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Default OCI spec does not mask /proc/acpi, so a container can change host hardware state","attack_vector":"Any tenant workload","remediation":"Upgrade Docker Engine or add explicit maskedPaths","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-10892"],"status":"curated","published":"2018-07-06"},{"cvss_vector":"CVSS:3.0/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:L/A:N","cwe":["CWE-20"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2018-10995","cve":"CVE-2018-10995","aliases":[],"title":"Slurm (user_name / gid field handling): Slurm trusts the user_name and gid fields carried in job RPCs instead of","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Slurm (user_name / gid field handling)","year":"2018","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Slurm trusts the user_name and gid fields carried in job RPCs instead of resolving identity itself, so the identity a job runs under can be steered by whoever crafts the RPC. On a shared cluster that is the first half of running work under someone else's account.","attack_vector":"A tenant who can submit jobs, or anyone who can talk to slurmctld or slurmd on the cluster network.","remediation":"Upgrade to Slurm 17.02.11 or 17.11.7 and restart slurmctld and slurmd. Pair the upgrade with a check that MUNGE is actually enforcing authentication on every node - this class of bug is only dangerous when the RPC path is not independently authenticated.","references":["https://lists.schedmd.com/pipermail/slurm-announce/2018/000008.html","https://nvd.nist.gov/vuln/detail/CVE-2018-10995","https://www.debian.org/security/2018/dsa-4254"],"status":"curated"},{"cvss_vector":"CVSS:3.0/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:L/A:L","cwe":["CWE-732"],"fleet":{"pain_class":"hot-patch"},"id":"CVE-2018-1724","cve":"CVE-2018-1724","aliases":["IBM X-Force 147439"],"title":"IBM Spectrum LSF (job submission, file permissions): Weak file permissions in the LSF install let a local user change","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"IBM Spectrum LSF (job submission, file permissions)","year":"2018","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Weak file permissions in the LSF install let a local user change the user a job is submitted as. On a shared GPU cluster that means running your workload under someone else's identity, charged to their account and with access to their data.","attack_vector":"A local user on an LSF submission host, affecting LSF 9.1.1, 9.1.2, 9.1.3 and 10.1.","remediation":"Apply IBM's LSF fix pack and, separately, tighten the permissions on the LSF configuration and binary directories - the underlying issue is filesystem mode, so the correct modes need to be verified after every LSF upgrade, not just once.","references":["https://www.ibm.com/support/pages/node/734767","https://nvd.nist.gov/vuln/detail/CVE-2018-1724"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2019-9836","cve":"CVE-2019-9836","aliases":[],"title":"AMD Platform Security Processor - SEV key derivation (PSP firmware <= 0.17 build 11): The SEV implementation in PSP","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Platform Security Processor - SEV key derivation (PSP firmware <= 0.17 build 11)","year":"2019","cvss_score":5.3,"severity":"medium","kev":false,"impact":"The SEV implementation in PSP firmware used a broken elliptic-curve parameter check, so an attacker with host privileges can run an invalid-curve attack against the platform Diffie-Hellman exchange and recover the SEV endorsement key. With that key the host can decrypt a guest's launch secret and read the encrypted VM's memory outright. If you sold confidential computing on top of SEV on affected firmware, the guarantee was not there.","attack_vector":"Local, requires hypervisor/host administrator privilege - which is exactly the party SEV is supposed to defend the guest against. No guest cooperation needed.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. Because this sits inside the SEV-SNP trust boundary, the update also moves the platform's reported TCB version: after patching you must refresh VCEK certificates from AMD's KDS and update whatever attestation policy your tenants (or your own confidential-VM control plane) pin against, or every guest launch will start failing validation. Fixed in PSP/SEV firmware 0.17 build 22 and later. Any guest launched or attested on older firmware should be treated as having had no confidentiality guarantee - rotate the secrets those guests held rather than just patching forward.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-9836","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2019-06-25"},{"id":"CVE-2020-29662","cve":"CVE-2020-29662","aliases":[],"title":"Harbor: Catalog registry API exposed on an unauthenticated path","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Harbor","year":"2020","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Catalog registry API exposed on an unauthenticated path; full image inventory disclosure","attack_vector":"Unauthenticated network","remediation":"Upgrade Harbor; put the registry behind auth","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-29662"],"status":"curated","published":"2021-02-02"},{"id":"CVE-2020-7202","cve":"CVE-2020-7202","aliases":["HPESBHF04069"],"title":"HPE iLO 4 / iLO 5 (unauthenticated information disclosure): An unauthenticated remote request pulls back the server","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HPE iLO 4 / iLO 5 (unauthenticated information disclosure)","year":"2020","cvss_score":5.3,"severity":"medium","kev":false,"impact":"An unauthenticated remote request pulls back the server serial number and other identifying detail from the iLO. Low impact on its own, high value as reconnaissance: it lets an attacker who can reach a management network enumerate exactly what hardware sits behind each iLO address, fingerprint generations, and pick which nodes are worth a real exploit - all without a single failed login to show up in an audit log. Covers ProLiant, Apollo, Synergy compute modules and Converged Systems, which is most of the HPE fleet shape a GPU operator would run.","attack_vector":"Anything routable to the iLO on the out-of-band management VLAN, unauthenticated. If any iLO is inadvertently internet-exposed, this is what a mass scanner harvests first.","remediation":"Flash iLO 5 to v2.31 or later and iLO 4 to v2.76 or later. Out-of-band, per-node, no host reboot and no drain. Given the low direct impact, most operators should fold this into the next scheduled iLO firmware campaign rather than running a dedicated one - but do treat any internet-reachable iLO as an emergency independent of this CVE.","references":["https://support.hpe.com/hpsc/doc/public/display?docLocale=en_US&docId=emr_na-hpesbhf04069en_us","https://nvd.nist.gov/vuln/detail/CVE-2020-7202"],"status":"curated","published":"2021-01-05"},{"id":"CVE-2020-8552","cve":"CVE-2020-8552","aliases":[],"title":"Kubernetes (kube-apiserver): Successful API requests can DoS the apiserver","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-apiserver)","year":"2020","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Successful API requests can DoS the apiserver","attack_vector":"Any authenticated cluster user","remediation":"Rolling control-plane upgrade; no GPU drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-8552"],"status":"curated","published":"2020-03-27"},{"id":"CVE-2021-1055","cve":"CVE-2021-1055","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): Improper access control in the escape handler leaks information","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2021","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Improper access control in the escape handler leaks information and lets an unprivileged caller crash the driver.","attack_vector":"Any local user with GPU device access.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1055"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-01-08"},{"id":"CVE-2021-22815","cve":"CVE-2021-22815","aliases":["SEVD-2021-313-03"],"title":"APC/Schneider Electric UPS, PDU, and cooling products using NMC2/NMC3 (Smart-UPS, Symmetra, Galaxy, rack PDUs, InRow","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"APC/Schneider Electric UPS, PDU, and cooling products using NMC2/NMC3","year":"2021","cvss_score":5.3,"severity":"medium","kev":false,"impact":"The card's troubleshooting archive — a diagnostic bundle that can contain configuration and log details — can be pulled off the device by someone who shouldn't have access to it, giving an attacker reconnaissance data useful for planning further attacks on that power/cooling unit.","attack_vector":"Network access to the card's web interface; the advisory describes this as an information-exposure issue reachable without full administrative rights.","remediation":"Firmware upgrade per Schneider's SEVD-2021-313-03 advisory (fixed AOS versions vary by card generation — NMC2 vs NMC3). This spans a very wide product line (UPS, rack PDUs, cooling, NetBotz), so treat it as a fleet-wide inventory-and-patch exercise rather than a one-off fix.","references":["https://download.schneider-electric.com/files?p_Doc_Ref=SEVD-2021-313-03"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-01-28"},{"id":"CVE-2021-45925","cve":"CVE-2021-45925","aliases":["AMI-SA-2022001","Nozomi Labs BMC firmware research"],"title":"AMI MegaRAC SPx 12 / SPx 13 (BMC login): The login flow answers differently for real and fake usernames, so","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx 12 / SPx 13 (BMC login)","year":"2021","cvss_score":5.3,"severity":"medium","kev":false,"impact":"The login flow answers differently for real and fake usernames, so an unauthenticated attacker can enumerate every valid BMC account. The value to an attacker is targeting: sweep the management range, learn which nodes still carry the ODM's default account or the provisioning template's service account, and aim credential-stuffing only at those. It turns a noisy brute-force into a quiet, low-attempt campaign that will not trip lockout thresholds.","attack_vector":"Unauthenticated network access to the BMC web login. Anything that can reach the BMC's HTTP/HTTPS port on the management VLAN.","remediation":"Firmware flash to SPx_12-update-7.00 / SPx_13-update-5.00 or later; low urgency on its own, fold it into whatever BMC flash campaign you are already running. The config-only work carries most of the value and costs nothing: delete vendor default accounts, avoid a fleet-wide shared username in the provisioning template, and enable BMC account lockout plus authentication logging to your SIEM so the enumeration sweep itself becomes visible.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2022001.pdf","https://www.nozominetworks.com/labs/vulnerability-advisories/cve-2021-45925/","https://nvd.nist.gov/vuln/detail/CVE-2021-45925"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-10-24"},{"id":"CVE-2022-27652","cve":"CVE-2022-27652","aliases":[],"title":"CRI-O: Containers started with non-empty default inheritable capabilities","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"CRI-O","year":"2022","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Containers started with non-empty default inheritable capabilities","attack_vector":"Any tenant workload","remediation":"Upgrade CRI-O; drain node","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-27652"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2022-04-18"},{"id":"CVE-2022-34684","cve":"CVE-2022-34684","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): An off-by-one error in nvidia.ko permits data tampering","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":5.3,"severity":"medium","kev":false,"impact":"An off-by-one error in nvidia.ko permits data tampering or information disclosure across the ioctl boundary. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5415. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34684","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:L/A:L","cwe":["CWE-125"],"fleet":{"pain_class":"node-drain"},"published":"2022-12-30"},{"id":"CVE-2022-36109","cve":"CVE-2022-36109","aliases":[],"title":"Docker / moby: Supplementary groups not set up properly","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2022","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Supplementary groups not set up properly; access to group-readable files inside container","attack_vector":"Any tenant workload","remediation":"Upgrade Docker Engine; restart containers","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-36109"],"status":"curated","published":"2022-09-09"},{"id":"CVE-2022-40258","cve":"CVE-2022-40258","aliases":[],"title":"AMI MegaRAC: Weak MD5 password hashing for BMC accounts","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC","year":"2022","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Weak MD5 password hashing for BMC accounts; offline cracking of captured hashes","attack_vector":"Local/offline after hash disclosure","remediation":"BMC firmware update plus credential rotation, since previously-hashed passwords must be considered recoverable","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-40258"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-01-31"},{"id":"CVE-2022-4122","cve":"CVE-2022-4122","aliases":[],"title":"Buildah: Symlink following when reading .containerignore/.dockerignore discloses host files","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Buildah","year":"2022","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Symlink following when reading .containerignore/.dockerignore discloses host files","attack_vector":"Malicious build context","remediation":"Upgrade Buildah on build hosts","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-4122"],"status":"curated","published":"2022-12-08"},{"id":"CVE-2022-42254","cve":"CVE-2022-42254","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): An out-of-bounds array access in nvidia.ko gives","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":5.3,"severity":"medium","kev":false,"impact":"An out-of-bounds array access in nvidia.ko gives an unprivileged local user a crash, a memory leak, or data corruption. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5415. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42254","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:L/A:L","cwe":["CWE-125"],"fleet":{"pain_class":"node-drain"},"published":"2022-12-30"},{"id":"CVE-2022-42255","cve":"CVE-2022-42255","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): A second out-of-bounds array access path in nvidia.ko","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":5.3,"severity":"medium","kev":false,"impact":"A second out-of-bounds array access path in nvidia.ko with the same unprivileged-local reach. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5415. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42255","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:L/A:L","cwe":["CWE-787"],"fleet":{"pain_class":"node-drain"},"published":"2022-12-30"},{"id":"CVE-2022-42256","cve":"CVE-2022-42256","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): An integer overflow in index validation defeats","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":5.3,"severity":"medium","kev":false,"impact":"An integer overflow in index validation defeats the driver's own bounds check, giving an unprivileged user an out-of-bounds access in kernel context. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5415. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42256","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:L/A:L","cwe":["CWE-190"],"fleet":{"pain_class":"node-drain"},"published":"2022-12-30"},{"id":"CVE-2022-42257","cve":"CVE-2022-42257","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): An integer overflow in nvidia.ko produces information","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":5.3,"severity":"medium","kev":false,"impact":"An integer overflow in nvidia.ko produces information disclosure, data tampering or a node crash. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5415. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42257","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:L/A:L","cwe":["CWE-190"],"fleet":{"pain_class":"node-drain"},"published":"2022-12-30"},{"id":"CVE-2022-42258","cve":"CVE-2022-42258","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): Another integer overflow path in nvidia.ko reachable","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Another integer overflow path in nvidia.ko reachable by any local GPU user. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5415. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42258","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:L/A:L","cwe":["CWE-190"],"fleet":{"pain_class":"node-drain"},"published":"2022-12-30"},{"id":"CVE-2022-42265","cve":"CVE-2022-42265","aliases":[],"title":"GPU Display Driver / vGPU guest driver: DoS / info disclosure (integer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver / vGPU guest driver","year":"2022","cvss_score":5.3,"severity":"medium","kev":false,"impact":"DoS / info disclosure (integer overflow)","attack_vector":"Any tenant with a container or vGPU guest","remediation":"Driver + vGPU Manager upgrade; rolling node reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42265","https://github.com/NVIDIA/product-security/tree/main/2024/5520"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:L/A:L","cwe":["CWE-190"],"fleet":{"pain_class":"node-reboot"},"published":"2022-12-30"},{"id":"CVE-2022-42288","cve":"CVE-2022-42288","aliases":[],"title":"DGX servers BMC: Info exposure from BMC","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX servers BMC","year":"2022","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Info exposure from BMC","attack_vector":"Network-adjacent","remediation":"Flash BMC 2.09.00+","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42288","https://github.com/NVIDIA/product-security/tree/main/2024/5513"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:L/I:N/A:N","cwe":["CWE-208"],"published":"2023-01-13"},{"id":"CVE-2022-48698","cve":"CVE-2022-48698","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A memory or reference-count leak in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2022","cvss_score":5.3,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu display core (DC/DM). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amd/display: fix memory leak when using debugfs_lookup()","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-48698","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-03"},{"id":"CVE-2023-20582","cve":"CVE-2023-20582","aliases":[],"title":"AMD IOMMU - nested page table entry faults bypass SEV-SNP RMP checks: The IOMMU mishandles invalid nested page table","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD IOMMU - nested page table entry faults bypass SEV-SNP RMP checks","year":"2023","cvss_score":5.3,"severity":"medium","kev":false,"impact":"The IOMMU mishandles invalid nested page table entries, letting a privileged attacker induce PTE faults that bypass SEV-SNP RMP enforcement and tamper with confidential guest memory. Same failure shape as the DTE variant and disclosed alongside it - the IOMMU's error paths are where RMP enforcement leaks.","attack_vector":"Hypervisor-privileged attacker driving DMA through the IOMMU.","remediation":"Fixed in AMD SEV firmware / AGESA and reaches you as an OEM SBIOS package - AMD hands AGESA to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before shipping BIOS. **Budget one to six months of OEM lag**, longer on older platforms and sometimes never on end-of-support SKUs. Applying it means draining the host and doing a full power cycle. Because the fix moves the platform's reported SEV-SNP TCB version, you must also pull fresh VCEK certificates from AMD's Key Distribution Service and update any attestation policy your tenants pin - otherwise guests will start failing launch validation the moment the BIOS lands. Some SEV firmware can alternatively be staged from linux-firmware (amd/amd_sev_*.sbin) and committed via the ccp driver at boot, which is faster than waiting on BIOS - check whether your platform supports firmware hot-load before assuming the OEM is the only route. Patch alongside the DTE variant; they ship together.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-20582","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2025-02-11"},{"id":"CVE-2023-20584","cve":"CVE-2023-20584","aliases":[],"title":"AMD IOMMU - invalid device table entries bypass SEV-SNP RMP checks: The IOMMU mishandles certain special address ranges","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD IOMMU - invalid device table entries bypass SEV-SNP RMP checks","year":"2023","cvss_score":5.3,"severity":"medium","kev":false,"impact":"The IOMMU mishandles certain special address ranges when the device table entry is invalid, and a privileged attacker with a compromised hypervisor can induce those DTE faults to skip the reverse-map table checks that enforce SEV-SNP page ownership. RMP checks are the mechanism that stops a device DMA from landing in a confidential guest's memory; bypassing them via the IOMMU means DMA-based guest tampering.","attack_vector":"Requires hypervisor privilege plus the ability to drive DMA - so a compromised host with a device (or a device it controls) it can point at guest pages.","remediation":"Fixed in AMD SEV firmware / AGESA and reaches you as an OEM SBIOS package - AMD hands AGESA to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before shipping BIOS. **Budget one to six months of OEM lag**, longer on older platforms and sometimes never on end-of-support SKUs. Applying it means draining the host and doing a full power cycle. Because the fix moves the platform's reported SEV-SNP TCB version, you must also pull fresh VCEK certificates from AMD's Key Distribution Service and update any attestation policy your tenants pin - otherwise guests will start failing launch validation the moment the BIOS lands. Some SEV firmware can alternatively be staged from linux-firmware (amd/amd_sev_*.sbin) and committed via the ccp driver at boot, which is faster than waiting on BIOS - check whether your platform supports firmware hot-load before assuming the OEM is the only route. Particularly relevant on GPU nodes, where large numbers of devices do high-rate DMA and the IOMMU is doing real work rather than sitting idle.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-20584","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2024-08-13"},{"id":"CVE-2023-20585","cve":"CVE-2023-20585","aliases":[],"title":"AMD IOMMU host buffer access - insufficient RMP checks (AMD-SB-3016): Insufficient RMP checking on IOMMU host buffer","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD IOMMU host buffer access - insufficient RMP checks (AMD-SB-3016)","year":"2023","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Insufficient RMP checking on IOMMU host buffer access produces an out-of-bounds write. This is the most operationally expensive item in the RMP family, because of what fixing it costs rather than what it does.","attack_vector":"Privileged attacker with hypervisor control, via IOMMU host buffer operations.","remediation":"Fixed in AMD PI/AGESA firmware and delivered only as an OEM SBIOS package - AMD ships the PI drop to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before releasing BIOS. **Budget one to six months of OEM lag**, and note that several CVEs in this batch are marked 'no fix planned' on Naples (EPYC 7001) - for those the only remediation is retiring the hardware. Applying it means cordon, drain and a full power cycle per node; there is no driver reload, no live patch and no VBIOS step. Because this touches the SEV-SNP trust boundary, the update moves the platform TCB version: refresh VCEK certificates from AMD's KDS and update tenant attestation policy, or confidential guest launches will fail immediately after the BIOS lands. **This one needs more than a BIOS flash**: AMD requires the firmware update, an OS update, *and* a full SNP guest shutdown and platform re-initialization. In practice that means draining every confidential VM off the host, tearing down SNP, updating, re-initializing and re-admitting - a materially longer maintenance window than the rest of the batch, and one you cannot overlap with normal rolling reboots. Plan capacity for it explicitly.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-20585","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3016.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"tags":["tenant-isolation"],"published":"2026-04-16"},{"id":"CVE-2023-20592","cve":"CVE-2023-20592","aliases":[],"title":"AMD SEV-ES (CacheWarp): CacheWarp: INVD lets a malicious hypervisor revert SEV-ES guest memory writes, breaking guest","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"AMD SEV-ES (CacheWarp)","year":"2023","cvss_score":5.3,"severity":"medium","kev":false,"impact":"CacheWarp: INVD lets a malicious hypervisor revert SEV-ES guest memory writes, breaking guest integrity and enabling auth bypass inside the VM","attack_vector":"Malicious/compromised host against a tenant confidential VM","remediation":"Microcode/AGESA update + reboot. Undermines the trust story of any SEV-based confidential GPU-VM offering - must be reflected in attestation policy, not just patching","references":["https://access.redhat.com/security/cve/CVE-2023-20592"],"status":"curated","fleet":{"ubiquity":"common - SEV-SNP is the CPU-side TEE that anchors \"confidential GPU\" offerings on EPYC Naples/Rome/Milan hosts","remediation_pain":"microcode+reboot for Milan; Naples/Rome are effectively unpatchable-mitigate-only, so affected nodes must be retired from any confidential-compute SKU","pain_class":"unpatchable / mitigate-only","why_fleet_wide":"Breaks the integrity guarantee of confidential VMs from a malicious hypervisor - the exact threat model a neocloud invokes when it tells a customer their weights are safe from the operator."},"published":"2023-11-14"},{"id":"CVE-2023-25173","cve":"CVE-2023-25173","aliases":[],"title":"containerd: Supplementary groups not set up correctly inside containers","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2023","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Supplementary groups not set up correctly inside containers; unexpected access to group-readable files","attack_vector":"Any tenant workload","remediation":"Rolling containerd upgrade with node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25173"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2023-02-16"},{"id":"CVE-2023-25192","cve":"CVE-2023-25192","aliases":[],"title":"AMI MegaRAC SPX (Redfish): User enumeration through Redfish","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPX (Redfish)","year":"2023","cvss_score":5.3,"severity":"medium","kev":false,"impact":"User enumeration through Redfish — maps valid BMC accounts fleet-wide ahead of a credential attack","attack_vector":"Network / Redfish, unauthenticated","remediation":"BMC firmware update to SPx12-update-7.00 / SPx13-update-5.00; low severity alone, but it is the reconnaissance leg of the MegaRAC chain","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25192"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-02-15"},{"id":"CVE-2023-25512","cve":"CVE-2023-25512","aliases":[],"title":"NVIDIA CUDA Toolkit - cuobjdump: An out-of-bounds read on a malformed input file reaches limited code execution","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA CUDA Toolkit - cuobjdump","year":"2023","cvss_score":5.3,"severity":"medium","kev":false,"impact":"An out-of-bounds read on a malformed input file reaches limited code execution and information disclosure. The realistic exposure is your build and profiling pipeline, not your runtime fleet: anything that automatically disassembles third-party fatbins, vendor kernels or model artifacts is running this parser on attacker-influenced input.","attack_vector":"Local, and requires a user or an automated job to run cuobjdump over an attacker-supplied file. CI jobs that inspect third-party CUDA binaries are the usual path.","remediation":"Update the CUDA Toolkit package (bulletin 5456). Cost: effectively zero - userspace SDK only, no driver reload, no node drain, no running-job impact. Rebuild build/CI images and move on.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25512","https://github.com/NVIDIA/product-security/tree/main/2023/5456"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:L/I:L/A:L","cwe":["CWE-125"],"published":"2023-04-22"},{"id":"CVE-2023-25513","cve":"CVE-2023-25513","aliases":[],"title":"NVIDIA CUDA Toolkit - cuobjdump: A second out-of-bounds read path on malformed input with the same limited","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA CUDA Toolkit - cuobjdump","year":"2023","cvss_score":5.3,"severity":"medium","kev":false,"impact":"A second out-of-bounds read path on malformed input with the same limited code-execution reach. The realistic exposure is your build and profiling pipeline, not your runtime fleet: anything that automatically disassembles third-party fatbins, vendor kernels or model artifacts is running this parser on attacker-influenced input.","attack_vector":"Local, and requires a user or an automated job to run cuobjdump over an attacker-supplied file. CI jobs that inspect third-party CUDA binaries are the usual path.","remediation":"Update the CUDA Toolkit package (bulletin 5456). Cost: effectively zero - userspace SDK only, no driver reload, no node drain, no running-job impact. Rebuild build/CI images and move on.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25513","https://github.com/NVIDIA/product-security/tree/main/2023/5456"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:L/I:L/A:L","cwe":["CWE-125"],"published":"2023-04-22"},{"id":"CVE-2023-25514","cve":"CVE-2023-25514","aliases":[],"title":"NVIDIA CUDA Toolkit - cuobjdump: A third out-of-bounds read path on malformed input, fixed in the same bulletin","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA CUDA Toolkit - cuobjdump","year":"2023","cvss_score":5.3,"severity":"medium","kev":false,"impact":"A third out-of-bounds read path on malformed input, fixed in the same bulletin. The realistic exposure is your build and profiling pipeline, not your runtime fleet: anything that automatically disassembles third-party fatbins, vendor kernels or model artifacts is running this parser on attacker-influenced input.","attack_vector":"Local, and requires a user or an automated job to run cuobjdump over an attacker-supplied file. CI jobs that inspect third-party CUDA binaries are the usual path.","remediation":"Update the CUDA Toolkit package (bulletin 5456). Cost: effectively zero - userspace SDK only, no driver reload, no node drain, no running-job impact. Rebuild build/CI images and move on.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25514","https://github.com/NVIDIA/product-security/tree/main/2023/5456"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:L/I:L/A:L","cwe":["CWE-125"],"published":"2023-04-22"},{"id":"CVE-2023-30633","cve":"CVE-2023-30633","aliases":["INSYDE-SA-2023045"],"title":"Insyde InsydeH2O (TrEEConfigDriver, TPM PCR reporting): Low CVSS, high operational consequence","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (TrEEConfigDriver, TPM PCR reporting)","year":"2023","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Low CVSS, high operational consequence. The driver can report false TPM Platform Configuration Register values, which means the measurements a node presents during remote attestation do not reflect what actually booted. Every control an operator layers on top of measured boot - proving to a customer that their node runs the firmware and image you claim, gating access to model weights or key material on an attestation quote, detecting a bootkit left by the previous tenant - silently returns a pass on a compromised node. This is the bug class that turns every other firmware CVE in this list from detectable into invisible.","attack_vector":"Local attacker on the host able to influence the platform configuration the driver reports, then any subsequent attestation. The attack is against the verifier's trust, not against the node's availability.","remediation":"OEM BIOS update on the fixed Insyde kernel; flash + reboot per node. No config workaround, because the whole point is that the reported state is wrong. Until patched, do not treat PCR-based attestation from affected platforms as authoritative for tenant-isolation or key-release decisions, and cross-check firmware integrity out of band (BMC-side SPI measurement, offline flash comparison) rather than trusting the node's own quote.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-30633","https://www.insyde.com/security-pledge/SA-2023045"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-10-19"},{"id":"CVE-2023-32280","cve":"CVE-2023-32280","aliases":["INTEL-SA-00922"],"title":"Intel Server OpenBMC firmware (before egs-1.05) - credential storage: Credentials are insufficiently protected","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Server OpenBMC firmware (before egs-1.05) - credential storage","year":"2023","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Credentials are insufficiently protected and an unauthenticated caller on the network can read them. Scored only 5.3 because what leaks is partial, but credential material recovered without authentication is worth more than the score suggests on a fleet that reuses BMC credentials across nodes - the standard practice for anyone who provisioned racks from a template. Read once, reuse everywhere.","attack_vector":"Unauthenticated, over the network, to the BMC management interface.","remediation":"Fixed in Intel Server OpenBMC egs-1.05 and later - per-node out-of-band BMC firmware update. The compensating control is per-node unique BMC credentials, which is config-only, costs a provisioning change, and blunts every credential-disclosure bug in this cluster rather than just this one. If your fleet currently shares one BMC password, fixing that is higher leverage than this specific patch.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-32280","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00922.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-02-14"},{"id":"CVE-2023-34344","cve":"CVE-2023-34344","aliases":["AMI-SA-2023005","NVIDIA OSR review"],"title":"AMI MegaRAC SPx (IPMI handler): Timing and response differences in the IPMI handler let an unauthenticated attacker","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx (IPMI handler)","year":"2023","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Timing and response differences in the IPMI handler let an unauthenticated attacker confirm which usernames exist on the BMC. On its own it leaks nothing but names; in a fleet it is the reconnaissance step that makes credential-stuffing efficient, because it tells the attacker which nodes still carry the vendor default account or the provisioning template's service account before they spend any attempts.","attack_vector":"Network-reachable IPMI service, no credentials, no interaction. Any host that can reach UDP/623 on the BMC can enumerate accounts across the whole management range in a single sweep.","remediation":"Firmware flash to SPx_12.7 / SPx_13.5, out-of-band per node, ODM-gated. Genuinely low urgency for the flash itself. The config-only work is what matters: remove or rename vendor default accounts, ensure no username is shared across the fleet by the provisioning template, and disable IPMI-over-LAN where your tooling allows it - that closes the enumeration surface outright with no reboot.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2023005.pdf","https://nvd.nist.gov/vuln/detail/CVE-2023-34344"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-06-12"},{"id":"CVE-2023-36844","cve":"CVE-2023-36844","aliases":[],"title":"Juniper Junos OS J-Web (EX): PHP external variable modification","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Juniper Junos OS J-Web (EX)","year":"2023","cvss_score":5.3,"severity":"medium","kev":true,"impact":"PHP external variable modification; part of the exploited J-Web chain","attack_vector":"Network, unauthenticated","remediation":"Junos upgrade with switch failover","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-36844"],"status":"curated","published":"2023-08-17"},{"id":"CVE-2023-36847","cve":"CVE-2023-36847","aliases":[],"title":"Juniper Junos OS J-Web (EX): Missing authentication on `installAppPackage.php` — unauthenticated file upload to the","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Juniper Junos OS J-Web (EX)","year":"2023","cvss_score":5.3,"severity":"medium","kev":true,"impact":"Missing authentication on `installAppPackage.php` — unauthenticated file upload to the switch filesystem","attack_vector":"Network, unauthenticated","remediation":"Junos upgrade; combined with CVE-2023-36845 this is a pre-auth RCE chain against management switches","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-36847"],"status":"curated","published":"2023-08-17"},{"id":"CVE-2023-38958","cve":"CVE-2023-38958","aliases":["CVE-2023-38954","CVE-2023-38955","CVE-2023-38956"],"title":"ZKTeco BioAccess IVS v3.3.1 access control platform: An unauthenticated attacker can open and close any door","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"ZKTeco BioAccess IVS v3.3.1 access control platform","year":"2023","cvss_score":5.3,"severity":"medium","kev":false,"impact":"An unauthenticated attacker can open and close any door the platform manages by sending a crafted web request. The CVSS score badly understates this: a 5.3 labelled 'access control issue' is, in the physical world, a remote door-open primitive with no credentials. The same disclosure set adds unauthenticated device enumeration (IP addresses and names of every reader and controller), unauthenticated path traversal for arbitrary file read, and SQL injection - so the attacker can map the site, find the door serving the GPU cage specifically, and then open it. ZKTeco gear is common in cost-sensitive and fast-built deployments, which describes a lot of newer GPU hosting sites, and it is often installed by a general contractor rather than a security integrator. What follows from an open cage door is the usual list: drives with model weights and customer data walk out, a console or USB device gets attached to a running node, or someone plugs into the out-of-band switch and reaches every BMC in the row.","attack_vector":"Unauthenticated HTTP request to the BioAccess IVS platform. It is a web-managed server; where it is reachable from the corporate network or, worse, published for remote administration, the attack is a single request from anywhere. Device enumeration first means the attacker does not need prior knowledge of your site layout.","remediation":"Upgrade past 3.3.1 to a fixed BioAccess IVS release - ZKTeco's patch cadence and advisory quality are weak, so verify with the vendor that the specific door-control endpoint is fixed rather than trusting a version bump. Given the vendor track record, the stronger recommendation for a datacenter is to treat this product as unsuitable for a hall or cage boundary and plan replacement with a platform that has a real security program. Meanwhile: take the platform entirely off any network reachable from outside the security VLAN, put it behind a jump host, and add a compensating physical control on the doors that actually matter - a mechanical lock, a mantrap with a second independent system, or a guard - because a remote unlock primitive with no authentication is not something a network ACL fully mitigates if the ACL is ever wrong.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-38958","https://claroty.com/team82/disclosure-dashboard/cve-2023-38958","https://nvd.nist.gov/vuln/detail/CVE-2023-38956"],"status":"curated"},{"id":"CVE-2023-4155","cve":"CVE-2023-4155","aliases":[],"title":"Linux KVM - SEV-ES/SEV-SNP VMGEXIT double-fetch race: A KVM guest running SEV-ES or SEV-SNP with several vCPUs can","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux KVM - SEV-ES/SEV-SNP VMGEXIT double-fetch race","year":"2023","cvss_score":5.3,"severity":"medium","kev":false,"impact":"A KVM guest running SEV-ES or SEV-SNP with several vCPUs can trigger a double-fetch race in the host's VMGEXIT handler and drive it into recursion. Repeated invocation exhausts the host kernel stack and panics the hypervisor - a confidential guest taking down the host it runs on, and with it every other tenant on that machine. This is guest-to-host denial of service, which is the direction operators care about.","attack_vector":"From inside a guest VM using SEV-ES or SEV-SNP with multiple vCPUs. Tenant-reachable - no host privilege needed.","remediation":"Fixed in the Linux kernel - KVM/x86 SEV code or the ccp/PSP driver. Take the distro kernel update (RHEL/Rocky, Ubuntu, SLES) and **reboot the host**; SEV/SNP hypervisor paths cannot be live-patched in any meaningful way, and SNP platform init/shutdown is not safe to cycle under running guests. Drain confidential-VM tenants, reboot, then re-admit. No firmware, VBIOS or AGESA step needed, which makes this one of the cheaper classes of SEV fix to roll out. Prioritise on any host that admits tenant-controlled confidential VMs; the attacker prerequisite is just 'has a VM here'.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-4155","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2023-09-13"},{"id":"CVE-2023-4693","cve":"CVE-2023-4693","aliases":[],"title":"GRUB2 (NTFS filesystem parser): Out-of-bounds read in the same NTFS path leaks GRUB heap memory","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (NTFS filesystem parser)","year":"2023","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Out-of-bounds read in the same NTFS path leaks GRUB heap memory. On its own it is an information leak, but it is the ASLR-defeating half that makes the paired write bug reliably exploitable.","attack_vector":"Attacker-supplied NTFS volume, physical or via BMC virtual media.","remediation":"grub2 package update + reboot. Same as its sibling - dropping the NTFS module from your build is the durable answer on a Linux-only fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-4693","https://access.redhat.com/security/cve/CVE-2023-4693"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2023-10-25"},{"id":"CVE-2023-48299","cve":"CVE-2023-48299","aliases":[],"title":"TorchServe (model/workflow API): Information disclosure of files on the serving host","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"TorchServe (model/workflow API)","year":"2023","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Information disclosure of files on the serving host","attack_vector":"Network to the management API","remediation":"Patch to 0.9.0+; firewall the management plane","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-48299"],"status":"curated","published":"2023-11-21"},{"id":"CVE-2023-52738","cve":"CVE-2023-52738","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/fence): A NULL pointer dereference in the amdgpu RAS","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/fence)","year":"2023","cvss_score":5.3,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu RAS / GPU reset and recovery path. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu/fence: Fix oops due to non-matching drm_sched init/fini","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52738","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-21"},{"cwe":["CWE-20","CWE-401"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-53814","cve":"CVE-2023-53814","aliases":[],"title":"Linux kernel (drivers/pci): When the kernel coalesces two adjacent host-bridge apertures it invalidates the absorbed","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/pci)","year":"2023","cvss_score":5.3,"severity":"medium","kev":false,"impact":"When the kernel coalesces two adjacent host-bridge apertures it invalidates the absorbed resource by zeroing its fields, but the validity test only looks at the end address - so a legitimate resource that happens to end at zero is silently discarded. A root-bus resource that is dropped is never registered, meaning the kernel no longer knows that range is spoken for and can hand it out again during later BAR assignment, on top of an aperture that is genuinely in use.","attack_vector":"Not attacker-triggered - it fires at enumeration time on any host bridge whose firmware or device tree describes contiguous apertures that get coalesced, and it is deterministic for a given platform rather than racy. There is no tenant or fabric path to it. It earns a place here because the consequence is in resource assignment: a bus range the kernel forgot about is a range it may reassign, and overlapping windows are how one device ends up decoding another's addresses. The reported instance is an ARM board, but the flawed check is in generic drivers/pci/probe.c and applies to any bridge that coalesces.","remediation":"Boot a kernel where the invalid-resource check tests the full resource, not just .end. Interim: on affected kernels, compare the root bus resources the kernel prints at boot against what firmware advertises and treat a missing aperture as a reason not to place multi-tenant workloads on that platform.","references":["https://git.kernel.org/stable/c/e4af080f3ef6a65b0d702988c2471a47c9ae2cc0","https://git.kernel.org/stable/c/fe6a1fbe83f5b23d7db93596b793561230f06b40","https://nvd.nist.gov/vuln/detail/CVE-2023-53814"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2023-6032","cve":"CVE-2023-6032","aliases":["SEVD-2023-318-03"],"title":"Schneider Electric Galaxy VS / VL / VXL three-phase UPS, Network Management Card over HTTPS: Path traversal lets","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Schneider Electric Galaxy VS / VL / VXL three-phase UPS, Network Management Card over HTTPS","year":"2023","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Path traversal lets an attacker enumerate and download files from the management card on a large three-phase UPS - the class of unit that sits between utility power and an entire GPU hall, not a single rack. What leaks is configuration and credential material for the power estate. Treat this as reconnaissance that precedes a physical availability attack rather than as a data-loss event in itself.","attack_vector":"Anyone who can reach the NMC's HTTPS interface. Galaxy-class UPS management cards live on the facility network, which in a leased colo is usually the landlord's network, not yours.","remediation":"Firmware update to the card, per SEVD-2023-318-03. Non-disruptive to the load, but on a leased site you may not own the equipment - in that case the real remediation is contractual: require the landlord to evidence the patch level of every UPS management card that feeds your halls, and require the facility network to be segmented from anything you run.","references":["https://download.schneider-electric.com/files?p_Doc_Ref=SEVD-2023-318-03&p_enDocType=Security+and+Safety+Notice&p_File_Name=SEVD-2023-318-03.pdf"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2023-11-15"},{"id":"CVE-2024-0171","cve":"CVE-2024-0171","aliases":[],"title":"Dell PowerEdge Server BIOS (AMD platforms, TOCTOU race): A time-of-check/time-of-use race in BIOS gives","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell PowerEdge Server BIOS (AMD platforms, TOCTOU race)","year":"2024","cvss_score":5.3,"severity":"medium","kev":false,"impact":"A time-of-check/time-of-use race in BIOS gives a low-privileged local attacker access to resources outside its authorization. Relevant to the AMD EPYC-based GPU platforms.","attack_vector":"Local low-privilege user on an AMD-based PowerEdge.","remediation":"Flash the fixed PowerEdge BIOS. A BIOS update is a cold reboot per node and cannot be done live - on a GPU fleet that means draining jobs and taking the box out of the scheduler, so batch it with other firmware work rather than doing a standalone pass.","references":["https://www.dell.com/support/kbdoc/en-us/000226253/dsa-2024-039-security-update-for-dell-amd-based-poweredge-server-vulnerability"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-21944","cve":"CVE-2024-21944","aliases":[],"title":"AMD SEV-SNP (BadRAM): BadRAM: improper validation of DIMM SPD metadata lets an attacker with physical access or ring0","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"AMD SEV-SNP (BadRAM)","year":"2024","cvss_score":5.3,"severity":"medium","kev":false,"impact":"BadRAM: improper validation of DIMM SPD metadata lets an attacker with physical access or ring0 on a non-compliant DIMM overwrite guest memory and forge SNP attestation","attack_vector":"Physical access / compromised host firmware against a confidential tenant VM","remediation":"AGESA/BIOS firmware update + reboot, plus DIMM SPD lockdown at the supply-chain level. Cannot be fixed in software - a genuine constraint on any \"we cannot see your data\" confidential-GPU claim","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21944"],"status":"curated","fleet":{"ubiquity":"common - 3rd/4th-gen EPYC (Milan, Milan-X, Genoa, Bergamo, Genoa-X, Siena) hosts under GPU nodes","remediation_pain":"firmware-flash - AMD's fix validates SPD metadata at boot, so it is a BIOS/AGESA update per node; the underlying attack needs ~$10 of hardware and physical access, which colo and bare-metal-rental models do not exclude","pain_class":"physical access","why_fleet_wide":"Forges SEV-SNP attestation reports and inserts undetectable backdoors into confidential VMs, i.e. every attestation a customer verified on affected hosts is retroactively meaningless."},"published":"2026-06-10"},{"id":"CVE-2024-23650","cve":"CVE-2024-23650","aliases":[],"title":"BuildKit: Malicious client or frontend crashes the BuildKit daemon","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"BuildKit","year":"2024","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Malicious client or frontend crashes the BuildKit daemon","attack_vector":"Anyone with build API access","remediation":"Upgrade BuildKit","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-23650"],"status":"curated","published":"2024-01-31"},{"id":"CVE-2024-27891","cve":"CVE-2024-27891","aliases":["Arista Security Advisory 0102"],"title":"Arista EOS (MACsec with egress ACLs): On interfaces with both MACsec and egress ACLs configured, the egress ACL is not","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (MACsec with egress ACLs)","year":"2024","cvss_score":5.3,"severity":"medium","kev":false,"impact":"On interfaces with both MACsec and egress ACLs configured, the egress ACL is not enforced for packets leaving those ports. The combination — link encryption plus egress filtering — is exactly what you deploy on inter-site or inter-pod links carrying multiple tenants, so the failure lands on the highest-trust links in the build.","attack_vector":"Traffic egressing an interface configured with both MACsec and an egress ACL. No attacker capability needed.","remediation":"EOS upgrade plus reload. Interim: move the filtering to the ingress direction on the far side of the link, which is a live config change and restores enforcement without touching MACsec.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-27891","https://www.arista.com/en/support/advisories-notices/security-advisory/19908-security-advisory-0102"],"status":"curated","tags":["tenant-isolation"],"published":"2026-06-04"},{"id":"CVE-2024-31157","cve":"CVE-2024-31157","aliases":[],"title":"Intel UEFI firmware (OutOfBandXML module): Improper initialisation in the OutOfBandXML UEFI module allows a privileged","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel UEFI firmware (OutOfBandXML module)","year":"2024","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Improper initialisation in the OutOfBandXML UEFI module allows a privileged user to disclose information from firmware. The out-of-band XML path is part of remote platform configuration, so it is reachable in the management workflows operators actually automate.","attack_vector":"Privileged local access on the host.","remediation":"Fixed in platform BIOS/UEFI firmware. That means an OEM release, a per-node drain, a flash and a cold reboot - and OEM availability commonly lags the Intel advisory by quarters on server boards. There is no microcode or OS-level shortcut for this class; budget it as a fleet-wide maintenance campaign, not a patch.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-31157","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01139.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-02-12"},{"id":"CVE-2024-37152","cve":"CVE-2024-37152","aliases":[],"title":"Argo CD: /api/v1/settings exposes sensitive settings without authentication","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2024","cvss_score":5.3,"severity":"medium","kev":false,"impact":"/api/v1/settings exposes sensitive settings without authentication","attack_vector":"Unauthenticated network","remediation":"Rolling Argo CD upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-37152"],"status":"curated","published":"2024-06-06"},{"id":"CVE-2024-38303","cve":"CVE-2024-38303","aliases":[],"title":"Dell PowerEdge 14G Intel BIOS (improper input validation): A high-privileged local attacker extracts information","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell PowerEdge 14G Intel BIOS (improper input validation)","year":"2024","cvss_score":5.3,"severity":"medium","kev":false,"impact":"A high-privileged local attacker extracts information from the firmware layer with scope change - firmware-resident secrets and platform configuration leaking to the OS.","attack_vector":"Local high-privilege user on a 14G Intel PowerEdge.","remediation":"Update 14G Intel BIOS to 2.22.x or later. Flash the fixed PowerEdge BIOS. A BIOS update is a cold reboot per node and cannot be done live - on a GPU fleet that means draining jobs and taking the box out of the scheduler, so batch it with other firmware work rather than doing a standalone pass.","references":["https://www.dell.com/support/kbdoc/en-us/000228135/dsa-2024-309-security-update-for-dell-poweredge-server-for-improper-input-validation-vulnerability"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-39707","cve":"CVE-2024-39707","aliases":["INSYDE-SA-2024007"],"title":"Insyde InsydeH2O (IHISI function 0x49, UEFI variable factory reset): IHISI function 0x49 restores certain UEFI","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (IHISI function 0x49, UEFI variable factory reset)","year":"2024","cvss_score":5.3,"severity":"medium","kev":false,"impact":"IHISI function 0x49 restores certain UEFI variables to factory defaults with no authentication by default. That is a rollback primitive: an attacker resets security-relevant firmware settings that an operator had hardened - and on affected platforms can wind protections back to a weaker known state without needing to exploit anything. For a fleet where node hardening is applied once at provisioning and then assumed, this quietly undoes it, and nothing in the OS logs the change.","attack_vector":"Local attacker on the host OS able to invoke the IHISI interface. No authentication required on affected platforms.","remediation":"OEM BIOS update carrying the Insyde fix, which makes the function require authentication. Firmware flash, reboot per node. Because the impact is settings rollback rather than code execution, the practical operator control is detection: baseline your BIOS settings at provisioning and re-verify them (via the OEM's redfish/BIOS-attribute API) on every node before it re-enters the tenant pool, rather than assuming a hardened node stays hardened.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-39707","https://www.insyde.com/security-pledge/SA-2024007"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-11-14"},{"id":"CVE-2024-42478","cve":"CVE-2024-42478","aliases":[],"title":"llama.cpp (RPC backend): Arbitrary address read via `rpc_tensor.data`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"llama.cpp (RPC backend)","year":"2024","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Arbitrary address read via `rpc_tensor.data`","attack_vector":"Unauthenticated network to the RPC port","remediation":"Rebuild; network-isolate","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-42478"],"status":"curated","published":"2024-08-12"},{"id":"CVE-2024-43101","cve":"CVE-2024-43101","aliases":[],"title":"Intel Data Center GPU Flex Series - Windows driver software: Improper access control in the Flex Series Windows driver","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel Data Center GPU Flex Series - Windows driver software","year":"2024","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Improper access control in the Flex Series Windows driver software allows an authenticated user to cause denial of service.","attack_vector":"Local, authenticated, on a Windows host running Flex Series driver software before 31.0.101.4255.","remediation":"Update to 31.0.101.4255 or later. Cost: node reboot after driver replacement.","references":["https://www.intel.com/content/www/us/en/security-center/default.html","https://nvd.nist.gov/vuln/detail/CVE-2024-43101"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-05-13"},{"id":"CVE-2024-45783","cve":"CVE-2024-45783","aliases":["GRUB2 2025 batch"],"title":"GRUB2 (HFS+ filesystem parser): A reference count can be decremented twice, producing a use-after-free","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (HFS+ filesystem parser)","year":"2024","cvss_score":5.3,"severity":"medium","kev":false,"impact":"A reference count can be decremented twice, producing a use-after-free. Lower severity on its own but a usable link in a chain with the write primitives in the same batch.","attack_vector":"Attacker-supplied HFS+ volume.","remediation":"grub2 package update + reboot; or drop the module.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-45783","https://www.openwall.com/lists/oss-security/2025/02/18/3"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-02-18"},{"id":"CVE-2025-23028","cve":"CVE-2025-23028","aliases":[],"title":"Cilium: Denial of service in the Cilium dataplane","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2025","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Denial of service in the Cilium dataplane","attack_vector":"Any pod on the cluster network","remediation":"Rolling Cilium upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23028"],"status":"curated","published":"2025-01-22"},{"id":"CVE-2025-29934","cve":"CVE-2025-29934","aliases":[],"title":"AMD CPU - stale TLB entries in SEV-SNP guests: A silicon bug lets a local admin-privileged attacker run an SEV-SNP","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD CPU - stale TLB entries in SEV-SNP guests","year":"2025","cvss_score":5.3,"severity":"medium","kev":false,"impact":"A silicon bug lets a local admin-privileged attacker run an SEV-SNP guest against stale TLB entries. The guest executes with translations that no longer reflect the real page mappings, which the host can steer - a data-integrity attack on a confidential VM that leaves no trace in the guest, since from inside the VM the memory simply reads wrong.","attack_vector":"Local, admin-privileged host attacker targeting a confidential guest.","remediation":"Fixed in AMD SEV firmware / AGESA and reaches you as an OEM SBIOS package - AMD hands AGESA to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before shipping BIOS. **Budget one to six months of OEM lag**, longer on older platforms and sometimes never on end-of-support SKUs. Applying it means draining the host and doing a full power cycle. Because the fix moves the platform's reported SEV-SNP TCB version, you must also pull fresh VCEK certificates from AMD's Key Distribution Service and update any attestation policy your tenants pin - otherwise guests will start failing launch validation the moment the BIOS lands. Some SEV firmware can alternatively be staged from linux-firmware (amd/amd_sev_*.sbin) and committed via the ccp driver at boot, which is faster than waiting on BIOS - check whether your platform supports firmware hot-load before assuming the OEM is the only route.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-29934","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2025-11-21"},{"id":"CVE-2025-33185","cve":"CVE-2025-33185","aliases":[],"title":"NVIDIA AIStore - AuthN: An unauthenticated user extracts information from the AIStore authentication component","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA AIStore - AuthN","year":"2025","cvss_score":5.3,"severity":"medium","kev":false,"impact":"An unauthenticated user extracts information from the AIStore authentication component. AIStore fronts training datasets, so what leaks here is metadata about your data estate.","attack_vector":"Network, unauthenticated, no user interaction. Anyone who can reach the AuthN endpoint.","remediation":"Upgrade AIStore per bulletin 5724 and roll the AuthN pods. Cost: rolling restart of the storage control plane; data path is unaffected.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33185","https://github.com/NVIDIA/product-security/tree/main/2025/5724"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:L/I:N/A:N","cwe":["CWE-862"],"fleet":{"pain_class":"daemon-restart"},"published":"2025-11-11"},{"id":"CVE-2025-38742","cve":"CVE-2025-38742","aliases":[],"title":"Dell iDRAC Service Module (iSM, incorrect permissions): Incorrect permission assignment on a critical resource lets","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC Service Module (iSM, incorrect permissions)","year":"2025","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Incorrect permission assignment on a critical resource lets a low-privileged local user reach code execution via the iSM agent.","attack_vector":"Low-privilege local account on the host OS.","remediation":"Upgrade iSM to 6.0.3.0. Package update and service restart only.","references":["https://www.dell.com/support/kbdoc/en-us/000359617/dsa-2025-311-security-update-for-dell-idrac-service-module-vulnerabilities"],"status":"curated","fleet":{"pain_class":"daemon-restart"}},{"id":"CVE-2025-5197","cve":"CVE-2025-5197","aliases":[],"title":"HuggingFace transformers: ReDoS in `convert_tf_weight_name_to_pt_weight_name`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"HuggingFace transformers","year":"2025","cvss_score":5.3,"severity":"medium","kev":false,"impact":"ReDoS in `convert_tf_weight_name_to_pt_weight_name`","attack_vector":"Customer-supplied TF checkpoint converted at load","remediation":"Upgrade transformers","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-5197"],"status":"curated","published":"2025-08-06"},{"id":"CVE-2025-52534","cve":"CVE-2025-52534","aliases":[],"title":"AMD CPU microcode - bound check: An improper bound check inside AMD CPU microcode lets a malicious **guest** write into","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD CPU microcode - bound check","year":"2025","cvss_score":5.3,"severity":"medium","kev":false,"impact":"An improper bound check inside AMD CPU microcode lets a malicious **guest** write into host memory. This is a straight VM-escape primitive expressed in silicon microcode rather than in the hypervisor: the tenant does not need a QEMU or KVM bug, the CPU itself lets the write through. Once a guest can write host memory the host is compromised, and from there every other VM on the box.","attack_vector":"From inside a guest VM - no host privilege required. That makes this materially more dangerous operationally than the local-admin microcode issues: your tenants are the attackers in the threat model.","remediation":"Fixed by an AMD microcode patch. Two delivery routes, and the difference matters: the linux-firmware amd-ucode blobs load early at boot (initramfs) and need only a reboot, while the durable fix is the microcode embedded in the OEM SBIOS/AGESA package, which carries the usual one-to-six-month OEM lag and a full power cycle. **For confidential computing you need the SBIOS route**: microcode late-loaded by the OS is not part of what SEV-SNP attests, so a guest checking the attestation report cannot tell the fix is present. AMD does not support late-loading microcode on a running EPYC host - treat this as reboot-required. After patching, expect the reported TCB version to change and plan the VCEK certificate refresh accordingly. Prioritise this above the local-admin microcode issues on any node that runs untrusted guest VMs - the attacker prerequisite here is simply 'is a tenant'.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-52534","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2026-02-10"},{"id":"CVE-2025-54511","cve":"CVE-2025-54511","aliases":[],"title":"AMD Secure Processor - privilege check on write path: The ASP accepts an input value and performs a write","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor - privilege check on write path","year":"2025","cvss_score":5.3,"severity":"medium","kev":false,"impact":"The ASP accepts an input value and performs a write without confirming the caller had sufficient privilege to ask for it. The result is an integrity loss inside the secure processor - a caller that should have been refused gets its write.","attack_vector":"Local, through the ASP's callable interface.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-54511","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2026-05-15"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:L/A:L","cwe":["CWE-77","CWE-74"],"fleet":{"pain_class":"unpatchable / mitigate-only"},"id":"CVE-2026-15035","cve":"CVE-2026-15035","aliases":[],"title":"BentoML OpenLLM 0.6.30 (async_run_command in src/openllm/common.py): A model repository directory name flows unescaped","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"BentoML OpenLLM 0.6.30 (async_run_command in src/openllm/common.py)","year":"2026","cvss_score":5.3,"severity":"medium","kev":false,"impact":"A model repository directory name flows unescaped into a shell command, so a crafted repository name executes attacker commands as the user running OpenLLM. The exploit is public and, as of the advisory, the project had not responded to the report.","attack_vector":"A local user of OpenLLM who adds or uses a model repository with an attacker-chosen directory name. Requires local access and low privileges.","remediation":"No vendor fix is confirmed. Do not pass untrusted repository paths to OpenLLM, run it as an unprivileged user in a container, and track the upstream issue before treating it as remediated.","references":["https://github.com/bentoml/OpenLLM/issues/1229","https://nvd.nist.gov/vuln/detail/CVE-2026-15035"],"status":"curated"},{"id":"CVE-2026-24208","cve":"CVE-2026-24208","aliases":[],"title":"Triton Inference Server: Info disclosure via path traversal on model files","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2026","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Info disclosure via path traversal on model files","attack_vector":"Network client","remediation":"Upgrade Triton; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24208","https://github.com/NVIDIA/product-security/tree/main/2026/5828"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:L","cwe":["CWE-22"],"published":"2026-05-20"},{"id":"CVE-2026-24227","cve":"CVE-2026-24227","aliases":[],"title":"TensorRT: DoS via resource exhaustion","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT","year":"2026","cvss_score":5.3,"severity":"medium","kev":false,"impact":"DoS via resource exhaustion","attack_vector":"Malicious model/engine input","remediation":"Bump TensorRT; rebuild serving images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24227","https://github.com/NVIDIA/product-security/tree/main/2026/5855"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:L","cwe":["CWE-502"],"published":"2026-07-14"},{"id":"CVE-2026-47262","cve":"CVE-2026-47262","aliases":[],"title":"containerd: Crafted image causes DoS during container creation","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2026","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Crafted image causes DoS during container creation","attack_vector":"Malicious image","remediation":"Rolling containerd upgrade with node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47262"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2026-07-01"},{"id":"CVE-2026-47622","cve":"CVE-2026-47622","aliases":[],"title":"NVIDIA Dynamo: Error messages from Dynamo leak sensitive information to an unauthenticated network caller","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Dynamo","year":"2026","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Error messages from Dynamo leak sensitive information to an unauthenticated network caller - typically internal paths, versions and configuration that shorten the next stage of an attack.","attack_vector":"Network, unauthenticated. Anyone who can reach the Dynamo serving endpoint and provoke an error.","remediation":"Upgrade Dynamo per bulletin 5842 and roll the deployment. Cost: rolling restart of the serving tier.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47622","https://github.com/NVIDIA/product-security/tree/main/2026/5842"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:L/I:N/A:N","cwe":["CWE-209"],"fleet":{"pain_class":"daemon-restart"},"published":"2026-08-04"},{"id":"CVE-2026-55686","cve":"CVE-2026-55686","aliases":[],"title":"Podman: Malicious image WORKDIR symlink creates directories or changes ownership on the host filesystem","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Podman","year":"2026","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Malicious image WORKDIR symlink creates directories or changes ownership on the host filesystem","attack_vector":"Malicious image","remediation":"Upgrade Podman to 5.7.1+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-55686"],"status":"curated","published":"2026-06-26"},{"id":"NCVD-2020-002-rdma-fabric-remote-dram-bank-con","cve":null,"aliases":["Bankrupt","RDMA memory-bank covert channel","Ustiugov et al., arXiv:2006.03854"],"title":"RDMA fabric + remote DRAM bank contention (cross-node covert channel): Bankrupt establishes a 74 Kb/s covert channel","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"RDMA fabric + remote DRAM bank contention (cross-node covert channel)","year":"2020","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Bankrupt establishes a 74 Kb/s covert channel between two processes on different machines in an RDMA network, by steering RDMA packets to addresses that map to a single DRAM bank on a shared intermediary node and timing the resulting queuing. It remained undetectable to existing monitoring - CPU and NIC performance counters showed nothing. For an operator, this defeats the assumption that network segmentation between tenants prevents exfiltration: a compromised process inside an isolated enclave can signal out to a colluding receiver anywhere on the same RDMA fabric, without opening a connection between them. Any data-loss-prevention story that relies on egress controls at the IP layer is bypassed.","attack_vector":"Spy and receiver each allocate their own private memory region on a common intermediary machine - a normal thing for any RDMA tenant to do. The spy issues RDMA operations to a chosen set of remote addresses, causing deep queuing at one memory bank; the receiver probes addresses mapped to the same bank in its own region and reads the timing. Both sides only ever touch memory they legitimately own, which is why nothing flags it.","remediation":"No patch. Mitigation is placement and monitoring: avoid a shared intermediary that both a sensitive tenant and an untrusted tenant can target (scheduler policy change), and where memory-bank interleaving is configurable, randomise the physical-address-to-bank mapping per tenant (BIOS/firmware setting, requires a host reboot). Detection needs per-QP latency-distribution telemetry rather than counters - a monitoring build-out, not a config toggle. For genuinely sensitive workloads, a dedicated fabric is the only reliable answer.","references":["https://arxiv.org/abs/2006.03854"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"tags":["tenant-isolation"]},{"id":"NCVD-2020-004-rdma-fabric-remote-dram-bank-con","cve":null,"aliases":["Bankrupt","RDMA memory-bank covert channel","Ustiugov et al., arXiv:2006.03854"],"title":"RDMA fabric + remote DRAM bank contention (cross-node covert channel): Bankrupt establishes a 74 Kb/s covert channel","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"RDMA fabric + remote DRAM bank contention (cross-node covert channel)","year":"2020","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Bankrupt establishes a 74 Kb/s covert channel between two processes on different machines in an RDMA network, by steering RDMA packets to addresses that map to a single DRAM bank on a shared intermediary node and timing the resulting queuing. It remained undetectable to existing monitoring - CPU and NIC performance counters showed nothing. For an operator, this defeats the assumption that network segmentation between tenants prevents exfiltration: a compromised process inside an isolated enclave can signal out to a colluding receiver anywhere on the same RDMA fabric, without opening a connection between them. Any data-loss-prevention story that relies on egress controls at the IP layer is bypassed.","attack_vector":"Spy and receiver each allocate their own private memory region on a common intermediary machine - a normal thing for any RDMA tenant to do. The spy issues RDMA operations to a chosen set of remote addresses, causing deep queuing at one memory bank; the receiver probes addresses mapped to the same bank in its own region and reads the timing. Both sides only ever touch memory they legitimately own, which is why nothing flags it.","remediation":"No patch. Mitigation is placement and monitoring: avoid a shared intermediary that both a sensitive tenant and an untrusted tenant can target (scheduler policy change), and where memory-bank interleaving is configurable, randomise the physical-address-to-bank mapping per tenant (BIOS/firmware setting, requires a host reboot). Detection needs per-QP latency-distribution telemetry rather than counters - a monitoring build-out, not a config toggle. For genuinely sensitive workloads, a dedicated fabric is the only reliable answer.","references":["https://arxiv.org/abs/2006.03854"],"status":"curated","tags":["tenant-isolation"]},{"id":"NCVD-2026-006-dmtf-libspdm-csr-generation-unde","cve":null,"aliases":["GHSA-j54w-759w-xj3m","DMTF-2026-0002"],"title":"DMTF libspdm CSR generation under the mbedTLS crypto backend (cryptlib_mbedtls): Stack corruption inside the firmware","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"DMTF libspdm CSR generation under the mbedTLS crypto backend (cryptlib_mbedtls)","year":"2026","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Stack corruption inside the firmware component that generates certificate signing requests - meaning the code path that establishes a device's cryptographic identity is the one that can be smashed. At minimum this is a denial of service against device provisioning; at worst, stack corruption in a firmware context of this privilege is a candidate for control-flow hijack. On a GPU fleet the affected components are the accelerators, NICs and management controllers that participate in device identity provisioning, and a device whose identity provisioning can be attacked is a device whose attestation claims an operator should not rely on. An oversized Common Name in the RequesterInfo field of a GET_CSR request writes past the end of a stack array. Reachable only where the responder advertises CSR_CAP and uses libspdm_gen_x509_csr().","attack_vector":"An SPDM requester able to send GET_CSR to a responder that has CSR_CAP enabled and uses the mbedTLS backend. That is a narrow configuration, but it is the configuration used by embedded firmware that cannot carry OpenSSL - which describes a lot of BMC and device firmware.","remediation":"Update libspdm and rebuild affected firmware; there is no CVE, so this will not surface through NVD-based scanning and you have to track the DMTF advisory series directly. The narrowing conditions are useful operationally: ask vendors whether their SPDM responder is built with CSR_CAP and against mbedTLS, and if CSR generation is not a capability you use, having it compiled out removes the path entirely. As with everything in libspdm, the actual rollout is per-vendor firmware images and per-device flashes.","references":["https://github.com/DMTF/libspdm/security/advisories/GHSA-j54w-759w-xj3m"],"status":"curated"},{"id":"NCVD-2026-007-dmtf-libspdm-responder-handling","cve":null,"aliases":["GHSA-m4wc-xmvg-369f","DMTF-2026-0001"],"title":"DMTF libspdm responder handling of GET_MEASUREMENT_EXTENSION_LOG: A requester reads memory it was never authorised","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"DMTF libspdm responder handling of GET_MEASUREMENT_EXTENSION_LOG","year":"2026","cvss_score":5.3,"severity":"medium","kev":false,"impact":"A requester reads memory it was never authorised to read, out of the responder - which in a datacenter is typically the firmware of a GPU, NIC, or the BMC acting as an SPDM responder. What leaks is whatever sits adjacent to the measurement extension log in that component's address space: potentially keys, measurement state, or other firmware data. Because the responder is a device that multiple hosts or tenants may talk to over the life of a node, this is a cross-boundary read from a component that is supposed to be the root of the attestation story rather than a target of it. The Offset and Length fields are added with wrapping arithmetic before the bounds check, so a crafted pair overflows and passes validation. Requires the responder to advertise MEL_CAP and CHUNK_CAP.","attack_vector":"Anything that can act as an SPDM requester to the affected responder - a host driver, a management controller, or a peer device on the fabric. On bare metal that includes the tenant's own host software if the platform lets host drivers speak SPDM to accelerators, which most do.","remediation":"This one has no CVE assigned, only a GHSA and a DMTF advisory number, so it will not appear in NVD-driven scanning at all - operators tracking firmware risk off CVE feeds alone will simply never see it. Remediation is a libspdm update followed by per-vendor firmware rebuilds and per-device flashes, on each vendor's own timeline. If your responders do not need the measurement extension log, having the vendor build with MEL_CAP disabled removes the reachable path, but that is a firmware build option rather than something an operator can toggle.","references":["https://github.com/DMTF/libspdm/security/advisories/GHSA-m4wc-xmvg-369f"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:L/I:N/A:N","cwe":["CWE-306"],"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2026-041-rclone-rc-server-debug-pprof-han","cve":null,"aliases":["GHSA-mfvx-7rcj-9m5g"],"title":"rclone (rc server, /debug/pprof handler): The pprof debug handler is mounted as its own route on the rclone","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"rclone (rc server, /debug/pprof handler)","year":"2026","cvss_score":5.3,"severity":"medium","kev":false,"impact":"The pprof debug handler is mounted as its own route on the rclone remote-control server, outside the fail-closed authentication check that guards everything else. An unauthenticated caller fetches /debug/pprof/cmdline and gets the full argv of the rclone process - which for a data mover means the object-storage access keys, bucket names and endpoints passed on the command line. The same gap also lets an unauthenticated caller enumerate the configured remote names.","attack_vector":"Anyone who can reach the rclone rc port. Data-staging jobs on GPU clusters routinely run rclone with --rc bound to the node interface so a controller can drive it, which puts this in reach of any co-tenant on the cluster network.","remediation":"Upgrade rclone to the release carrying this fix and restart the rc server. Independently, stop passing credentials as command-line arguments - move them to the rclone config file or environment - and bind --rc-addr to localhost with --rc-user/--rc-pass set. No CVE ID has been assigned; track it by the GHSA.","references":["https://github.com/rclone/rclone/security/advisories/GHSA-mfvx-7rcj-9m5g"],"status":"curated"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:L","cwe":["CWE-400"],"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2026-044-vllm-deepstream-video-backend-vi","cve":null,"aliases":["GHSA-cqm8-jxg6-fqfq"],"title":"vLLM (DeepStream video backend, VideoMediaIO backend selection): A performance feature merged past two existing","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (DeepStream video backend, VideoMediaIO backend selection)","year":"2026","cvss_score":5.3,"severity":"medium","kev":false,"impact":"A performance feature merged past two existing security boundaries and reopened both. vLLM guards against request-level selection of GPU video backends by asking backend_requires_gpu(), but that function returns false for any name it does not know — and DeepStream was never registered, so it reads as a non-GPU backend and the guard waves it through. Its decode path also skips VLLM_MAX_IMAGE_PIXELS, the frame-dimension limit added specifically to stop compressed media from exhausting memory. An unauthenticated client therefore activates NVDEC/GStreamer at request time, initialises the process-wide GPU decode pool outside the startup memory reservation, and submits video every other backend would have rejected. Measured effect was half of lightweight canary requests timing out. For a multi-tenant serving fleet this is one caller degrading GPU decode capacity for everyone on the replica, from an unauthenticated position.","attack_vector":"Network, unauthenticated, against a vLLM deployment on 0.26.x that accepts video input. The attacker names the deepstream backend in the request to bypass the GPU-backend restriction and then submits oversized media.","remediation":"Upgrade to vLLM 0.27.0 or later and restart the servers. In the interim, reject request-level video_backend and backend parameters at your gateway rather than relying on the in-process check, and cap request body size and video dimensions upstream. Verify GPU memory reservation covers any decode pool your deployment can actually reach.","references":["https://github.com/vllm-project/vllm/security/advisories/GHSA-cqm8-jxg6-fqfq"],"status":"curated"},{"id":"CVE-2020-15257","cve":"CVE-2020-15257","aliases":[],"title":"containerd: containerd-shim abstract-socket API exposed to host-network containers","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2020","cvss_score":5.2,"severity":"medium","kev":false,"impact":"containerd-shim abstract-socket API exposed to host-network containers; container escape to host root","attack_vector":"Any tenant pod running with hostNetwork:true","remediation":"Upgrade containerd; drain node to restart shims. Also ban hostNetwork for tenant pods via policy","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-15257"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2020-12-01"},{"id":"CVE-2021-46746","cve":"CVE-2021-46746","aliases":[],"title":"AMD Secure Processor TEE - Secure OS stack overrun (AMD-SB-3003): A stack overrun in the ASP Secure OS trusted","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor TEE - Secure OS stack overrun (AMD-SB-3003)","year":"2021","cvss_score":5.2,"severity":"medium","kev":false,"impact":"A stack overrun in the ASP Secure OS trusted execution environment, denying service to the secure processor. With the ASP down, the platform loses fTPM services, SEV key operations and attestation - so on a confidential-computing host this is not a cosmetic crash, it takes your CVM capacity offline until the node is power-cycled.","attack_vector":"Local, through the ASP TEE interface.","remediation":"Fixed in AMD PI/AGESA firmware and delivered only as an OEM SBIOS package - AMD ships the PI drop to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before releasing BIOS. **Budget one to six months of OEM lag**, and note that several CVEs in this batch are marked 'no fix planned' on Naples (EPYC 7001) - for those the only remediation is retiring the hardware. Applying it means cordon, drain and a full power cycle per node; there is no driver reload, no live patch and no VBIOS step.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-46746","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3003.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"published":"2024-08-13"},{"id":"CVE-2023-31011","cve":"CVE-2023-31011","aliases":[],"title":"DGX H100 BMC (REST): DoS","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX H100 BMC (REST)","year":"2023","cvss_score":5.2,"severity":"medium","kev":false,"impact":"DoS","attack_vector":"Network-adjacent BMC REST client","remediation":"Flash BMC 23.08.18 out-of-band","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5473/5473.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:H/UI:N/S:U/C:L/I:H/A:N","cwe":["CWE-20"],"published":"2023-09-20"},{"id":"CVE-2023-31189","cve":"CVE-2023-31189","aliases":["INTEL-SA-00922"],"title":"Intel Server OpenBMC firmware (before egs-1.09) - authentication logic: An authenticated low-privilege user escalates","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Server OpenBMC firmware (before egs-1.09) - authentication logic","year":"2023","cvss_score":5.2,"severity":"medium","kev":false,"impact":"An authenticated low-privilege user escalates privilege on the BMC, with a scope change - the escalation crosses a security boundary rather than staying inside the account model. On fleets that hand out constrained BMC accounts to tenants, support staff or monitoring systems (read-only sensor scraping is a common one), this converts any of those accounts into something with more control over the node. The lesson for a GPU operator: a read-only BMC account issued to a customer or a monitoring vendor is not a safe thing to hand out on this firmware.","attack_vector":"Local access with a low-privileged authenticated BMC account. Needs an existing credential, so exposure tracks how widely you distribute BMC accounts.","remediation":"Fixed in Intel Server OpenBMC egs-1.09 and later; delivery is a per-node out-of-band BMC firmware update through Intel's platform packages, with the usual OEM rebase lag on non-Intel-badged boards using the same base. Config-only compensation: stop issuing BMC accounts to third parties, proxy sensor and telemetry reads through your own collector rather than giving monitoring vendors direct BMC credentials.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31189","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00922.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-02-14"},{"id":"CVE-2024-20397","cve":"CVE-2024-20397","aliases":[],"title":"Cisco NX-OS (bootloader / image signature verification): Secure boot on the switch is defeatable: an attacker","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco NX-OS (bootloader / image signature verification)","year":"2024","cvss_score":5.2,"severity":"medium","kev":false,"impact":"Secure boot on the switch is defeatable: an attacker with physical access or admin credentials can make a Nexus load an unsigned NX-OS image. That is how a fabric compromise becomes permanent — a modified image keeps root across reloads and reflashes, and nothing in your config management will notice. Affects Nexus 3000/7000/9000, MDS 9000 and UCS 6400/6500 fabric interconnects, so it covers both the Ethernet and the storage fabric.","attack_vector":"Physical access to the switch (console/bootloader prompt) or an existing administrative account. Realistic threat model for colocation, shared cages, and any switch that has passed through a supply chain or an RMA.","remediation":"BIOS update on both the primary and the alternate BIOS bank — either through `install all` with a fixed NX-OS release or Cisco's release-independent BIOS upgrade script. This is a firmware flash, needs a reload, and must be applied per-device; you cannot fix it with config. Pair it with physical access control on the console ports.","references":["https://sec.cloudapps.cisco.com/security/center/content/CiscoSecurityAdvisory/cisco-sa-nxos-image-sig-bypas-pQDRQvjL","https://nvd.nist.gov/vuln/detail/CVE-2024-20397"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-12-04"},{"id":"CVE-2024-33660","cve":"CVE-2024-33660","aliases":["AMI-SA-2024004"],"title":"AMI AptioV UEFI BIOS (SPI flash integrity verification): An actor with physical access can modify the SPI flash","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI AptioV UEFI BIOS (SPI flash integrity verification)","year":"2024","cvss_score":5.2,"severity":"medium","kev":false,"impact":"An actor with physical access can modify the SPI flash without the modification being detected. The score is modest because it needs hands on the hardware, but the operator consequence is that your firmware integrity story has no floor: a node that passed through an untrusted physical environment can carry an undetectable implant. That matters concretely for GPU fleets in leased colo where remote hands are third-party staff, for hardware shipped internationally, for anything bought on the secondary market during a supply crunch, and for RMA units returning from a vendor depot.","attack_vector":"Physical access to the machine, no credentials needed. Anyone who can open the chassis and reach the SPI flash - datacenter remote-hands staff, shipping and logistics handling, a vendor's repair depot, or a hosting provider's own technicians.","remediation":"BIOS update to BKC_5.37 or later - firmware flash plus reboot per node, vendor-rebase-gated. Because the threat model is physical rather than network, patching is only part of it: the process controls are what actually help. Take a firmware measurement baseline per node at commissioning, re-measure after any physical service event or RMA return, use chassis intrusion detection and seal logging, and treat any node that came back from third-party hands without a verified measurement as needing a reflash before it rejoins the pool.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/2024/AMI-SA-2024004.pdf","https://nvd.nist.gov/vuln/detail/CVE-2024-33660"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-11-12"},{"cvss_vector":"CVSS:3.1/AV:P/AC:L/PR:L/UI:N/S:C/C:H/I:N/A:N","cwe":["CWE-501"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-24153","cve":"CVE-2026-24153","aliases":[],"title":"NVIDIA Jetson Linux (initrd, nvluks trusted application): The nvluks trusted application is left enabled after initrd","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Jetson Linux (initrd, nvluks trusted application)","year":"2026","cvss_score":5.2,"severity":"medium","kev":false,"impact":"The nvluks trusted application is left enabled after initrd has finished with it, so the component that unwraps disk-encryption keys stays reachable once the system is up. An attacker with the device in hand and a low-privileged account can ask it to do its job and recover the contents of the encrypted rootfs. In practical terms full-disk encryption stops protecting a Jetson device that leaves your custody, which is the one scenario it was deployed for.","attack_vector":"Physical possession of the device plus a low-privileged local foothold. Not reachable over the network, so this is a lost, stolen, seized or RMA'd hardware problem rather than a remote fleet problem - which makes it a real concern for anything deployed outside a controlled facility.","remediation":"Update Jetson Linux to 35.6.4, 36.5 or 38.4 depending on branch, and reboot for the new initrd to take effect. Any device that has been outside your physical control while running an affected version should be treated as key-compromised: re-key the LUKS volumes and rotate every secret the device held.","references":["https://github.com/NVIDIA/product-security/tree/main/2026/5797","https://nvd.nist.gov/vuln/detail/CVE-2026-24153"],"status":"curated"},{"id":"CVE-2021-47389","cve":"CVE-2021-47389","aliases":[],"title":"Linux KVM/SVM - missing sev_decommission in sev_receive_start: KVM failed to DECOMMISSION the current SEV context","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux KVM/SVM - missing sev_decommission in sev_receive_start","year":"2021","cvss_score":5.1,"severity":"medium","kev":false,"impact":"KVM failed to DECOMMISSION the current SEV context when binding an ASID fails after RECEIVE_START. The firmware-side SEV context is left allocated with nothing owning it, exhausting the limited pool of SEV contexts the platform supports. Repeat the failure enough times and the host can no longer launch confidential VMs at all - a resource-exhaustion denial of service against your confidential-computing capacity, reachable through the guest-import path.","attack_vector":"Through the SEV guest receive/import path - so a tenant or control-plane action that fails repeatedly, deliberately or otherwise.","remediation":"Fixed in the Linux kernel. Distro kernel update plus a host reboot. Note that recovering exhausted SEV contexts on an unpatched host generally means an SNP platform shutdown/init cycle, which requires draining every confidential guest anyway.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-47389"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-21"},{"id":"CVE-2022-3172","cve":"CVE-2022-3172","aliases":[],"title":"Kubernetes (kube-apiserver): Aggregated API server can redirect apiserver clients","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-apiserver)","year":"2022","cvss_score":5.1,"severity":"medium","kev":false,"impact":"Aggregated API server can redirect apiserver clients; SSRF from the control plane","attack_vector":"Whoever controls an aggregated APIService backend","remediation":"Rolling control-plane upgrade; audit registered APIServices","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated","published":"2023-11-03"},{"id":"CVE-2023-40551","cve":"CVE-2023-40551","aliases":["shim 15.8 batch"],"title":"shim (MZ/PE header parser): Out-of-bounds read parsing MZ binaries","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"shim (MZ/PE header parser)","year":"2023","cvss_score":5.1,"severity":"medium","kev":false,"impact":"Out-of-bounds read parsing MZ binaries. Leaks memory contents from the pre-boot environment, which can include key material the firmware has not yet cleared.","attack_vector":"Crafted MZ binary loaded via shim.","remediation":"shim package update + reboot. Bundle with the rest of the shim 15.8 batch.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-40551","https://www.openwall.com/lists/oss-security/2024/01/26/1"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-01-29"},{"id":"CVE-2024-47973","cve":"CVE-2024-47973","aliases":["Solidigm SA-000563","over-provisioning data disclosure"],"title":"Solidigm DC SSDs (D3-S4510/S4520/S4610/S4620, D5-P5316, D7-P5520/P5620, DC S4500/S4600)","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Solidigm DC SSDs (D3-S4510/S4520/S4610/S4620, D5-P5316, D7-P5520/P5620, DC S4500/S4600) - over-provisioned NAND not…","year":"2024","cvss_score":5.1,"severity":"medium","kev":false,"impact":"A defect in how the drive handles its over-provisioned capacity leaks data to an attacker. Over-provisioned blocks are the spare NAND the FTL keeps outside the host-visible LBA range - they hold real tenant data that was written and then remapped, and they are invisible to every host-side wipe. BREAKS TENANT HANDOFF: dd-over-the-whole-device, blkdiscard, mkfs and any 'overwrite the visible address space' routine cannot reach these blocks by construction, so a previous tenant's data survives a reclaim that looks completely thorough from the host. This is the concrete, CVE'd instance of the general wear-levelling problem operators are usually told about only in the abstract.","attack_vector":"A tenant with local/root access on the bare-metal host after reclaim, reading back data the previous tenant wrote. Low privilege required on the host; no physical access needed.","remediation":"Firmware flash, per SKU, drive offline: ACV10340 for D5-P5316, XCV10151/XC311151 for D3-S4510/S4610, 7CV10111 for D3-S4520/S4620, 9CV10410 for D7-P5520/P5620, YCV10200 for D5-P5530, applied with Solidigm Storage Tool. Note the sting in the advisory: for DC S4500 and DC S4600 Solidigm states it has NO plans to ship a fix unless a customer explicitly requests one - on those SKUs this is effectively UNPATCHABLE, and the only safe reclaim is physical destruction or accepting that host-side wipes do not cover OP blocks. Because no host-side tool can reach over-provisioned NAND, the durable control is again encryption you own: LUKS/dm-crypt per tenant with the key in your KMS, so remapped ciphertext in OP blocks is worthless after you delete the key. Flashing a 10,000-drive fleet is a rolling drain across weeks with per-SKU tooling and per-SKU firmware images; expect the inventory step (working out which of five Solidigm families each node actually has) to take as long as the flashing.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-47973","https://www.solidigm.com/support-page/support-security.html","https://www.solidigm.com/content/dam/solidigm/en/site/support/support-community/cve-(security)/documents/public-security-advisory-v2.pdf"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"published":"2024-10-07"},{"id":"CVE-2024-7881","cve":"CVE-2024-7881","aliases":["TFV-13","data memory-dependent prefetch leak"],"title":"Arm Neoverse V2 / V3 / V3AE, Cortex-X3 / X4 / X925, C1-series","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arm Neoverse V2 / V3 / V3AE, Cortex-X3 / X4 / X925, C1-series; mitigated in Trusted Firmware-A v2.2-v2.12 and LTS…","year":"2024","cvss_score":5.1,"severity":"medium","kev":false,"impact":"Unprivileged code can steer a data-memory-dependent prefetcher into loading a privileged address and then dereferencing its contents, turning the prefetcher into a read oracle for kernel, hypervisor or secure-world memory. On a shared training or inference node this is a slow but real cross-boundary leak - the kind of thing that gets you key material and pointers for a follow-on exploit rather than an instant escape. Neoverse V2 is the Grace core, so GH200 and GB200 head nodes are in scope.","attack_vector":"Any unprivileged process on the host, or unprivileged code inside a guest. No devices, no network, no physical access. A container tenant on a shared Arm node is enough.","remediation":"EL3 firmware sets CPUACTLR6_EL1[41]=1 (or IMP_CPUECTLR_EL1[49]=1 on C1-Pro) to disable the offending prefetcher, and exposes it via SMCCC_ARCH_WORKAROUND_4. It is enabled by default in fixed TF-A on vulnerable cores, so the operator task is 'take the OEM firmware build and flash it', with the usual reboot and drain. Disabling a prefetcher is not free - expect single-digit percent regression on pointer-chasing workloads. There is no software-only mitigation you can apply without the firmware update.","references":["https://trustedfirmware-a.readthedocs.io/en/latest/security_advisories/security-advisory-tfv-13.html","https://nvd.nist.gov/vuln/detail/CVE-2024-7881"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-01-28"},{"cvss_vector":"CVSS:4.0/AV:N/AC:L/AT:N/PR:H/UI:N/VC:L/VI:N/VA:N/SC:N/SI:N/SA:N","cwe":["CWE-532"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-2327","cve":"CVE-2025-2327","aliases":[],"title":"Pure Storage FlashArray key rotation logging (Rapid Data Locking): The Key Encryption Key is written to logs during","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Pure Storage FlashArray key rotation logging (Rapid Data Locking)","year":"2025","cvss_score":5.1,"severity":"medium","kev":false,"impact":"The Key Encryption Key is written to logs during rotation when Rapid Data Locking is configured. Anyone who can read those logs - including whoever receives a support bundle - holds the key that protects data at rest on the array.","attack_vector":"Read access to FlashArray logs, or possession of a support bundle collected from an affected array with RDL enabled.","remediation":"Upgrade Purity//FA to the fixed release, then rotate the KEK again on a patched version and destroy or recall any log set or support bundle collected while the flaw was present.","references":["https://support.purestorage.com/category/m_pure_storage_product_security","https://nvd.nist.gov/vuln/detail/CVE-2025-2327"],"status":"curated"},{"id":"CVE-2025-54388","cve":"CVE-2025-54388","aliases":[],"title":"Docker / moby: On firewalld reload, published container ports become reachable from outside despite the intended","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2025","cvss_score":5.1,"severity":"medium","kev":false,"impact":"On firewalld reload, published container ports become reachable from outside despite the intended restriction","attack_vector":"Unauthenticated network","remediation":"Upgrade Docker Engine; verify iptables/nftables rules after any firewalld reload","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-54388"],"status":"curated","published":"2025-07-30"},{"id":"CVE-2019-14891","cve":"CVE-2019-14891","aliases":[],"title":"CRI-O: All pod processes share one memory cgroup, so a workload OOM kills conmon and destabilises the node","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"CRI-O","year":"2019","cvss_score":5,"severity":"medium","kev":false,"impact":"All pod processes share one memory cgroup, so a workload OOM kills conmon and destabilises the node","attack_vector":"Any tenant workload","remediation":"Upgrade CRI-O; drain node","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-14891"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2019-11-25"},{"id":"CVE-2021-32760","cve":"CVE-2021-32760","aliases":[],"title":"containerd: Crafted image can change Unix file permissions of existing host files during extraction","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2021","cvss_score":5,"severity":"medium","kev":false,"impact":"Crafted image can change Unix file permissions of existing host files during extraction","attack_vector":"Malicious image","remediation":"Rolling containerd upgrade with node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-32760"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2021-07-19"},{"id":"CVE-2022-21701","cve":"CVE-2022-21701","aliases":[],"title":"Istio: A user with CREATE on Gateway API resources escalates privilege in istiod","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Istio","year":"2022","cvss_score":5,"severity":"medium","kev":false,"impact":"A user with CREATE on Gateway API resources escalates privilege in istiod","attack_vector":"Cluster user with namespace access and Gateway API rights","remediation":"Rolling istiod upgrade; restrict Gateway creation","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-21701"],"status":"curated","published":"2022-01-19"},{"id":"CVE-2022-42292","cve":"CVE-2022-42292","aliases":[],"title":"GeForce Experience installer: Local privesc via symlink following","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GeForce Experience installer","year":"2022","cvss_score":5,"severity":"medium","kev":false,"impact":"Local privesc via symlink following","attack_vector":"Local user","remediation":"Consumer-only; no DC action","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42292","https://github.com/NVIDIA/product-security/tree/main/2023/5384"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:R/S:U/C:N/I:L/A:H","cwe":["CWE-59"],"published":"2023-02-12"},{"cwe":["CWE-401","CWE-696"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-48904","cve":"CVE-2022-48904","aliases":[],"title":"Linux kernel (drivers/iommu/amd): AMD-Vi updated the domain's I/O page-table mode before running the code that frees","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/amd)","year":"2022","cvss_score":5,"severity":"medium","kev":false,"impact":"AMD-Vi updated the domain's I/O page-table mode before running the code that frees the old page table, so the teardown walked with the wrong mode and the page-table memory was never released. Upstream observed it directly when launching VMs with passthrough devices - so on a node doing normal tenant VM churn, host kernel memory bleeds away on every passthrough VM start until the node runs out and takes every tenant on it down with it.","attack_vector":"Triggered by the routine act of attaching a device to a domain and changing its page-table mode, which is what happens each time a passthrough VM is launched on an AMD-Vi host. The tenant drives the rate through ordinary start/stop of its own instances via the control plane; no host root is needed to cause the churn, though the VMM performs the attach. Conditional on AMD-Vi and PCI passthrough being in use.","remediation":"No fixed release is listed in this record; apply the linked stable commits or run a current stable/LTS kernel on AMD passthrough nodes. Interim: monitor host slab/page usage across passthrough VM churn and drain nodes that trend upward, and avoid rapid create/destroy cycles of passthrough instances until patched.","references":["https://git.kernel.org/stable/c/378e2fe1eb58d5c2ed55c8fe5e11f9db5033cdd6","https://git.kernel.org/stable/c/c78627f757e37c2cf386b59c700c4e1574988597","https://nvd.nist.gov/vuln/detail/CVE-2022-48904"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2023-0203","cve":"CVE-2023-0203","aliases":[],"title":"ConnectX-5/6/6-DX NIC firmware: NIC DoS (insufficient access-control granularity)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"ConnectX-5/6/6-DX NIC firmware","year":"2023","cvss_score":5,"severity":"medium","kev":false,"impact":"NIC DoS (insufficient access-control granularity)","attack_vector":"Any unprivileged tenant with a VF","remediation":"Flash NIC firmware 35.1012+; node reboot","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5459/5459.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:C/C:N/I:N/A:L","cwe":["CWE-1220"],"fleet":{"pain_class":"node-reboot"},"published":"2023-04-22"},{"id":"CVE-2023-0205","cve":"CVE-2023-0205","aliases":[],"title":"ConnectX-5/6/6-DX NIC firmware: NIC DoS","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"ConnectX-5/6/6-DX NIC firmware","year":"2023","cvss_score":5,"severity":"medium","kev":false,"impact":"NIC DoS","attack_vector":"Any unprivileged tenant with a VF","remediation":"Flash NIC firmware 35.1012+; node reboot","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5459/5459.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:C/C:N/I:N/A:L","cwe":["CWE-1220"],"fleet":{"pain_class":"node-reboot"},"published":"2023-04-22"},{"id":"CVE-2023-0264","cve":"CVE-2023-0264","aliases":[],"title":"Keycloak: OIDC authentication flaw - attacker reusing data from a same-realm request impersonates a user","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Keycloak","year":"2023","cvss_score":5,"severity":"medium","kev":false,"impact":"OIDC authentication flaw - attacker reusing data from a same-realm request impersonates a user and mints session tokens","attack_vector":"Network (remote)","remediation":"Control-plane: URGENT - Keycloak fronts the tenant portal; upgrade + invalidate all sessions","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0264"],"status":"curated","published":"2023-08-04"},{"id":"CVE-2023-25809","cve":"CVE-2023-25809","aliases":[],"title":"runc: Rootless runc leaves /sys/fs/cgroup writable inside the container","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"runc","year":"2023","cvss_score":5,"severity":"medium","kev":false,"impact":"Rootless runc leaves /sys/fs/cgroup writable inside the container","attack_vector":"Any tenant workload under rootless runc","remediation":"Replace runc binary; drain node","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25809"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2023-03-29"},{"cwe":["CWE-362","CWE-416"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-53501","cve":"CVE-2023-53501","aliases":[],"title":"Linux kernel (drivers/iommu/amd): Unbinding a PASID races the I/O page-fault (PPR) notifications still in flight","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/amd)","year":"2023","cvss_score":5,"severity":"medium","kev":false,"impact":"Unbinding a PASID races the I/O page-fault (PPR) notifications still in flight against it. The refcount that is supposed to keep the PASID state alive until outstanding faults drain hits zero on the unbind path, so the state object can be torn down while the fault handler is still using it - a PASID/SVA lifetime break on AMD-Vi. Upstream's observable is a refcount warning plus leaked state rather than a demonstrated use-after-free.","attack_vector":"A process using AMD SVA/PASID that unbinds while its device still has page-fault requests outstanding - i.e. a tenant workload that programs a PASID-capable accelerator and exits with DMA still pending. No host root. Conditional on AMD-Vi with the iommu_v2 PASID/PPR path in use.","remediation":"The record lists no fixed release; boot a kernel carrying the stable fix commits below. Interim control: do not enable SVA/PASID for tenant workloads on unpatched AMD-Vi hosts.","references":["https://git.kernel.org/stable/c/a50d60b8f2aff46dd7c7edb4a5835cdc4d432c22","https://git.kernel.org/stable/c/13ed255248dfbbb7f23f9170c7a537fb9ca22c73","https://nvd.nist.gov/vuln/detail/CVE-2023-53501"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-200","CWE-908"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-54034","cve":"CVE-2023-54034","aliases":[],"title":"Linux kernel (drivers/iommu/iommufd): The vfio type1 info structure is not zeroed before being filled and copied out","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/iommufd)","year":"2023","cvss_score":5,"severity":"medium","kev":false,"impact":"The vfio type1 info structure is not zeroed before being filled and copied out, and the copy-in covers fewer bytes than the struct, so padding bytes of kernel stack are handed to the caller. Kernel memory disclosure to a tenant through iommufd's vfio compatibility ioctl.","attack_vector":"A tenant or VMM holding /dev/iommu, or a vfio container backed by iommufd, calling the type1 GET_INFO ioctl through the compat layer. One ioctl, no race, no host root.","remediation":"The record lists no fixed release; boot a kernel carrying the stable fix commits below. Interim control: keep /dev/iommu out of tenant containers.","references":["https://git.kernel.org/stable/c/7adcec686e4d699c169d34c722132b2bce5232cb","https://git.kernel.org/stable/c/b3551ead616318ea155558cdbe7e91495b8d9b33","https://nvd.nist.gov/vuln/detail/CVE-2023-54034"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-200","CWE-908"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-54137","cve":"CVE-2023-54137","aliases":[],"title":"Linux kernel (drivers/vfio): Uninitialized kernel stack bytes sitting in a structure hole are copied out to userspace","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/vfio)","year":"2023","cvss_score":5,"severity":"medium","kev":false,"impact":"Uninitialized kernel stack bytes sitting in a structure hole are copied out to userspace through the VFIO container's info ioctl. A tenant reads back kernel stack contents it was never meant to see - small on its own, but exactly the kind of leak used to infer layout or recover a pointer before a heavier bug is fired.","attack_vector":"Any tenant holding /dev/vfio/vfio calling VFIO_IOMMU_GET_INFO on its own container. One ioctl, no race, no host root, no hardware precondition beyond the legacy type1 container being in use.","remediation":"The record lists no fixed release; boot a kernel carrying the stable fix commits below. Interim control: drop /dev/vfio from containers that do not need passthrough - there is no way to filter a single capability out of the info ioctl.","references":["https://git.kernel.org/stable/c/ad83d83dd891244de0d07678b257dc976db7c132","https://git.kernel.org/stable/c/13fd667db999bffb557c5de7adb3c14f1713dd51","https://nvd.nist.gov/vuln/detail/CVE-2023-54137"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-42934","cve":"CVE-2024-42934","aliases":[],"title":"OpenIPMI before 2.0.36: Where this bites an operator is in test and CI infrastructure rather than production nodes","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"OpenIPMI before 2.0.36","year":"2024","cvss_score":5,"severity":"medium","kev":false,"impact":"Where this bites an operator is in test and CI infrastructure rather than production nodes: ipmi_sim is what teams use to emulate BMCs when developing provisioning automation and firmware tooling. An attacker who can reach a simulator instance crashes it or, with luck, bypasses its authentication - and a compromised CI runner that builds and signs your provisioning images is a supply-chain foothold into the real fleet. The low probability of code execution is the honest read; the availability impact on a build pipeline is the likely one. An out-of-bounds array access on the authentication type in the ipmi_sim BMC simulator, with denial of service the likely outcome and authentication bypass or code execution possible at low probability.","attack_vector":"Network access to a running ipmi_sim instance. In most environments that is a developer workstation or a CI runner rather than the production management VLAN, but CI runners are frequently more reachable than people assume.","remediation":"Package update to OpenIPMI 2.0.36 or later on any host running ipmi_sim - a distribution package update, no firmware and no reboot. The broader hygiene point: BMC simulators used in CI should not be reachable from anything but the test harness, and should not run on a host that also holds production BMC credentials or signing keys.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-42934","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2024/42xxx/CVE-2024-42934.json"],"status":"curated"},{"id":"CVE-2024-48936","cve":"CVE-2024-48936","aliases":[],"title":"Slurm: Authentication-handling mistake in stepmgr lets an attacker execute processes under other users' jobs","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Slurm","year":"2024","cvss_score":5,"severity":"medium","kev":false,"impact":"Authentication-handling mistake in stepmgr lets an attacker execute processes under other users' jobs; cross-tenant job compromise on a shared GPU cluster","attack_vector":"Any user who can submit a job","remediation":"Upgrade Slurm to 24.05.4+; disable --stepmgr if not needed","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-48936"],"status":"curated","published":"2024-10-28"},{"cwe":["CWE-682","CWE-190"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-21724","cve":"CVE-2025-21724","aliases":[],"title":"Linux kernel (drivers/iommu/iommufd): The iommufd dirty-tracking bitmap computed an index by shifting a 32-bit constant","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/iommufd)","year":"2025","cvss_score":5,"severity":"medium","kev":false,"impact":"The iommufd dirty-tracking bitmap computed an index by shifting a 32-bit constant by a user-supplied page shift, so a large shift produces undefined behaviour and a garbage index into the bitmap that tracks which IOVAs a passthrough device has written. Garbage indexing into that structure is the wrong kind of wrong for a mapping-tracking data structure - treat it as untrusted input reaching IOVA bookkeeping.","attack_vector":"A process holding /dev/iommu supplies an out-of-range page shift (upstream cites 63) on the dirty-tracking bitmap path. Plain ioctl input validation on the fd a passthrough tenant already holds; no host root, no device required. Conditional on iommufd being in use on the node. Same input surface as CVE-2025-40293, which turns the same overflow into a divide-by-zero.","remediation":"No fixed release is listed in this record; apply the linked stable commits or run a current stable/LTS kernel, and take it together with the CVE-2025-40293 fix since they harden the same path. Interim: keep /dev/iommu out of containers that do not perform passthrough and validate page-size arguments in the VMM layer.","references":["https://git.kernel.org/stable/c/44d9c94b7a3f29a3e07c4753603a35e9b28842a3","https://git.kernel.org/stable/c/38ac76fc06bc6826a3e4b12a98efbe98432380a9","https://nvd.nist.gov/vuln/detail/CVE-2025-21724"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-23260","cve":"CVE-2025-23260","aliases":[],"title":"AIS Operator (AIStore): Improper access control on cluster storage","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"AIS Operator (AIStore)","year":"2025","cvss_score":5,"severity":"medium","kev":false,"impact":"Improper access control on cluster storage","attack_vector":"Tenant with cluster network access","remediation":"Upgrade the operator Helm chart","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23260","https://github.com/NVIDIA/product-security/tree/main/2025/5660"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:C/C:L/I:N/A:N","cwe":["CWE-266"],"published":"2025-06-24"},{"id":"CVE-2025-23332","cve":"CVE-2025-23332","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): A null-pointer dereference in a Linux driver kernel","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2025","cvss_score":5,"severity":"medium","kev":false,"impact":"A null-pointer dereference in a Linux driver kernel module, reachable locally, panics the node. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5703. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23332","https://github.com/NVIDIA/product-security/tree/main/2025/5703"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-drain"},"published":"2025-10-23"},{"id":"CVE-2026-41413","cve":"CVE-2026-41413","aliases":[],"title":"Istio: A RequestAuthentication jwksUri pointed at an internal service makes istiod issue an unauthenticated request","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Istio","year":"2026","cvss_score":5,"severity":"medium","kev":false,"impact":"A RequestAuthentication jwksUri pointed at an internal service makes istiod issue an unauthenticated request; SSRF from the control plane","attack_vector":"Cluster user with namespace access","remediation":"Rolling istiod upgrade to 1.28.6/1.29.2+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-41413"],"status":"curated","published":"2026-05-07"},{"cwe":["CWE-416","CWE-672"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-64473","cve":"CVE-2026-64473","aliases":[],"title":"Linux kernel (drivers/vfio): Vfio deleted the device before removing its debugfs tree, so debugfs files stay visible","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/vfio)","year":"2026","cvss_score":5,"severity":"medium","kev":false,"impact":"Vfio deleted the device before removing its debugfs tree, so debugfs files stay visible while the devres-allocated state behind them has already been released. Anything that opens those files during the unregister window - which lasts as long as userspace still holds references to the device - reads through a stale inode private pointer into freed memory.","attack_vector":"Needs host root or whatever principal can read /sys/kernel/debug/vfio, and needs the read to land inside the unregister window of a device that a tenant is still holding open. Not tenant-reachable in a normal container (debugfs is not mounted there); the realistic trigger is host monitoring or debugging tooling that walks vfio debugfs while devices are being torn down.","remediation":"Update to a stable kernel carrying commits 6cc60b41 / a53109ff. Interim: do not mount debugfs on production tenant nodes, or keep monitoring agents off /sys/kernel/debug/vfio while devices are unbinding.","references":["https://git.kernel.org/stable/c/6cc60b41d61657dc469893d14e8e55d160056ff1","https://git.kernel.org/stable/c/a53109ffb6b5148e11a27fb7670355b92db12dd3","https://nvd.nist.gov/vuln/detail/CVE-2026-64473"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2018-9279","cve":"CVE-2018-9279","aliases":[],"title":"Eaton UPS 9PX 8000 SP web interface: The device's own web page contains the user password in cleartext in the page","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Eaton UPS 9PX 8000 SP web interface","year":"2018","cvss_score":4.9,"severity":"medium","kev":false,"impact":"The device's own web page contains the user password in cleartext in the page source. Anyone who gets a single authenticated view - or a screenshot, or a saved page in a support ticket - has the credential. The companion issue CVE-2018-9280 does the same for the SNMPv3 read and write user passwords, which is worse, because the SNMP write community is a control channel.","attack_vector":"Anyone who can load the UPS web interface, or who obtains a saved copy of the page.","remediation":"Firmware update where available. Rotate the UPS and SNMPv3 credentials, and specifically rotate the SNMP write credential - and then ask whether SNMP write needs to be enabled at all, because on most UPS deployments it does not.","references":["https://www.bishopfox.com/news/2018/10/eaton-ups-9px-8000-sp-multiple-vulnerabilities/"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2018-10-24"},{"id":"CVE-2019-11245","cve":"CVE-2019-11245","aliases":[],"title":"Kubernetes (kubelet): Container restart runs as uid 0 despite mustRunAsNonRoot","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kubelet)","year":"2019","cvss_score":4.9,"severity":"medium","kev":false,"impact":"Container restart runs as uid 0 despite mustRunAsNonRoot","attack_vector":"Any tenant workload","remediation":"Rolling kubelet upgrade with node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-11245"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2019-08-29"},{"id":"CVE-2020-11484","cve":"CVE-2020-11484","aliases":[],"title":"NVIDIA DGX BMC (AMI firmware): An administrative BMC user can pull the hash of the BMC/IPMI user password","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVIDIA DGX BMC (AMI firmware)","year":"2020","cvss_score":4.9,"severity":"medium","kev":false,"impact":"An administrative BMC user can pull the hash of the BMC/IPMI user password. In practice that means one compromised BMC admin session yields offline-crackable credentials that are frequently reused across an entire DGX fleet - a lateral movement multiplier across every node in the rack. DGX-1 before BMC 3.38.30.","attack_vector":"An attacker who already has administrative access to one BMC, including via the hard-coded credentials in the same advisory.","remediation":"Flash the DGX BMC firmware from NVIDIA's DGX firmware update container (DGX-1 to 3.38.30 or later, DGX-2 to 1.06.06 or later; DGX A100 per the bulletin's table). A BMC flash does not require the host OS to reboot but drops out-of-band management for several minutes and NVIDIA recommends a host power cycle afterwards, so treat it as a per-node maintenance window. Rotate every BMC and IPMI credential after the flash - flashing does not invalidate secrets an attacker already pulled. Keep BMCs on an isolated management VLAN with no route from tenant or job networks.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-11484"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2020-10-29"},{"id":"CVE-2021-44769","cve":"CVE-2021-44769","aliases":["AMI-SA-2022001","Nozomi Labs BMC firmware research"],"title":"AMI MegaRAC SPx 12 / SPx 13 (BMC TLS certificate generation): Malformed input to the BMC's certificate-generation","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx 12 / SPx 13 (BMC TLS certificate generation)","year":"2021","cvss_score":4.9,"severity":"medium","kev":false,"impact":"Malformed input to the BMC's certificate-generation function permanently wedges the controller, and the only documented recovery is a factory reset. On a GPU fleet that is worse than a normal DoS: you lose out-of-band power control, console and virtual media on the affected nodes, and getting them back means an on-site factory reset that also wipes your BMC configuration - accounts, certificates, network settings, alert destinations - so every affected node has to be re-provisioned by hand. A malicious insider or a compromised management account can do this across a rack in minutes and cost you days of remote-hands work.","attack_vector":"Network access to the BMC with high privileges - an administrative BMC account. Because BMC admin credentials are so often cloned across a fleet by the provisioning system, one leaked credential scales this to every node that shares it.","remediation":"Firmware flash to SPx_12-update-5.00 / SPx_13-update-3.00 or later, out-of-band per node, ODM-gated. Config-only risk reduction: unique per-node BMC admin credentials so one leak cannot sweep the fleet, and - critically - keep an exported, version-controlled copy of every BMC's configuration so that if you do have to factory-reset, restoring is automated rather than a per-node manual rebuild.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2022001.pdf","https://www.nozominetworks.com/labs/vulnerability-advisories/cve-2021-44769/","https://nvd.nist.gov/vuln/detail/CVE-2021-44769"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-10-24"},{"id":"CVE-2022-21417","cve":"CVE-2022-21417","aliases":[],"title":"MySQL Server: InnoDB flaw allowing a high-privileged network attacker to cause a repeatable DoS","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"MySQL Server","year":"2022","cvss_score":4.9,"severity":"medium","kev":false,"impact":"InnoDB flaw allowing a high-privileged network attacker to cause a repeatable DoS","attack_vector":"Network (remote)","remediation":"Control-plane: quarterly CPU patch","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-21417"],"status":"curated","published":"2022-04-19"},{"id":"CVE-2022-22488","cve":"CVE-2022-22488","aliases":["IBM X-Force 226337"],"title":"IBM OpenBMC OP910 / OP940 certificate handling (phosphor-certificate-manager lineage): A privileged BMC user","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"IBM OpenBMC OP910 / OP940 certificate handling (phosphor-certificate-manager lineage)","year":"2022","cvss_score":4.9,"severity":"medium","kev":false,"impact":"A privileged BMC user who uploads or deletes CA certificates rapidly enough takes the BMC down. Low severity because it needs an admin account, but the fleet-relevant version is not malice: it is your own certificate-rotation automation. An operator scripting CA distribution across a few hundred BMCs can trip this and take out out-of-band management fleet-wide during what was supposed to be a routine hygiene job.","attack_vector":"Authenticated BMC administrator over the network - including your own automation holding admin credentials.","remediation":"Fixed in later OP910/OP940 firmware; per-node system firmware update with a maintenance window, and not worth a dedicated campaign at this severity. The operational fix is free: rate-limit and serialize certificate operations in your BMC automation, and stagger fleet-wide certificate pushes rather than fanning out at full concurrency.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-22488","https://www.ibm.com/support/pages/node/6840155","https://exchange.xforce.ibmcloud.com/vulnerabilities/226337"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-12-12"},{"id":"CVE-2023-22084","cve":"CVE-2023-22084","aliases":[],"title":"MySQL Server: InnoDB flaw - a high-privileged network attacker can hang or repeatedly crash the server","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"MySQL Server","year":"2023","cvss_score":4.9,"severity":"medium","kev":false,"impact":"InnoDB flaw - a high-privileged network attacker can hang or repeatedly crash the server","attack_vector":"Network (remote)","remediation":"Control-plane: quarterly CPU patch on managed MySQL","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-22084"],"status":"curated","published":"2023-10-17"},{"id":"CVE-2023-29153","cve":"CVE-2023-29153","aliases":[],"title":"Intel SPS firmware: Uncontrolled resource consumption in SPS firmware lets a privileged user deny service","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SPS firmware","year":"2023","cvss_score":4.9,"severity":"medium","kev":false,"impact":"Uncontrolled resource consumption in SPS firmware lets a privileged user deny service to the management engine over the network. Losing the management engine means losing out-of-band recovery on that node - which is exactly when you need it.","attack_vector":"Privileged user with network access to the affected interface.","remediation":"Fixed in Intel CSME/SPS firmware, which reaches you as an OEM BIOS or firmware package - not as a microcode or OS update. That means: wait for your server vendor to ship it, drain the node, flash, and reboot. OEM availability is the long pole and routinely lags the Intel advisory by one or more quarters on server platforms. Track it per platform SKU, because vendors ship these unevenly across their own product lines.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-29153","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01003.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-02-14"},{"id":"CVE-2023-46118","cve":"CVE-2023-46118","aliases":[],"title":"RabbitMQ: HTTP API enforces no request body limit","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"RabbitMQ","year":"2023","cvss_score":4.9,"severity":"medium","kev":false,"impact":"HTTP API enforces no request body limit -> authenticated user exhausts node memory (DoS)","attack_vector":"Network (remote)","remediation":"Control-plane: broker upgrade; set max_message_size","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-46118"],"status":"curated","published":"2023-10-25"},{"id":"CVE-2024-0116","cve":"CVE-2024-0116","aliases":[],"title":"NVIDIA Triton Inference Server: An out-of-bounds read triggered by releasing a shared memory region while it is still","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2024","cvss_score":4.9,"severity":"medium","kev":false,"impact":"An out-of-bounds read triggered by releasing a shared memory region while it is still in use crashes the server. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5565. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0116","https://github.com/NVIDIA/product-security/tree/main/2024/5565"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:H/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-125"],"published":"2024-10-01"},{"id":"CVE-2024-23444","cve":"CVE-2024-23444","aliases":[],"title":"Elasticsearch: elasticsearch-certutil --csr writes the private key to disk unencrypted despite --pass","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Elasticsearch","year":"2024","cvss_score":4.9,"severity":"medium","kev":false,"impact":"elasticsearch-certutil --csr writes the private key to disk unencrypted despite --pass","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade + reissue any cert whose key was generated this way","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-23444"],"status":"curated","published":"2024-07-31"},{"id":"CVE-2024-53880","cve":"CVE-2024-53880","aliases":[],"title":"Triton Inference Server: DoS (integer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2024","cvss_score":4.9,"severity":"medium","kev":false,"impact":"DoS (integer overflow)","attack_vector":"Client of the inference endpoint","remediation":"Upgrade Triton; redeploy serving images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53880","https://github.com/NVIDIA/product-security/tree/main/2025/5612"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:H/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-190"],"published":"2025-02-12"},{"id":"CVE-2025-26482","cve":"CVE-2025-26482","aliases":["DSA-2025-046"],"title":"Dell PowerEdge Server BIOS + iDRAC9 (information disclosure): Information disclosure spanning both the BIOS and iDRAC9","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell PowerEdge Server BIOS + iDRAC9 (information disclosure)","year":"2025","cvss_score":4.9,"severity":"medium","kev":false,"impact":"Information disclosure spanning both the BIOS and iDRAC9 on a very large PowerEdge model list - and that list explicitly includes the GPU platforms: XE9680, XE9680L, XE9640 and XE8640. Low severity in isolation; what makes it worth tracking is coverage. If you run Dell GPU nodes, this one almost certainly applies to your exact SKUs, and disclosed platform/firmware detail is the reconnaissance that makes a later BMC or BIOS exploit reliable rather than a guess.","attack_vector":"A high-privilege attacker with remote access - i.e. someone who already holds an administrative iDRAC credential. This is a post-compromise information-leak rather than an entry point, which is why the score is moderate.","remediation":"Two separate rollouts. The iDRAC9 side is an out-of-band firmware flash, per-node, no host reboot, no drain. The BIOS side needs a System BIOS update that applies only on the next reboot, so it costs a drain of running training jobs - for XE9680-class nodes that is a real scheduling problem, since those are the machines you least want to take down. Sequence the iDRAC flash immediately and batch the BIOS update into the next planned maintenance window. Per-model version floors are in the advisory (e.g. 2.5.4 for the R660/R760 family, 1.2.6 for the R470/R570/R670/R770 family).","references":["https://www.dell.com/support/kbdoc/en-us/000370138/dsa-2025-046-security-update-for-dell-poweredge-server-and-dell-idrac9-for-information-disclosure-vulnerability","https://nvd.nist.gov/vuln/detail/CVE-2025-26482"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2025-09-25"},{"id":"CVE-2019-11251","cve":"CVE-2019-11251","aliases":[],"title":"Kubernetes (kubectl): Double-symlink in tar output escapes the kubectl cp destination","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kubectl)","year":"2019","cvss_score":4.8,"severity":"medium","kev":false,"impact":"Double-symlink in tar output escapes the kubectl cp destination","attack_vector":"Malicious image","remediation":"Upgrade kubectl on operator and CI machines","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-11251"],"status":"curated","published":"2020-02-03"},{"id":"CVE-2019-6195","cve":"CVE-2019-6195","aliases":[],"title":"Lenovo XClarity Controller (XCC): Authorization bypass","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lenovo XClarity Controller (XCC)","year":"2019","cvss_score":4.8,"severity":"medium","kev":false,"impact":"Authorization bypass — a low-privilege authenticated user gains read access beyond their role","attack_vector":"Network / XCC web","remediation":"XCC firmware update; low CVSS but relevant where BMC access is delegated to tenants or remote-hands staff","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-6195"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2020-02-14"},{"id":"CVE-2020-15211","cve":"CVE-2020-15211","aliases":[],"title":"TensorFlow Lite (flatbuffer models): Out-of-bounds via duplicate tensor indices in flatbuffer models","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"TensorFlow Lite (flatbuffer models)","year":"2020","cvss_score":4.8,"severity":"medium","kev":false,"impact":"Out-of-bounds via duplicate tensor indices in flatbuffer models","attack_vector":"Customer-supplied TFLite model","remediation":"Patch; edge/CPU inference tiers only","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-15211"],"status":"curated","published":"2020-09-25"},{"id":"CVE-2022-3466","cve":"CVE-2022-3466","aliases":[],"title":"CRI-O: Shipped OpenShift CRI-O builds regressed the CVE-2022-2995 fix","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"CRI-O","year":"2022","cvss_score":4.8,"severity":"medium","kev":false,"impact":"Shipped OpenShift CRI-O builds regressed the CVE-2022-2995 fix; supply-chain style reintroduction of a fixed bug","attack_vector":"Any tenant workload on an affected OCP build","remediation":"Verify CRI-O build version explicitly, not just the OCP version; upgrade and drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-3466"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2023-09-15"},{"id":"CVE-2023-31339","cve":"CVE-2023-31339","aliases":[],"title":"ARM Trusted Firmware in AMD Zynq UltraScale+ MPSoC/RFSoC: Improper input validation in the ARM Trusted Firmware used","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ARM Trusted Firmware in AMD Zynq UltraScale+ MPSoC/RFSoC","year":"2023","cvss_score":4.8,"severity":"medium","kev":false,"impact":"Improper input validation in the ARM Trusted Firmware used on AMD's Zynq UltraScale+ parts allows out-of-bounds reads and data leakage. Relevant to datacenter operators through the side door: Zynq and Versal parts show up as SmartNIC, DPU, storage-controller and management-plane silicon inside servers, so this is firmware running on your network path rather than on your compute path.","attack_vector":"Local to the device, requires privileged access to the ATF interface on the Zynq part.","remediation":"Fixed in updated ARM Trusted Firmware from AMD/Xilinx. Delivery depends entirely on who integrated the part - a SmartNIC vendor, a storage OEM, your own board team - so tracing the update path is often harder than applying it. Requires a device firmware update and a reset of the affected card. Inventory which AMD/Xilinx adaptive SoCs are in your servers; most operators cannot answer that question, which is the real finding.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31339","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-08-13"},{"id":"CVE-2023-7258","cve":"CVE-2023-7258","aliases":[],"title":"gVisor: Reference-counting bug in mount-point tracking panics the sandbox","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"gVisor","year":"2023","cvss_score":4.8,"severity":"medium","kev":false,"impact":"Reference-counting bug in mount-point tracking panics the sandbox","attack_vector":"A tenant running as root inside the sandbox with mount permission","remediation":"Upgrade runsc; restart sandboxed pods","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-7258"],"status":"curated","published":"2024-05-15"},{"id":"CVE-2025-24513","cve":"CVE-2025-24513","aliases":[],"title":"ingress-nginx: auth-secret file path traversal in the controller","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2025","cvss_score":4.8,"severity":"medium","kev":false,"impact":"auth-secret file path traversal in the controller","attack_vector":"Cluster user who can create Ingress objects","remediation":"Rolling controller upgrade","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated","published":"2025-03-25"},{"id":"CVE-2025-29949","cve":"CVE-2025-29949","aliases":[],"title":"AMD Secure Processor bootloader - legacy recovery mode: Insufficient input sanitisation in the ASP bootloader's legacy","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor bootloader - legacy recovery mode","year":"2025","cvss_score":4.8,"severity":"medium","kev":false,"impact":"Insufficient input sanitisation in the ASP bootloader's legacy recovery path lets an attacker write out of bounds and corrupt Secure DRAM, bricking the boot flow. The outcome is denial of service at the firmware level - a node that will not come back up, which on a GPU fleet means an RMA-shaped hole rather than a reboot.","attack_vector":"Local, and only through legacy recovery mode - so it needs an attacker who can force the platform into recovery, which usually means firmware-level or physical access.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. Where the platform allows it, disabling legacy ASP recovery mode in SBIOS removes the reachable path without waiting for the OEM.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-29949","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2026-02-10"},{"id":"CVE-2026-24147","cve":"CVE-2026-24147","aliases":[],"title":"Triton Inference Server: Unauthorized file access via path traversal","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2026","cvss_score":4.8,"severity":"medium","kev":false,"impact":"Unauthorized file access via path traversal","attack_vector":"Authenticated inference client","remediation":"Upgrade Triton; redeploy serving images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24147","https://github.com/NVIDIA/product-security/tree/main/2026/5816"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:L/I:N/A:L","cwe":["CWE-22"],"published":"2026-04-07"},{"id":"CVE-2026-41174","cve":"CVE-2026-41174","aliases":[],"title":"Traefik: Cross-namespace isolation not enforced in the Kubernetes CRD provider","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Traefik","year":"2026","cvss_score":4.8,"severity":"medium","kev":false,"impact":"Cross-namespace isolation not enforced in the Kubernetes CRD provider","attack_vector":"Cluster user with namespace access","remediation":"Rolling Traefik upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-41174"],"status":"curated","published":"2026-04-30"},{"id":"CVE-2018-3626","cve":"CVE-2018-3626","aliases":[],"title":"Intel SGX SDK (Edger8r generated code, side channel): Edger8r generated bridge code that was susceptible to a side","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SGX SDK (Edger8r generated code, side channel)","year":"2018","cvss_score":4.7,"severity":"medium","kev":false,"impact":"Edger8r generated bridge code that was susceptible to a side channel, so a local user could recover information crossing the enclave boundary. Earliest member of the generated-bridge-code family.","attack_vector":"Local user interacting with a vulnerable enclave.","remediation":"Rebuild enclaves with SGX SDK 2.1.2 (Linux) / 1.9.6 (Windows) or later and re-attest.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-3626","https://security-center.intel.com/advisory.aspx?intelid=INTEL-SA-00117&languageid=en-fr"],"status":"curated","published":"2018-03-20"},{"id":"CVE-2020-5967","cve":"CVE-2020-5967","aliases":[],"title":"NVIDIA Linux GPU Display Driver, UVM driver (nvidia-uvm.ko): A race in the Unified Virtual Memory kernel module lets","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Linux GPU Display Driver, UVM driver (nvidia-uvm.ko)","year":"2020","cvss_score":4.7,"severity":"medium","kev":false,"impact":"A race in the Unified Virtual Memory kernel module lets a local process wedge or crash the driver. UVM is on the hot path for managed-memory CUDA workloads, so on a Linux training node this is a way for any job with GPU access to take out the whole host's GPUs. Ubuntu shipped it as a security update, so it is real on distro-packaged fleets.","attack_vector":"Any local user or container that has /dev/nvidia-uvm mapped in - which is every GPU container under the standard container toolkit configuration.","remediation":"Install the fixed Linux GPU Display Driver branch. nvidia.ko / nvidia-uvm.ko cannot be replaced while any process holds a GPU, so plan a node drain: cordon the node, stop every CUDA job and GPU container, unload the modules or reboot, install, reload. Container runtimes that bind-mount the driver libraries (nvidia-container-toolkit) need restarting so running pods pick up the new userspace. No firmware flash.","references":["https://usn.ubuntu.com/4404-1/","https://usn.ubuntu.com/4404-2/","https://nvd.nist.gov/vuln/detail/CVE-2020-5967"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2020-06-25"},{"id":"CVE-2020-8563","cve":"CVE-2020-8563","aliases":[],"title":"Kubernetes (cloud-controller-manager): vSphere cloud credentials leaked into logs at verbosity 4+","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (cloud-controller-manager)","year":"2020","cvss_score":4.7,"severity":"medium","kev":false,"impact":"vSphere cloud credentials leaked into logs at verbosity 4+","attack_vector":"Anyone with log-pipeline read access","remediation":"Reduce log verbosity; rotate cloud credentials; scrub logs","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-8563"],"status":"curated","published":"2020-12-07"},{"id":"CVE-2020-8564","cve":"CVE-2020-8564","aliases":[],"title":"Kubernetes: Malformed docker config leaks registry pull secrets into logs","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes","year":"2020","cvss_score":4.7,"severity":"medium","kev":false,"impact":"Malformed docker config leaks registry pull secrets into logs","attack_vector":"Anyone with log read access","remediation":"Reduce verbosity; rotate registry pull secrets","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-8564"],"status":"curated","published":"2020-12-07"},{"id":"CVE-2020-8565","cve":"CVE-2020-8565","aliases":[],"title":"Kubernetes: Authorization and bearer tokens written to logs at verbosity 9","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes","year":"2020","cvss_score":4.7,"severity":"medium","kev":false,"impact":"Authorization and bearer tokens written to logs at verbosity 9","attack_vector":"Anyone with log read access","remediation":"Reduce verbosity; rotate tokens","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-8565"],"status":"curated","published":"2020-12-07"},{"id":"CVE-2020-8566","cve":"CVE-2020-8566","aliases":[],"title":"Kubernetes (kube-controller-manager): Ceph RBD admin secrets written to controller-manager logs","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-controller-manager)","year":"2020","cvss_score":4.7,"severity":"medium","kev":false,"impact":"Ceph RBD admin secrets written to controller-manager logs","attack_vector":"Anyone with log read access","remediation":"Reduce verbosity; rotate Ceph credentials","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-8566"],"status":"curated","published":"2020-12-07"},{"id":"CVE-2021-1117","cve":"CVE-2021-1117","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): Improper input validation in the escape handler under specific","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2021","cvss_score":4.7,"severity":"medium","kev":false,"impact":"Improper input validation in the escape handler under specific configurations, giving an unprivileged local user a denial of service against the driver.","attack_vector":"Any local unprivileged user with GPU device access on a Windows host in the affected configuration.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1117"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-10-27"},{"id":"CVE-2021-26318","cve":"CVE-2021-26318","aliases":[],"title":"AMD processors - PREFETCH instruction timing and power side channel: Timing and power measurements around the x86","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD processors - PREFETCH instruction timing and power side channel","year":"2021","cvss_score":4.7,"severity":"medium","kev":false,"impact":"Timing and power measurements around the x86 PREFETCH instructions leak kernel address space layout on some AMD CPUs. Defeating KASLR is not itself a compromise, but it is the step that converts an unreliable kernel memory-corruption bug - and the amdgpu/amdkfd driver long tail is full of them - into a reliable exploit. Treat it as an exploitability multiplier for everything else in this database.","attack_vector":"Local, unprivileged. Reachable from inside a container.","remediation":"Mitigated by AMD microcode plus, on most of these, a kernel-side change - and the durable delivery vehicle is the OEM SBIOS/AGESA package, which carries **one to six months of OEM lag** and needs a drained node and a full power cycle. The linux-firmware amd-ucode blobs get you the microcode sooner via initramfs early-load and a reboot, but AMD does not support late-loading microcode on a running EPYC host, so either way this is reboot-required, not a live patch. Kernel-side mitigation also exists. Because the value of this bug is chaining, the practical defence is to keep the kernel memory-safety patches current rather than to treat KASLR as a real boundary.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26318","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2021-10-13"},{"id":"CVE-2022-27672","cve":"CVE-2022-27672","aliases":["Cross-Thread Return Address Predictions"],"title":"AMD processors with SMT - speculative execution across SMT mode switch: With SMT enabled, certain AMD processors","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD processors with SMT - speculative execution across SMT mode switch","year":"2022","cvss_score":4.7,"severity":"medium","kev":false,"impact":"With SMT enabled, certain AMD processors speculatively execute using a branch target supplied by the sibling thread after an SMT mode switch. One hardware thread therefore steers the other's speculation - a Spectre-v2-shaped attack that crosses the thread boundary rather than the process boundary, so it defeats isolation between two tenants sharing a physical core even when they are in different VMs.","attack_vector":"Local, requires SMT enabled and attacker/victim on sibling threads of the same core - the default arrangement on a bin-packed cluster.","remediation":"Mitigated by AMD microcode plus, on most of these, a kernel-side change - and the durable delivery vehicle is the OEM SBIOS/AGESA package, which carries **one to six months of OEM lag** and needs a drained node and a full power cycle. The linux-firmware amd-ucode blobs get you the microcode sooner via initramfs early-load and a reboot, but AMD does not support late-loading microcode on a running EPYC host, so either way this is reboot-required, not a live patch. The zero-cost mitigation available today is scheduling policy rather than patching: these attacks need the attacker and victim co-resident on sibling SMT threads, so either disable SMT (costing roughly 10-25% throughput on most inference and training workloads) or enforce core isolation so no two tenants ever share a physical core. On a GPU fleet the CPU is rarely the bottleneck, which makes disabling SMT a cheaper trade than it looks on paper.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-27672","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2023-03-01"},{"id":"CVE-2022-28693","cve":"CVE-2022-28693","aliases":["RSBA","Return Stack Buffer Alternate"],"title":"Intel processors (return stack buffer alternate prediction): When the return stack buffer underflows, the processor","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (return stack buffer alternate prediction)","year":"2022","cvss_score":4.7,"severity":"medium","kev":false,"impact":"When the return stack buffer underflows, the processor falls back to an alternate predictor that can be influenced by another context, leaking information. Same structural theme as PBRSB: the return predictor is not as well isolated as the ISA implies.","attack_vector":"Local authorised code on the host.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28693","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00707.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"published":"2025-02-14"},{"id":"CVE-2022-49781","cve":"CVE-2022-49781","aliases":[],"title":"Linux perf/x86/amd - race between amd_pmu_enable_all, perf NMI and throttling: A race between AMD PMU enablement","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux perf/x86/amd - race between amd_pmu_enable_all, perf NMI and throttling","year":"2022","cvss_score":4.7,"severity":"medium","kev":false,"impact":"A race between AMD PMU enablement, the perf NMI handler and event throttling crashes the host. Performance counters are enabled by every profiling and observability agent, and throttling kicks in precisely when counters are busy - so the crash window opens under exactly the monitoring load an AI cluster runs continuously.","attack_vector":"Local, through perf event enablement. Reachable by node observability agents and, where perf_event_paranoid is relaxed, by tenants.","remediation":"Distro kernel update plus reboot; no firmware step.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49781"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-05-01"},{"cwe":["CWE-754"],"fleet":{"pain_class":"node-drain"},"id":"CVE-2022-50636","cve":"CVE-2022-50636","aliases":[],"title":"Linux kernel (drivers/pci): Pci_device_is_present() read the Vendor/Device ID directly, which always reads as all-ones","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/pci)","year":"2022","cvss_score":4.7,"severity":"medium","kev":false,"impact":"Pci_device_is_present() read the Vendor/Device ID directly, which always reads as all-ones on a virtual function, so the kernel concluded every VF was absent. Drivers acted on that false answer - the reported case marks the device broken and drops the completion that would have finished in-flight I/O, wedging the request queue permanently. The practical consequence is a stuck VF teardown, not memory corruption: the process disabling VFs hangs unkillably in D state and the VF's resources cannot be reclaimed.","attack_vector":"PF/VF state confusion in core PCI code, so it applies wherever VFs exist - which in a GPU cloud is the mechanism for handing a slice of a NIC or GPU to a tenant. The hang is hit on the operator's reclaim path (unbinding a VF driver, or writing 0 to sriov_numvfs) while I/O is still in flight, and a tenant keeping its device continuously busy makes that the normal case rather than a rare race. Requires SR-IOV in use; no tenant privilege is needed to keep I/O outstanding.","remediation":"Update to a kernel carrying the fix (no fixed_in published; stable commits below). Interim: quiesce and stop the tenant workload, then confirm the VF is idle, before unbinding its driver or reducing sriov_numvfs - do not tear down VFs under live I/O.","references":["https://git.kernel.org/stable/c/f4b44c7766dae2b8681f621941cabe9f14066d59","https://git.kernel.org/stable/c/643d77fda08d06f863af35e80a7e517ea61d9629","https://nvd.nist.gov/vuln/detail/CVE-2022-50636"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2023-0192","cve":"CVE-2023-0192","aliases":[],"title":"GPU Display Driver: Improper privilege management","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2023","cvss_score":4.7,"severity":"medium","kev":false,"impact":"Improper privilege management","attack_vector":"Local operator / tenant","remediation":"Driver upgrade at next maintenance window","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0192","https://github.com/NVIDIA/product-security/tree/main/2023/5452"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:U/C:H/I:N/A:L","cwe":["CWE-269"],"published":"2023-04-01"},{"id":"CVE-2023-20583","cve":"CVE-2023-20583","aliases":["Collide+Power"],"title":"AMD processors - power side channel on cache line data changes: An authenticated attacker who can read CPU power","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD processors - power side channel on cache line data changes","year":"2023","cvss_score":4.7,"severity":"medium","kev":false,"impact":"An authenticated attacker who can read CPU power consumption can infer data as it moves through a shared cache line, because the energy cost of a write depends on how many bits flip. Collide+Power generalised this into a technique that works against any victim sharing the cache hierarchy, without needing the victim to have any specific code pattern. Leakage rates are low - this is not a fast exfiltration channel - but it is a channel that exists between tenants on the same socket regardless of what your isolation configuration says.","attack_vector":"Local, authenticated, needs access to power/energy telemetry (RAPL-style interfaces) and cache co-residency with the victim. Reachable from a container if you expose power telemetry into it, which some GPU and CPU monitoring stacks do.","remediation":"Mitigated primarily by restricting who can read fine-grained power telemetry: on Linux, the RAPL/energy interfaces should be root-only (this was tightened upstream), and you should not be passing power monitoring into tenant containers. Check what your DCGM-equivalent AMD telemetry stack exposes and to whom - operators often mount host telemetry paths into monitoring sidecars that tenants can reach. Kernel/driver-level fix plus a config review; no firmware flash and no reboot if you are only tightening permissions. The zero-cost mitigation available today is scheduling policy rather than patching: these attacks need the attacker and victim co-resident on sibling SMT threads, so either disable SMT (costing roughly 10-25% throughput on most inference and training workloads) or enforce core isolation so no two tenants ever share a physical core. On a GPU fleet the CPU is rarely the bottleneck, which makes disabling SMT a cheaper trade than it looks on paper.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-20583","https://collidepower.com/","https://www.amd.com/en/resources/product-security.html"],"status":"curated","tags":["tenant-isolation"],"published":"2023-08-01"},{"id":"CVE-2023-4641","cve":"CVE-2023-4641","aliases":[],"title":"shadow-utils: Possible password leak during passwd(1) change (uninitialised memory)","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"shadow-utils","year":"2023","cvss_score":4.7,"severity":"medium","kev":false,"impact":"Possible password leak during passwd(1) change (uninitialised memory)","attack_vector":"Local user","remediation":"Package update; no reboot","references":["https://access.redhat.com/security/cve/CVE-2023-4641"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2023-12-27"},{"cwe":["CWE-1254","CWE-125"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-52746","cve":"CVE-2023-52746","aliases":[],"title":"Linux kernel (net/xfrm): The 32-bit compat translation of xfrm netlink attributes uses the attacker-supplied attribute","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/xfrm)","year":"2023","cvss_score":4.7,"severity":"medium","kev":false,"impact":"The 32-bit compat translation of xfrm netlink attributes uses the attacker-supplied attribute type as an array index after a bounds check that the CPU can speculate past. This is a Spectre v1 gadget in the xfrm netlink parser: architecturally safe, but speculatively it reads kernel memory outside the policy table and can leak it through a cache side channel.","attack_vector":"A 32-bit process sending xfrm netlink messages, needing CAP_NET_ADMIN in the network namespace - which a container granted NET_ADMIN with its own netns has, and CONFIG_COMPAT plus a 32-bit or compat-capable tenant binary. Turning the gadget into an actual leak requires a working speculation side channel, so treat this as a hardening gap rather than a directly weaponizable read.","remediation":"Boot a kernel carrying the linked stable commits (which add array_index_nospec). Interim: drop CAP_NET_ADMIN from tenant containers, or build/boot without CONFIG_COMPAT on nodes that never run 32-bit workloads.","references":["https://git.kernel.org/stable/c/a893cc644812728e86e9aff517fd5698812ecef0","https://git.kernel.org/stable/c/5dc688fae6b7be9dbbf5304a3d2520d038e06db5","https://nvd.nist.gov/vuln/detail/CVE-2023-52746"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-21792","cve":"CVE-2024-21792","aliases":[],"title":"Intel Neural Compressor (TOCTOU): A time-of-check/time-of-use race in Neural Compressor lets an authenticated local","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Intel Neural Compressor (TOCTOU)","year":"2024","cvss_score":4.7,"severity":"medium","kev":false,"impact":"A time-of-check/time-of-use race in Neural Compressor lets an authenticated local user win a window and read information they should not. Low severity on its own; useful as a step in a chain on a shared optimisation node.","attack_vector":"Authenticated local user on the node running Neural Compressor.","remediation":"Upgrade to 2.5.0 or later. Package update, service restart.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21792","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01109.html"],"status":"curated","fleet":{"pain_class":"daemon-restart"},"published":"2024-05-16"},{"id":"CVE-2024-2201","cve":"CVE-2024-2201","aliases":[],"title":"Intel CPU (Native BHI): Native Branch History Injection - unprivileged user leaks kernel memory despite eIBRS","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel CPU (Native BHI)","year":"2024","cvss_score":4.7,"severity":"medium","kev":false,"impact":"Native Branch History Injection - unprivileged user leaks kernel memory despite eIBRS; also XSA-456 for Xen","attack_vector":"Any tenant process in a container; tenant VM guest","remediation":"Kernel mitigation (BHI_DIS_S or software sequence) + reboot; standing perf cost. Affects Intel Xeon hosts under GPU nodes","references":["https://access.redhat.com/security/cve/CVE-2024-2201"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-12-19"},{"id":"CVE-2024-27040","cve":"CVE-2024-27040","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":4.7,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add 'replay' NULL check in 'edp_set_replay_allow_active()'","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-27040","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-01"},{"id":"CVE-2024-32473","cve":"CVE-2024-32473","aliases":[],"title":"Docker / moby: IPv6 not disabled on interfaces where it should be","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2024","cvss_score":4.7,"severity":"medium","kev":false,"impact":"IPv6 not disabled on interfaces where it should be; unexpected reachability of containers over IPv6","attack_vector":"Any pod on the cluster network","remediation":"Upgrade moby; audit IPv6 firewall posture","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-32473"],"status":"curated","published":"2024-04-18"},{"id":"CVE-2024-36024","cve":"CVE-2024-36024","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A race condition or locking defect in the amdgpu display","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":4.7,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdgpu display core (DC/DM). Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amd/display: Disable idle reallow as part of command/gpint execution","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36024","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-30"},{"cwe":["CWE-401"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-38632","cve":"CVE-2024-38632","aliases":[],"title":"Linux kernel (drivers/vfio/pci): A failed interrupt-context allocation while enabling INTx leaks the IRQ name string.","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/vfio/pci)","year":"2024","cvss_score":4.7,"severity":"medium","kev":false,"impact":"A failed interrupt-context allocation while enabling INTx leaks the IRQ name string. Individually trivial, but a tenant can drive the failing path repeatedly, and unreclaimable kernel allocations on a shared node are a slow noisy-neighbour outage that eventually reaches every other tenant.","attack_vector":"A tenant holding a vfio-pci device fd calling VFIO_DEVICE_SET_IRQS to enable INTx while the interrupt-context allocation fails. Requires memory pressure to make the allocation fail, so this is an OOM-adjacent amplifier rather than a clean standalone primitive. No host root.","remediation":"Update to 5.15.168, 6.1.113, or 6.6.33 or later. Interim control: cap tenant memory with cgroups so they cannot easily drive the host into the allocation-failure regime.","references":["https://git.kernel.org/stable/c/a6d810554d7d9d07041f14c5fcd453f3d3fed594","https://git.kernel.org/stable/c/91ced077db2062604ec270b1046f8337e9090079","https://nvd.nist.gov/vuln/detail/CVE-2024-38632"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-42227","cve":"CVE-2024-42227","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":4.7,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Fix overlapping copy within dml_core_mode_programming","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-42227","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-07-30"},{"id":"CVE-2024-46851","cve":"CVE-2024-46851","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":4.7,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Avoid race between dcn10_set_drr() and dc_state_destruct()","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46851","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-09-27"},{"id":"CVE-2024-46870","cve":"CVE-2024-46870","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A race condition or locking defect in the amdgpu display","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":4.7,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdgpu display core (DC/DM). Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amd/display: Disable DMCUB timeout for DCN35","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46870","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-10-09"},{"cwe":["CWE-401"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-56742","cve":"CVE-2024-56742","aliases":[],"title":"Linux kernel (drivers/vfio/pci/mlx5): Pages allocated for a device migration buffer are not freed when adding them to","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel (drivers/vfio/pci/mlx5)","year":"2024","cvss_score":4.7,"severity":"medium","kev":false,"impact":"Pages allocated for a device migration buffer are not freed when adding them to the scatter-gather table fails, leaking host memory on every failed attempt. On a fleet that live-migrates SR-IOV NIC functions between tenants, repeated failures bleed host RAM that never comes back.","attack_vector":"The mlx5 VFIO variant driver's migration buffer path, exercised by the host's migration control plane with a size influenced by the tenant device's state. Requires the SG-table add to fail, i.e. memory pressure. Conditional on ConnectX VFs being passed through with live migration enabled - common on ConnectX-backed neoclouds. Host-side path, not a direct tenant ioctl.","remediation":"The record lists no fixed release; boot a kernel carrying the stable fix commits below. Interim control: monitor host memory across migration failures and drain nodes that show unexplained kernel memory growth.","references":["https://git.kernel.org/stable/c/769fe4ce444b646b0bf6ac308de80686c730c7df","https://git.kernel.org/stable/c/c44f1b2ddfa81c8d7f8e9b6bc76c427bc00e69d5","https://nvd.nist.gov/vuln/detail/CVE-2024-56742"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-401"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-21770","cve":"CVE-2025-21770","aliases":[],"title":"Linux kernel (drivers/iommu): Removing a device from the per-IOMMU page-fault queue responds to outstanding faults but","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu)","year":"2025","cvss_score":4.7,"severity":"medium","kev":false,"impact":"Removing a device from the per-IOMMU page-fault queue responds to outstanding faults but never frees the group structures holding them, so every teardown that happens with faults still pending leaks host kernel memory. The tenant controls both halves - it decides how many faults its assigned device has in flight and when PRI gets disabled - so repeated attach/detach cycles bleed the host.","attack_vector":"A tenant with an ATS/PRI-capable assigned device (SVM-capable GPU, PRI-capable NIC, DSA/IAA) drives the device to generate page requests and then disables PRI or releases the device while requests are still outstanding, repeatedly. No host root needed. Conditional on PRI/IOPF being enabled for tenant-assigned devices.","remediation":"No fixed release is listed in this record; apply the linked stable commits or run a current stable/LTS kernel. Interim: disable PRI/ATS on tenant-assigned devices that do not need demand paging, and rate-limit device attach/detach cycles per tenant.","references":["https://git.kernel.org/stable/c/db60d2d896a17decd58d143eef92cf22eb0a0176","https://git.kernel.org/stable/c/90d5429cd2921ca2714684ed525898d431bb9283","https://nvd.nist.gov/vuln/detail/CVE-2025-21770"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-662"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-23161","cve":"CVE-2025-23161","aliases":[],"title":"Linux kernel (drivers/pci/controller): The Intel VMD driver guarded config-space access with a lock type that becomes a","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/pci/controller)","year":"2025","cvss_score":4.7,"severity":"medium","kev":false,"impact":"The Intel VMD driver guarded config-space access with a lock type that becomes a sleeping lock under PREEMPT_RT, while the PCI core calls into it with interrupts disabled. Reading config space then sleeps in atomic context - a BUG splat and a wedged CPU on the shared node, reached from an ordinary sysfs read rather than anything privileged.","attack_vector":"The reported call chain starts in sysfs: pci_read_config -> pci_user_read_config_byte -> vmd_pci_read, i.e. a read of /sys/bus/pci/devices/<dev>/config for a device behind Intel VMD. Any local process that can open that file reaches it, including a tenant container with the default sysfs mount. Conditional on two things and inert without both: a PREEMPT_RT kernel, and Intel VMD enabled in BIOS (common on Intel server platforms fronting NVMe).","remediation":"Update to a kernel carrying the fix (no fixed_in published; stable commits below). Interim: on PREEMPT_RT nodes, either disable VMD in BIOS or mask sysfs config-space files from tenant containers; on non-RT kernels no action is needed.","references":["https://git.kernel.org/stable/c/c250262d6485ca333e9821f85b07eb383ec546b1","https://git.kernel.org/stable/c/c2968c812339593ac6e2bdd5cc3adabe3f05fa53","https://nvd.nist.gov/vuln/detail/CVE-2025-23161"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-23269","cve":"CVE-2025-23269","aliases":[],"title":"Jetson Xavier / Orin: Info disclosure via side channel","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Jetson Xavier / Orin","year":"2025","cvss_score":4.7,"severity":"medium","kev":false,"impact":"Info disclosure via side channel","attack_vector":"Local attacker on the device","remediation":"Flash JetPack; edge fleet","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23269","https://github.com/NVIDIA/product-security/tree/main/2025/5662"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-1423"],"published":"2025-07-17"},{"id":"CVE-2025-37899","cve":"CVE-2025-37899","aliases":[],"title":"Linux kernel (ksmbd): Use-after-free in ksmbd session logoff (found by an LLM-assisted audit)","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (ksmbd)","year":"2025","cvss_score":4.7,"severity":"medium","kev":false,"impact":"Use-after-free in ksmbd session logoff (found by an LLM-assisted audit)","attack_vector":"Authenticated network to an exposed ksmbd share","remediation":"Livepatchable; otherwise drain + reboot. Best answer remains not shipping ksmbd","references":["https://access.redhat.com/security/cve/CVE-2025-37899"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-05-20"},{"id":"CVE-2025-38104","cve":"CVE-2025-38104","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu)","year":"2025","cvss_score":4.7,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu power management (SMU/powerplay). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdgpu: Replace Mutex with Spinlock for RLCG register access to avoid Priority Inversion in SRIOV","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38104","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-04-18"},{"id":"CVE-2025-64432","cve":"CVE-2025-64432","aliases":[],"title":"KubeVirt: Flawed aggregation-layer authentication flow enables RBAC bypass","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"KubeVirt","year":"2025","cvss_score":4.7,"severity":"medium","kev":false,"impact":"Flawed aggregation-layer authentication flow enables RBAC bypass","attack_vector":"Cluster user with namespace access","remediation":"Upgrade KubeVirt","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-64432"],"status":"curated","published":"2025-11-07"},{"id":"CVE-2026-24199","cve":"CVE-2026-24199","aliases":[],"title":"GPU Display Driver: DoS (race in GPU resource allocation)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2026","cvss_score":4.7,"severity":"medium","kev":false,"impact":"DoS (race in GPU resource allocation)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; rolling reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24199","https://github.com/NVIDIA/product-security/tree/main/2026/5821"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-362"],"fleet":{"pain_class":"node-reboot"},"published":"2026-05-26"},{"cwe":["CWE-415","CWE-672"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-43097","cve":"CVE-2026-43097","aliases":[],"title":"Linux kernel (drivers/pci/controller): The Hyper-V PCI front-end frees its PCI domain number twice on a probe failure","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/pci/controller)","year":"2026","cvss_score":4.7,"severity":"medium","kev":false,"impact":"The Hyper-V PCI front-end frees its PCI domain number twice on a probe failure - once itself and once again through the bridge release callback. The second free returns an ID that may already have been reissued to a different bus, so two PCI domains can end up carrying the same domain number. Duplicate domain numbering means sysfs paths and device lookups that are supposed to distinguish two hierarchies stop doing so, on the very driver that presents passthrough devices to a guest.","attack_vector":"Guest-side, inside a Linux VM on Hyper-V/Azure that is being given a passthrough device - pci-hyperv is that paravirtual front-end. It requires hv_pci_probe() to fail after the domain number is stored, which is a host- or fabric-side condition (device offer withdrawn, channel setup failing) rather than something guest userspace triggers; a tenant with no control over device offers cannot force it. Nodes not running under Hyper-V never load the driver.","remediation":"Boot a kernel where pci-hyperv leaves domain_nr release to the PCI core. Interim: on Hyper-V hosts, watch for the 'ida_free called for id=... which is not allocated' warning as the marker that a domain ID has been double-freed, and restart the affected VM rather than letting it continue with ambiguous domain numbering.","references":["https://git.kernel.org/stable/c/21bc8e0ba5c2a081b0a2808c976d4c9dbddf1e48","https://git.kernel.org/stable/c/b6422dff0e518245019233432b6bccfc30b73e2f","https://nvd.nist.gov/vuln/detail/CVE-2026-43097"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2021-33082","cve":"CVE-2021-33082","aliases":["INTEL-SA-00563","Solidigm SA-000563"],"title":"Intel / Solidigm SSD, SSD DC and Optane SSD firmware","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel / Solidigm SSD, SSD DC and Optane SSD firmware - NVMe Sanitize (Block Erase) leaves prior data recoverable","year":"2021","cvss_score":4.6,"severity":"medium","kev":false,"impact":"Sensitive data is not removed before the media is reused. The drive accepts an NVMe Sanitize with Block Erase, reports completion, and leaves the previous contents recoverable. BREAKS TENANT HANDOFF in the most direct way in this whole category: this IS the command most bare-metal reclaim pipelines issue between customers. Your automation records a successful sanitize, the node goes back into the pool, and the next tenant can carve the previous tenant's datasets, model checkpoints and any credentials that touched local scratch out of the flash. Worse, it fails clean - there is no error to alert on, so a fleet can run in this state for years and every audit log says the erase succeeded.","attack_vector":"The next tenant with root on the reclaimed bare-metal host, or anyone who obtains the physical drive later (RMA, decommission, resale). The attacker does not need to break anything - they read what the sanitize left behind.","remediation":"Two moves, do both. (1) Immediate policy/workaround, no downtime: stop using Sanitize with Block Erase (SANACT=04h in the Sanitize command's block-erase form) in your reclaim path and switch to Sanitize with Crypto Erase, or Format NVM with Crypto Erase / User Data Erase - this is the vendor's own prescribed workaround and it is a one-line change in most reclaim scripts, so ship it today. (2) Flash affected drives to fixed firmware; this needs the drive quiesced, usually a node drain, and Solidigm Storage Tool / Intel MAS with per-SKU firmware images, so plan it as a rolling maintenance campaign. Note that several older SKUs in the same advisory family are explicitly end-of-support with no fix planned unless a customer asks - for those, the workaround IS the remediation. Independently: layer LUKS/dm-crypt with an operator-held key on all tenant-visible local NVMe, so reclaim means destroying your key rather than trusting the drive's report. Verifying erase across a 10,000-drive fleet by actually reading back raw blocks is a weeks-long operation and essentially no operator does it - which is precisely why a firmware that lies about sanitize success goes undetected.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-33082","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00563.html","https://www.solidigm.com/support-page/support-security.html","https://www.solidigm.com/content/dam/solidigm/en/site/support/support-community/cve-(security)/documents/public-security-advisory-v2.pdf"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"published":"2022-05-12"},{"id":"CVE-2022-40633","cve":"CVE-2022-40633","aliases":["ICSA-23-061-03"],"title":"Rittal CMC III cabinet lock / access-card system: The access cards used to open control cabinets secured with Rittal","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Rittal CMC III cabinet lock / access-card system","year":"2022","cvss_score":4.6,"severity":"medium","kev":false,"impact":"The access cards used to open control cabinets secured with Rittal CMC III locks can be cloned. In a datacenter this is the rack-level and cabinet-level access boundary - the CMC III system is exactly what many operators use to electronically lock individual racks and cabinets and to log who opened them, and it is often the control that lets a provider tell a customer 'only your staff can open your rack'. Cloning a card defeats that, and it defeats it invisibly: the lock logs a legitimate card opening a legitimate cabinet. Someone at an opened rack can pull NVMe drives holding model weights and customer data, attach a console to a node's serial or VGA port, plug into the out-of-band management switch that fronts every BMC in the row, or insert a hardware implant on a management link. The CVSS of 4.6 reflects the physical-access precondition, not the consequence - for a bare-metal GPU provider whose entire isolation story is physical, this is a boundary failure, and one that also breaks tenant handoff because the same credentials and the same locks carry across tenancies.","attack_vector":"Physical proximity to a valid card, then physical presence at the cabinet. No network access is involved. The precondition that matters is that the attacker must already be inside the hall - so this is the second stage after tailgating, a compromised hall-door credential, or legitimate access as a contractor, landlord technician, or another tenant's staff in a shared hall.","remediation":"Not fixable by patching - it is the credential technology. Rittal's guidance is to move to a more secure card technology where the hardware supports it; in practice that means replacing readers and reissuing cards, a per-cabinet hardware cost. Compensating controls that work today: tamper alarms on cabinet doors wired into a system separate from the lock itself, camera coverage of aisles with retention long enough to review, and a policy that any cabinet-open event is reconciled against a work order rather than just logged. In a shared hall, treat the cabinet lock as a deterrent rather than a boundary, and put anything that genuinely requires isolation behind a full cage with a second access factor.","references":["https://www.cisa.gov/news-events/ics-advisories/icsa-23-061-03","https://nvd.nist.gov/vuln/detail/CVE-2022-40633"],"status":"curated"},{"id":"CVE-2023-6134","cve":"CVE-2023-6134","aliases":[],"title":"Keycloak: Redirect scheme filtering bypassed by appending a wildcard","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Keycloak","year":"2023","cvss_score":4.6,"severity":"medium","kev":false,"impact":"Redirect scheme filtering bypassed by appending a wildcard -> XSS and follow-on attacks","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; tighten allowed redirect URIs","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-6134"],"status":"curated","published":"2023-12-14"},{"id":"CVE-2023-6927","cve":"CVE-2023-6927","aliases":[],"title":"Keycloak: Wildcard in the JARM form_post.jwt response mode","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Keycloak","year":"2023","cvss_score":4.6,"severity":"medium","kev":false,"impact":"Wildcard in the JARM form_post.jwt response mode -> steal authorization codes and tokens (bypasses the CVE-2023-6134 fix)","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade + remove wildcards from client redirect URIs","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-6927"],"status":"curated","published":"2023-12-18"},{"id":"CVE-2024-40635","cve":"CVE-2024-40635","aliases":[],"title":"containerd: UID:GID larger than 32-bit signed max wraps to 0, silently running the container as root","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2024","cvss_score":4.6,"severity":"medium","kev":false,"impact":"UID:GID larger than 32-bit signed max wraps to 0, silently running the container as root","attack_vector":"Malicious image setting a huge numeric USER","remediation":"Rolling containerd upgrade with node drain; add admission check on runAsUser","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-40635"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2025-03-17"},{"id":"CVE-2025-0031","cve":"CVE-2025-0031","aliases":[],"title":"AMD SEV firmware - use-after-free allowing a SINGLE_SOCKET guest to activate on the wrong socket (AMD-SB-3023): A","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV firmware - use-after-free allowing a SINGLE_SOCKET guest to activate on the wrong socket (AMD-SB-3023)","year":"2025","cvss_score":4.6,"severity":"medium","kev":false,"impact":"A use-after-free in SEV firmware lets a migrated guest whose policy specifies SINGLE_SOCKET be activated on a different socket than its migration agent. The guest chose that policy to constrain where its keys and memory live; violating it silently means a tenant's stated confidential-computing constraint is not being enforced, and they have no way to tell.","attack_vector":"Malicious hypervisor driving guest migration.","remediation":"Fixed in AMD PI/AGESA firmware and delivered only as an OEM SBIOS package - AMD ships the PI drop to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before releasing BIOS. **Budget one to six months of OEM lag**, and note that several CVEs in this batch are marked 'no fix planned' on Naples (EPYC 7001) - for those the only remediation is retiring the hardware. Applying it means cordon, drain and a full power cycle per node; there is no driver reload, no live patch and no VBIOS step. Because this touches the SEV-SNP trust boundary, the update moves the platform TCB version: refresh VCEK certificates from AMD's KDS and update tenant attestation policy, or confidential guest launches will fail immediately after the BIOS lands. Worth flagging to any tenant who sets socket-scoped SNP guest policies - if you sell policy enforcement as a feature, this is a period during which it was not enforced.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0031","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3023.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"tags":["tenant-isolation"],"published":"2026-02-10"},{"id":"CVE-2025-23292","cve":"CVE-2025-23292","aliases":[],"title":"NVIDIA License System - Delegated Licensing Service (DLS): SQL injection in the DLS appliance reaching a high integrity","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA License System - Delegated Licensing Service (DLS)","year":"2025","cvss_score":4.6,"severity":"medium","kev":false,"impact":"SQL injection in the DLS appliance reaching a high integrity impact and partial denial of service in the UI - an authenticated attacker can alter licensing records.","attack_vector":"Adjacent network, high privileges, user interaction. An admin-level account on the licensing appliance.","remediation":"Update the DLS appliance per bulletin 5705 and review licensing records for tampering. Cost: appliance restart.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23292","https://github.com/NVIDIA/product-security/tree/main/2025/5705"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:H/PR:H/UI:R/S:U/C:N/I:H/A:L","cwe":["CWE-943"],"published":"2025-09-30"},{"cvss_vector":"CVSS:4.0/AV:N/AC:L/AT:P/PR:N/UI:N/VC:N/VI:H/VA:N/SC:N/SI:N/SA:N/E:U","cwe":["CWE-287"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2025-27414","cve":"CVE-2025-27414","aliases":[],"title":"MinIO (SFTP gateway): The SFTP frontend trusts an SSH public key it should not, letting an attacker authenticate as","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"MinIO (SFTP gateway)","year":"2025","cvss_score":4.6,"severity":"medium","kev":false,"impact":"The SFTP frontend trusts an SSH public key it should not, letting an attacker authenticate as another user without that user's private key. Whatever buckets that identity can reach are now readable and writable by the attacker over SFTP.","attack_vector":"Any client that can reach the MinIO SFTP port on a deployment with SFTP enabled and public-key auth configured.","remediation":"Upgrade to the release named in GHSA-wc79-7x8x-2p58 and restart the SFTP listener. If SFTP is not a requirement, disable it entirely - it is a second authentication surface on the same data.","references":["https://github.com/minio/minio/security/advisories/GHSA-wc79-7x8x-2p58","https://nvd.nist.gov/vuln/detail/CVE-2025-27414"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-29943","cve":"CVE-2025-29943","aliases":[],"title":"AMD CPU pipeline configuration - SEV-SNP guest stack pointer corruption: A write-what-where condition in CPU pipeline","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD CPU pipeline configuration - SEV-SNP guest stack pointer corruption","year":"2025","cvss_score":4.6,"severity":"medium","kev":false,"impact":"A write-what-where condition in CPU pipeline configuration lets an admin-privileged attacker corrupt the stack pointer inside an SEV-SNP guest. Corrupting a guest's stack pointer from outside is a control-flow attack on a VM whose memory the host is not supposed to be able to touch - a route to steering execution inside a confidential workload.","attack_vector":"Local, admin-privileged host attacker.","remediation":"Fixed in AMD SEV firmware / AGESA and reaches you as an OEM SBIOS package - AMD hands AGESA to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before shipping BIOS. **Budget one to six months of OEM lag**, longer on older platforms and sometimes never on end-of-support SKUs. Applying it means draining the host and doing a full power cycle. Because the fix moves the platform's reported SEV-SNP TCB version, you must also pull fresh VCEK certificates from AMD's Key Distribution Service and update any attestation policy your tenants pin - otherwise guests will start failing launch validation the moment the BIOS lands. Some SEV firmware can alternatively be staged from linux-firmware (amd/amd_sev_*.sbin) and committed via the ccp driver at boot, which is faster than waiting on BIOS - check whether your platform supports firmware hot-load before assuming the OEM is the only route.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-29943","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2026-01-16"},{"id":"CVE-2025-47291","cve":"CVE-2025-47291","aliases":[],"title":"containerd: User-namespaced containers not placed under the Kubernetes cgroup, defeating resource limits","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2025","cvss_score":4.6,"severity":"medium","kev":false,"impact":"User-namespaced containers not placed under the Kubernetes cgroup, defeating resource limits; noisy-neighbour / node DoS","attack_vector":"Any tenant workload using user namespaces","remediation":"Rolling containerd upgrade with node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-47291"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2025-05-21"},{"id":"CVE-2025-48517","cve":"CVE-2025-48517","aliases":[],"title":"AMD SEV firmware - ASID range enforcement between SEV-ES and SEV-SNP guests: A malicious hypervisor can launch a SEV-ES","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV firmware - ASID range enforcement between SEV-ES and SEV-SNP guests","year":"2025","cvss_score":4.6,"severity":"medium","kev":false,"impact":"A malicious hypervisor can launch a SEV-ES guest using an ASID from the range reserved for SEV-SNP guests. ASIDs key the memory encryption, so overlapping the ranges lets a weaker-protected ES guest sit where an SNP guest's protections were assumed - a partial confidentiality loss for the SNP tenant. The interesting part is that the attack uses a legitimate hypervisor operation with an out-of-range parameter rather than any memory-safety bug.","attack_vector":"Requires hypervisor privilege and the ability to launch guests - i.e. the cloud operator or anyone who compromises the control plane.","remediation":"Fixed in AMD SEV firmware / AGESA and reaches you as an OEM SBIOS package - AMD hands AGESA to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before shipping BIOS. **Budget one to six months of OEM lag**, longer on older platforms and sometimes never on end-of-support SKUs. Applying it means draining the host and doing a full power cycle. Because the fix moves the platform's reported SEV-SNP TCB version, you must also pull fresh VCEK certificates from AMD's Key Distribution Service and update any attestation policy your tenants pin - otherwise guests will start failing launch validation the moment the BIOS lands. Some SEV firmware can alternatively be staged from linux-firmware (amd/amd_sev_*.sbin) and committed via the ccp driver at boot, which is faster than waiting on BIOS - check whether your platform supports firmware hot-load before assuming the OEM is the only route.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-48517","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2026-02-10"},{"id":"CVE-2025-56548","cve":"CVE-2025-56548","aliases":["PT-2025-171"],"title":"Broadcom NetXtreme-E network adapter firmware: The lower-severity half of the same Positive Technologies NetXtreme-E","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Broadcom NetXtreme-E network adapter firmware","year":"2025","cvss_score":4.6,"severity":"medium","kev":false,"impact":"The lower-severity half of the same Positive Technologies NetXtreme-E firmware disclosure. Carried here because it is fixed by the same firmware image — an operator who patches for the 8.2 issue gets this for free, and one who tracks only high-severity CVEs will never see it listed.","attack_vector":"Adapter firmware interface; same exposure surface as the companion issue.","remediation":"Same NetXtreme-E firmware update; NIC flash and cold power cycle. No separate action needed once the fixed image is deployed.","references":["https://global.ptsecurity.com/en/about/news/pt-expert-helped-patch-vulnerabilities-broadcom-network-adapter-firmware/","https://nvd.nist.gov/vuln/detail/CVE-2025-56548"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-66664","cve":"CVE-2025-66664","aliases":[],"title":"AMD Secure Processor TEE SOC driver - SR-IOV GFX firmware load command: A malformed DRV_SOC_CMD_ID_LOAD_GFX_IP_FW","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor TEE SOC driver - SR-IOV GFX firmware load command","year":"2025","cvss_score":4.6,"severity":"medium","kev":false,"impact":"A malformed DRV_SOC_CMD_ID_LOAD_GFX_IP_FW SR-IOV command causes an out-of-bounds read in the ASP's TEE SOC driver. This one matters specifically because it sits on the SR-IOV path - the mechanism by which a single GPU is carved up between virtual functions belonging to different tenants. A guest VF driver reaching a firmware-load command handler is precisely the boundary GPU virtualisation is supposed to hold.","attack_vector":"Reachable from an SR-IOV virtual function, i.e. from a guest VM that has been assigned a GPU VF - so this is guest-to-platform, not merely host-local. Only applies where GPU SR-IOV virtualisation is actually enabled.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. If you run GPU SR-IOV with untrusted tenants in the VFs, this is worth prioritising over its 4.6 score; if you do passthrough of whole physical GPUs instead, the path is not exercised.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-66664","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2026-05-15"},{"id":"CVE-2021-3473","cve":"CVE-2021-3473","aliases":[],"title":"Lenovo XClarity Controller: Backup/restore password written to an internal XCC log buffer","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lenovo XClarity Controller","year":"2021","cvss_score":4.5,"severity":"medium","kev":false,"impact":"Backup/restore password written to an internal XCC log buffer — credential leak to anyone who can pull BMC logs","attack_vector":"Local/network, authenticated","remediation":"XCC firmware update plus rotation of any XCC backup passwords used during fleet provisioning","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-3473"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2021-04-13"},{"id":"CVE-2025-23252","cve":"CVE-2025-23252","aliases":[],"title":"NVDebug tool: Info disclosure / privesc via diagnostic tool","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVDebug tool","year":"2025","cvss_score":4.5,"severity":"medium","kev":false,"impact":"Info disclosure / privesc via diagnostic tool","attack_vector":"Local operator","remediation":"Upgrade NVDebug on nodes","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23252","https://github.com/NVIDIA/product-security/tree/main/2025/5651"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:H/UI:R/S:U/C:H/I:N/A:N","cwe":["CWE-1244"],"published":"2025-06-18"},{"id":"CVE-2025-23274","cve":"CVE-2025-23274","aliases":[],"title":"CUDA Toolkit: Info disclosure (buffer over-read)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2025","cvss_score":4.5,"severity":"medium","kev":false,"impact":"Info disclosure (buffer over-read)","attack_vector":"Malicious artifact","remediation":"Bump CUDA Toolkit; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23274","https://github.com/NVIDIA/product-security/tree/main/2025/5661"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:L/I:L/A:L","cwe":["CWE-125"],"published":"2025-09-24"},{"id":"CVE-2025-4166","cve":"CVE-2025-4166","aliases":[],"title":"HashiCorp Vault: KV v2 leaks sensitive payload content into server and audit logs on malformed requests","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HashiCorp Vault","year":"2025","cvss_score":4.5,"severity":"medium","kev":false,"impact":"KV v2 leaks sensitive payload content into server and audit logs on malformed requests","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; scrub and re-secure the audit log store","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-4166"],"status":"curated","published":"2025-05-02"},{"id":"CVE-2017-5698","cve":"CVE-2017-5698","aliases":["INTEL-SA-00082"],"title":"Intel AMT / ISM / SBT firmware anti-rollback, ME 11.0.25.3001 and 11.0.26.3000: The patched ME firmware does","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel AMT / ISM / SBT firmware anti-rollback, ME 11.0.25.3001 and 11.0.26.3000","year":"2017","cvss_score":4.4,"severity":"medium","kev":false,"impact":"The patched ME firmware does not enforce anti-rollback, so a local administrator can downgrade the ME back to an 11.6.x image that is vulnerable to the KEV-listed AMT auth bypass. Operationally this means a node you have already remediated can be silently un-remediated: your fleet inventory says 'patched', the ME says otherwise, and the auth bypass is live again. Because the downgrade lives in the ME region, it survives a host reimage and therefore survives tenant handoff.","attack_vector":"Local root or administrator on the host, using the normal Intel ME firmware update path (HECI/MEI device). No physical access, no ME exploit required - just the vendor's own update tool. Any tenant who has been given real root on a bare-metal node can do this before handing the node back.","remediation":"Flash to an ME build that enforces the rollback floor - again an OEM BIOS/ME bundle from Dell/HPE/Lenovo/Supermicro/Gigabyte/Quanta, with a reboot. The durable control is process, not firmware: read back and attest the actual ME firmware version at node reclaim time rather than trusting an inventory record, and block host-side ME update tooling (MEI device access, Intel MEInfo/FWUpdate binaries) inside tenant images.","references":["https://nvd.nist.gov/vuln/detail/CVE-2017-5698","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00082.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2019-0117","cve":"CVE-2019-0117","aliases":[],"title":"Intel SGX protected memory subsystem: Insufficient access control in the SGX protected-memory subsystem allows","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SGX protected memory subsystem","year":"2019","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Insufficient access control in the SGX protected-memory subsystem allows information disclosure from enclave memory to a privileged local user. Part of the long tail of SGX hardware issues that each force a TCB recovery.","attack_vector":"Privileged local access on the host.","remediation":"Microcode/platform firmware update and re-attestation of all enclaves. Where the fix ships in microcode it can be late-loaded at boot; where it ships in the platform BIOS, expect an OEM release and a per-node drain and reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-0117","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00219.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2019-11-14"},{"id":"CVE-2019-5698","cve":"CVE-2019-5698","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): An index value from the guest is not validated by the vGPU plugin, letting one","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2019","cvss_score":4.4,"severity":"medium","kev":false,"impact":"An index value from the guest is not validated by the vGPU plugin, letting one tenant VM crash the host-side plugin and deny service to every co-tenant on that GPU.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-5698"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2019-11-09"},{"id":"CVE-2020-24497","cve":"CVE-2020-24497","aliases":[],"title":"Intel E810 Ethernet controller firmware (NVM < 1.4.1.13): Early-generation E810 firmware flaw (an access-control","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel E810 Ethernet controller firmware (NVM < 1.4.1.13)","year":"2020","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Early-generation E810 firmware flaw (an access-control failure) resulting in denial of service on the adapter. Relevant to any E810 adapter still on shipping-era NVM, which is common when NIC firmware was never brought into the patch pipeline.","attack_vector":"Varies by issue - the unauthenticated variant is reachable from the network, the others need privileged host access.","remediation":"Fixed in the E810 NVM (adapter firmware) image. Deploy with Intel's NVM Update Utility, which needs a driver reload and a power cycle - not just a warm reboot - for the new image to take effect. Drain the node. Distinct from the ice driver updates: you need both, and they ship on different schedules. Target NVM 1.4.1.13 or later, but jump straight to a current image rather than the minimum fixed version.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-24497","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00456.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-02-17"},{"id":"CVE-2020-24498","cve":"CVE-2020-24498","aliases":[],"title":"Intel E810 Ethernet controller firmware (NVM < 1.4.1.13): Early-generation E810 firmware flaw (a buffer overflow","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel E810 Ethernet controller firmware (NVM < 1.4.1.13)","year":"2020","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Early-generation E810 firmware flaw (a buffer overflow reachable by a privileged user) resulting in denial of service on the adapter. Relevant to any E810 adapter still on shipping-era NVM, which is common when NIC firmware was never brought into the patch pipeline.","attack_vector":"Varies by issue - the unauthenticated variant is reachable from the network, the others need privileged host access.","remediation":"Fixed in the E810 NVM (adapter firmware) image. Deploy with Intel's NVM Update Utility, which needs a driver reload and a power cycle - not just a warm reboot - for the new image to take effect. Drain the node. Distinct from the ice driver updates: you need both, and they ship on different schedules. Target NVM 1.4.1.13 or later, but jump straight to a current image rather than the minimum fixed version.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-24498","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00456.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-02-17"},{"id":"CVE-2020-24500","cve":"CVE-2020-24500","aliases":[],"title":"Intel E810 Ethernet controller firmware (NVM < 1.4.1.13): Early-generation E810 firmware flaw (a second buffer overflow","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel E810 Ethernet controller firmware (NVM < 1.4.1.13)","year":"2020","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Early-generation E810 firmware flaw (a second buffer overflow reachable by a privileged user) resulting in denial of service on the adapter. Relevant to any E810 adapter still on shipping-era NVM, which is common when NIC firmware was never brought into the patch pipeline.","attack_vector":"Varies by issue - the unauthenticated variant is reachable from the network, the others need privileged host access.","remediation":"Fixed in the E810 NVM (adapter firmware) image. Deploy with Intel's NVM Update Utility, which needs a driver reload and a power cycle - not just a warm reboot - for the new image to take effect. Drain the node. Distinct from the ice driver updates: you need both, and they ship on different schedules. Target NVM 1.4.1.13 or later, but jump straight to a current image rather than the minimum fixed version.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-24500","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00456.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-02-17"},{"id":"CVE-2020-5973","cve":"CVE-2020-5973","aliases":[],"title":"NVIDIA vGPU software (guest kernel-mode driver + vGPU plugin): The vGPU plugin lets guest code reach privileged","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software (guest kernel-mode driver + vGPU plugin)","year":"2020","cvss_score":4.4,"severity":"medium","kev":false,"impact":"The vGPU plugin lets guest code reach privileged operations it should not have, which a tenant uses to deny service on the shared GPU. vGPU 8.x before 8.4, 9.x before 9.4, 10.x before 10.3.","attack_vector":"A user inside a guest VM with a vGPU, using the guest kernel-mode driver.","remediation":"Upgrade the vGPU Manager on the host and the vGPU guest driver inside each tenant VM to the fixed release. Host side is a node drain plus reboot; guest side is a per-VM driver install and reboot. Because the guest driver is inside tenant-controlled VMs, in a multi-tenant estate you cannot fully remediate the guest half yourself - the host-side upgrade is the control you own. No VBIOS flash.","references":["https://usn.ubuntu.com/4404-1/","https://usn.ubuntu.com/4404-2/","https://nvd.nist.gov/vuln/detail/CVE-2020-5973"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2020-06-30"},{"id":"CVE-2020-5982","cve":"CVE-2020-5982","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): No rate limiting in the kernel-mode scheduler: a local process floods","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2020","cvss_score":4.4,"severity":"medium","kev":false,"impact":"No rate limiting in the kernel-mode scheduler: a local process floods it with requests and starves the GPU. On a shared node, one noisy tenant is a denial-of-service tool against everything else on the card.","attack_vector":"Any local user or container with GPU access.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5982"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2020-10-02"},{"id":"CVE-2021-0197","cve":"CVE-2021-0197","aliases":[],"title":"Intel E810 Ethernet controller firmware (NVM): Firmware-level flaw in the E810 network controller allowing a privileged","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel E810 Ethernet controller firmware (NVM)","year":"2021","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Firmware-level flaw in the E810 network controller allowing a privileged local user to cause denial of service on the adapter. Individually minor; collectively they are the reason NIC firmware needs to be in your patch pipeline rather than frozen at whatever the OEM shipped.","attack_vector":"Privileged local access on the host.","remediation":"Fixed in the E810 NVM (adapter firmware) image. Deploy with Intel's NVM Update Utility, which needs a driver reload and a power cycle - not just a warm reboot - for the new image to take effect. Drain the node. Distinct from the ice driver updates: you need both, and they ship on different schedules.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-0197","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00554.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-11-17"},{"id":"CVE-2021-0198","cve":"CVE-2021-0198","aliases":[],"title":"Intel E810 Ethernet controller firmware (NVM): Firmware-level flaw in the E810 network controller allowing a privileged","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel E810 Ethernet controller firmware (NVM)","year":"2021","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Firmware-level flaw in the E810 network controller allowing a privileged local user to cause denial of service on the adapter. Individually minor; collectively they are the reason NIC firmware needs to be in your patch pipeline rather than frozen at whatever the OEM shipped.","attack_vector":"Privileged local access on the host.","remediation":"Fixed in the E810 NVM (adapter firmware) image. Deploy with Intel's NVM Update Utility, which needs a driver reload and a power cycle - not just a warm reboot - for the new image to take effect. Drain the node. Distinct from the ice driver updates: you need both, and they ship on different schedules.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-0198","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00554.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-11-17"},{"id":"CVE-2021-0199","cve":"CVE-2021-0199","aliases":[],"title":"Intel E810 Ethernet controller firmware (NVM): Firmware-level flaw in the E810 network controller allowing a privileged","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel E810 Ethernet controller firmware (NVM)","year":"2021","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Firmware-level flaw in the E810 network controller allowing a privileged local user to cause denial of service on the adapter. Individually minor; collectively they are the reason NIC firmware needs to be in your patch pipeline rather than frozen at whatever the OEM shipped.","attack_vector":"Privileged local access on the host.","remediation":"Fixed in the E810 NVM (adapter firmware) image. Deploy with Intel's NVM Update Utility, which needs a driver reload and a power cycle - not just a warm reboot - for the new image to take effect. Drain the node. Distinct from the ice driver updates: you need both, and they ship on different schedules.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-0199","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00554.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-11-17"},{"id":"CVE-2021-1103","cve":"CVE-2021-1103","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): Another guest-reachable NULL dereference in the vGPU plugin causing denial of","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Another guest-reachable NULL dereference in the vGPU plugin causing denial of service to co-tenants. vGPU 12.x before 12.3, 11.x before 11.5, 8.x before 8.8.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1103"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2021-07-21"},{"id":"CVE-2021-27363","cve":"CVE-2021-27363","aliases":[],"title":"Linux iSCSI: Kernel pointer leak - iscsi_transport handle exposed to unprivileged users via sysfs","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux iSCSI","year":"2021","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Kernel pointer leak - iscsi_transport handle exposed to unprivileged users via sysfs","attack_vector":"Local","remediation":"Data-plane: kernel patch, batch with a reboot window","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-27363"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2021-03-07"},{"id":"CVE-2021-33128","cve":"CVE-2021-33128","aliases":[],"title":"Intel E810 Ethernet controller firmware (NVM): Firmware-level flaw in the E810 network controller allowing a privileged","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel E810 Ethernet controller firmware (NVM)","year":"2021","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Firmware-level flaw in the E810 network controller allowing a privileged local user to cause denial of service on the adapter. Individually minor; collectively they are the reason NIC firmware needs to be in your patch pipeline rather than frozen at whatever the OEM shipped.","attack_vector":"Privileged local access on the host.","remediation":"Fixed in the E810 NVM (adapter firmware) image. Deploy with Intel's NVM Update Utility, which needs a driver reload and a power cycle - not just a warm reboot - for the new image to take effect. Drain the node. Distinct from the ice driver updates: you need both, and they ship on different schedules.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-33128","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00593.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2022-08-18"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:N/I:H/A:N","fleet":{"pain_class":"daemon-restart"},"id":"CVE-2021-38882","cve":"CVE-2021-38882","aliases":[],"title":"IBM Spectrum Scale file audit logging retention: A privileged administrator deletes audit records before their","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"IBM Spectrum Scale file audit logging retention","year":"2021","cvss_score":4.4,"severity":"medium","kev":false,"impact":"A privileged administrator deletes audit records before their retention period expires, so an insider with admin rights can erase evidence of their own access to tenant data.","attack_vector":"Administrative access to a Spectrum Scale 5.1.0 through 5.1.1.1 cluster with file audit logging configured.","remediation":"Upgrade to 5.1.1.2 or later. Independently, ship audit records off the cluster to append-only storage that cluster admins cannot reach, so the retention guarantee does not depend on the storage system policing its own administrators.","references":["https://www.ibm.com/support/pages/node/6516426","https://nvd.nist.gov/vuln/detail/CVE-2021-38882"],"status":"curated"},{"id":"CVE-2022-21894","cve":"CVE-2022-21894","aliases":["BlackLotus"],"title":"Windows Boot Manager: Secure Boot bypass exploited in the wild by the BlackLotus UEFI bootkit","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Windows Boot Manager","year":"2022","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Secure Boot bypass exploited in the wild by the BlackLotus UEFI bootkit; the bootkit survives OS reinstall and disk replacement because it lives in the ESP with a revoked-but-still-trusted bootloader","attack_vector":"Local, high privilege","remediation":"dbx revocation and the phased Microsoft boot-manager revocation rollout. Low CVSS badly understates it: the score reflects the local-privilege precondition, not the below-OS persistence that follows","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-21894"],"status":"curated","published":"2022-01-11"},{"id":"CVE-2022-28709","cve":"CVE-2022-28709","aliases":[],"title":"Intel E810 Ethernet controller firmware (NVM): Firmware-level flaw in the E810 network controller allowing a privileged","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel E810 Ethernet controller firmware (NVM)","year":"2022","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Firmware-level flaw in the E810 network controller allowing a privileged local user to cause denial of service on the adapter. Individually minor; collectively they are the reason NIC firmware needs to be in your patch pipeline rather than frozen at whatever the OEM shipped.","attack_vector":"Privileged local access on the host.","remediation":"Fixed in the E810 NVM (adapter firmware) image. Deploy with Intel's NVM Update Utility, which needs a driver reload and a power cycle - not just a warm reboot - for the new image to take effect. Drain the node. Distinct from the ice driver updates: you need both, and they ship on different schedules.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28709","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00593.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2022-08-18"},{"id":"CVE-2022-34667","cve":"CVE-2022-34667","aliases":[],"title":"NVIDIA CUDA Toolkit - cuobjdump: A stack-based buffer overflow on a malformed input file yields limited denial","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA CUDA Toolkit - cuobjdump","year":"2022","cvss_score":4.4,"severity":"medium","kev":false,"impact":"A stack-based buffer overflow on a malformed input file yields limited denial of service and data-integrity loss for the invoking user. The realistic exposure is your build and profiling pipeline, not your runtime fleet: anything that automatically disassembles third-party fatbins, vendor kernels or model artifacts is running this parser on attacker-influenced input.","attack_vector":"Local, and requires a user or an automated job to run cuobjdump over an attacker-supplied file. CI jobs that inspect third-party CUDA binaries are the usual path.","remediation":"Update the CUDA Toolkit package (bulletin 5373). Cost: effectively zero - userspace SDK only, no driver reload, no node drain, no running-job impact. Rebuild build/CI images and move on.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34667","https://github.com/NVIDIA/product-security/tree/main/2022/5373"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:L/A:L","cwe":["CWE-121"],"published":"2022-11-19"},{"id":"CVE-2022-34673","cve":"CVE-2022-34673","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): An out-of-bounds array access in nvidia.ko yields denial","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":4.4,"severity":"medium","kev":false,"impact":"An out-of-bounds array access in nvidia.ko yields denial of service, information disclosure or data tampering from an unprivileged local account. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5415. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34673","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:L/A:L","cwe":["CWE-190"],"fleet":{"pain_class":"node-drain"},"published":"2022-12-30"},{"id":"CVE-2022-42259","cve":"CVE-2022-42259","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): An integer overflow in nvidia.ko crashes the driver","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":4.4,"severity":"medium","kev":false,"impact":"An integer overflow in nvidia.ko crashes the driver and takes the node's GPUs with it. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5415. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42259","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:L/A:L","cwe":["CWE-190"],"fleet":{"pain_class":"node-drain"},"published":"2022-12-30"},{"cwe":["CWE-404"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-48916","cve":"CVE-2022-48916","aliases":[],"title":"Linux kernel (drivers/iommu/intel): On VT-d scalable mode with VMD enabled, RID2PASID setup fails for devices behind","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/intel)","year":"2022","cvss_score":4.4,"severity":"medium","kev":false,"impact":"On VT-d scalable mode with VMD enabled, RID2PASID setup fails for devices behind the VMD bridge, the device fails to join its IOMMU group, and the failure path adds the device info to a list it is already on. The result is a kernel BUG and a panic during boot or PCI enumeration. Practically this means a node configured for both NVMe VMD and the scalable-mode IOMMU that passthrough needs does not stay up - and scalable mode is exactly what you enable to do PASID-based device assignment.","attack_vector":"Not tenant-reachable. Triggered by host configuration: Intel VT-d in scalable mode plus VMD enabled in BIOS, on Sapphire Rapids class platforms. The panic happens during device enumeration, so it hits at boot or on PCI rescan, and needs host/firmware-level access to set up. Worth tracking because it is the combination a GPU node with NVMe VMD and PASID passthrough naturally lands on.","remediation":"Fixed in 5.13 / 5.14 per this record; run a kernel at or beyond those on Intel nodes, or apply the linked stable commits. Interim: do not enable VMD and IOMMU scalable mode together on the same host - disable VMD in BIOS on nodes that need scalable-mode passthrough.","references":["https://git.kernel.org/stable/c/2aaa085bd012a83be7104356301828585a2253ed","https://git.kernel.org/stable/c/d5ad4214d9c6c6e465c192789020a091282dfee7","https://nvd.nist.gov/vuln/detail/CVE-2022-48916"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2023-0193","cve":"CVE-2023-0193","aliases":[],"title":"CUDA Toolkit: DoS / info disclosure (buffer over-read in cuobjdump)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2023","cvss_score":4.4,"severity":"medium","kev":false,"impact":"DoS / info disclosure (buffer over-read in cuobjdump)","attack_vector":"Malicious cubin/model artifact fed to build tooling","remediation":"Bump CUDA Toolkit in base images; rebuild and republish tenant base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0193","https://github.com/NVIDIA/product-security/tree/main/2023/5446"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:L/I:N/A:L","cwe":["CWE-125"],"published":"2023-03-10"},{"id":"CVE-2023-25520","cve":"CVE-2023-25520","aliases":[],"title":"Jetson TX2 / AGX Xavier (nvbootctrl): DoS via invalid boot config","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Jetson TX2 / AGX Xavier (nvbootctrl)","year":"2023","cvss_score":4.4,"severity":"medium","kev":false,"impact":"DoS via invalid boot config","attack_vector":"Local privileged attacker","remediation":"Flash JetPack 32.7.4+","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5466/5466.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-20"],"published":"2023-06-23"},{"id":"CVE-2023-27593","cve":"CVE-2023-27593","aliases":[],"title":"Cilium: Agent pod hostPath allows writing to /opt/cni/bin, replacing the CNI binary on the host","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2023","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Agent pod hostPath allows writing to /opt/cni/bin, replacing the CNI binary on the host","attack_vector":"Attacker with access to a Cilium agent pod","remediation":"Rolling Cilium upgrade; restrict who can exec into kube-system","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-27593"],"status":"curated","published":"2023-03-17"},{"id":"CVE-2023-31356","cve":"CVE-2023-31356","aliases":[],"title":"AMD SEV firmware - incomplete memory cleanup (AMD-SB-3003): Incomplete memory cleanup in the SEV firmware allows","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV firmware - incomplete memory cleanup (AMD-SB-3003)","year":"2023","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Incomplete memory cleanup in the SEV firmware allows corruption of guest private memory. The recurring pattern in this database - firmware that does not scrub or fully release state between uses - applied to the memory of confidential guests, where the whole product claim is that nobody but the guest can touch it.","attack_vector":"Local, privileged, on a host running SEV guests.","remediation":"Fixed in AMD PI/AGESA firmware and delivered only as an OEM SBIOS package - AMD ships the PI drop to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before releasing BIOS. **Budget one to six months of OEM lag**, and note that several CVEs in this batch are marked 'no fix planned' on Naples (EPYC 7001) - for those the only remediation is retiring the hardware. Applying it means cordon, drain and a full power cycle per node; there is no driver reload, no live patch and no VBIOS step. Because this touches the SEV-SNP trust boundary, the update moves the platform TCB version: refresh VCEK certificates from AMD's KDS and update tenant attestation policy, or confidential guest launches will fail immediately after the BIOS lands.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31356","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3003.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"tags":["tenant-isolation"],"published":"2024-08-13"},{"id":"CVE-2023-49100","cve":"CVE-2023-49100","aliases":["TFV-11","SDEI interrupt bind out-of-bounds read"],"title":"Arm Trusted Firmware-A before v2.10, SDEI service (sdei_interrupt_bind SMC handler): An SMC argument from the normal","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arm Trusted Firmware-A before v2.10, SDEI service (sdei_interrupt_bind SMC handler)","year":"2023","cvss_score":4.4,"severity":"medium","kev":false,"impact":"An SMC argument from the normal world reaches plat_ic_get_interrupt_type without adequate validation, giving an out-of-bounds read inside EL3. It is a small primitive on its own, but EL3 is the most privileged code on the part - it owns the secure world, the PSCI power state machine and the root of trust - so any memory-safety defect there is a stepping stone toward full platform compromise from a host kernel that is otherwise contained.","attack_vector":"Host kernel or hypervisor code (EL1/EL2) issuing SDEI SMCs. A guest cannot reach it directly unless the hypervisor forwards SDEI, so the realistic path is a tenant who has already got kernel on the host, or a compromised host agent.","remediation":"Upgrade to TF-A v2.10 or later with the TFV-11 fix, via an OEM platform firmware build. Flash + reboot + drain. If SDEI is not used by your platform, having it compiled out of BL31 is the cleaner answer - ask the OEM whether it is even enabled before assuming you are exposed.","references":["https://trustedfirmware-a.readthedocs.io/en/latest/security_advisories/security-advisory-tfv-11.html","https://nvd.nist.gov/vuln/detail/CVE-2023-49100"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-02-21"},{"cwe":["CWE-476","CWE-252"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-52570","cve":"CVE-2023-52570","aliases":[],"title":"Linux kernel (drivers/vfio/mdev): If creating an mdev type's sysfs entries partially fails, the parent still registers","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/vfio/mdev)","year":"2023","cvss_score":4.4,"severity":"medium","kev":false,"impact":"If creating an mdev type's sysfs entries partially fails, the parent still registers as successful, and later unregistration walks uninitialized type pointers and dereferences garbage. mdev is the mechanism vGPU partitioning is built on, so this sits on the code that carves one physical GPU into per-tenant devices.","attack_vector":"Needs host root: the failure is reached by an allocation failing during mdev parent registration at module load, then unloading the module. Upstream found it with fault injection. Not tenant-reachable - flagged because it is on the mdev/vGPU registration path that defines tenant device boundaries, not because a tenant can drive it.","remediation":"The record lists no fixed release; boot a kernel carrying the stable fix commits below. Interim control: do not load/unload mdev parent drivers on a node carrying live tenants.","references":["https://git.kernel.org/stable/c/c01b2e0ee22ef8b4dd7509a93aecc0ac0826bae4","https://git.kernel.org/stable/c/52093779b1830ac184a23848d971f06404cf513e","https://nvd.nist.gov/vuln/detail/CVE-2023-52570"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-0110","cve":"CVE-2024-0110","aliases":[],"title":"CUDA Toolkit: OOB write","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2024","cvss_score":4.4,"severity":"medium","kev":false,"impact":"OOB write -> possible code exec in build tooling","attack_vector":"Malicious cubin/model artifact","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0110","https://github.com/NVIDIA/product-security/tree/main/2024/5564"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:L/I:L/A:N","cwe":["CWE-787"],"published":"2024-08-31"},{"id":"CVE-2024-0131","cve":"CVE-2024-0131","aliases":[],"title":"GPU Display Driver: Info disclosure (OOB kernel memory access)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2024","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Info disclosure (OOB kernel memory access)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; rolling reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0131","https://github.com/NVIDIA/product-security/tree/main/2025/5614"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-805"],"fleet":{"pain_class":"node-reboot"},"published":"2025-02-02"},{"id":"CVE-2024-0139","cve":"CVE-2024-0139","aliases":[],"title":"Base Command Manager (Linux): Local privesc via insecure temporary file handling","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Base Command Manager (Linux)","year":"2024","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Local privesc via insecure temporary file handling","attack_vector":"Local user on a managed node","remediation":"Patch Base Command Manager on head and compute nodes","references":["https://github.com/NVIDIA/product-security/tree/main/2024/5600","https://nvd.nist.gov/vuln/detail/CVE-2024-0139"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:R/S:U/C:N/I:N/A:H","cwe":["CWE-377"],"published":"2024-12-06"},{"id":"CVE-2024-21970","cve":"CVE-2024-21970","aliases":[],"title":"AMD Power Management Firmware (SMU) - array index validation: An unvalidated array index in AMD's power management","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Power Management Firmware (SMU) - array index validation","year":"2024","cvss_score":4.4,"severity":"medium","kev":false,"impact":"An unvalidated array index in AMD's power management firmware lets a privileged attacker corrupt AGESA memory. The SMU is a live, always-running microcontroller with broad platform reach, so memory corruption there is an integrity problem for the whole node's power and clock management - including, on GPU nodes, the mechanisms that keep accelerators inside their thermal envelope.","attack_vector":"Local, privileged. Reachable through the SMU mailbox interface, which requires root.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. SMU firmware ships inside the SBIOS/AGESA bundle - there is no separate SMU update channel you can drive yourself.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21970","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2025-09-06"},{"id":"CVE-2024-28085","cve":"CVE-2024-28085","aliases":[],"title":"util-linux (wall): WallEscape: escape-sequence injection via wall(1)","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"util-linux (wall)","year":"2024","cvss_score":4.4,"severity":"medium","kev":false,"impact":"WallEscape: escape-sequence injection via wall(1) - can spoof a sudo prompt and steal a password on a shared host","attack_vector":"Local user on a shared login/head node","remediation":"Package update; no reboot. Only matters where multiple tenants share a login node","references":["https://access.redhat.com/security/cve/CVE-2024-28085"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2024-03-27"},{"id":"CVE-2024-51741","cve":"CVE-2024-51741","aliases":[],"title":"Redis: Malformed ACL selector triggers a server panic","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Redis","year":"2024","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Malformed ACL selector triggers a server panic -> denial of service","attack_vector":"Local","remediation":"Control-plane: upgrade to 7.2.7/7.4.2","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-51741"],"status":"curated","published":"2025-01-06"},{"id":"CVE-2025-1118","cve":"CVE-2025-1118","aliases":["GRUB2 2025 batch"],"title":"GRUB2 (dump command lockdown): The dump command was not disabled under Secure Boot lockdown, letting a privileged user","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (dump command lockdown)","year":"2025","cvss_score":4.4,"severity":"medium","kev":false,"impact":"The dump command was not disabled under Secure Boot lockdown, letting a privileged user read arbitrary memory at boot. That includes anything the firmware left in RAM - most usefully, key material and Secure Boot state. Low CVSS, high value as a reconnaissance primitive before a real bypass.","attack_vector":"Local privileged user at the GRUB shell.","remediation":"grub2 package update + reboot. A GRUB password limits access to the shell in the meantime, but does not fix the lockdown gap.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-1118","https://www.openwall.com/lists/oss-security/2025/02/18/3"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-02-19"},{"id":"CVE-2025-12896","cve":"CVE-2025-12896","aliases":["CVE-2025-12902","Solidigm locked-drive bypass"],"title":"Solidigm DC SSD firmware - unauthorized access to a LOCKED storage device via improper resource management: An attacker","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Solidigm DC SSD firmware - unauthorized access to a LOCKED storage device via improper resource management","year":"2025","cvss_score":4.4,"severity":"medium","kev":false,"impact":"An attacker with local or physical access gains unauthorized access to a drive that is in the locked state; the paired issue additionally allows denial of service. 'Locked' is the state your decommission and reclaim runbooks rely on - the drive is supposed to be an inert brick until someone presents the credential. BREAKS TENANT HANDOFF: a drive locked at the end of a tenancy, or locked before shipping back on RMA, can be read anyway. The DoS variant additionally gives a tenant a way to take a drive out from under the host, which on a shared bare-metal node is a self-service outage. These are the most recent public confirmations that the locked-drive guarantee keeps failing on datacenter NVMe - four separate CVEs across two advisory cycles on the same product family.","attack_vector":"A tenant with local (host-level, elevated) access on the bare-metal machine, or anyone with physical access to the drive after it leaves the rack - RMA courier, decommission handler, resale buyer.","remediation":"Firmware update from Solidigm's security page for the affected DC SKUs; drive offline, node drained, Solidigm Storage Tool per SKU. Check your own inventory against the advisory rather than assuming coverage, because Solidigm's DC line spans many families with different firmware trains and older SKUs in this same advisory family have previously been marked no-fix. The structural lesson for an operator: this is now the fourth-plus distinct 'locked drive is not actually locked' finding on datacenter NVMe in two years, so budget for it as a recurring class rather than a one-off patch - assume drive-level locking will fail again and keep a software encryption layer (LUKS/dm-crypt, operator-held key) as the control that actually enforces tenant separation.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-12896","https://nvd.nist.gov/vuln/detail/CVE-2025-12902","https://www.solidigm.com/support-page/support-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2025-11-07"},{"id":"CVE-2025-21590","cve":"CVE-2025-21590","aliases":[],"title":"Juniper Junos OS kernel: Improper isolation in the Junos kernel lets a local attacker with shell access inject","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Juniper Junos OS kernel","year":"2025","cvss_score":4.4,"severity":"medium","kev":true,"impact":"Improper isolation in the Junos kernel lets a local attacker with shell access inject arbitrary code — exploited in the wild to implant persistent backdoors on routers","attack_vector":"Local, shell access","remediation":"Junos upgrade fleet-wide with routing failover; the in-the-wild usage means affected devices need forensic verification, not just patching","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21590"],"status":"curated","published":"2025-03-12"},{"cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-22031","cve":"CVE-2025-22031","aliases":[],"title":"Linux kernel (drivers/pci/pcie): PCIe bandwidth control dereferences a bridge's subordinate bus pointer without","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/pci/pcie)","year":"2025","cvss_score":4.4,"severity":"medium","kev":false,"impact":"PCIe bandwidth control dereferences a bridge's subordinate bus pointer without checking it, and that pointer stays NULL when the kernel runs out of bus numbers to assign. The node panics during PCI enumeration - it does not come up, so the capacity is simply gone rather than degraded.","attack_vector":"Not attacker-driven: the precondition is a firmware/topology condition - BIOS leaves bridges unnumbered and the kernel exhausts the bus-number space while fixing it, which happens on deep or dense bridge hierarchies. That makes it a fleet-availability item for GPU boxes with large PCIe switch fabrics or many hotplug-capable ports, and for any node whose BIOS was just updated. No tenant reachability; include it in kernel-baseline hygiene rather than in the tenant threat model.","remediation":"Update to a kernel carrying the fix (no fixed_in published; stable commits below). Interim: check dmesg on new or re-flashed hardware for 'cannot be assigned' bus-number messages before putting a node into service, and update BIOS so bridges are correctly numbered.","references":["https://git.kernel.org/stable/c/d93d309013e89631630a12b1770d27e4be78362a","https://git.kernel.org/stable/c/1181924af78e5299ddec6e457789c02dd5966559","https://nvd.nist.gov/vuln/detail/CVE-2025-22031"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-23247","cve":"CVE-2025-23247","aliases":[],"title":"CUDA Toolkit: DoS via malformed string parsing","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2025","cvss_score":4.4,"severity":"medium","kev":false,"impact":"DoS via malformed string parsing","attack_vector":"Malicious binary artifact","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23247","https://github.com/NVIDIA/product-security/tree/main/2025/5643"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:L/I:L/A:N","cwe":["CWE-130"],"published":"2025-05-27"},{"id":"CVE-2025-23286","cve":"CVE-2025-23286","aliases":[],"title":"GPU Display Driver: Info disclosure (buffer over-read)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2025","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Info disclosure (buffer over-read)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; rolling reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23286","https://github.com/NVIDIA/product-security/tree/main/2025/5670"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-125"],"fleet":{"pain_class":"node-reboot"},"published":"2025-08-02"},{"id":"CVE-2025-23335","cve":"CVE-2025-23335","aliases":[],"title":"NVIDIA Triton Inference Server: A specific model configuration plus a specific input causes an underflow","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2025","cvss_score":4.4,"severity":"medium","kev":false,"impact":"A specific model configuration plus a specific input causes an underflow in the TensorRT backend and kills the server. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5687. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23335","https://github.com/NVIDIA/product-security/tree/main/2025/5687"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:H/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-191"],"published":"2025-08-06"},{"id":"CVE-2025-23336","cve":"CVE-2025-23336","aliases":[],"title":"NVIDIA Triton Inference Server: Loading a misconfigured model causes a denial of service","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2025","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Loading a misconfigured model causes a denial of service - relevant where tenants can register their own models. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5691. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23336","https://github.com/NVIDIA/product-security/tree/main/2025/5691"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:H/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-20"],"published":"2025-09-17"},{"id":"CVE-2025-23345","cve":"CVE-2025-23345","aliases":[],"title":"NVIDIA GPU Display Driver - kernel mode layer (Windows nvlddmkm.sys and Linux nvidia.ko): An out-of-bounds read in the","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - kernel mode layer (Windows nvlddmkm.sys and Linux nvidia.ko)","year":"2025","cvss_score":4.4,"severity":"medium","kev":false,"impact":"An out-of-bounds read in the driver's video decoder leaks memory or crashes the node. Relevant to transcode and video-analytics fleets that feed tenant-supplied streams straight into NVDEC. Both the Windows and Linux datacenter drivers are affected, so a mixed fleet needs two separate rollouts.","attack_vector":"Local and unprivileged on either OS. On Linux it is reachable from any GPU container via /dev/nvidia*; on Windows from any session holding a GPU handle.","remediation":"Upgrade both the Linux and the Windows datacenter driver branches listed in bulletin 5703. Cost: Linux needs a drain and nvidia.ko reload per node; Windows needs a reboot per node. Two change windows unless your fleet is homogeneous.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23345","https://github.com/NVIDIA/product-security/tree/main/2025/5703"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:L/A:L","cwe":["CWE-125"],"fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-10-23"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-284"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2025-24313","cve":"CVE-2025-24313","aliases":["INTEL-SA-01329"],"title":"Intel Device Plugins for Kubernetes (GPU/accelerator device plugin, access control): Improper access control in Intel's","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Intel Device Plugins for Kubernetes (GPU/accelerator device plugin, access control)","year":"2025","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Improper access control in Intel's Kubernetes device plugins lets a privileged local user deny service on the node. Because the device plugin is what advertises and allocates accelerators to the kubelet, knocking it over means the node stops offering its accelerators and scheduled workloads lose their allocation.","attack_vector":"A privileged user with local access to a node running Intel Device Plugins for Kubernetes before 0.32.0.","remediation":"Upgrade the Intel device plugin DaemonSet to 0.32.0 or later and let it roll across the nodes. The DaemonSet restart is enough - no drain or reboot - but confirm accelerator capacity re-registers on each node after the rollout.","references":["https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01329.html","https://nvd.nist.gov/vuln/detail/CVE-2025-24313"],"status":"curated"},{"id":"CVE-2025-33195","cve":"CVE-2025-33195","aliases":[],"title":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware: Unexpected memory buffer operations in SROOT firmware","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware","year":"2025","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Unexpected memory buffer operations in SROOT firmware reach data tampering and privilege escalation. These sit in the GB10 root-of-trust chain, so a successful exploit undermines the platform's own attestation and secure-boot story rather than just the OS above it.","attack_vector":"Local access to the DGX Spark. Several of the set need no privileges at all; the rest need host root. This is a desk-side developer box, so physical and local access assumptions are much weaker than for a racked DGX.","remediation":"Apply the DGX Spark firmware update from bulletin 5720. Cost: flash plus reboot, low drain cost given the form factor, but not live-patchable and root-of-trust firmware cannot be rolled back once applied.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33195","https://github.com/NVIDIA/product-security/tree/main/2025/5720"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:L/A:L","cwe":["CWE-119"],"fleet":{"pain_class":"firmware-flash"},"published":"2025-11-25"},{"id":"CVE-2025-33196","cve":"CVE-2025-33196","aliases":[],"title":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware: SROOT firmware reuses a resource without clearing","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware","year":"2025","cvss_score":4.4,"severity":"medium","kev":false,"impact":"SROOT firmware reuses a resource without clearing it, leaking its previous contents to the next consumer - residual-data exposure inside the root of trust. These sit in the GB10 root-of-trust chain, so a successful exploit undermines the platform's own attestation and secure-boot story rather than just the OS above it.","attack_vector":"Local access to the DGX Spark. Several of the set need no privileges at all; the rest need host root. This is a desk-side developer box, so physical and local access assumptions are much weaker than for a racked DGX.","remediation":"Apply the DGX Spark firmware update from bulletin 5720. Cost: flash plus reboot, low drain cost given the form factor, but not live-patchable and root-of-trust firmware cannot be rolled back once applied.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33196","https://github.com/NVIDIA/product-security/tree/main/2025/5720"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-226"],"fleet":{"pain_class":"firmware-flash"},"published":"2025-11-25"},{"id":"CVE-2025-33221","cve":"CVE-2025-33221","aliases":[],"title":"GPU Display Driver: DoS (input validation)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2025","cvss_score":4.4,"severity":"medium","kev":false,"impact":"DoS (input validation)","attack_vector":"Privileged local user","remediation":"Driver upgrade; rolling reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33221","https://github.com/NVIDIA/product-security/tree/main/2026/5821"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-20"],"fleet":{"pain_class":"node-reboot"},"published":"2026-05-26"},{"cwe":["CWE-476","CWE-252"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-71313","cve":"CVE-2025-71313","aliases":[],"title":"Linux kernel (drivers/pci/endpoint/functions): The NTB endpoint function drivers never checked whether their workqueue","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/pci/endpoint/functions)","year":"2025","cvss_score":4.4,"severity":"medium","kev":false,"impact":"The NTB endpoint function drivers never checked whether their workqueue was actually created, so a failed allocation leaves a NULL pointer that is later handed to queue_work() during endpoint controller init - a NULL dereference that panics the machine at link-up time rather than failing the bind cleanly.","attack_vector":"Endpoint mode with the NTB (non-transparent bridge) function driver bound, plus an allocation failure at bind time - so realistically a memory-pressured endpoint device, not a targeted attack. The dereference itself lands in epf_ntb_epc_init(), which runs when the connected host brings the link up. Inert on a conventional GPU server; relevant to NTB-based host-to-host interconnect hardware.","remediation":"Update to a kernel carrying the fix (no fixed_in published; stable commits below). Interim: do not bind the NTB endpoint function drivers on memory-constrained endpoint devices, and keep headroom so the allocation succeeds.","references":["https://git.kernel.org/stable/c/314eab6740bcda504ef978be599f805de05ce6de","https://git.kernel.org/stable/c/03f336a869b3a3f119d3ae52ac9723739c7fb7b6","https://nvd.nist.gov/vuln/detail/CVE-2025-71313"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-47487","cve":"CVE-2026-47487","aliases":[],"title":"Triton Inference Server: Path traversal in model file operations","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2026","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Path traversal in model file operations","attack_vector":"Authenticated client","remediation":"Upgrade Triton; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47487","https://github.com/NVIDIA/product-security/tree/main/2026/5860"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:N/A:L","cwe":["CWE-22"],"published":"2026-08-04"},{"id":"CVE-2020-13788","cve":"CVE-2020-13788","aliases":[],"title":"Harbor: SSRF: a user who can edit projects scans the Harbor host's intranet","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Harbor","year":"2020","cvss_score":4.3,"severity":"medium","kev":false,"impact":"SSRF: a user who can edit projects scans the Harbor host's intranet","attack_vector":"Authenticated registry user","remediation":"Upgrade Harbor; egress-restrict the registry pod","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-13788"],"status":"curated","published":"2020-07-15"},{"id":"CVE-2020-8551","cve":"CVE-2020-8551","aliases":[],"title":"Kubernetes (kubelet): Kubelet API DoS, including via the unauthenticated read-only port","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kubelet)","year":"2020","cvss_score":4.3,"severity":"medium","kev":false,"impact":"Kubelet API DoS, including via the unauthenticated read-only port","attack_vector":"Any pod on the cluster network","remediation":"Rolling kubelet upgrade with node drain; disable the read-only port (10255)","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-8551"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2020-03-27"},{"id":"CVE-2021-3956","cve":"CVE-2021-3956","aliases":[],"title":"Lenovo XClarity Controller (LDAP mode): Read-only authentication bypass when XCC is in LDAP-only authentication mode","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lenovo XClarity Controller (LDAP mode)","year":"2021","cvss_score":4.3,"severity":"medium","kev":false,"impact":"Read-only authentication bypass when XCC is in LDAP-only authentication mode","attack_vector":"Network","remediation":"XCC firmware update; also affects the common neocloud pattern of centralising BMC auth in LDAP/AD","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-3956"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2022-05-18"},{"id":"CVE-2022-39229","cve":"CVE-2022-39229","aliases":[],"title":"Grafana: A user can block another user's login by registering their email address as a username","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Grafana","year":"2022","cvss_score":4.3,"severity":"medium","kev":false,"impact":"A user can block another user's login by registering their email address as a username","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; low severity but a real ops denial vector","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-39229"],"status":"curated","published":"2022-10-13"},{"id":"CVE-2022-41354","cve":"CVE-2022-41354","aliases":[],"title":"Argo CD: Unauthenticated attackers can enumerate existing applications","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2022","cvss_score":4.3,"severity":"medium","kev":false,"impact":"Unauthenticated attackers can enumerate existing applications","attack_vector":"Unauthenticated network reaching the Argo CD API","remediation":"Rolling Argo CD upgrade; put the API behind auth","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-41354"],"status":"curated","published":"2023-03-27"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:L","fleet":{"pain_class":"node-reboot"},"id":"CVE-2023-31042","cve":"CVE-2023-31042","aliases":[],"title":"Pure Storage FlashBlade object store protocol: An authenticated object-store user degrades both data access and","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Pure Storage FlashBlade object store protocol","year":"2023","cvss_score":4.3,"severity":"medium","kev":false,"impact":"An authenticated object-store user degrades both data access and replication for the whole array. One tenant's S3 client can therefore interrupt everyone else's reads and break the copy that was supposed to be the recovery point.","attack_vector":"Any authenticated client of the FlashBlade object store protocol - the same access a tenant needs to use their own bucket.","remediation":"Upgrade Purity//FB to the fixed release named in Pure's bulletin. Monitor replication lag as a signal while unpatched, since the replication impact is the part that quietly breaks recovery.","references":["https://support.purestorage.com/Employee_Handbooks/Technical_Services/PSIRT/Security_Bulletin_for_FlashBlade_Object_Store_Protocol_CVE-2023-31042","https://nvd.nist.gov/vuln/detail/CVE-2023-31042"],"status":"curated"},{"id":"CVE-2023-31203","cve":"CVE-2023-31203","aliases":[],"title":"OpenVINO Model Server: Input-validation flaw in OpenVINO Model Server reachable without authentication","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"OpenVINO Model Server","year":"2023","cvss_score":4.3,"severity":"medium","kev":false,"impact":"Input-validation flaw in OpenVINO Model Server reachable without authentication. Same exposure profile as the later Model Server issues - it sits on the inference request path.","attack_vector":"Anything that can reach the serving endpoint.","remediation":"Upgrade to the 2022.3 or later Model Server build shipped with OpenVINO 2023. Container update and rolling restart.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31203","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00901.html"],"status":"curated","fleet":{"pain_class":"daemon-restart"},"published":"2023-11-14"},{"id":"CVE-2024-29040","cve":"CVE-2024-29040","aliases":[],"title":"tpm2-tss (FAPI quote verification): The JSON quote info returned by Fapi_Quote accepts an arbitrary TPM2_GENERATED","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"tpm2-tss (FAPI quote verification)","year":"2024","cvss_score":4.3,"severity":"medium","kev":false,"impact":"The JSON quote info returned by Fapi_Quote accepts an arbitrary TPM2_GENERATED magic value, so a malicious device can hand back a quote that the library accepts but that the TPM never produced. That is attestation forgery: a compromised node convinces the verifier it booted a measured, clean image. For any operator selling verified bare metal or confidential GPU compute, this breaks the assertion the whole product rests on - and it breaks it silently, since a forged quote validates.","attack_vector":"A malicious or compromised endpoint being attested. The attacker is the machine claiming to be healthy, not a third party on the wire.","remediation":"Update tpm2-tss to 4.1.0 or later wherever your attestation verifier runs and restart the service - package-level, no firmware, no reboot. Then re-run attestation across the fleet, because any quote validated by the old library proves nothing. Worth auditing whether your verifier does its own magic-value check rather than trusting the library.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-29040","https://github.com/tpm2-software/tpm2-tss/security/advisories"],"status":"curated","fleet":{"pain_class":"daemon-restart"},"published":"2024-06-28"},{"id":"CVE-2024-8059","cve":"CVE-2024-8059","aliases":["LEN-172051"],"title":"Lenovo XClarity Controller (XCC) - audit log: When an account username is exactly 16 characters, XCC writes the IPMI","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lenovo XClarity Controller (XCC) - audit log","year":"2024","cvss_score":4.3,"severity":"medium","kev":false,"impact":"When an account username is exactly 16 characters, XCC writes the IPMI credentials into its own audit log entries in the clear. The consequence is that anyone allowed to read BMC audit logs - a broader set than anyone allowed to administer the BMC, since logs get shipped to SIEMs, ticketing systems and shared dashboards - picks up working IPMI credentials for the node. Those credentials give power control and boot manipulation. It is a small bug with an awkward blast radius, because log data usually flows to systems with much weaker access control than the BMC itself. Affects a long list of ThinkSystem and ThinkAgile models.","attack_vector":"Anyone with read access to XCC audit logs, or to whatever downstream system those logs are forwarded into. Requires that at least one account on the node has a 16-character username, which is common where naming conventions produce fixed-length service account names.","remediation":"Flash XCC to the per-model version in LEN-172051 - out-of-band, per-node, no host reboot and no drain. Two things the flash does not do, and you must: purge or re-scope any already-collected XCC audit logs sitting in your log pipeline, and rotate the IPMI credentials that were exposed. A quick config-only check meanwhile: look for 16-character usernames across the fleet, since only those trigger the leak.","references":["https://support.lenovo.com/us/en/product_security/LEN-172051","https://nvd.nist.gov/vuln/detail/CVE-2024-8059"],"status":"curated","published":"2024-09-13"},{"id":"CVE-2025-25012","cve":"CVE-2025-25012","aliases":[],"title":"Kibana: Open redirect leading to SSRF via a specially crafted URL","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Kibana","year":"2025","cvss_score":4.3,"severity":"medium","kev":false,"impact":"Open redirect leading to SSRF via a specially crafted URL","attack_vector":"Network (remote)","remediation":"Control-plane: Kibana upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-25012"],"status":"curated","published":"2025-06-25"},{"id":"CVE-2025-33197","cve":"CVE-2025-33197","aliases":[],"title":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware: A null-pointer dereference in SROOT firmware crashes","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware","year":"2025","cvss_score":4.3,"severity":"medium","kev":false,"impact":"A null-pointer dereference in SROOT firmware crashes the platform. These sit in the GB10 root-of-trust chain, so a successful exploit undermines the platform's own attestation and secure-boot story rather than just the OS above it.","attack_vector":"Local access to the DGX Spark. Several of the set need no privileges at all; the rest need host root. This is a desk-side developer box, so physical and local access assumptions are much weaker than for a racked DGX.","remediation":"Apply the DGX Spark firmware update from bulletin 5720. Cost: flash plus reboot, low drain cost given the form factor, but not live-patchable and root-of-trust firmware cannot be rolled back once applied.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33197","https://github.com/NVIDIA/product-security/tree/main/2025/5720"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:C/C:N/I:N/A:L","cwe":["CWE-476"],"fleet":{"pain_class":"firmware-flash"},"published":"2025-11-25"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:L/I:N/A:N","cwe":["CWE-209"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2026-22052","cve":"CVE-2026-22052","aliases":[],"title":"NetApp ONTAP S3 NAS bucket directory listing: An authenticated S3 user lists the contents of directories they have no","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"NetApp ONTAP S3 NAS bucket directory listing","year":"2026","cvss_score":4.3,"severity":"medium","kev":false,"impact":"An authenticated S3 user lists the contents of directories they have no rights to. Where ONTAP S3 is the object endpoint feeding a training pipeline, that exposes another tenant's dataset layout and object keys.","attack_vector":"Any authenticated S3 client of an ONTAP 9.12.1 or later system with S3 NAS buckets configured.","remediation":"Upgrade to the fixed ONTAP release. In the interim, avoid mapping S3 buckets onto NAS paths that are shared across tenants, and prefer per-tenant buckets rooted at separate volumes.","references":["https://security.netapp.com/advisory/NTAP-20260304-0001","https://nvd.nist.gov/vuln/detail/CVE-2026-22052"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2026-24176","cve":"CVE-2026-24176","aliases":[],"title":"KAI Scheduler: Improper access control in resource allocation (cross-tenant quota abuse)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"KAI Scheduler","year":"2026","cvss_score":4.3,"severity":"medium","kev":false,"impact":"Improper access control in resource allocation (cross-tenant quota abuse)","attack_vector":"Any tenant with cluster API access","remediation":"Upgrade the KAI Scheduler Helm chart; rolling control-plane update, no tenant eviction","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24176","https://github.com/NVIDIA/product-security/tree/main/2026/5818"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:L/A:N","cwe":["CWE-863"],"published":"2026-04-21"},{"id":"CVE-2026-24232","cve":"CVE-2026-24232","aliases":[],"title":"Transformers4Rec: Code exec via insecure deserialization in the pipeline","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Transformers4Rec","year":"2026","cvss_score":4.3,"severity":"medium","kev":false,"impact":"Code exec via insecure deserialization in the pipeline","attack_vector":"Malicious dataset/model","remediation":"Bump the package; rebuild recsys images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24232","https://github.com/NVIDIA/product-security/tree/main/2026/5869"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:C/C:N/I:N/A:L","cwe":["CWE-502"],"published":"2026-07-21"},{"id":"CVE-2026-24241","cve":"CVE-2026-24241","aliases":[],"title":"NVIDIA License System - Delegated Licensing Service (DLS): Improper authentication in the DLS lets an unauthenticated","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA License System - Delegated Licensing Service (DLS)","year":"2026","cvss_score":4.3,"severity":"medium","kev":false,"impact":"Improper authentication in the DLS lets an unauthenticated attacker on an adjacent network read information from the licensing service.","attack_vector":"Adjacent network, no privileges, no user interaction - the weakest precondition of the DLS set.","remediation":"Update the DLS appliance per bulletin 5789. Cost: appliance restart, no tenant impact.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24241","https://github.com/NVIDIA/product-security/tree/main/2026/5789"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:N/UI:N/S:U/C:L/I:N/A:N","cwe":["CWE-287"],"published":"2026-02-24"},{"id":"CVE-2026-39395","cve":"CVE-2026-39395","aliases":[],"title":"cosign / sigstore: verify-blob-attestation reports \"Verified OK\" for malformed or mismatched payloads","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"cosign / sigstore","year":"2026","cvss_score":4.3,"severity":"medium","kev":false,"impact":"verify-blob-attestation reports \"Verified OK\" for malformed or mismatched payloads","attack_vector":"Malicious artifact","remediation":"Upgrade cosign to 3.0.6/2.6.3+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-39395"],"status":"curated","published":"2026-04-07"},{"id":"CVE-2026-65920","cve":"CVE-2026-65920","aliases":[],"title":"diffusers (shard file loader): Path traversal in `_get_checkpoint_shard_files`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"diffusers (shard file loader)","year":"2026","cvss_score":4.3,"severity":"medium","kev":false,"impact":"Path traversal in `_get_checkpoint_shard_files`","attack_vector":"Customer-supplied sharded checkpoint (including safetensors shards)","remediation":"Upgrade past 0.39.0. Note safetensors' safety guarantee covers the tensor payload, not the shard-index filenames","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-65920"],"status":"curated","published":"2026-07-23"},{"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:L/A:N","fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2022-004-cilium-agent-network-policy-name","cve":null,"aliases":["GHSA-pfhr-pccp-hwmh"],"title":"Cilium agent (network policy namespace label selectors): A tenant chooses which network policy applies to their own","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium agent (network policy namespace label selectors)","year":"2022","cvss_score":4.3,"severity":"medium","kev":false,"impact":"A tenant chooses which network policy applies to their own pods. Where policies use namespace label selectors, an attacker with rights to deploy pods — directly or through a Deployment, DaemonSet or any higher-level controller — crafts extra pod labels so their pod is matched by a different, more permissive policy than the one intended for it. Pod-deploy rights are the baseline grant in any shared Kubernetes cluster, so this turns the standard tenant permission into a policy-selection primitive and undermines namespace-label-based separation, which is the usual way multi-tenant GPU clusters express 'tenant A cannot talk to tenant B'.","attack_vector":"Network / in-cluster. Requires Kubernetes pod-deploy rights in the cluster, plus policies (CiliumNetworkPolicy, CiliumClusterwideNetworkPolicy or standard NetworkPolicy) that select on namespace labels.","remediation":"Upgrade the Cilium agent to 1.10.14, 1.11.8, 1.12.1 or later and roll the DaemonSet. No workaround exists for affected versions. Independently, constrain what labels tenants can set with an admission policy so label-driven selection cannot be steered from a tenant workload manifest.","references":["https://github.com/cilium/cilium/security/advisories/GHSA-pfhr-pccp-hwmh"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2015-7267","cve":"CVE-2015-7267","aliases":["CVE-2015-7268","CVE-2015-7269","Hot Plug attack","Forced Restart attack","Hot Unplug attack"],"title":"Self-encrypting drives in TCG Opal / eDrive mode","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Self-encrypting drives in TCG Opal / eDrive mode - Samsung 850 Pro, Samsung PM851, Seagate ST500LT015, ST500LT025 on…","year":"2015","cvss_score":4.2,"severity":"medium","kev":false,"impact":"Three related bypasses that all exploit the same design assumption: an SED stays UNLOCKED once the platform has authenticated it, and only re-locks on a true power cycle. If the drive never sees power drop - hot-plugging a SATA drive while the machine sleeps, forcing a soft reset and booting an alternative OS, or attaching a second SATA connector and an alternate power source before pulling the data cable - the DEK stays live and the attacker reads plaintext with no credential. Included here because this is the CLASS of failure operators keep re-encountering: the newer Solidigm 'locked drive is not locked' CVEs are the same assumption failing again a decade later. BREAKS TENANT HANDOFF for any workflow that assumes a drive is safe because it is locked - locked-while-powered is not locked, and a running or sleeping node is a readable node.","attack_vector":"A physically proximate attacker with access to the running or sleeping machine and its drive cabling - a colo neighbour, a rack tech, or anyone who reaches a node between tenancies while it is still powered. No password required.","remediation":"Mostly a policy and platform-configuration fix rather than a drive flash - the drives behaved as the Opal model specified, so there is no universal firmware patch. Concretely: disable sleep/suspend states on bare-metal hosts so the drive is never in the powered-but-unattended state that makes hot-plug work, require full power cycles rather than soft resets in your reclaim path, and verify that the platform re-authenticates the drive after any reset. Between tenants, do not hand over a node that has merely been rebooted - power it fully down as part of reclaim. The durable answer is the same as everywhere else in this category: software encryption with an operator-held key, so that a drive left unlocked still yields only ciphertext to whoever gets to it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2015-7267","https://nvd.nist.gov/vuln/detail/CVE-2015-7268","https://nvd.nist.gov/vuln/detail/CVE-2015-7269"],"status":"curated","published":"2017-11-27"},{"id":"CVE-2018-12038","cve":"CVE-2018-12038","aliases":["Self-Encrypting Deception","VU#395981"],"title":"Samsung 840 EVO SSD - disk encryption key exposed through wear-levelled NAND and vendor-specific commands: The drive","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Samsung 840 EVO SSD - disk encryption key exposed through wear-levelled NAND and vendor-specific commands","year":"2018","cvss_score":4.2,"severity":"medium","kev":false,"impact":"The drive stores key material in ordinary wear-levelled flash. Because the FTL never overwrites in place, changing the password leaves the OLD key version sitting in a stale physical block that vendor-specific commands can still reach. So even a drive whose password was rotated - the exact hygiene step an operator performs between customers - still hands over a key that decrypts the previous tenant's data. BREAKS TENANT HANDOFF, and it breaks it in the specific way that defeats the mitigation most operators would reach for: rotating the credential is not enough, because the old key is physically still there.","attack_vector":"Anyone holding the physical drive who can issue vendor-specific (undocumented) ATA commands to the controller - the next bare-metal tenant, an RMA handler, or a buyer of decommissioned hardware. Requires no knowledge of any current or previous password.","remediation":"No firmware fix restores the guarantee, because the exposure is old key material already committed to NAND; flashing new firmware does not scrub blocks the FTL has retired. Physically destroy any 840 EVO that ever held tenant data. Going forward, do not use drive-managed encryption as the tenant boundary - use LUKS/dm-crypt with an operator-held key so that 'erase' is a key-deletion event in your KMS, not a request to the drive. This is also the canonical argument for why a wear-levelled device can never prove sanitization to you: the blocks that matter are the ones the drive has already hidden from the host address space.","references":["https://kb.cert.org/vuls/id/395981","https://nvd.nist.gov/vuln/detail/CVE-2018-12038","https://msrc.microsoft.com/update-guide/vulnerability/ADV180028","https://security.netapp.com/advisory/ntap-20181112-0001/"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2018-11-20"},{"id":"CVE-2022-0532","cve":"CVE-2022-0532","aliases":[],"title":"CRI-O: \"Safe\" sysctls applied to the host when a pod uses host IPC/network","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"CRI-O","year":"2022","cvss_score":4.2,"severity":"medium","kev":false,"impact":"\"Safe\" sysctls applied to the host when a pod uses host IPC/network","attack_vector":"Cluster user able to create a hostIPC pod","remediation":"Upgrade CRI-O; block hostIPC/hostNetwork for tenants","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-0532"],"status":"curated","published":"2022-02-09"},{"id":"CVE-2023-31014","cve":"CVE-2023-31014","aliases":[],"title":"GeForce NOW Android app: Info disclosure / code exec via implicit intent","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GeForce NOW Android app","year":"2023","cvss_score":4.2,"severity":"medium","kev":false,"impact":"Info disclosure / code exec via implicit intent","attack_vector":"Malicious app on the same device","remediation":"Not applicable to server fleets","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5476/5476.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:R/S:U/C:L/I:L/A:L","cwe":["CWE-927"],"published":"2023-09-20"},{"id":"CVE-2023-31031","cve":"CVE-2023-31031","aliases":[],"title":"DGX A100 SBIOS: Buffer overflow in SBIOS","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX A100 SBIOS","year":"2023","cvss_score":4.2,"severity":"medium","kev":false,"impact":"Buffer overflow in SBIOS","attack_vector":"Local operator","remediation":"Flash SBIOS 1.25+; node power cycle","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31031","https://github.com/NVIDIA/product-security/tree/main/2024/5513"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:L/I:L/A:L","cwe":["CWE-122"],"fleet":{"pain_class":"firmware-flash"},"published":"2024-01-12"},{"id":"CVE-2024-0104","cve":"CVE-2024-0104","aliases":[],"title":"Mellanox OS / MetroX / Onyx / Skyway: Improper access control on switch mgmt","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Mellanox OS / MetroX / Onyx / Skyway","year":"2024","cvss_score":4.2,"severity":"medium","kev":false,"impact":"Improper access control on switch mgmt","attack_vector":"Authenticated switch user","remediation":"Upgrade switch OS image","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0104","https://github.com/NVIDIA/product-security/tree/main/2024/5559"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:L/UI:N/S:U/C:L/I:L/A:N","cwe":["CWE-284"],"published":"2024-08-08"},{"id":"CVE-2024-29902","cve":"CVE-2024-29902","aliases":[],"title":"cosign / sigstore: Remote image with a malicious attachment DoSes the machine running cosign","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"cosign / sigstore","year":"2024","cvss_score":4.2,"severity":"medium","kev":false,"impact":"Remote image with a malicious attachment DoSes the machine running cosign","attack_vector":"Malicious image","remediation":"Upgrade cosign to 2.2.4+","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-29902"],"status":"curated","published":"2024-04-10"},{"id":"CVE-2024-29903","cve":"CVE-2024-29903","aliases":[],"title":"cosign / sigstore: Crafted software artifacts DoS the cosign host, affecting all colocated services","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"cosign / sigstore","year":"2024","cvss_score":4.2,"severity":"medium","kev":false,"impact":"Crafted software artifacts DoS the cosign host, affecting all colocated services","attack_vector":"Malicious artifact","remediation":"Upgrade cosign to 2.2.4+; isolate the verifier","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-29903"],"status":"curated","published":"2024-04-10"},{"id":"CVE-2024-45678","cve":"CVE-2024-45678","aliases":["EUCLEAK"],"title":"Infineon cryptographic library (ECDSA) in security microcontrollers: Electromagnetic side channel in Infineon's ECDSA","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Infineon cryptographic library (ECDSA) in security microcontrollers","year":"2024","cvss_score":4.2,"severity":"medium","kev":false,"impact":"Electromagnetic side channel in Infineon's ECDSA implementation allows secret key extraction from affected security chips. The library is used well beyond the headline YubiKey case, across Infineon security microcontrollers that also serve as TPMs and platform root-of-trust devices. Where such a chip anchors a fleet's attestation or admin authentication, a cloned credential is indistinguishable from the real one.","attack_vector":"Physical access plus specialised equipment and time with the device. In a datacenter this is not automatically out of scope: a colo cage, an RMA path, a decommissioning contractor, or remote-hands staff all supply that access, and hardware in transit is the classic exposure window.","remediation":"Chip firmware cannot be updated in the affected devices - the fix ships only in new hardware revisions. So the answer is inventory, then replacement or acceptance, plus rotating any credential the affected chip holds. For an operator, the practical control is chain-of-custody: tamper-evident sealing, tracked RMA handling, and never returning a root-of-trust device to a pool without re-provisioning.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-45678","https://ninjalab.io/eucleak/"],"status":"curated","published":"2024-09-03"},{"id":"CVE-2025-23275","cve":"CVE-2025-23275","aliases":[],"title":"CUDA Toolkit: Code exec potential (buffer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2025","cvss_score":4.2,"severity":"medium","kev":false,"impact":"Code exec potential (buffer overflow)","attack_vector":"Malicious artifact","remediation":"Bump CUDA Toolkit; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23275","https://github.com/NVIDIA/product-security/tree/main/2025/5661"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:R/S:U/C:L/I:L/A:L","cwe":["CWE-787"],"published":"2025-09-24"},{"id":"CVE-2025-23301","cve":"CVE-2025-23301","aliases":[],"title":"NVIDIA HGX / DGX (Hopper and Blackwell) - GPU VBIOS: A VBIOS misconfiguration lets an attacker set an unsafe GPU debug","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA HGX / DGX (Hopper and Blackwell) - GPU VBIOS","year":"2025","cvss_score":4.2,"severity":"medium","kev":false,"impact":"A VBIOS misconfiguration lets an attacker set an unsafe GPU debug access level. Debug access on a datacenter GPU is the setting that gates whether host-side tooling can inspect GPU internal state - raising it on a shared or confidential-computing node undercuts the platform's own claim that tenant state is opaque. NVIDIA scores the direct impact as denial of service with a changed scope, but the interesting property is that the security posture of the GPU is settable from software.","attack_vector":"Local, low privileges, with high attack complexity. The attacker needs code on the host, not inside a guest - so this matters most where you run tenant containers on bare metal rather than behind a hypervisor.","remediation":"Apply the VBIOS update in NVIDIA bulletin 5674. Cost: a VBIOS flash on an HGX baseboard covers all eight GPUs on the board and requires a full node drain plus power cycle - you cannot flash per-GPU while jobs run. Fold it into your next planned firmware-bundle window rather than treating it as a hotfix; NVIDIA ships HGX firmware as a bundle (VBIOS + NVSwitch + ERoT) and mixing versions is unsupported.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23301","https://github.com/NVIDIA/product-security/tree/main/2025/5674"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:C/C:N/I:L/A:L","cwe":["CWE-1244"],"fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2025-09-04"},{"id":"CVE-2025-23302","cve":"CVE-2025-23302","aliases":[],"title":"NVIDIA HGX / DGX (Hopper and Blackwell) - NVSwitch LS10 firmware: A misconfiguration of the LS10 NVSwitch lets an","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA HGX / DGX (Hopper and Blackwell) - NVSwitch LS10 firmware","year":"2025","cvss_score":4.2,"severity":"medium","kev":false,"impact":"A misconfiguration of the LS10 NVSwitch lets an attacker set an unsafe debug access level on the fabric switch itself. The NVSwitch is the component that enforces which GPU can address which peer over NVLink; anything that loosens its debug posture is a question mark over the multi-GPU partitioning that a shared HGX node depends on. Direct scored impact is denial of service with a changed scope.","attack_vector":"Local, low privileges, high complexity, from the host that manages the fabric. Fabric Manager runs here, so the practical prerequisite is code on the baseboard host - not inside a tenant VM.","remediation":"Apply the NVSwitch firmware update from bulletin 5674 as part of the HGX firmware bundle. Cost: full node drain and power cycle across the whole baseboard; NVLink topology is re-established by Fabric Manager on restart, so verify fabric health and NVLink link counts after the flash before returning the node to the pool.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23302","https://github.com/NVIDIA/product-security/tree/main/2025/5674"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:C/C:N/I:L/A:L","cwe":["CWE-1244"],"fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2025-09-04"},{"cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:L/UI:N/S:U/C:L/I:L/A:N","cwe":["CWE-863"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2025-43904","cve":"CVE-2025-43904","aliases":[],"title":"Slurm (slurmdbd accounting, Coordinator role): A Coordinator - the delegated role a site gives a team lead over their","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Slurm (slurmdbd accounting, Coordinator role)","year":"2025","cvss_score":4.2,"severity":"medium","kev":false,"impact":"A Coordinator - the delegated role a site gives a team lead over their own account - can promote a user to full Slurm Administrator. That collapses the boundary between per-tenant account admin and cluster-wide admin, which on a multi-tenant GPU cluster means one team's lead can rewrite every other tenant's limits, QOS and fairshare.","attack_vector":"A user holding Coordinator on any account in the accounting database. Coordinators are commonly handed out to per-project leads, so the population that can exploit this is larger than it sounds.","remediation":"Upgrade to Slurm 23.11.11, 24.05.8 or 24.11.5 and restart slurmdbd. Then audit the admin level of every user in sacctmgr and revoke any Administrator grant you did not issue yourself.","references":["https://lists.schedmd.com/mailman3/hyperkitty/list/slurm-announce@lists.schedmd.com/message/B73QHKW6TKE2T5KDWVPIWNE5H4KWX667/","https://nvd.nist.gov/vuln/detail/CVE-2025-43904"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:C/C:L/I:L/A:N","cwe":["CWE-863"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2025-66433","cve":"CVE-2025-66433","aliases":["HTCONDOR-2025-0002"],"title":"HTCondor (condor_schedd / Access Point): A user plants a specially crafted job that lies dormant, then runs as a","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HTCondor (condor_schedd / Access Point)","year":"2025","cvss_score":4.2,"severity":"medium","kev":false,"impact":"A user plants a specially crafted job that lies dormant, then runs as a different non-root user of the attacker's choosing once the Access Point is upgraded to an affected version. The upgrade itself is the trigger, which makes this an unusually nasty one to reason about - the exploit is armed before you install the vulnerable code.","attack_vector":"A user with WRITE access to the schedd, i.e. anyone allowed to submit jobs, on an AP running 24.7.3 or later.","remediation":"Upgrade to HTCondor 24.12.14, 25.0.3 or 25.3.1. Before and after the upgrade, hunt for pre-planted jobs with: condor_q -all -constraint 'OsUser != Owner' and condor_rm anything suspicious. An AP already running a vulnerable version cannot have a new attack initiated against it, so the priority order is: scan the queue, then upgrade.","references":["https://htcondor.org/security/vulnerabilities/HTCONDOR-2025-0002.html","https://nvd.nist.gov/vuln/detail/CVE-2025-66433"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2020-8561","cve":"CVE-2020-8561","aliases":[],"title":"Kubernetes (kube-apiserver): Admission webhook responses redirect apiserver requests into private networks","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-apiserver)","year":"2020","cvss_score":4.1,"severity":"medium","kev":false,"impact":"Admission webhook responses redirect apiserver requests into private networks; SSRF","attack_vector":"Whoever controls a registered webhook backend","remediation":"Rolling control-plane upgrade; audit webhook configurations","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-8561"],"status":"curated","published":"2021-09-20"},{"id":"CVE-2021-1088","cve":"CVE-2021-1088","aliases":[],"title":"NVIDIA GPU firmware microcontroller (Falcon): Debug mechanisms in the GPU's internal microcontroller are reachable with","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU firmware microcontroller (Falcon)","year":"2021","cvss_score":4.1,"severity":"medium","kev":false,"impact":"Debug mechanisms in the GPU's internal microcontroller are reachable with insufficient access control, leaking information from below the driver. The microcontroller sits underneath every context on the card, so what leaks is not bounded by process or VM. Applies across the datacenter line including DGX-1, DGX-2 and DGX Station A100.","attack_vector":"A user with elevated privileges on the host. In a bare-metal GPU-rental model that is the tenant themselves, which makes 'requires root' a much weaker precondition than it sounds.","remediation":"NVIDIA shipped the fix in GPU firmware/microcode delivered with the R470 and R450 driver branches and, on some SKUs, in an updated VBIOS. On most datacenter parts the microcontroller image is loaded by the driver at GPU init, so a driver upgrade plus a node reboot applies it; check the bulletin's product table, because a subset of boards also needs an out-of-band VBIOS/InfoROM update, which is an offline per-node flash with the GPU idle. Either way the node has to be drained.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1088"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2021-11-20"},{"id":"CVE-2021-1105","cve":"CVE-2021-1105","aliases":[],"title":"NVIDIA GPU firmware microcontroller (Falcon): Debug registers on the GPU's internal microcontroller are readable at","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU firmware microcontroller (Falcon)","year":"2021","cvss_score":4.1,"severity":"medium","kev":false,"impact":"Debug registers on the GPU's internal microcontroller are readable at runtime, leaking state from beneath the driver. Listed across the DGX-1, DGX-2 and DGX Station A100 lines as well as the GeForce/Quadro/Tesla range.","attack_vector":"A user with elevated privileges on the GPU host - which in bare-metal GPU rental is the tenant.","remediation":"NVIDIA shipped the fix in GPU firmware/microcode delivered with the R470 and R450 driver branches and, on some SKUs, in an updated VBIOS. On most datacenter parts the microcontroller image is loaded by the driver at GPU init, so a driver upgrade plus a node reboot applies it; check the bulletin's product table, because a subset of boards also needs an out-of-band VBIOS/InfoROM update, which is an offline per-node flash with the GPU idle. Either way the node has to be drained.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1105"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2021-11-20"},{"id":"CVE-2021-1125","cve":"CVE-2021-1125","aliases":[],"title":"NVIDIA GPU firmware microcontroller (Falcon): Program data in the GPU's internal microcontroller can be corrupted by a","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU firmware microcontroller (Falcon)","year":"2021","cvss_score":4.1,"severity":"medium","kev":false,"impact":"Program data in the GPU's internal microcontroller can be corrupted by a privileged user. Corrupting firmware-level state affects the whole card, not one context, and the corruption is not visible to anything running above the driver. Listed for DGX-1, DGX-2 and DGX Station A100 among others.","attack_vector":"A user with elevated privileges on the GPU host - the tenant themselves on rented bare metal.","remediation":"NVIDIA shipped the fix in GPU firmware/microcode delivered with the R470 and R450 driver branches and, on some SKUs, in an updated VBIOS. On most datacenter parts the microcontroller image is loaded by the driver at GPU init, so a driver upgrade plus a node reboot applies it; check the bulletin's product table, because a subset of boards also needs an out-of-band VBIOS/InfoROM update, which is an offline per-node flash with the GPU idle. Either way the node has to be drained.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1125"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2021-11-20"},{"id":"CVE-2021-23219","cve":"CVE-2021-23219","aliases":[],"title":"NVIDIA GPU firmware microcontroller (Falcon): Loading crafted microcode gives a privileged user access to information","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU firmware microcontroller (Falcon)","year":"2021","cvss_score":4.1,"severity":"medium","kev":false,"impact":"Loading crafted microcode gives a privileged user access to information the microcontroller is supposed to protect. Same family as the other 2021 microcontroller bugs and listed for DGX-1, DGX-2 and DGX Station A100.","attack_vector":"A user with elevated privileges on the GPU host.","remediation":"NVIDIA shipped the fix in GPU firmware/microcode delivered with the R470 and R450 driver branches and, on some SKUs, in an updated VBIOS. On most datacenter parts the microcontroller image is loaded by the driver at GPU init, so a driver upgrade plus a node reboot applies it; check the bulletin's product table, because a subset of boards also needs an out-of-band VBIOS/InfoROM update, which is an offline per-node flash with the GPU idle. Either way the node has to be drained.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-23219"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2021-11-20"},{"id":"CVE-2021-34399","cve":"CVE-2021-34399","aliases":[],"title":"NVIDIA GPU firmware microcontroller (Falcon): Registers in the GPU's internal microcontroller are not scrubbed, so a","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU firmware microcontroller (Falcon)","year":"2021","cvss_score":4.1,"severity":"medium","kev":false,"impact":"Registers in the GPU's internal microcontroller are not scrubbed, so a privileged user reads leftover state from whatever ran before them. On a shared or sequentially rented GPU this is residue from the previous tenant's workload.","attack_vector":"A user with elevated privileges on the GPU host - the incoming tenant on rented bare metal.","remediation":"NVIDIA shipped the fix in GPU firmware/microcode delivered with the R470 and R450 driver branches and, on some SKUs, in an updated VBIOS. On most datacenter parts the microcontroller image is loaded by the driver at GPU init, so a driver upgrade plus a node reboot applies it; check the bulletin's product table, because a subset of boards also needs an out-of-band VBIOS/InfoROM update, which is an offline per-node flash with the GPU idle. Either way the node has to be drained.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-34399"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2021-11-20"},{"id":"CVE-2021-34400","cve":"CVE-2021-34400","aliases":[],"title":"NVIDIA GPU firmware microcontroller (Falcon): Unscrubbed microcontroller memory leaks data to a privileged user. Same","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU firmware microcontroller (Falcon)","year":"2021","cvss_score":4.1,"severity":"medium","kev":false,"impact":"Unscrubbed microcontroller memory leaks data to a privileged user. Same residue problem as the unscrubbed-registers issue but with a bigger surface, and it is exactly the failure mode that makes bare-metal GPU handoff between tenants risky: memory that should have been wiped at teardown is still readable by whoever gets the card next.","attack_vector":"A user with elevated privileges on the GPU host, typically the next tenant to be scheduled onto the card.","remediation":"NVIDIA shipped the fix in GPU firmware/microcode delivered with the R470 and R450 driver branches and, on some SKUs, in an updated VBIOS. On most datacenter parts the microcontroller image is loaded by the driver at GPU init, so a driver upgrade plus a node reboot applies it; check the bulletin's product table, because a subset of boards also needs an out-of-band VBIOS/InfoROM update, which is an offline per-node flash with the GPU idle. Either way the node has to be drained. On any fleet that reassigns bare-metal GPU nodes between tenants, pair the driver/firmware update with an explicit GPU reset and memory-scrub step in the reprovisioning pipeline; do not rely on the card clearing itself.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-34400"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2021-11-20"},{"id":"CVE-2022-28192","cve":"CVE-2022-28192","aliases":[],"title":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko): A use-after-free in the host vGPU","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko)","year":"2022","cvss_score":4.1,"severity":"medium","kev":false,"impact":"A use-after-free in the host vGPU Manager, reachable when host-side resources are freed out of sequence, crashes the hypervisor's GPU stack. The attacker sits inside a tenant guest VM and the blast radius is the hypervisor host, which is running every other tenant's vGPU on the same physical GPU. This is precisely the boundary a vGPU-based multi-tenant offering is sold on. Rated hard to exploit because it needs elevated control over freeing host resources, so treat it as a chain component rather than a standalone break.","attack_vector":"A tenant inside their own guest VM, driving the paravirtualised vGPU control interface. Several of these need only an unprivileged process in the guest; the rest need guest root, which a tenant already has on a VM they rented. No host credentials are involved at any point.","remediation":"Patch the vGPU Manager on the hypervisor host per bulletin 5353. Cost: the highest of any class here. The host driver cannot be reloaded while vGPUs are attached, so every tenant VM on that hypervisor must be live-migrated or powered off - a full host drain. NVIDIA also enforces a supported host/guest driver skew, so budget a matching guest-driver campaign in the same window or tenants lose their vGPU on next boot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28192","https://github.com/NVIDIA/product-security/tree/main/2022/5353"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-drain"},"tags":["tenant-isolation"],"published":"2022-05-17"},{"id":"CVE-2023-52862","cve":"CVE-2023-52862","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2023","cvss_score":4.1,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Fix null pointer dereference in error message","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52862","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-05-21"},{"id":"CVE-2024-0133","cve":"CVE-2024-0133","aliases":[],"title":"Container Toolkit: Unauthorized empty-file creation on the host","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Container Toolkit","year":"2024","cvss_score":4.1,"severity":"medium","kev":false,"impact":"Unauthorized empty-file creation on the host","attack_vector":"Any tenant with a crafted container image","remediation":"Bump toolkit + restart runtime","references":["https://github.com/NVIDIA/product-security/tree/main/2024/5582","https://nvd.nist.gov/vuln/detail/CVE-2024-0133"],"status":"curated","fleet":{"ubiquity":"Universal - same package, same install base","remediation_pain":"`daemon-restart` (1.16.2)","pain_class":"daemon-restart","why_fleet_wide":"Sibling TOCTOU allowing creation of arbitrary files on the host from inside a container; low severity alone but a stepping stone to host compromise on shared nodes"},"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:R/S:C/C:N/I:L/A:N","cwe":["CWE-367"],"published":"2024-09-26"},{"id":"CVE-2024-0134","cve":"CVE-2024-0134","aliases":[],"title":"Container Toolkit / GPU Operator: Unauthorized file creation on the host (data tampering)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Container Toolkit / GPU Operator","year":"2024","cvss_score":4.1,"severity":"medium","kev":false,"impact":"Unauthorized file creation on the host (data tampering)","attack_vector":"Any tenant with a crafted container image","remediation":"Bump toolkit; upgrade GPU Operator Helm chart","references":["https://github.com/NVIDIA/product-security/tree/main/2024/5585","https://nvd.nist.gov/vuln/detail/CVE-2024-0134"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:R/S:C/C:N/I:L/A:N","cwe":["CWE-61"],"published":"2024-11-05"},{"id":"CVE-2025-20044","cve":"CVE-2025-20044","aliases":[],"title":"Intel TDX module: The TDX module is the software that stands between the host/VMM and every confidential VM on the box","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel TDX module","year":"2025","cvss_score":4.1,"severity":"medium","kev":false,"impact":"The TDX module is the software that stands between the host/VMM and every confidential VM on the box; a privilege escalation inside it is a break of the boundary that separates a tenant's trust domain from the operator and from other TDs. Specific flaw: improper locking, letting a privileged host user escalate.","attack_vector":"A privileged user on the host - which in the TDX threat model is the adversary the whole design exists to exclude, so 'requires host privilege' is not a mitigating factor here.","remediation":"Update the Intel TDX module. The TDX module is loaded by the SEAM loader at boot, so the practical rollout is: stage the new module, drain every trust domain off the node, and reboot. It is not a live-patchable component and running TDs cannot be migrated through it. After the update, every TD must re-attest because the TDX module SVN is part of the attestation report - so anything that pinned the old measurement will fail until you update your attestation policy too. No OEM BIOS release needed for the module itself, which makes this materially faster than a platform firmware update.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-20044","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01245.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2025-08-12"},{"id":"CVE-2018-12037","cve":"CVE-2018-12037","aliases":["Self-Encrypting Deception","VU#395981","Radboud SED research"],"title":"Crucial/Micron MX100, MX200, MX300; Samsung 840 EVO and 850 EVO (ATA-high mode)","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Crucial/Micron MX100, MX200, MX300; Samsung 840 EVO and 850 EVO (ATA-high mode); Samsung T3 and T5 portable SSDs…","year":"2018","cvss_score":4,"severity":"medium","kev":false,"impact":"The drive's data-encryption key is not derived from the password at all. Set an ATA password or an Opal credential, and the drive still encrypts with a key that sits in the controller independent of that secret - so anyone who can talk to the firmware recovers the plaintext without ever knowing the password. BREAKS TENANT HANDOFF: an operator who 'securely wipes' a node by setting/resetting a drive password, or who relies on the SED to make old data unreadable, has done nothing. The next tenant who gets that bare-metal box, or anyone who receives the drive on an RMA or decommission pallet, reads the prior tenant's training data, checkpoints, SSH keys and cloud credentials in the clear. It also means your 'crypto erase' is not a crypto erase - the key you thought you threw away was never the key protecting the data.","attack_vector":"Anyone who obtains the drive and can issue vendor/debug commands to its controller: the next tenant on the same bare-metal host, an RMA return path, a decommissioned node in a resale channel, or a rack tech with five minutes and a SATA cable. No password, no user credential, and no network access needed - physical or low-level bus access to the drive is the whole requirement.","remediation":"Treat as UNPATCHABLE in practice on the affected SKUs - these are consumer/prosumer drives, several vendors never shipped a firmware fix, and the flaw is in how the key hierarchy was designed rather than a bounds check. The real fix is a policy change: stop trusting hardware encryption as your tenant-separation boundary and layer software encryption you control (LUKS/dm-crypt with a key held in your KMS, or per-tenant filesystem encryption) on top of every local NVMe/SATA device. Then 'crypto erase between tenants' means destroying a key you own, which is verifiable, instead of asking the drive to forget something. Also disable Windows eDrive/BitLocker hardware offload fleet-wide (see the ADV180028 entry). For drives already in service: re-provision with software encryption requires a full re-image and re-encrypt of every affected node, and any drive that held sensitive data before the policy change should be physically destroyed rather than resold, because you cannot retroactively prove it was erased.","references":["https://kb.cert.org/vuls/id/395981","https://nvd.nist.gov/vuln/detail/CVE-2018-12037","https://www.ru.nl/en/research/research-news/radboud-university-researchers-discover-security-flaw-in-ssd-hard-drives","https://www.dell.com/support/kbdoc/en-us/000139235/self-encrypting-drives-vulnerabilities-cve-2018-12037-and-cve-2018-12038-impact-on-dell-emc-server-dell-storage-networking-and-dell-clients","https://msrc.microsoft.com/update-guide/vulnerability/ADV180028"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"published":"2018-11-20"},{"id":"CVE-2021-26400","cve":"CVE-2021-26400","aliases":[],"title":"AMD processors - speculative reordering of loads on shared memory: AMD processors may speculatively reorder load","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD processors - speculative reordering of loads on shared memory","year":"2021","cvss_score":4,"severity":"medium","kev":false,"impact":"AMD processors may speculatively reorder load instructions such that stale data is observed when several processors operate on shared memory. Where that shared memory spans a trust boundary - a shared page between a guest and the host, or between containers - stale reads become a disclosure channel, and worse, code that relies on memory ordering for its own security checks can be made to see the wrong value.","attack_vector":"Local, requires shared memory between attacker and victim and multiple processors operating on it concurrently.","remediation":"Mitigated by AMD microcode plus, on most of these, a kernel-side change - and the durable delivery vehicle is the OEM SBIOS/AGESA package, which carries **one to six months of OEM lag** and needs a drained node and a full power cycle. The linux-firmware amd-ucode blobs get you the microcode sooner via initramfs early-load and a reboot, but AMD does not support late-loading microcode on a running EPYC host, so either way this is reboot-required, not a live patch.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26400","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2022-05-11"},{"cwe":["CWE-911"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2022-50505","cve":"CVE-2022-50505","aliases":[],"title":"Linux kernel (drivers/iommu/amd): The AMD-Vi PPR (peripheral page request) notifier looked up the faulting PCI device","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/iommu/amd)","year":"2022","cvss_score":4,"severity":"medium","kev":false,"impact":"The AMD-Vi PPR (peripheral page request) notifier looked up the faulting PCI device and never dropped the reference it took, so every page-request event a device generates permanently pins one more reference to that device. The device can then never be cleanly released back to the pool, and the refcount is driven by a tenant's own workload - reclaiming a GPU or accelerator after a tenant is done with it stops working.","attack_vector":"Device-driven: a tenant running an SVA/PRI workload on an AMD-Vi host makes the assigned device emit page requests, and each one leaks a pci_dev reference in host kernel context. No host privilege needed - the tenant just uses demand paging on its own assigned device. Conditional on AMD-Vi with the amd_iommu_v2 / PPR path active and a PRI-capable device assigned to the tenant.","remediation":"No fixed release is listed in this record; apply the linked stable commits or run a current stable/LTS kernel on AMD nodes. Interim: disable PRI/ATS on tenant-assigned devices where demand paging is not required, and watch for devices that refuse to unbind after a tenant releases them.","references":["https://git.kernel.org/stable/c/bdb2113dd8f17a3cc84a2b4be4968a849f69ec72","https://git.kernel.org/stable/c/efd50c65fd1cdef63eb58825f3fe72496443764c","https://nvd.nist.gov/vuln/detail/CVE-2022-50505"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2023-25524","cve":"CVE-2023-25524","aliases":[],"title":"Omniverse Launcher: Access-token disclosure","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Omniverse Launcher","year":"2023","cvss_score":4,"severity":"medium","kev":false,"impact":"Access-token disclosure -> user impersonation","attack_vector":"Local user / browser","remediation":"Upgrade Launcher; low relevance to headless DC fleets","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5472/5472.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:U/C:L/I:N/A:N","cwe":["CWE-598"],"published":"2023-08-03"},{"cwe":["CWE-401","CWE-911"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-35908","cve":"CVE-2024-35908","aliases":[],"title":"Linux kernel (net/tls): Tls_sw_recvmsg takes a psock reference before acquiring the reader lock and returns without","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/tls)","year":"2024","cvss_score":4,"severity":"medium","kev":false,"impact":"Tls_sw_recvmsg takes a psock reference before acquiring the reader lock and returns without dropping it if the lock fails, so every failed receive pins a psock forever. A tenant that loops interrupted receives on a kTLS+sockmap socket leaks kernel objects until the node runs out of memory.","attack_vector":"Local and unprivileged on the socket side - the reader lock fails on signal interruption or timeout, which the tenant controls. Requires the socket to also carry a BPF psock (sockmap), so it applies on nodes running a service mesh or CNI that combines sockmap with kTLS, not on plain kTLS sockets.","remediation":"Boot a kernel carrying the linked stable commits. Interim: do not run sockmap/sk_msg policy over kTLS sockets, and cap per-tenant memory so a leak is bounded by the cgroup rather than the node.","references":["https://git.kernel.org/stable/c/30fabe50a7ace3e9d57cf7f9288f33ea408491c8","https://git.kernel.org/stable/c/f1b7f14130d782433bc98c1e1e41ce6b4d4c3096","https://nvd.nist.gov/vuln/detail/CVE-2024-35908"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2024-47825","cve":"CVE-2024-47825","aliases":[],"title":"Cilium: Deny rules for prefixes broader than /32 can be ignored","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2024","cvss_score":4,"severity":"medium","kev":false,"impact":"Deny rules for prefixes broader than /32 can be ignored; egress restrictions silently fail","attack_vector":"Any tenant workload","remediation":"Rolling Cilium upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-47825"],"status":"curated","published":"2024-10-21"},{"cwe":["CWE-401"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2024-56745","cve":"CVE-2024-56745","aliases":[],"title":"Linux kernel (drivers/pci): Every write to a device's reset_method sysfs attribute that contains no space leaks the","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/pci)","year":"2024","cvss_score":4,"severity":"medium","kev":false,"impact":"Every write to a device's reset_method sysfs attribute that contains no space leaks the buffer it allocated, because strsep() has already nulled the pointer that the free uses. The leak is unbounded and driven entirely by the writer, so a loop over that attribute grows kernel memory until the node is under memory pressure - a slow noisy-neighbour outage on a shared host, on the knob that decides how a device is reset between tenants.","attack_vector":"Needs write access to /sys/bus/pci/devices/<dev>/reset_method, which is root-owned - so host root, or a privileged container with /sys mounted writable. Not reachable by a plain tenant container, not reachable from a guest, and not reachable over the fabric. The reason it is worth carrying is the surface rather than the severity: reset_method is the control that determines which reset a device gets when it is reclaimed from one tenant and handed to the next, and it should not be writable from anything a tenant runs.","remediation":"Boot a kernel where reset_method_store() iterates over a separate temporary pointer so the original allocation is still freed. Interim: ensure /sys is mounted read-only in containers and that no tenant workload runs with privileges to write PCI sysfs attributes.","references":["https://git.kernel.org/stable/c/403efb4457c0c8f8f51e904cc57d39193780c6bd","https://git.kernel.org/stable/c/931d07ccffcc3614f20aaf602b31e89754e21c59","https://nvd.nist.gov/vuln/detail/CVE-2024-56745"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-32793","cve":"CVE-2025-32793","aliases":[],"title":"Cilium: WireGuard encryption gap in a specific Cilium configuration","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2025","cvss_score":4,"severity":"medium","kev":false,"impact":"WireGuard encryption gap in a specific Cilium configuration","attack_vector":"Anyone on the underlay network","remediation":"Rolling Cilium upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-32793"],"status":"curated","published":"2025-04-21"},{"cwe":["CWE-404"],"fleet":{"pain_class":"node-reboot"},"id":"CVE-2025-38625","cve":"CVE-2025-38625","aliases":[],"title":"Linux kernel (drivers/vfio/pci/pds): The pds VFIO variant driver shipped without a detach_ioas operation, so it had no","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (drivers/vfio/pci/pds)","year":"2025","cvss_score":4,"severity":"medium","kev":false,"impact":"The pds VFIO variant driver shipped without a detach_ioas operation, so it had no way to detach the device from the iommufd address space it was attached to. VFIO core's registration check catches this and refuses to load the driver, which is the fail-closed outcome - but the hazard the check exists to prevent is precisely a passthrough device left attached to an IOAS after the tenant that owned it is gone, i.e. a DMA window that outlives its owner. As shipped, the operator-visible symptom is that pds passthrough simply does not work when iommufd is enabled.","attack_vector":"Not directly attacker-triggered in the released form: the missing op makes probe fail with a WARN when CONFIG_IOMMUFD is enabled and a device is bound to pds_vfio_pci. Worth carrying because it is a completeness gap in the IOAS attach/detach contract on a passthrough driver, and because it silently removes AMD Pensando DPU/SmartNIC passthrough from any node running an iommufd-enabled kernel. Conditional on CONFIG_IOMMUFD and pds_vfio_pci being bound.","remediation":"No fixed release is listed in this record; apply the linked stable commits or run a current stable/LTS kernel on nodes with Pensando/pds SR-IOV VFs. Interim: if pds passthrough is required, keep those nodes on the legacy VFIO type1 container path rather than iommufd until the kernel is patched.","references":["https://git.kernel.org/stable/c/7dbfae90c5a33f6b694e7068bc9522cc2655373d","https://git.kernel.org/stable/c/1df8150ab4cc422bddfbd312d6758c50b688a971","https://nvd.nist.gov/vuln/detail/CVE-2025-38625"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-48514","cve":"CVE-2025-48514","aliases":[],"title":"AMD SEV firmware - SEV-ES guest attacking an SNP guest: Coarse access-control granularity in SEV firmware lets a","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV firmware - SEV-ES guest attacking an SNP guest","year":"2025","cvss_score":4,"severity":"medium","kev":false,"impact":"Coarse access-control granularity in SEV firmware lets a privileged attacker create a SEV-ES guest positioned to attack an SNP guest, costing the SNP guest confidentiality. The pattern is worth internalising: mixing SEV generations on one host means the weakest guest type in the mix can become the attack platform against the strongest.","attack_vector":"Requires the ability to launch guests with chosen SEV parameters - the operator or a compromised control plane.","remediation":"Fixed in AMD reference firmware (AGESA / PSP / SEV firmware) and delivered only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo and the ODMs each rebuild and requalify AMD's AGESA drop before shipping. **Expect one to six months of OEM lag**, and on end-of-support platforms expect nothing. Applying it is a drain plus full power cycle, not a driver reload. Verify by reading back the PSP/SMU firmware version afterwards rather than trusting the BIOS version string. This sits inside the SEV-SNP trust boundary, so the update moves the platform's reported TCB version: refresh VCEK certificates from AMD's KDS and update any attestation policy your tenants pin, or confidential guest launches will start failing right after the BIOS lands. Interim control: do not co-schedule legacy SEV/SEV-ES guests alongside SEV-SNP guests on the same host. Pinning confidential-tier workloads to SNP-only hosts removes the attack platform entirely and costs you scheduling flexibility rather than a maintenance window.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-48514","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2026-02-10"},{"id":"CVE-2025-54509","cve":"CVE-2025-54509","aliases":[],"title":"AMD IOMMU register interface - ASP coherency: Improper access control on the IOMMU register interface lets a privileged","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD IOMMU register interface - ASP coherency","year":"2025","cvss_score":4,"severity":"medium","kev":false,"impact":"Improper access control on the IOMMU register interface lets a privileged attacker force non-coherent accesses by the AMD Secure Processor. Incoherent reads by the security engine mean it can be shown stale or inconsistent data - a subtle way to make the ASP act on something other than what is actually in memory.","attack_vector":"Local, privileged, via the IOMMU register interface.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-54509","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2026-06-09"},{"id":"CVE-2025-64715","cve":"CVE-2025-64715","aliases":[],"title":"Cilium: Egress policies referencing AWS security group IDs are misapplied","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2025","cvss_score":4,"severity":"medium","kev":false,"impact":"Egress policies referencing AWS security group IDs are misapplied","attack_vector":"Any tenant workload","remediation":"Rolling Cilium upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-64715"],"status":"curated","published":"2025-11-29"},{"id":"CVE-2021-26387","cve":"CVE-2021-26387","aliases":[],"title":"AMD Secure Processor kernel - DRAM mapping into protected areas (AMD-SB-3003): An access-control gap in the ASP kernel","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor kernel - DRAM mapping into protected areas (AMD-SB-3003)","year":"2021","cvss_score":3.9,"severity":"low","kev":false,"impact":"An access-control gap in the ASP kernel allows DRAM to be mapped into areas the secure processor treats as protected. Mapping attacker-influenced DRAM into a protected region is how you get the secure processor to operate on data it believes is trustworthy - low score, but it is a building block rather than an endpoint.","attack_vector":"Local, privileged, through the ASP interface.","remediation":"Fixed in AMD PI/AGESA firmware and delivered only as an OEM SBIOS package - AMD ships the PI drop to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before releasing BIOS. **Budget one to six months of OEM lag**, and note that several CVEs in this batch are marked 'no fix planned' on Naples (EPYC 7001) - for those the only remediation is retiring the hardware. Applying it means cordon, drain and a full power cycle per node; there is no driver reload, no live patch and no VBIOS step. Marked 'no fix planned' on Naples (EPYC 7001) - on that generation the remediation is hardware retirement, which for 2017-era EPYC in an AI fleet is likely already overdue on performance grounds.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26387","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3003.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"published":"2024-08-13"},{"id":"CVE-2021-46772","cve":"CVE-2021-46772","aliases":[],"title":"AGESA Boot Loader (ABL) - SPI ROM header input validation (AMD-SB-3003): The AGESA Boot Loader does not properly","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AGESA Boot Loader (ABL) - SPI ROM header input validation (AMD-SB-3003)","year":"2021","cvss_score":3.9,"severity":"low","kev":false,"impact":"The AGESA Boot Loader does not properly validate SPI ROM headers, so malformed header content is acted on during early boot. Anything that runs before signature enforcement is fully established is disproportionately valuable to an attacker regardless of its CVSS.","attack_vector":"Local, requires SPI ROM write access - root plus flash, a compromised BMC, or supply-chain access.","remediation":"Fixed in AMD PI/AGESA firmware and delivered only as an OEM SBIOS package - AMD ships the PI drop to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before releasing BIOS. **Budget one to six months of OEM lag**, and note that several CVEs in this batch are marked 'no fix planned' on Naples (EPYC 7001) - for those the only remediation is retiring the hardware. Applying it means cordon, drain and a full power cycle per node; there is no driver reload, no live patch and no VBIOS step. Enable platform SPI write protection as the compensating control; boot-time parsers cannot be defended from the OS.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-46772","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3003.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"published":"2024-08-13"},{"id":"CVE-2022-24735","cve":"CVE-2022-24735","aliases":[],"title":"Redis: Lua environment weakness lets a user inject code that runs with another Redis user's privileges","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Redis","year":"2022","cvss_score":3.9,"severity":"low","kev":false,"impact":"Lua environment weakness lets a user inject code that runs with another Redis user's privileges","attack_vector":"Local","remediation":"Control-plane: upgrade + ACL review","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-24735"],"status":"curated","published":"2022-04-27"},{"id":"CVE-2023-20867","cve":"CVE-2023-20867","aliases":[],"title":"VMware Tools: A fully compromised ESXi host can force VMware Tools to skip host-to-guest authentication","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"VMware Tools","year":"2023","cvss_score":3.9,"severity":"low","kev":true,"impact":"A fully compromised ESXi host can force VMware Tools to skip host-to-guest authentication - used by UNC3886 for stealthy guest access [KEV]","attack_vector":"Compromised hypervisor against tenant guests","remediation":"VMware Tools update inside every guest image - tenant-side action a neocloud can only mandate, not perform. Low CVSS, high real-world significance for post-escape persistence","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-20867"],"status":"curated","published":"2023-06-13"},{"id":"CVE-2025-32004","cve":"CVE-2025-32004","aliases":[],"title":"Intel SGX SDK (Edger8r code generator): The Edger8r tool generates the trusted/untrusted bridge code for enclaves","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SGX SDK (Edger8r code generator)","year":"2025","cvss_score":3.9,"severity":"low","kev":false,"impact":"The Edger8r tool generates the trusted/untrusted bridge code for enclaves; an input-validation flaw here means the generated bridge itself can be unsafe. Every enclave built with the affected SDK inherits the problem.","attack_vector":"Local authenticated user against an enclave built with the affected generator.","remediation":"Rebuild enclaves with a fixed SGX SDK and re-attest. Vendor-side fix; no operator reboot but also nothing you can patch yourself.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-32004","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01383.html"],"status":"curated","published":"2025-08-12"},{"cwe":["CWE-285"],"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2021-015-argo-workflows-argo-server-auth","cve":null,"aliases":["GHSA-prqf-xr2j-xf65"],"title":"Argo Workflows (Argo Server, --auth-mode=client on Kubernetes >= 1.19): PRIVILEGE ESCALATION TO THE SERVER'S IDENTITY","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo Workflows (Argo Server, --auth-mode=client on Kubernetes >= 1.19)","year":"2021","cvss_score":3.9,"severity":"low","kev":false,"impact":"PRIVILEGE ESCALATION TO THE SERVER'S IDENTITY: in a specific configuration the client's authentication is silently ignored and the server's own credentials are used instead, so every caller gets whatever the Argo Server account can do. The conditions are narrow — Kubernetes 1.19 or newer, Argo Server running outside a Kubernetes pod (bare metal or a VM), --auth-mode=client without --auth-mode=server, clients authenticating with a client key, and the server holding more permissions than the connecting account. But the reason this matters to an operator is that it inverts the intended posture: client mode is what you switch to in order to make callers act as themselves, so the deployment that took the hardening step is the one that silently loses it. The maintainers describe it as a proactive fix with no known exploits.","attack_vector":"Network, low privileges: any client that can authenticate to an Argo Server running off-cluster in the configuration above. The escalation happens automatically rather than requiring a crafted request.","remediation":"Upgrade Argo Workflows to a release carrying the fix and restart the server. Prefer running Argo Server inside a Kubernetes pod, which takes the affected configuration off the table. Where it must run off-cluster, keep the server's own service account scoped no wider than the least-privileged caller you accept, so an ignored client identity does not hand out more than the caller already had.","references":["https://github.com/argoproj/argo-workflows/security/advisories/GHSA-prqf-xr2j-xf65","https://github.com/argoproj/argo-workflows/pull/6506"],"status":"curated"},{"id":"CVE-2023-42776","cve":"CVE-2023-42776","aliases":[],"title":"Intel SGX DCAP for Windows: Input-validation flaw in the Windows DCAP components allowing local information disclosure","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SGX DCAP for Windows","year":"2023","cvss_score":3.8,"severity":"low","kev":false,"impact":"Input-validation flaw in the Windows DCAP components allowing local information disclosure. Low severity, but DCAP is the attestation plumbing - anything that touches it deserves a look in a confidential-compute deployment.","attack_vector":"Local authenticated user on a Windows host running DCAP.","remediation":"Update SGX DCAP for Windows to 1.19.100.3 or later. Userspace, service restart.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-42776","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01014.html"],"status":"curated","fleet":{"pain_class":"daemon-restart"},"published":"2024-02-14"},{"id":"CVE-2024-36348","cve":"CVE-2024-36348","aliases":["TSA","Transient Scheduler Attack"],"title":"AMD processors - speculative inference of control registers despite UMIP: Part of the Transient Scheduler Attacks batch","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD processors - speculative inference of control registers despite UMIP","year":"2024","cvss_score":3.8,"severity":"low","kev":false,"impact":"Part of the Transient Scheduler Attacks batch AMD disclosed in July 2025. A user process can speculatively infer the contents of control registers even when UMIP - the feature specifically added to stop userspace reading them - is enabled. Control register contents leak kernel configuration and address-layout information, which is the reconnaissance step that makes a subsequent kernel exploit reliable. Low score, real utility to an attacker chaining it.","attack_vector":"Local, unprivileged user process. Part of the TSA family that also covers cross-thread and cross-privilege leakage on affected Zen parts.","remediation":"Mitigated by AMD microcode plus, on most of these, a kernel-side change - and the durable delivery vehicle is the OEM SBIOS/AGESA package, which carries **one to six months of OEM lag** and needs a drained node and a full power cycle. The linux-firmware amd-ucode blobs get you the microcode sooner via initramfs early-load and a reboot, but AMD does not support late-loading microcode on a running EPYC host, so either way this is reboot-required, not a live patch. AMD shipped the TSA mitigations in microcode plus a kernel change (VERW-based clearing on transitions) in the July 2025 wave. Patch the whole TSA batch together - the siblings covering store-queue and L1 leakage carry the higher scores. Some TSA mitigations cost measurable performance on context-switch-heavy workloads, so benchmark before you assume the fix is free.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36348","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2025-07-08"},{"id":"CVE-2024-36349","cve":"CVE-2024-36349","aliases":["TSA","Transient Scheduler Attack"],"title":"AMD processors - speculative inference of TSC_AUX when reads are disabled: Sibling of the other Transient Scheduler","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD processors - speculative inference of TSC_AUX when reads are disabled","year":"2024","cvss_score":3.8,"severity":"low","kev":false,"impact":"Sibling of the other Transient Scheduler Attack disclosures. A user process can speculatively infer TSC_AUX even when the platform has disabled that read. TSC_AUX carries the CPU and NUMA node identity, so leaking it tells an attacker exactly where they are running - which is the prerequisite for arranging co-residency with a target tenant and then mounting a cross-core or cross-thread channel against them.","attack_vector":"Local, unprivileged user process on affected AMD parts.","remediation":"Mitigated by AMD microcode plus, on most of these, a kernel-side change - and the durable delivery vehicle is the OEM SBIOS/AGESA package, which carries **one to six months of OEM lag** and needs a drained node and a full power cycle. The linux-firmware amd-ucode blobs get you the microcode sooner via initramfs early-load and a reboot, but AMD does not support late-loading microcode on a running EPYC host, so either way this is reboot-required, not a live patch. Ships with the rest of the July 2025 TSA batch; do not cherry-pick individual CVEs out of it. Benchmark after applying - the TSA mitigations add work on privilege transitions.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36349","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2025-07-08"},{"cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:L/I:N/A:N","cwe":["CWE-362"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2020-27746","cve":"CVE-2020-27746","aliases":[],"title":"Slurm (X11 forwarding, xauth magic-cookie setup): Slurm shells out to xauth to install a user's X11 magic cookie, and","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Slurm (X11 forwarding, xauth magic-cookie setup)","year":"2020","cvss_score":3.7,"severity":"low","kev":false,"impact":"Slurm shells out to xauth to install a user's X11 magic cookie, and there is a window where the cookie is visible under /proc. A co-tenant polling /proc on the same node steals the cookie and attaches to that user's X11 session - keystrokes, screen contents, and the ability to inject input.","attack_vector":"A co-tenant with a job or shell on the same compute node as a victim who submitted with --x11. Only jobs that requested X11 forwarding are exposed.","remediation":"Upgrade to Slurm 19.05.8 or 20.02.6 and restart slurmd. If you cannot upgrade, disable X11 forwarding (PrologFlags without X11) - on a GPU training cluster X11 forwarding is almost never load-bearing and turning it off is cheap.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-27746","https://www.debian.org/security/2021/dsa-4841"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2022-23134","cve":"CVE-2022-23134","aliases":[],"title":"Zabbix: Some setup.php steps reachable by unauthenticated users","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Zabbix","year":"2022","cvss_score":3.7,"severity":"low","kev":true,"impact":"Some setup.php steps reachable by unauthenticated users -> configuration change","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; restrict frontend network access","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-23134"],"status":"curated","published":"2022-01-13"},{"cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:N/I:L/A:N","cwe":["CWE-327","CWE-328"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2026-13482","cve":"CVE-2026-13482","aliases":[],"title":"SkyPilot (sky/users/server.py, user ID derivation from username): User IDs are derived with a weak hash of the","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"SkyPilot (sky/users/server.py, user ID derivation from username)","year":"2026","cvss_score":3.7,"severity":"low","kev":false,"impact":"User IDs are derived with a weak hash of the username, so a remote attacker can work toward a collision and get a user ID that belongs to someone else. Practical exploitation is difficult, but the failure mode is identity confusion in the layer that decides whose clusters and quota a request touches.","attack_vector":"Remote, unauthenticated, but high attack complexity. Affects SkyPilot up to 0.12.0. The exploit has been published.","remediation":"Upgrade SkyPilot past 0.12.0 once the maintainers ship the fix tracked in issue 9194, and restart the API server. This is a VulDB-sourced report - confirm against the SkyPilot release notes before scheduling a maintenance window on it alone.","references":["https://github.com/skypilot-org/skypilot/issues/9194","https://nvd.nist.gov/vuln/detail/CVE-2026-13482"],"status":"curated"},{"id":"CVE-2026-24122","cve":"CVE-2026-24122","aliases":[],"title":"cosign / sigstore: Expired issuing certificate treated as valid during verification","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"cosign / sigstore","year":"2026","cvss_score":3.7,"severity":"low","kev":false,"impact":"Expired issuing certificate treated as valid during verification","attack_vector":"Malicious image","remediation":"Upgrade cosign","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24122"],"status":"curated","published":"2026-02-19"},{"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2021-017-nats-server-tls-ciphersuite-sele","cve":null,"aliases":["GHSA-jj54-5q2m-q7pj","CVE-2021-32026 (reserved)"],"title":"NATS server (TLS ciphersuite selection via CLI flags): A configuration footgun in the cluster message bus: NATS","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"NATS server (TLS ciphersuite selection via CLI flags)","year":"2021","cvss_score":3.7,"severity":"low","kev":false,"impact":"A configuration footgun in the cluster message bus: NATS defaults to a restricted modern ciphersuite set (RSA/ECDSA with AES-GCM plus SHA2, or ChaCha20/Poly1305), but if an administrator sets the TLS key and certificate through command-line options rather than a configuration file, those restrictions are dropped and every ciphersuite Go supports is enabled. Clients then negotiate suites the operator never intended to offer. The vendor is explicit that none of the additional suites are broken, so this is a hardening regression rather than a break — no embargo was applied and no rushed release was made. It is worth carrying for an AI-cluster operator because NATS commonly carries job dispatch and control-plane messaging between components, CLI-flag configuration is exactly what container entrypoints and Helm charts generate, and the failure is invisible: the deployment looks TLS-enabled and the negotiated posture is silently weaker than the documented default.","attack_vector":"Network, against a nats-server started with TLS parameters supplied as command-line options rather than in a configuration file. A client chooses among the unexpectedly widened ciphersuite set during the handshake.","remediation":"Upgrade nats-server to 2.2.3 or later and restart. Workaround without upgrading: move TLS parameters out of CLI flags into a configuration file, which preserves the restricted default set. Audit container entrypoints and Helm values for TLS flags passed on the command line, since that is where this configuration shape originates.","references":["https://github.com/nats-io/nats-server/security/advisories/GHSA-jj54-5q2m-q7pj","https://advisories.nats.io/CVE/CVE-2021-32026.txt"],"status":"curated"},{"id":"CVE-2024-45310","cve":"CVE-2024-45310","aliases":[],"title":"runc: runc can be tricked into creating empty files/directories at arbitrary host locations","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"runc","year":"2024","cvss_score":3.6,"severity":"low","kev":false,"impact":"runc can be tricked into creating empty files/directories at arbitrary host locations","attack_vector":"Any tenant workload with control over the pod's mount config","remediation":"Replace runc binary; drain node","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-45310"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2024-09-03"},{"id":"CVE-2026-0995","cve":"CVE-2026-0995","aliases":["TFV-16","SME TLBI erratum"],"title":"Arm C1-Pro before r1p2; Trusted Firmware-A v2.10 and later on multi-core configurations with the CME complex enabled","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arm C1-Pro before r1p2; Trusted Firmware-A v2.10 and later on multi-core configurations with the CME complex enabled","year":"2026","cvss_score":3.6,"severity":"low","kev":false,"impact":"SME / SVE / SIMD memory accesses on one core can outlive another core's TLB invalidate and barrier, so vector loads and stores complete against a translation that was supposed to be dead. The consequence is memory touched outside the expected translation or privilege boundary. Because SME and SVE are exactly what ML kernels use, this is a fault that fires most readily under the workload you actually run, not under a synthetic test.","attack_vector":"Requires code on multiple cores of an affected C1-Pro part - a guest or host process issuing wide vector memory operations while another core performs TLB maintenance. Local only.","remediation":"The TF-A mitigation is heavy: EL3 coordinates a secure-SGI rendezvous across cores using atomic counters, and the OS must call into that SMC interface during affected TLB maintenance. So you need both an OEM firmware build with WORKAROUND_CVE_2026_0995=1 and a patched kernel; neither alone is sufficient. Flash + reboot + drain, plus a kernel roll. Silicon revision r1p2 and later does not need it, so on a fleet refresh this is a spec item to demand from the vendor rather than a patch to carry forever.","references":["https://trustedfirmware-a.readthedocs.io/en/latest/security_advisories/security-advisory-tfv-16.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-03-02"},{"cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:L/UI:N/S:U/C:L/I:N/A:N","fleet":{"pain_class":"node-reboot"},"id":"CVE-2020-8588","cve":"CVE-2020-8588","aliases":[],"title":"NetApp Clustered Data ONTAP Storage Virtual Machine boundary: A user in one SVM determines whether data exists on a","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"NetApp Clustered Data ONTAP Storage Virtual Machine boundary","year":"2020","cvss_score":3.5,"severity":"low","kev":false,"impact":"A user in one SVM determines whether data exists on a different SVM. The SVM is the tenant boundary in ONTAP, so this leaks the shape of another customer's namespace across it.","attack_vector":"An authenticated user on an adjacent network with access to any SVM on a Clustered Data ONTAP system earlier than 9.3P20 or 9.5P15.","remediation":"Upgrade to 9.3P20 / 9.5P15 or later. If SVMs are being used as a hard tenant boundary, treat namespace metadata as having been observable until the upgrade lands.","references":["https://security.netapp.com/advisory/ntap-20210201-0001/","https://nvd.nist.gov/vuln/detail/CVE-2020-8588"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:L/UI:N/S:U/C:L/I:N/A:N","fleet":{"pain_class":"node-reboot"},"id":"CVE-2020-8589","cve":"CVE-2020-8589","aliases":[],"title":"NetApp Clustered Data ONTAP Storage Virtual Machine boundary: A user in one SVM enumerates the names of other SVMs and","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"NetApp Clustered Data ONTAP Storage Virtual Machine boundary","year":"2020","cvss_score":3.5,"severity":"low","kev":false,"impact":"A user in one SVM enumerates the names of other SVMs and the filenames inside them. Dataset and checkpoint names alone tell a competitor what another tenant is training.","attack_vector":"An authenticated user on an adjacent network with access to any SVM on a Clustered Data ONTAP system earlier than 9.3P20 or 9.5P15.","remediation":"Upgrade to 9.3P20 / 9.5P15 or later. Where SVM naming itself is sensitive, rename after patching, since the old names are already exposed.","references":["https://security.netapp.com/advisory/ntap-20210201-0002/","https://nvd.nist.gov/vuln/detail/CVE-2020-8589"],"status":"curated","tags":["tenant-isolation"]},{"cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:N/UI:R/S:U/C:L/I:N/A:N","cwe":["CWE-89"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2023-41891","cve":"CVE-2023-41891","aliases":["GHSA-r847-6w6h-r8g4"],"title":"FlyteAdmin (list endpoints, SQL injection through list filters): FlyteAdmin's list endpoints interpolate filter","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"FlyteAdmin (list endpoints, SQL injection through list filters)","year":"2023","cvss_score":3.5,"severity":"low","kev":false,"impact":"FlyteAdmin's list endpoints interpolate filter parameters into SQL, so a crafted REST request runs attacker-chosen statements against the control-plane database. That database holds every project's and every tenant's workflow, execution and launch-plan records, so the read boundary between projects collapses.","attack_vector":"A user who can reach the FlyteAdmin API. In most deployments that means someone already behind the VPN or holding a valid login.","remediation":"Upgrade FlyteAdmin to 1.1.124 or later and restart. Keep FlyteAdmin off the public internet regardless, and review database audit logs for unexpected queries from the admin service account.","references":["https://github.com/flyteorg/flyteadmin/security/advisories/GHSA-r847-6w6h-r8g4","https://nvd.nist.gov/vuln/detail/CVE-2023-41891"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2025-2295","cve":"CVE-2025-2295","aliases":["GHSA-8522-69fh-w74x"],"title":"EDK II NetworkPkg (IScsiDxe, Ready-To-Transfer PDU handling): A malicious iSCSI target sends a crafted R2T PDU","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"EDK II NetworkPkg (IScsiDxe, Ready-To-Transfer PDU handling)","year":"2025","cvss_score":3.5,"severity":"low","kev":false,"impact":"A malicious iSCSI target sends a crafted R2T PDU with a bogus offset and length, and the booting firmware obligingly transmits chunks of its own memory back to the target. Low severity on paper, and it is a read-only leak, but what leaks is DXE-phase memory - boot secrets, variable contents, buffer addresses - which is exactly the reconnaissance an attacker needs before firing one of the higher-severity overflow bugs in the same stack.","attack_vector":"An attacker controlling or impersonating the iSCSI target a node boots from. Unauthenticated, pre-OS, from the storage network.","remediation":"OEM BIOS update, flash + reboot per node - but given the low score, do not expect OEMs to ship it urgently or to call it out prominently in release notes. The config workaround is the better first move: turn off the UEFI iSCSI initiator on nodes that boot locally, require mutual CHAP where you do boot from SAN, and segment the storage fabric.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-2295","https://github.com/tianocore/edk2/security/advisories/GHSA-8522-69fh-w74x"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2025-03-14"},{"id":"CVE-2026-70412","cve":"CVE-2026-70412","aliases":["DSA-2026-348"],"title":"Dell iDRAC9 / iDRAC10 (memory erase, data remanence): Data survives an iDRAC memory erase and stays readable","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC9 / iDRAC10 (memory erase, data remanence)","year":"2026","cvss_score":3.5,"severity":"low","kev":false,"impact":"Data survives an iDRAC memory erase and stays readable afterwards. The CVSS is low, but the operational meaning for a bare-metal GPU cloud is not: the erase step in your node-reprovisioning pipeline does not actually erase, so a low-privilege user on the next tenancy can read remnants left by the previous one. This is exactly the cross-tenant handoff failure that bare-metal operators promise does not happen, and it fails quietly - the wipe reports success. Affects both the iDRAC9 and iDRAC10 generations.","attack_vector":"A low-privilege account with remote access to the iDRAC - which on a bare-metal cloud can be the next tenant, if your product hands tenants any BMC-adjacent access at all, or anyone reaching the management VLAN.","remediation":"Flash iDRAC9 to 7.20.30.50 or iDRAC10 to 1.20.60.50 or later. Out-of-band, per-node, no host reboot and no job drain. Beyond the flash, treat this as a pipeline bug rather than a node bug: if your reprovisioning runbook relies on the iDRAC erase as the cross-tenant boundary, add an independent verification step, and consider re-checking nodes that were recycled between tenants on unpatched firmware.","references":["https://www.dell.com/support/kbdoc/en-us/000497902/dsa-2026-348-security-update-for-dell-idrac9-and-idrac10-vulnerability","https://nvd.nist.gov/vuln/detail/CVE-2026-70412"],"status":"curated","published":"2026-08-17"},{"id":"CVE-2023-2431","cve":"CVE-2023-2431","aliases":[],"title":"Kubernetes (kubelet): Pods with an empty localhost seccomp profile field silently bypass seccomp enforcement","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kubelet)","year":"2023","cvss_score":3.4,"severity":"low","kev":false,"impact":"Pods with an empty localhost seccomp profile field silently bypass seccomp enforcement","attack_vector":"Any tenant workload","remediation":"Rolling kubelet upgrade with node drain; add an admission check on seccompProfile","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-2431"],"status":"curated","fleet":{"pain_class":"node-drain"},"published":"2023-06-16"},{"cvss_vector":"CVSS:3.0/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:N/A:N","cwe":["CWE-200"],"fleet":{"pain_class":"node-drain"},"id":"CVE-2018-1993","cve":"CVE-2018-1993","aliases":[],"title":"IBM Spectrum Scale Local Read Only Cache (LROC): With LROC enabled, a read of one file can silently return the contents","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"IBM Spectrum Scale Local Read Only Cache (LROC)","year":"2018","cvss_score":3.3,"severity":"low","kev":false,"impact":"With LROC enabled, a read of one file can silently return the contents of a different file. A tenant reading its own dataset gets back bytes belonging to someone else, and a training job can ingest another tenant's data without anyone noticing.","attack_vector":"Any user able to read files on a node with LROC enabled. This is a correctness bug in the cache rather than an exploit chain, so it fires during ordinary I/O.","remediation":"Upgrade to the fixed Spectrum Scale level. If an upgrade cannot happen right away, disable LROC on affected nodes - the performance loss is far cheaper than cross-tenant data bleed, and any data read while LROC was active should be treated as suspect.","references":["https://www.ibm.com/support/docview.wss?uid=ibm10793719","https://nvd.nist.gov/vuln/detail/CVE-2018-1993"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2018-20855","cve":"CVE-2018-20855","aliases":["RDMA/mlx5 uninitialized mlx5_ib_create_qp_resp"],"title":"Linux kernel mlx5_ib (create QP response): mlx5_ib_create_qp_resp is never initialized in create_qp_common, so creating","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_ib (create QP response)","year":"2018","cvss_score":3.3,"severity":"low","kev":false,"impact":"mlx5_ib_create_qp_resp is never initialized in create_qp_common, so creating a queue pair returns uninitialized kernel stack memory to the calling process. Low severity on its own, but it is a free kernel-stack read for any tenant with RDMA access, useful for defeating address-space layout randomization before a heavier exploit.","attack_vector":"Local, low-privileged - any user able to create an RDMA queue pair through libibverbs, which on a GPU cluster is every workload using RDMA collectives.","remediation":"Upgrade the host kernel past 4.18.7 or take the distro backport. Any modern kernel already carries this; the value here is checking that legacy long-lived nodes in the fleet are not still on pre-4.18 kernels. Host reboot to apply.","references":["https://ubuntu.com/security/CVE-2018-20855","https://nvd.nist.gov/vuln/detail/CVE-2018-20855"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2019-07-26"},{"id":"CVE-2019-0174","cve":"CVE-2019-0174","aliases":["RAMBleed","INTEL-SA-00247"],"title":"DDR3 and DDR4 DRAM, including ECC modules; tracked by Intel as a partial-physical-address disclosure issue: Turns","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"DDR3 and DDR4 DRAM, including ECC modules; tracked by Intel as a partial-physical-address disclosure issue","year":"2019","cvss_score":3.3,"severity":"low","kev":false,"impact":"Turns Rowhammer from a write primitive into a read primitive. The attacker does not corrupt the victim's data - they observe whether their own bits flip, which is data-dependent on the neighbouring row's contents, and read the victim's memory out that way. The published result was extraction of a 2048-bit OpenSSH RSA host key from another process at about 0.3 bits per second. Because it is read-only, ECC does not stop it and nothing in the victim's process ever misbehaves, so there is no detection surface at all. The low CVSS is misleading for a multi-tenant operator: the practical claim is cross-tenant key theft with no artefact.","attack_vector":"Unprivileged local code sharing DRAM with the victim - a container, a VM, or a co-scheduled batch job. The attacker needs to get their pages physically adjacent to the victim's, which memory-massaging techniques make reliable on a busy host.","remediation":"No patch. ECC does not help - the attack never needs a flip to survive. Refresh-rate increases raise cost but do not close it. The controls that work: keep untrusted tenants off shared memory controllers, and at the application layer make the secrets worth less by rotating keys aggressively and using memory-hard or constantly-rekeyed representations for long-lived secrets on shared hosts. If you host customer inference on shared CPU memory, assume any long-lived key resident there is readable by a co-tenant given hours.","references":["https://rambleed.com/","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00247.html","https://nvd.nist.gov/vuln/detail/CVE-2019-0174"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"published":"2019-06-13"},{"id":"CVE-2021-26342","cve":"CVE-2021-26342","aliases":[],"title":"AMD SEV guest VMs - TLB flush after VMCB creation sequence: The CPU may fail to flush the TLB after a particular","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV guest VMs - TLB flush after VMCB creation sequence","year":"2021","cvss_score":3.3,"severity":"low","kev":false,"impact":"The CPU may fail to flush the TLB after a particular sequence involving creation of a new VMCB in SEV guest VMs. Stale translations across a VM boundary mean one guest's memory can be reached through another's cached mapping - the low CVSS reflects the difficulty of arranging the sequence, not the severity of what happens if you do.","attack_vector":"Requires a host able to arrange a specific VMCB creation sequence - hypervisor-privileged.","remediation":"Fixed in AMD reference firmware (AGESA / PSP / SEV firmware) and delivered only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo and the ODMs each rebuild and requalify AMD's AGESA drop before shipping. **Expect one to six months of OEM lag**, and on end-of-support platforms expect nothing. Applying it is a drain plus full power cycle, not a driver reload. Verify by reading back the PSP/SMU firmware version afterwards rather than trusting the BIOS version string. This sits inside the SEV-SNP trust boundary, so the update moves the platform's reported TCB version: refresh VCEK certificates from AMD's KDS and update any attestation policy your tenants pin, or confidential guest launches will start failing right after the BIOS lands.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26342","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2022-05-11"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:L/A:N","fleet":{"pain_class":"daemon-restart"},"id":"CVE-2021-29671","cve":"CVE-2021-29671","aliases":[],"title":"IBM Spectrum Scale file audit logging: A local user touches files without the access being recorded, so the audit trail","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"IBM Spectrum Scale file audit logging","year":"2021","cvss_score":3.3,"severity":"low","kev":false,"impact":"A local user touches files without the access being recorded, so the audit trail the operator relies on to prove who read what stops being complete. In a shared cluster this is the difference between detecting and not detecting a data-exfiltration incident.","attack_vector":"Local account on a Spectrum Scale 5.1.0.1 node with file audit logging enabled.","remediation":"Upgrade to the fixed level. Treat audit logs written by 5.1.0.1 as incomplete rather than authoritative for any investigation covering that window.","references":["https://www.ibm.com/support/pages/node/6441429","https://nvd.nist.gov/vuln/detail/CVE-2021-29671"],"status":"curated"},{"id":"CVE-2022-23649","cve":"CVE-2022-23649","aliases":[],"title":"cosign / sigstore: Cosign can be tricked into claiming a Rekor transparency-log entry exists when it does not","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"cosign / sigstore","year":"2022","cvss_score":3.3,"severity":"low","kev":false,"impact":"Cosign can be tricked into claiming a Rekor transparency-log entry exists when it does not","attack_vector":"Malicious image with a crafted signature","remediation":"Upgrade cosign in the admission and CI paths","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-23649"],"status":"curated","published":"2022-02-18"},{"id":"CVE-2022-24736","cve":"CVE-2022-24736","aliases":[],"title":"Redis: Crafted Lua script triggers a NULL pointer dereference","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Redis","year":"2022","cvss_score":3.3,"severity":"low","kev":false,"impact":"Crafted Lua script triggers a NULL pointer dereference -> redis-server crash","attack_vector":"Local","remediation":"Control-plane: rolling Redis upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-24736"],"status":"curated","published":"2022-04-27"},{"id":"CVE-2022-42336","cve":"CVE-2022-42336","aliases":[],"title":"Xen on AMD Family 17h / Hygon Family 18h - guest SSBD selection: Setting Speculative Store Bypass Disable on AMD Family","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen on AMD Family 17h / Hygon Family 18h - guest SSBD selection","year":"2022","cvss_score":3.3,"severity":"low","kev":false,"impact":"Setting Speculative Store Bypass Disable on AMD Family 17h and Hygon Family 18h has to be coordinated at the physical core level, and Xen's logic did not do that correctly. The consequence is that a guest which asked for SSBD protection may not actually get it - or may have it silently disabled by a sibling. A tenant hardening itself against Spectre-v4 gets a mitigation that is not in force, which is worse than knowing it is off.","attack_vector":"Cross-guest speculative execution between VMs sharing a physical core.","remediation":"Fixed in Xen (XSA-431). Hypervisor update plus host reboot. Verify per-guest that SSBD is genuinely active afterwards rather than trusting the requested setting.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42336","https://xenbits.xen.org/xsa/advisory-431.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2023-05-17"},{"id":"CVE-2023-0196","cve":"CVE-2023-0196","aliases":[],"title":"CUDA Toolkit: DoS (null deref)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2023","cvss_score":3.3,"severity":"low","kev":false,"impact":"DoS (null deref)","attack_vector":"Malicious ELF/cubin artifact","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0196","https://github.com/NVIDIA/product-security/tree/main/2023/5446"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-476"],"published":"2023-03-02"},{"id":"CVE-2023-20519","cve":"CVE-2023-20519","aliases":[],"title":"AMD SEV-SNP guest context page - use-after-free enabling migration-agent masquerade (AMD-SB-3002): A use-after-free in","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV-SNP guest context page - use-after-free enabling migration-agent masquerade (AMD-SB-3002)","year":"2023","cvss_score":3.3,"severity":"low","kev":false,"impact":"A use-after-free in the SNP guest context page lets a malicious hypervisor masquerade as the guest's migration agent. The guest then negotiates its migration with the attacker instead of a legitimate MA - which is to say it hands over the state that memory encryption existed to protect, voluntarily, to the party it was protecting itself from.","attack_vector":"Malicious or compromised hypervisor. Only exercised where SEV-SNP live migration is enabled.","remediation":"Fixed in AMD PI/AGESA firmware and delivered only as an OEM SBIOS package - AMD ships the PI drop to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before releasing BIOS. **Budget one to six months of OEM lag**, and note that several CVEs in this batch are marked 'no fix planned' on Naples (EPYC 7001) - for those the only remediation is retiring the hardware. Applying it means cordon, drain and a full power cycle per node; there is no driver reload, no live patch and no VBIOS step. Because this touches the SEV-SNP trust boundary, the update moves the platform TCB version: refresh VCEK certificates from AMD's KDS and update tenant attestation policy, or confidential guest launches will fail immediately after the BIOS lands. If you do not offer live migration for confidential VMs, the path is unreachable and this can wait for the next firmware wave. If you do, disable it until the fleet is patched - that is a scheduler policy change, not a maintenance window.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-20519","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3002.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"tags":["tenant-isolation"],"published":"2023-11-14"},{"id":"CVE-2023-25510","cve":"CVE-2023-25510","aliases":[],"title":"NVIDIA CUDA Toolkit - cuobjdump: A null-pointer dereference on a malformed binary crashes the tool","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA CUDA Toolkit - cuobjdump","year":"2023","cvss_score":3.3,"severity":"low","kev":false,"impact":"A null-pointer dereference on a malformed binary crashes the tool. The realistic exposure is your build and profiling pipeline, not your runtime fleet: anything that automatically disassembles third-party fatbins, vendor kernels or model artifacts is running this parser on attacker-influenced input.","attack_vector":"Local, and requires a user or an automated job to run cuobjdump over an attacker-supplied file. CI jobs that inspect third-party CUDA binaries are the usual path.","remediation":"Update the CUDA Toolkit package (bulletin 5456). Cost: effectively zero - userspace SDK only, no driver reload, no node drain, no running-job impact. Rebuild build/CI images and move on.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25510","https://github.com/NVIDIA/product-security/tree/main/2023/5456"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-476"],"published":"2023-04-22"},{"id":"CVE-2023-25511","cve":"CVE-2023-25511","aliases":[],"title":"NVIDIA CUDA Toolkit - cuobjdump: A division-by-zero on crafted input crashes the tool","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA CUDA Toolkit - cuobjdump","year":"2023","cvss_score":3.3,"severity":"low","kev":false,"impact":"A division-by-zero on crafted input crashes the tool. The realistic exposure is your build and profiling pipeline, not your runtime fleet: anything that automatically disassembles third-party fatbins, vendor kernels or model artifacts is running this parser on attacker-influenced input.","attack_vector":"Local, and requires a user or an automated job to run cuobjdump over an attacker-supplied file. CI jobs that inspect third-party CUDA binaries are the usual path.","remediation":"Update the CUDA Toolkit package (bulletin 5456). Cost: effectively zero - userspace SDK only, no driver reload, no node drain, no running-job impact. Rebuild build/CI images and move on.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25511","https://github.com/NVIDIA/product-security/tree/main/2023/5456"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-369"],"published":"2023-04-22"},{"id":"CVE-2023-25523","cve":"CVE-2023-25523","aliases":[],"title":"CUDA Toolkit (nvdisasm): DoS (null deref via malformed ELF)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit (nvdisasm)","year":"2023","cvss_score":3.3,"severity":"low","kev":false,"impact":"DoS (null deref via malformed ELF)","attack_vector":"Malicious model/binary artifact","remediation":"Upgrade to CUDA Toolkit 12.2+; rebuild base images","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5469/5469.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-476"],"published":"2023-07-04"},{"id":"CVE-2023-31306","cve":"CVE-2023-31306","aliases":[],"title":"AMD graphics driver - dynamic power management (DPM) array index validation: An unvalidated array index in the driver's","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"AMD graphics driver - dynamic power management (DPM) array index validation","year":"2023","cvss_score":3.3,"severity":"low","kev":false,"impact":"An unvalidated array index in the driver's dynamic power management functions produces an out-of-bounds access. DPM controls clocks and power states; on Instinct parts that is the machinery keeping accelerators inside their power and thermal envelope, so corruption here is worth more attention than the 3.3 score suggests even though the direct security impact is limited.","attack_vector":"Local, requires the ability to pass malformed arguments to DPM functions.","remediation":"Update the AMD graphics driver and reload or reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31306","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-09-06"},{"id":"CVE-2024-0072","cve":"CVE-2024-0072","aliases":[],"title":"CUDA Toolkit: DoS (null deref)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2024","cvss_score":3.3,"severity":"low","kev":false,"impact":"DoS (null deref)","attack_vector":"Malicious cubin/model artifact","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0072","https://github.com/NVIDIA/product-security/tree/main/2024/5517"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-476"],"published":"2024-04-05"},{"id":"CVE-2024-0076","cve":"CVE-2024-0076","aliases":[],"title":"CUDA Toolkit: Info disclosure (buffer over-read)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2024","cvss_score":3.3,"severity":"low","kev":false,"impact":"Info disclosure (buffer over-read)","attack_vector":"Malicious binary artifact","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0076","https://github.com/NVIDIA/product-security/tree/main/2024/5517"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-125"],"published":"2024-04-05"},{"id":"CVE-2024-0102","cve":"CVE-2024-0102","aliases":[],"title":"CUDA Toolkit: Info disclosure (OOB read)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2024","cvss_score":3.3,"severity":"low","kev":false,"impact":"Info disclosure (OOB read)","attack_vector":"Malicious binary artifact","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0102","https://github.com/NVIDIA/product-security/tree/main/2024/5548"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-125"],"published":"2024-08-08"},{"id":"CVE-2024-0109","cve":"CVE-2024-0109","aliases":[],"title":"CUDA Toolkit: Info disclosure (OOB read)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2024","cvss_score":3.3,"severity":"low","kev":false,"impact":"Info disclosure (OOB read)","attack_vector":"Malicious binary artifact","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0109","https://github.com/NVIDIA/product-security/tree/main/2024/5564"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-125"],"published":"2024-08-31"},{"id":"CVE-2024-0123","cve":"CVE-2024-0123","aliases":[],"title":"NVIDIA CUDA Toolkit - nvdisasm: Improper input validation on a malicious ELF crashes the disassembler","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA CUDA Toolkit - nvdisasm","year":"2024","cvss_score":3.3,"severity":"low","kev":false,"impact":"Improper input validation on a malicious ELF crashes the disassembler. The realistic exposure is your build and profiling pipeline, not your runtime fleet: anything that automatically disassembles third-party fatbins, vendor kernels or model artifacts is running this parser on attacker-influenced input.","attack_vector":"Local, and requires a user or an automated job to run nvdisasm over an attacker-supplied file. CI jobs that inspect third-party CUDA binaries are the usual path.","remediation":"Update the CUDA Toolkit package (bulletin 5577). Cost: effectively zero - userspace SDK only, no driver reload, no node drain, no running-job impact. Rebuild build/CI images and move on.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0123","https://github.com/NVIDIA/product-security/tree/main/2024/5577"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-1285"],"published":"2024-10-03"},{"id":"CVE-2024-0124","cve":"CVE-2024-0124","aliases":[],"title":"NVIDIA CUDA Toolkit - nvdisasm: A use-after-free on a malformed ELF causes a crash and potentially worse depending","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA CUDA Toolkit - nvdisasm","year":"2024","cvss_score":3.3,"severity":"low","kev":false,"impact":"A use-after-free on a malformed ELF causes a crash and potentially worse depending on heap state. The realistic exposure is your build and profiling pipeline, not your runtime fleet: anything that automatically disassembles third-party fatbins, vendor kernels or model artifacts is running this parser on attacker-influenced input.","attack_vector":"Local, and requires a user or an automated job to run nvdisasm over an attacker-supplied file. CI jobs that inspect third-party CUDA binaries are the usual path.","remediation":"Update the CUDA Toolkit package (bulletin 5577). Cost: effectively zero - userspace SDK only, no driver reload, no node drain, no running-job impact. Rebuild build/CI images and move on.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0124","https://github.com/NVIDIA/product-security/tree/main/2024/5577"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-416"],"published":"2024-10-03"},{"id":"CVE-2024-0125","cve":"CVE-2024-0125","aliases":[],"title":"NVIDIA CUDA Toolkit - nvdisasm: A null-pointer dereference on a malformed ELF crashes the disassembler","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA CUDA Toolkit - nvdisasm","year":"2024","cvss_score":3.3,"severity":"low","kev":false,"impact":"A null-pointer dereference on a malformed ELF crashes the disassembler. The realistic exposure is your build and profiling pipeline, not your runtime fleet: anything that automatically disassembles third-party fatbins, vendor kernels or model artifacts is running this parser on attacker-influenced input.","attack_vector":"Local, and requires a user or an automated job to run nvdisasm over an attacker-supplied file. CI jobs that inspect third-party CUDA binaries are the usual path.","remediation":"Update the CUDA Toolkit package (bulletin 5577). Cost: effectively zero - userspace SDK only, no driver reload, no node drain, no running-job impact. Rebuild build/CI images and move on.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0125","https://github.com/NVIDIA/product-security/tree/main/2024/5577"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-476"],"published":"2024-10-03"},{"id":"CVE-2024-0149","cve":"CVE-2024-0149","aliases":[],"title":"GPU Display Driver: Info disclosure (OOB read)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2024","cvss_score":3.3,"severity":"low","kev":false,"impact":"Info disclosure (OOB read)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; rolling reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0149","https://github.com/NVIDIA/product-security/tree/main/2025/5614"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:N/A:N","cwe":["CWE-125"],"fleet":{"pain_class":"node-reboot"},"published":"2025-01-28"},{"id":"CVE-2024-53870","cve":"CVE-2024-53870","aliases":[],"title":"CUDA Toolkit: Info disclosure (OOB read)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2024","cvss_score":3.3,"severity":"low","kev":false,"impact":"Info disclosure (OOB read)","attack_vector":"Malicious cubin/model artifact","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53870","https://github.com/NVIDIA/product-security/tree/main/2025/5594"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-125"],"published":"2025-02-25"},{"id":"CVE-2024-53871","cve":"CVE-2024-53871","aliases":[],"title":"CUDA Toolkit: Info disclosure (OOB read)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2024","cvss_score":3.3,"severity":"low","kev":false,"impact":"Info disclosure (OOB read)","attack_vector":"Malicious cubin/model artifact","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53871","https://github.com/NVIDIA/product-security/tree/main/2025/5594"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-125"],"published":"2025-02-25"},{"id":"CVE-2024-53872","cve":"CVE-2024-53872","aliases":[],"title":"CUDA Toolkit: Info disclosure (OOB read)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2024","cvss_score":3.3,"severity":"low","kev":false,"impact":"Info disclosure (OOB read)","attack_vector":"Malicious cubin/model artifact","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53872","https://github.com/NVIDIA/product-security/tree/main/2025/5594"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-125"],"published":"2025-02-25"},{"id":"CVE-2024-53873","cve":"CVE-2024-53873","aliases":[],"title":"CUDA Toolkit: Info disclosure (OOB read)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2024","cvss_score":3.3,"severity":"low","kev":false,"impact":"Info disclosure (OOB read)","attack_vector":"Malicious cubin/model artifact","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53873","https://github.com/NVIDIA/product-security/tree/main/2025/5594"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-125"],"published":"2025-02-25"},{"id":"CVE-2024-53874","cve":"CVE-2024-53874","aliases":[],"title":"CUDA Toolkit: Info disclosure (OOB read)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2024","cvss_score":3.3,"severity":"low","kev":false,"impact":"Info disclosure (OOB read)","attack_vector":"Malicious cubin/model artifact","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53874","https://github.com/NVIDIA/product-security/tree/main/2025/5594"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-125"],"published":"2025-02-25"},{"id":"CVE-2024-53875","cve":"CVE-2024-53875","aliases":[],"title":"CUDA Toolkit: Info disclosure (OOB read)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2024","cvss_score":3.3,"severity":"low","kev":false,"impact":"Info disclosure (OOB read)","attack_vector":"Malicious cubin/model artifact","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53875","https://github.com/NVIDIA/product-security/tree/main/2025/5594"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-125"],"published":"2025-02-25"},{"id":"CVE-2024-53876","cve":"CVE-2024-53876","aliases":[],"title":"CUDA Toolkit: Info disclosure (OOB read)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2024","cvss_score":3.3,"severity":"low","kev":false,"impact":"Info disclosure (OOB read)","attack_vector":"Malicious cubin/model artifact","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53876","https://github.com/NVIDIA/product-security/tree/main/2025/5594"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-125"],"published":"2025-02-25"},{"id":"CVE-2024-53877","cve":"CVE-2024-53877","aliases":[],"title":"CUDA Toolkit: DoS (null deref)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2024","cvss_score":3.3,"severity":"low","kev":false,"impact":"DoS (null deref)","attack_vector":"Malicious cubin/model artifact","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53877","https://github.com/NVIDIA/product-security/tree/main/2025/5594"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-476"],"published":"2025-02-25"},{"id":"CVE-2025-20613","cve":"CVE-2025-20613","aliases":[],"title":"Intel TDX firmware (PRNG seeding): A predictable seed in the TDX firmware's pseudo-random number generator. Predictable","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel TDX firmware (PRNG seeding)","year":"2025","cvss_score":3.3,"severity":"low","kev":false,"impact":"A predictable seed in the TDX firmware's pseudo-random number generator. Predictable randomness inside a confidential-compute TCB undermines whatever the module derived from it - key material, nonces, address-space randomisation inside the boundary - so the low CVSS understates the structural concern.","attack_vector":"An authenticated user on the host.","remediation":"Update the Intel TDX module. The TDX module is loaded by the SEAM loader at boot, so the practical rollout is: stage the new module, drain every trust domain off the node, and reboot. It is not a live-patchable component and running TDs cannot be migrated through it. After the update, every TD must re-attest because the TDX module SVN is part of the attestation report - so anything that pinned the old measurement will fail until you update your attestation policy too. No OEM BIOS release needed for the module itself, which makes this materially faster than a platform firmware update.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-20613","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01312.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2025-08-12"},{"id":"CVE-2025-23248","cve":"CVE-2025-23248","aliases":[],"title":"NVIDIA CUDA Toolkit - nvdisasm: An out-of-bounds read on a malformed ELF crashes the tool","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA CUDA Toolkit - nvdisasm","year":"2025","cvss_score":3.3,"severity":"low","kev":false,"impact":"An out-of-bounds read on a malformed ELF crashes the tool. The realistic exposure is your build and profiling pipeline, not your runtime fleet: anything that automatically disassembles third-party fatbins, vendor kernels or model artifacts is running this parser on attacker-influenced input.","attack_vector":"Local, and requires a user or an automated job to run nvdisasm over an attacker-supplied file. CI jobs that inspect third-party CUDA binaries are the usual path.","remediation":"Update the CUDA Toolkit package (bulletin 5661). Cost: effectively zero - userspace SDK only, no driver reload, no node drain, no running-job impact. Rebuild build/CI images and move on.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23248","https://github.com/NVIDIA/product-security/tree/main/2025/5661"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-125"],"published":"2025-09-24"},{"id":"CVE-2025-23255","cve":"CVE-2025-23255","aliases":[],"title":"NVIDIA CUDA Toolkit - cuobjdump: An out-of-bounds read on a malformed ELF crashes the tool","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA CUDA Toolkit - cuobjdump","year":"2025","cvss_score":3.3,"severity":"low","kev":false,"impact":"An out-of-bounds read on a malformed ELF crashes the tool. The realistic exposure is your build and profiling pipeline, not your runtime fleet: anything that automatically disassembles third-party fatbins, vendor kernels or model artifacts is running this parser on attacker-influenced input.","attack_vector":"Local, and requires a user or an automated job to run cuobjdump over an attacker-supplied file. CI jobs that inspect third-party CUDA binaries are the usual path.","remediation":"Update the CUDA Toolkit package (bulletin 5661). Cost: effectively zero - userspace SDK only, no driver reload, no node drain, no running-job impact. Rebuild build/CI images and move on.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23255","https://github.com/NVIDIA/product-security/tree/main/2025/5661"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-125"],"published":"2025-09-24"},{"id":"CVE-2025-23271","cve":"CVE-2025-23271","aliases":[],"title":"CUDA Toolkit: Info disclosure (buffer over-read)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2025","cvss_score":3.3,"severity":"low","kev":false,"impact":"Info disclosure (buffer over-read)","attack_vector":"Malicious artifact","remediation":"Bump CUDA Toolkit; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23271","https://github.com/NVIDIA/product-security/tree/main/2025/5661"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-125"],"published":"2025-09-24"},{"id":"CVE-2025-23287","cve":"CVE-2025-23287","aliases":[],"title":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys): An attacker with local access reads sensitive","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys)","year":"2025","cvss_score":3.3,"severity":"low","kev":false,"impact":"An attacker with local access reads sensitive system-level information through the Windows display driver - useful for fingerprinting the host before a heavier attack. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5670. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23287","https://github.com/NVIDIA/product-security/tree/main/2025/5670"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:N/A:N","cwe":["CWE-497"],"fleet":{"pain_class":"node-drain"},"published":"2025-08-02"},{"id":"CVE-2025-23288","cve":"CVE-2025-23288","aliases":[],"title":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys): Local unprivileged access to the Windows display","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys)","year":"2025","cvss_score":3.3,"severity":"low","kev":false,"impact":"Local unprivileged access to the Windows display driver exposes sensitive system information. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5670. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23288","https://github.com/NVIDIA/product-security/tree/main/2025/5670"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:N/A:N","cwe":["CWE-497"],"fleet":{"pain_class":"node-drain"},"published":"2025-08-02"},{"id":"CVE-2025-23308","cve":"CVE-2025-23308","aliases":[],"title":"NVIDIA CUDA Toolkit - nvdisasm: A heap-based buffer overflow on a malicious ELF gives arbitrary code execution","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA CUDA Toolkit - nvdisasm","year":"2025","cvss_score":3.3,"severity":"low","kev":false,"impact":"A heap-based buffer overflow on a malicious ELF gives arbitrary code execution at the privilege level of whoever ran nvdisasm - the most serious of the CUDA parser bugs, and CI service accounts are usually not low-privilege. The realistic exposure is your build and profiling pipeline, not your runtime fleet: anything that automatically disassembles third-party fatbins, vendor kernels or model artifacts is running this parser on attacker-influenced input.","attack_vector":"Local, and requires a user or an automated job to run nvdisasm over an attacker-supplied file. CI jobs that inspect third-party CUDA binaries are the usual path.","remediation":"Update the CUDA Toolkit package (bulletin 5661). Cost: effectively zero - userspace SDK only, no driver reload, no node drain, no running-job impact. Rebuild build/CI images and move on.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23308","https://github.com/NVIDIA/product-security/tree/main/2025/5661"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:L/I:N/A:N","cwe":["CWE-122"],"published":"2025-09-24"},{"id":"CVE-2025-23338","cve":"CVE-2025-23338","aliases":[],"title":"NVIDIA CUDA Toolkit - nvdisasm: An out-of-bounds write on a malicious ELF crashes the tool and corrupts heap state","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA CUDA Toolkit - nvdisasm","year":"2025","cvss_score":3.3,"severity":"low","kev":false,"impact":"An out-of-bounds write on a malicious ELF crashes the tool and corrupts heap state. The realistic exposure is your build and profiling pipeline, not your runtime fleet: anything that automatically disassembles third-party fatbins, vendor kernels or model artifacts is running this parser on attacker-influenced input.","attack_vector":"Local, and requires a user or an automated job to run nvdisasm over an attacker-supplied file. CI jobs that inspect third-party CUDA binaries are the usual path.","remediation":"Update the CUDA Toolkit package (bulletin 5661). Cost: effectively zero - userspace SDK only, no driver reload, no node drain, no running-job impact. Rebuild build/CI images and move on.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23338","https://github.com/NVIDIA/product-security/tree/main/2025/5661"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-129"],"published":"2025-09-24"},{"id":"CVE-2025-23339","cve":"CVE-2025-23339","aliases":[],"title":"NVIDIA CUDA Toolkit - cuobjdump: A stack-based buffer overflow on a malicious ELF gives arbitrary code execution","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA CUDA Toolkit - cuobjdump","year":"2025","cvss_score":3.3,"severity":"low","kev":false,"impact":"A stack-based buffer overflow on a malicious ELF gives arbitrary code execution as the invoking user. The realistic exposure is your build and profiling pipeline, not your runtime fleet: anything that automatically disassembles third-party fatbins, vendor kernels or model artifacts is running this parser on attacker-influenced input.","attack_vector":"Local, and requires a user or an automated job to run cuobjdump over an attacker-supplied file. CI jobs that inspect third-party CUDA binaries are the usual path.","remediation":"Update the CUDA Toolkit package (bulletin 5661). Cost: effectively zero - userspace SDK only, no driver reload, no node drain, no running-job impact. Rebuild build/CI images and move on.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23339","https://github.com/NVIDIA/product-security/tree/main/2025/5661"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:L/I:N/A:N","cwe":["CWE-121"],"published":"2025-09-24"},{"id":"CVE-2025-23340","cve":"CVE-2025-23340","aliases":[],"title":"NVIDIA CUDA Toolkit - nvdisasm: Another out-of-bounds read on a malformed ELF, fixed alongside the rest","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA CUDA Toolkit - nvdisasm","year":"2025","cvss_score":3.3,"severity":"low","kev":false,"impact":"Another out-of-bounds read on a malformed ELF, fixed alongside the rest. The realistic exposure is your build and profiling pipeline, not your runtime fleet: anything that automatically disassembles third-party fatbins, vendor kernels or model artifacts is running this parser on attacker-influenced input.","attack_vector":"Local, and requires a user or an automated job to run nvdisasm over an attacker-supplied file. CI jobs that inspect third-party CUDA binaries are the usual path.","remediation":"Update the CUDA Toolkit package (bulletin 5661). Cost: effectively zero - userspace SDK only, no driver reload, no node drain, no running-job impact. Rebuild build/CI images and move on.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23340","https://github.com/NVIDIA/product-security/tree/main/2025/5661"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-125"],"published":"2025-09-24"},{"id":"CVE-2025-23346","cve":"CVE-2025-23346","aliases":[],"title":"NVIDIA CUDA Toolkit - cuobjdump: A null-pointer dereference from an unprivileged user crashes the tool","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA CUDA Toolkit - cuobjdump","year":"2025","cvss_score":3.3,"severity":"low","kev":false,"impact":"A null-pointer dereference from an unprivileged user crashes the tool. The realistic exposure is your build and profiling pipeline, not your runtime fleet: anything that automatically disassembles third-party fatbins, vendor kernels or model artifacts is running this parser on attacker-influenced input.","attack_vector":"Local, and requires a user or an automated job to run cuobjdump over an attacker-supplied file. CI jobs that inspect third-party CUDA binaries are the usual path.","remediation":"Update the CUDA Toolkit package (bulletin 5661). Cost: effectively zero - userspace SDK only, no driver reload, no node drain, no running-job impact. Rebuild build/CI images and move on.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23346","https://github.com/NVIDIA/product-security/tree/main/2025/5661"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-476"],"published":"2025-09-24"},{"id":"CVE-2025-33198","cve":"CVE-2025-33198","aliases":[],"title":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware: A second resource-reuse path in SROOT firmware leaks","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware","year":"2025","cvss_score":3.3,"severity":"low","kev":false,"impact":"A second resource-reuse path in SROOT firmware leaks residual data. These sit in the GB10 root-of-trust chain, so a successful exploit undermines the platform's own attestation and secure-boot story rather than just the OS above it.","attack_vector":"Local access to the DGX Spark. Several of the set need no privileges at all; the rest need host root. This is a desk-side developer box, so physical and local access assumptions are much weaker than for a racked DGX.","remediation":"Apply the DGX Spark firmware update from bulletin 5720. Cost: flash plus reboot, low drain cost given the form factor, but not live-patchable and root-of-trust firmware cannot be rolled back once applied.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33198","https://github.com/NVIDIA/product-security/tree/main/2025/5720"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:N/A:N","cwe":["CWE-226"],"fleet":{"pain_class":"firmware-flash"},"published":"2025-11-25"},{"id":"CVE-2025-54410","cve":"CVE-2025-54410","aliases":[],"title":"Docker / moby: Related firewalld handling defect affecting Moby port exposure","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2025","cvss_score":3.3,"severity":"low","kev":false,"impact":"Related firewalld handling defect affecting Moby port exposure","attack_vector":"Unauthenticated network","remediation":"Upgrade Docker Engine","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-54410"],"status":"curated","published":"2025-07-30"},{"id":"CVE-2026-41579","cve":"CVE-2026-41579","aliases":[],"title":"runc: setupPtmx/setupDev rootfs setup flaw during container rootfs construction","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"runc","year":"2026","cvss_score":3.3,"severity":"low","kev":false,"impact":"setupPtmx/setupDev rootfs setup flaw during container rootfs construction","attack_vector":"Any tenant workload with crafted rootfs","remediation":"Replace runc binary at next maintenance window; low severity, batch with other node work","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-41579"],"status":"curated","published":"2026-07-01"},{"id":"CVE-2023-20573","cve":"CVE-2023-20573","aliases":[],"title":"AMD SEV-SNP - debug exception delivery to guests: A privileged attacker can suppress delivery of debug exceptions","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV-SNP - debug exception delivery to guests","year":"2023","cvss_score":3.2,"severity":"low","kev":false,"impact":"A privileged attacker can suppress delivery of debug exceptions to SEV-SNP guests. The guest does not get debug information it expects, which is mostly an availability and observability problem - but for a guest that relies on debug exceptions as part of a self-protection or integrity-checking scheme, silently swallowing them removes that check.","attack_vector":"Privileged host attacker against a confidential guest.","remediation":"Fixed in AMD SEV firmware / AGESA and reaches you as an OEM SBIOS package - AMD hands AGESA to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before shipping BIOS. **Budget one to six months of OEM lag**, longer on older platforms and sometimes never on end-of-support SKUs. Applying it means draining the host and doing a full power cycle. Because the fix moves the platform's reported SEV-SNP TCB version, you must also pull fresh VCEK certificates from AMD's Key Distribution Service and update any attestation policy your tenants pin - otherwise guests will start failing launch validation the moment the BIOS lands. Some SEV firmware can alternatively be staged from linux-firmware (amd/amd_sev_*.sbin) and committed via the ccp driver at boot, which is faster than waiting on BIOS - check whether your platform supports firmware hot-load before assuming the OEM is the only route. Low priority relative to the RMP-bypass and microcode issues; batch it into the next BIOS wave.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-20573","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-01-11"},{"id":"CVE-2024-21977","cve":"CVE-2024-21977","aliases":[],"title":"AMD CPU microcode - RDRAND entropy after patch load: Incomplete cleanup after loading a microcode patch degrades the","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD CPU microcode - RDRAND entropy after patch load","year":"2024","cvss_score":3.2,"severity":"low","kev":false,"impact":"Incomplete cleanup after loading a microcode patch degrades the entropy of RDRAND, so a privileged attacker can weaken the randomness that SEV-SNP guests draw on. Guests that seed keys or nonces from RDRAND get predictable material, which quietly breaks their crypto without breaking anything visible. The score is low; the failure mode - silent, undetectable, affects key generation - is not.","attack_vector":"Local, privileged attacker who can trigger microcode patch loading. Effect lands on SEV-SNP guests on that host.","remediation":"Fixed by an AMD microcode patch. Two delivery routes, and the difference matters: the linux-firmware amd-ucode blobs load early at boot (initramfs) and need only a reboot, while the durable fix is the microcode embedded in the OEM SBIOS/AGESA package, which carries the usual one-to-six-month OEM lag and a full power cycle. **For confidential computing you need the SBIOS route**: microcode late-loaded by the OS is not part of what SEV-SNP attests, so a guest checking the attestation report cannot tell the fix is present. AMD does not support late-loading microcode on a running EPYC host - treat this as reboot-required. After patching, expect the reported TCB version to change and plan the VCEK certificate refresh accordingly. Guests that generated long-lived keys on an affected host should rotate them; patching stops future weak output but does nothing about keys already derived from degraded entropy.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21977","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2025-09-05"},{"id":"CVE-2024-36331","cve":"CVE-2024-36331","aliases":[],"title":"AMD CPU cache initialization - SEV-SNP guest memory integrity: Improper initialization of CPU cache memory lets a","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD CPU cache initialization - SEV-SNP guest memory integrity","year":"2024","cvss_score":3.2,"severity":"low","kev":false,"impact":"Improper initialization of CPU cache memory lets a hypervisor-privileged attacker overwrite SEV-SNP guest memory, costing guest data integrity. Cache-state manipulation is a recurring theme in SEV attacks (the CacheWarp research works the same seam) because the encryption protects DRAM, not what the cache does on the way there.","attack_vector":"Hypervisor-privileged attacker.","remediation":"Fixed in AMD SEV firmware / AGESA and reaches you as an OEM SBIOS package - AMD hands AGESA to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before shipping BIOS. **Budget one to six months of OEM lag**, longer on older platforms and sometimes never on end-of-support SKUs. Applying it means draining the host and doing a full power cycle. Because the fix moves the platform's reported SEV-SNP TCB version, you must also pull fresh VCEK certificates from AMD's Key Distribution Service and update any attestation policy your tenants pin - otherwise guests will start failing launch validation the moment the BIOS lands. Some SEV firmware can alternatively be staged from linux-firmware (amd/amd_sev_*.sbin) and committed via the ccp driver at boot, which is faster than waiting on BIOS - check whether your platform supports firmware hot-load before assuming the OEM is the only route.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36331","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2025-09-06"},{"id":"CVE-2025-33199","cve":"CVE-2025-33199","aliases":[],"title":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware: Incorrect control-flow behaviour in SROOT firmware","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware","year":"2025","cvss_score":3.2,"severity":"low","kev":false,"impact":"Incorrect control-flow behaviour in SROOT firmware permits data tampering. These sit in the GB10 root-of-trust chain, so a successful exploit undermines the platform's own attestation and secure-boot story rather than just the OS above it.","attack_vector":"Local access to the DGX Spark. Several of the set need no privileges at all; the rest need host root. This is a desk-side developer box, so physical and local access assumptions are much weaker than for a racked DGX.","remediation":"Apply the DGX Spark firmware update from bulletin 5720. Cost: flash plus reboot, low drain cost given the form factor, but not live-patchable and root-of-trust firmware cannot be rolled back once applied.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33199","https://github.com/NVIDIA/product-security/tree/main/2025/5720"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:C/C:N/I:L/A:N","cwe":["CWE-670"],"fleet":{"pain_class":"firmware-flash"},"published":"2025-11-25"},{"id":"CVE-2021-25740","cve":"CVE-2021-25740","aliases":[],"title":"Kubernetes: Endpoint/EndpointSlice confused-deputy lets users reach networks they should not","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes","year":"2021","cvss_score":3.1,"severity":"low","kev":false,"impact":"Endpoint/EndpointSlice confused-deputy lets users reach networks they should not","attack_vector":"Cluster user with namespace access","remediation":"No complete upstream fix; restrict Endpoint creation via admission policy","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-25740"],"status":"curated","published":"2021-09-20"},{"id":"CVE-2021-32718","cve":"CVE-2021-32718","aliases":[],"title":"RabbitMQ: Unsanitized username rendered in the management UI","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"RabbitMQ","year":"2021","cvss_score":3.1,"severity":"low","kev":false,"impact":"Unsanitized username rendered in the management UI -> stored XSS against an admin","attack_vector":"Network (remote)","remediation":"Control-plane: management plugin upgrade; keep the UI off the public internet","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-32718"],"status":"curated","published":"2021-06-28"},{"id":"CVE-2023-32082","cve":"CVE-2023-32082","aliases":[],"title":"etcd: LeaseTimeToLive exposes key names to a user without read permission on those keys","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"etcd","year":"2023","cvss_score":3.1,"severity":"low","kev":false,"impact":"LeaseTimeToLive exposes key names to a user without read permission on those keys","attack_vector":"Network (remote)","remediation":"Control-plane: etcd upgrade; audit RBAC on the cluster store","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-32082"],"status":"curated","published":"2023-05-11"},{"id":"CVE-2023-46737","cve":"CVE-2023-46737","aliases":[],"title":"cosign / sigstore: Attacker-controlled registry returns unbounded attestations, DoSing the verifier","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"cosign / sigstore","year":"2023","cvss_score":3.1,"severity":"low","kev":false,"impact":"Attacker-controlled registry returns unbounded attestations, DoSing the verifier","attack_vector":"Malicious registry, e.g. a tenant-specified image source","remediation":"Upgrade cosign; restrict which registries the admission controller will contact","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-46737"],"status":"curated","published":"2023-11-07"},{"id":"CVE-2024-51744","cve":"CVE-2024-51744","aliases":[],"title":"Prometheus / Thanos (golang-jwt): Unclear ParseWithClaims error behavior","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Prometheus / Thanos (golang-jwt)","year":"2024","cvss_score":3.1,"severity":"low","kev":false,"impact":"Unclear ParseWithClaims error behavior -> callers may accept an expired-and-invalid token","attack_vector":"Network (remote)","remediation":"Control-plane: dependency bump and rebuild of Go control-plane services","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-51744"],"status":"curated","published":"2024-11-04"},{"id":"CVE-2024-7598","cve":"CVE-2024-7598","aliases":[],"title":"Kubernetes (kube-apiserver): NetworkPolicy is not applied during a race in namespace termination, so pods briefly run","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-apiserver)","year":"2024","cvss_score":3.1,"severity":"low","kev":false,"impact":"NetworkPolicy is not applied during a race in namespace termination, so pods briefly run unrestricted","attack_vector":"Cluster user who can create and delete namespaces","remediation":"Rolling control-plane upgrade; treat namespace churn as a policy-gap window","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated","published":"2025-03-20"},{"id":"CVE-2026-15605","cve":"CVE-2026-15605","aliases":[],"title":"wandb SDK (`ArtifactManifestEntry.download`): Hash-handling weakness in artifact download integrity","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"wandb SDK (`ArtifactManifestEntry.download`)","year":"2026","cvss_score":3.1,"severity":"low","kev":false,"impact":"Hash-handling weakness in artifact download integrity","attack_vector":"Poisoned artifact in the registry","remediation":"Upgrade the SDK in base images; weakens artifact-integrity guarantees for model supply chain","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-15605"],"status":"curated","published":"2026-07-13"},{"id":"CVE-2026-24513","cve":"CVE-2026-24513","aliases":[],"title":"ingress-nginx: auth-url protection bypass","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2026","cvss_score":3.1,"severity":"low","kev":false,"impact":"auth-url protection bypass; authentication in front of a tenant service can be skipped","attack_vector":"Unauthenticated network","remediation":"Rolling controller upgrade","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated","published":"2026-02-03"},{"id":"CVE-2021-25743","cve":"CVE-2021-25743","aliases":[],"title":"Kubernetes (kubectl): kubectl does not neutralise ANSI escape sequences in output","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kubectl)","year":"2021","cvss_score":3,"severity":"low","kev":false,"impact":"kubectl does not neutralise ANSI escape sequences in output; terminal injection on the operator's machine","attack_vector":"Any tenant who can set an object field the operator will print","remediation":"Upgrade kubectl on operator machines","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-25743"],"status":"curated","published":"2022-01-07"},{"id":"CVE-2021-41190","cve":"CVE-2021-41190","aliases":[],"title":"OCI Distribution Spec: Content-Type alone determines manifest type, so a manifest can be interpreted differently","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"OCI Distribution Spec","year":"2021","cvss_score":3,"severity":"low","kev":false,"impact":"Content-Type alone determines manifest type, so a manifest can be interpreted differently by different clients; signature and policy confusion","attack_vector":"Malicious image in any registry","remediation":"Upgrade registry and client tooling; pin by digest rather than tag","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-41190"],"status":"curated","published":"2021-11-17"},{"cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:L/UI:R/S:C/C:N/I:L/A:N","cwe":["CWE-843"],"fleet":{"pain_class":"node-drain"},"id":"NCVD-2021-013-containerd-oci-manifest-index-pa","cve":null,"aliases":["GHSA-5j5w-g665-5m35"],"title":"containerd (OCI manifest / index parsing, Content-Type handling): SUPPLY CHAIN: the image digest stops being an","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd (OCI manifest / index parsing, Content-Type handling)","year":"2021","cvss_score":3,"severity":"low","kev":false,"impact":"SUPPLY CHAIN: the image digest stops being an unambiguous identifier. Manifest and index documents in the OCI Distribution and Image specs are ambiguous without a Content-Type header, and affected containerd versions trust that header to decide how to deserialise the document. A registry that returns a different Content-Type across two pulls of the same digest gets that digest interpreted as two different images. Everything an operator builds on digest pinning — admission policies that pin by digest, signature verification bound to a digest, reproducible node images across a GPU fleet — quietly loses its guarantee, because the thing being verified and the thing being run can diverge while the digest matches.","attack_vector":"Network, requiring a malicious or compromised registry (or a MITM on registry traffic) able to vary the Content-Type it returns, plus a pull of the affected image. Low privileges and user interaction in the sense that someone has to pull.","remediation":"Upgrade containerd to 1.4.12 or 1.5.8 or later, which reject manifests containing a 'manifests' field and indices containing a 'layers' field; drain the node for the runtime restart. Until then, pull only from registries you control or trust, and treat digest pinning alone as insufficient assurance on affected versions.","references":["https://github.com/containerd/containerd/security/advisories/GHSA-5j5w-g665-5m35","https://github.com/opencontainers/distribution-spec/security/advisories/GHSA-mc8v-mgrf-8f4m","https://github.com/opencontainers/image-spec/security/advisories/GHSA-77vh-xpmg-72qh"],"status":"curated"},{"id":"CVE-2021-41089","cve":"CVE-2021-41089","aliases":[],"title":"Docker / moby: `docker cp` into a crafted container changes Unix permissions of existing host files","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2021","cvss_score":2.8,"severity":"low","kev":false,"impact":"`docker cp` into a crafted container changes Unix permissions of existing host files","attack_vector":"Any tenant workload on a node where operators run docker cp","remediation":"Upgrade Docker Engine","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-41089"],"status":"curated","published":"2021-10-04"},{"id":"CVE-2023-31028","cve":"CVE-2023-31028","aliases":[],"title":"NVIDIA nvJPEG2000 library: Improper input validation on a crafted JPEG2000 file causes a partial denial of service","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA nvJPEG2000 library","year":"2023","cvss_score":2.8,"severity":"low","kev":false,"impact":"Improper input validation on a crafted JPEG2000 file causes a partial denial of service in the decoding library. Low severity on its own, but nvJPEG2000 sits inside DALI and medical/geospatial imaging pipelines that ingest customer files by design, so the untrusted-input assumption is real.","attack_vector":"Local, requires the library to decode an attacker-supplied image. Any data-loading pipeline that accepts tenant or customer imagery is the delivery path.","remediation":"Update the nvJPEG2000 library per bulletin 5517 and rebuild the images that link it. Cost: package update and job restart only; no driver or firmware change.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31028","https://github.com/NVIDIA/product-security/tree/main/2024/5517"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-20"],"published":"2024-04-05"},{"id":"CVE-2024-0080","cve":"CVE-2024-0080","aliases":[],"title":"nvTIFF library: DoS via malformed image","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"nvTIFF library","year":"2024","cvss_score":2.8,"severity":"low","kev":false,"impact":"DoS via malformed image","attack_vector":"Malicious dataset input","remediation":"Bump nvTIFF in image-processing images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0080","https://github.com/NVIDIA/product-security/tree/main/2024/5517"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-20"],"published":"2024-04-05"},{"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:N/I:L/A:N","cwe":["CWE-1220"],"fleet":{"pain_class":"hot-patch"},"id":"CVE-2024-52814","cve":"CVE-2024-52814","aliases":["GHSA-h974-w8pg-cx73"],"title":"Argo Workflows Helm chart (argo-helm, workflow-role privileges on workflowtasksets / workflowartifactgctasks): The","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo Workflows Helm chart (argo-helm, workflow-role privileges on workflowtasksets / workflowartifactgctasks)","year":"2024","cvss_score":2.8,"severity":"low","kev":false,"impact":"The chart hands workflowtasksets and workflowartifactgctasks permissions to every workflow pod when only agent and artifact-GC pods need them. A tenant workload can tamper with status reporting for other pods and templates. Impact is limited to status integrity, not code execution.","attack_vector":"A user who can get a workflow executed in the namespace, on argo-helm charts below 0.45.0.","remediation":"Upgrade the argo-workflows Helm chart to 0.45.0 or later and apply. Roll it together with the 0.44.0 fix for CVE-2024-52799 - both are edits to the same workflow-role and land without a controller restart.","references":["https://github.com/argoproj/argo-helm/security/advisories/GHSA-h974-w8pg-cx73","https://nvd.nist.gov/vuln/detail/CVE-2024-52814"],"status":"curated"},{"id":"CVE-2024-53878","cve":"CVE-2024-53878","aliases":[],"title":"CUDA Toolkit: DoS (improper type checking)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2024","cvss_score":2.8,"severity":"low","kev":false,"impact":"DoS (improper type checking)","attack_vector":"Malicious artifact","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53878","https://github.com/NVIDIA/product-security/tree/main/2025/5594"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-1284"],"published":"2025-02-25"},{"id":"CVE-2024-53879","cve":"CVE-2024-53879","aliases":[],"title":"CUDA Toolkit: DoS (improper type checking)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2024","cvss_score":2.8,"severity":"low","kev":false,"impact":"DoS (improper type checking)","attack_vector":"Malicious artifact","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53879","https://github.com/NVIDIA/product-security/tree/main/2025/5594"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-1284"],"published":"2025-02-25"},{"id":"CVE-2021-25737","cve":"CVE-2021-25737","aliases":[],"title":"Kubernetes: Endpoint IPs can redirect pod traffic to private node networks","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes","year":"2021","cvss_score":2.7,"severity":"low","kev":false,"impact":"Endpoint IPs can redirect pod traffic to private node networks","attack_vector":"Cluster user able to create Endpoints","remediation":"Rolling control-plane upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-25737"],"status":"curated","published":"2021-09-06"},{"id":"CVE-2024-3177","cve":"CVE-2024-3177","aliases":[],"title":"Kubernetes (kube-apiserver): Init/ephemeral container envFrom bypasses the ServiceAccount mountable-secrets policy","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-apiserver)","year":"2024","cvss_score":2.7,"severity":"low","kev":false,"impact":"Init/ephemeral container envFrom bypasses the ServiceAccount mountable-secrets policy","attack_vector":"Cluster user with namespace access","remediation":"Rolling control-plane upgrade; no GPU drain","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated","fleet":{"ubiquity":"Universal component, low impact - every cluster runs the ServiceAccount admission plugin","remediation_pain":"`daemon-restart` - control-plane-only upgrade, no GPU node drain needed","pain_class":"node-drain","why_fleet_wide":"`envFrom` bypasses the mountable-secrets restriction, leaking secrets across a namespace boundary; control-plane-scoped and low severity, so not a fleet emergency"},"published":"2024-04-22"},{"id":"CVE-2025-4563","cve":"CVE-2025-4563","aliases":[],"title":"Kubernetes (kube-apiserver): Nodes can bypass DRA authorization checks","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-apiserver)","year":"2025","cvss_score":2.7,"severity":"low","kev":false,"impact":"Nodes can bypass DRA authorization checks. Directly relevant: DRA is how GPUs are allocated in modern clusters","attack_vector":"A compromised node / kubelet credential","remediation":"Rolling control-plane upgrade; no GPU drain, but re-audit DRA ResourceClaim allocations","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated","published":"2025-06-23"},{"id":"CVE-2025-25183","cve":"CVE-2025-25183","aliases":[],"title":"vLLM (prefix cache hash collisions): Crafted prompts collide hashes","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (prefix cache hash collisions)","year":"2025","cvss_score":2.6,"severity":"low","kev":false,"impact":"Crafted prompts collide hashes → cache reuse across requests","attack_vector":"Co-tenant on a shared instance","remediation":"Same as above — architectural, not patchable away","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-25183"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"published":"2025-02-07"},{"id":"CVE-2025-46570","cve":"CVE-2025-46570","aliases":[],"title":"vLLM (prefix cache): Prefix-cache timing side channel leaks other tenants' prompts","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (prefix cache)","year":"2025","cvss_score":2.6,"severity":"low","kev":false,"impact":"Prefix-cache timing side channel leaks other tenants' prompts","attack_vector":"Co-tenant issuing timed prompts against a shared serving instance","remediation":"No clean fix while prefix caching is shared. Do not share a vLLM instance across tenants — the cache is a cross-tenant channel","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-46570"],"status":"curated","published":"2025-05-29"},{"id":"CVE-2023-20581","cve":"CVE-2023-20581","aliases":[],"title":"AMD IOMMU access control - SEV-SNP RMP check bypass (AMD-SB-3009): An IOMMU access-control flaw lets a privileged","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD IOMMU access control - SEV-SNP RMP check bypass (AMD-SB-3009)","year":"2023","cvss_score":2.5,"severity":"low","kev":false,"impact":"An IOMMU access-control flaw lets a privileged attacker bypass reverse-map table checks, undermining SEV-SNP guest memory protection. One of a family of IOMMU-mediated RMP bypasses - the RMP guards CPU accesses well, and the recurring weakness is device accesses arriving through the IOMMU's error and edge-case paths.","attack_vector":"Privileged attacker with a compromised hypervisor, driving DMA.","remediation":"Fixed in AMD PI/AGESA firmware and delivered only as an OEM SBIOS package - AMD ships the PI drop to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before releasing BIOS. **Budget one to six months of OEM lag**, and note that several CVEs in this batch are marked 'no fix planned' on Naples (EPYC 7001) - for those the only remediation is retiring the hardware. Applying it means cordon, drain and a full power cycle per node; there is no driver reload, no live patch and no VBIOS step. Because this touches the SEV-SNP trust boundary, the update moves the platform TCB version: refresh VCEK certificates from AMD's KDS and update tenant attestation policy, or confidential guest launches will fail immediately after the BIOS lands. Patch as a set with the other IOMMU/RMP bypasses rather than individually - they share firmware releases and any one of them reopens the class.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-20581","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3009.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"tags":["tenant-isolation"],"published":"2025-02-11"},{"id":"CVE-2024-27457","cve":"CVE-2024-27457","aliases":[],"title":"Intel TDX module firmware: Missing check for an exceptional condition in the TDX module allows a privileged user","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel TDX module firmware","year":"2024","cvss_score":2.5,"severity":"low","kev":false,"impact":"Missing check for an exceptional condition in the TDX module allows a privileged user to reach information disclosure. Scored very low, but it is inside the TDX TCB, so it still triggers a module SVN bump and therefore a re-attestation cycle.","attack_vector":"Privileged host user.","remediation":"Update the Intel TDX module. The TDX module is loaded by the SEAM loader at boot, so the practical rollout is: stage the new module, drain every trust domain off the node, and reboot. It is not a live-patchable component and running TDs cannot be migrated through it. After the update, every TD must re-attest because the TDX module SVN is part of the attestation report - so anything that pinned the old measurement will fail until you update your attestation policy too. No OEM BIOS release needed for the module itself, which makes this materially faster than a platform firmware update.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-27457","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01099.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-10-08"},{"id":"CVE-2025-23273","cve":"CVE-2025-23273","aliases":[],"title":"CUDA Toolkit: DoS (division by zero)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2025","cvss_score":2.5,"severity":"low","kev":false,"impact":"DoS (division by zero)","attack_vector":"Malicious artifact","remediation":"Bump CUDA Toolkit; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23273","https://github.com/NVIDIA/product-security/tree/main/2025/5661"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:N/I:N/A:L","cwe":["CWE-369"],"published":"2025-09-24"},{"id":"CVE-2025-23290","cve":"CVE-2025-23290","aliases":[],"title":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko): A guest can read global GPU metrics","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko)","year":"2025","cvss_score":2.5,"severity":"low","kev":false,"impact":"A guest can read global GPU metrics that are influenced by work running in other tenants' VMs. The attacker sits inside a tenant guest VM and the blast radius is the hypervisor host, which is running every other tenant's vGPU on the same physical GPU. This is precisely the boundary a vGPU-based multi-tenant offering is sold on. Low CVSS (2.5) and no memory corruption, but this is a genuine cross-tenant side channel: utilisation, clock and memory-pressure telemetry that moves with a neighbour's workload leaks the shape of that workload. If you sell confidential or isolated GPU capacity, this is a claim you cannot make while it is unpatched.","attack_vector":"A tenant inside their own guest VM, driving the paravirtualised vGPU control interface. Several of these need only an unprivileged process in the guest; the rest need guest root, which a tenant already has on a VM they rented. No host credentials are involved at any point.","remediation":"Patch the vGPU Manager on the hypervisor host per bulletin 5670. Cost: the highest of any class here. The host driver cannot be reloaded while vGPUs are attached, so every tenant VM on that hypervisor must be live-migrated or powered off - a full host drain. NVIDIA also enforces a supported host/guest driver skew, so budget a matching guest-driver campaign in the same window or tenants lose their vGPU on next boot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23290","https://github.com/NVIDIA/product-security/tree/main/2025/5670"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:L/I:N/A:N","cwe":["CWE-200"],"fleet":{"pain_class":"node-drain"},"tags":["tenant-isolation"],"published":"2025-08-02"},{"id":"CVE-2025-55703","cve":"CVE-2025-55703","aliases":[],"title":"Sunbird Power IQ 9.2.0 API: Error-based SQL injection through an outdated API endpoint with missing input validation","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Sunbird Power IQ 9.2.0 API","year":"2025","cvss_score":2.5,"severity":"low","kev":false,"impact":"Error-based SQL injection through an outdated API endpoint with missing input validation. Low score, but Power IQ is the power-monitoring layer that holds PDU credentials and outlet-level topology for the estate, so any read primitive into its database is worth closing.","attack_vector":"Access to the Power IQ API.","remediation":"Apply the Sunbird fix. Additionally, disable legacy API endpoints you do not use - the root cause here is an old endpoint left enabled.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-55703"],"status":"curated","published":"2025-12-15"},{"id":"CVE-2025-23291","cve":"CVE-2025-23291","aliases":[],"title":"NVIDIA License System - Delegated Licensing Service (DLS): An authorised-looking action leads to information disclosure","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA License System - Delegated Licensing Service (DLS)","year":"2025","cvss_score":2.4,"severity":"low","kev":false,"impact":"An authorised-looking action leads to information disclosure from the licensing service. Low severity and high complexity, but it is inventory data about your entire vGPU estate.","attack_vector":"Adjacent network, high privileges and user interaction required - realistically an insider or a compromised admin session.","remediation":"Update the DLS appliance per bulletin 5705. Cost: appliance restart, no tenant impact.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23291","https://github.com/NVIDIA/product-security/tree/main/2025/5705"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:H/PR:H/UI:R/S:C/C:L/I:N/A:N","cwe":["CWE-312"],"published":"2025-09-30"},{"id":"CVE-2018-25103","cve":"CVE-2018-25103","aliases":["AMI-SA-2024002","originally CVE-2024-3708"],"title":"AMI MegaRAC SPx (embedded lighttpd web server): Use-after-free in the lighttpd request parser embedded in MegaRAC SPx","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx (embedded lighttpd web server)","year":"2018","cvss_score":2.3,"severity":"low","kev":false,"impact":"Use-after-free in the lighttpd request parser embedded in MegaRAC SPx. AMI rates the direct impact as minor - low confidentiality and availability effect via HTTP request smuggling. The reason it belongs on an operator's radar is not its score but what it proves: AMI was still shipping a 2018-era lighttpd in production BMC firmware in 2024, which tells you the embedded OSS stack in your BMCs (lighttpd, nginx, cURL, OpenSSL, busybox) is years behind and is not covered by whatever OS patching process you run on the host.","attack_vector":"Network access to the BMC's web server, unauthenticated but requiring a particular request shape and some user interaction. Reachable from anything that can hit the BMC's HTTP/HTTPS port.","remediation":"Firmware flash to SPx_12.7+ / SPx_13.6, out-of-band per node, ODM-gated. Do not schedule a fleet flash for this CVE alone - its real use is as an argument for building BMC firmware version inventory. The durable action is to start tracking the running BMC build per node and the OSS components inside it, so the next embedded-library CVE is an inventory query rather than a research project.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/2024/AMI-SA-2024002.pdf","https://www.runzero.com/blog/lighttpd/","https://nvd.nist.gov/vuln/detail/CVE-2018-25103"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-06-17"},{"id":"CVE-2025-22853","cve":"CVE-2025-22853","aliases":[],"title":"Intel TDX firmware: Improper synchronisation in TDX firmware, exploitable by a privileged host user to escalate","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel TDX firmware","year":"2025","cvss_score":2.3,"severity":"low","kev":false,"impact":"Improper synchronisation in TDX firmware, exploitable by a privileged host user to escalate. Race conditions in the module are hard to trigger but sit on the tenant boundary.","attack_vector":"Privileged host user, requires winning a race.","remediation":"Update the Intel TDX module. The TDX module is loaded by the SEAM loader at boot, so the practical rollout is: stage the new module, drain every trust domain off the node, and reboot. It is not a live-patchable component and running TDs cannot be migrated through it. After the update, every TD must re-attest because the TDX module SVN is part of the attestation report - so anything that pinned the old measurement will fail until you update your attestation policy too. No OEM BIOS release needed for the module itself, which makes this materially faster than a platform firmware update.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-22853","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01312.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2025-08-12"},{"id":"CVE-2025-33200","cve":"CVE-2025-33200","aliases":[],"title":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware: A third resource-reuse path in SROOT firmware leaks","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware","year":"2025","cvss_score":2.3,"severity":"low","kev":false,"impact":"A third resource-reuse path in SROOT firmware leaks information to a privileged local caller. These sit in the GB10 root-of-trust chain, so a successful exploit undermines the platform's own attestation and secure-boot story rather than just the OS above it.","attack_vector":"Local access to the DGX Spark. Several of the set need no privileges at all; the rest need host root. This is a desk-side developer box, so physical and local access assumptions are much weaker than for a racked DGX.","remediation":"Apply the DGX Spark firmware update from bulletin 5720. Cost: flash plus reboot, low drain cost given the form factor, but not live-patchable and root-of-trust firmware cannot be rolled back once applied.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33200","https://github.com/NVIDIA/product-security/tree/main/2025/5720"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:L/I:N/A:N","cwe":["CWE-226"],"fleet":{"pain_class":"firmware-flash"},"published":"2025-11-25"},{"cvss_vector":"CVSS:4.0/AV:N/AC:L/AT:P/PR:L/UI:N/VC:N/VI:N/VA:L/SC:N/SI:N/SA:N","cwe":["CWE-476"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2026-42183","cve":"CVE-2026-42183","aliases":["GHSA-p4gq-3vxj-f4jq"],"title":"Argo Workflows (Argo Server, SSO RBAC delegation gatekeeper): With SSO_DELEGATE_RBAC_TO_NAMESPACE enabled, an SSO user","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo Workflows (Argo Server, SSO RBAC delegation gatekeeper)","year":"2026","cvss_score":2.3,"severity":"low","kev":false,"impact":"With SSO_DELEGATE_RBAC_TO_NAMESPACE enabled, an SSO user whose claims match a namespace RBAC rule but no SSO-namespace rule triggers a nil dereference and panics the request path. Low impact on its own, but it is a login-time crash, so the users it affects cannot reach the workflow UI or API at all.","attack_vector":"An authenticated SSO user whose group or claim mapping exists in a tenant namespace but not in the SSO namespace. Requires the operator to have turned on RBAC delegation.","remediation":"Upgrade Argo Server to 4.0.5 and restart. As an interim measure either disable SSO_DELEGATE_RBAC_TO_NAMESPACE or ensure every delegated namespace rule has a matching rule in the SSO namespace.","references":["https://github.com/argoproj/argo-workflows/security/advisories/GHSA-p4gq-3vxj-f4jq","https://nvd.nist.gov/vuln/detail/CVE-2026-42183"],"status":"curated"},{"id":"CVE-2020-8562","cve":"CVE-2020-8562","aliases":[],"title":"Kubernetes (kube-apiserver): TOCTOU/DNS-rebinding bypass of the link-local and localhost proxy protections","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-apiserver)","year":"2020","cvss_score":2.2,"severity":"low","kev":false,"impact":"TOCTOU/DNS-rebinding bypass of the link-local and localhost proxy protections","attack_vector":"Cluster user with proxy rights","remediation":"Rolling control-plane upgrade; network-level egress controls on the control plane","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-8562"],"status":"curated","published":"2022-02-01"},{"id":"CVE-2015-3456","cve":"CVE-2015-3456","aliases":[],"title":"QEMU / KVM / Xen (VENOM): VENOM: out-of-bounds write in the virtual Floppy Disk Controller","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"QEMU / KVM / Xen (VENOM)","year":"2015","cvss_score":2,"severity":"low","kev":false,"impact":"VENOM: out-of-bounds write in the virtual Floppy Disk Controller - guest-to-host code execution; present even when the FDC is disabled in the guest config","attack_vector":"Tenant VM guest","remediation":"QEMU update + VM restart. The origin of the \"unused emulated device is still attack surface\" lesson - audit and strip emulated devices from tenant VM templates","references":["https://nvd.nist.gov/vuln/detail/CVE-2015-3456"],"status":"curated","published":"2015-05-13"},{"id":"CVE-2023-0194","cve":"CVE-2023-0194","aliases":[],"title":"GPU Display Driver: Physical memory access","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2023","cvss_score":2,"severity":"low","kev":false,"impact":"Physical memory access","attack_vector":"Local operator with physical access","remediation":"Driver upgrade; no tenant eviction needed","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0194","https://github.com/NVIDIA/product-security/tree/main/2023/5452"],"status":"curated","cvss_vector":"CVSS:3.1/AV:P/AC:H/PR:N/UI:N/S:U/C:N/I:N/A:L","cwe":["CWE-1284"],"published":"2023-04-01"},{"id":"CVE-2023-0195","cve":"CVE-2023-0195","aliases":[],"title":"GPU Display Driver: Physical memory disclosure","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2023","cvss_score":2,"severity":"low","kev":false,"impact":"Physical memory disclosure","attack_vector":"Local operator with physical access","remediation":"Driver upgrade at next window","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0195","https://github.com/NVIDIA/product-security/tree/main/2023/5452"],"status":"curated","cvss_vector":"CVSS:3.1/AV:P/AC:H/PR:N/UI:N/S:U/C:L/I:N/A:N","cwe":["CWE-1284"],"published":"2023-04-01"},{"id":"CVE-2024-23591","cve":"CVE-2024-23591","aliases":["LEN-150020"],"title":"Lenovo ThinkSystem SR670 V2 (shipped in Manufacturing Mode): SR670 V2 servers built between roughly June 2021 and July","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Lenovo ThinkSystem SR670 V2 (shipped in Manufacturing Mode)","year":"2024","cvss_score":2,"severity":"low","kev":false,"impact":"SR670 V2 servers built between roughly June 2021 and July 2023 left the factory still in Manufacturing Mode, which means Intel Boot Guard firmware-integrity enforcement and Intel SPS security settings can be modified or disabled. The CVSS is 2.0 and that number is misleading for a GPU operator: the SR670 V2 is a four-GPU A100/H100-class node, so the affected units are precisely the accelerator fleet, and what is broken is the hardware root of trust that is supposed to stop firmware tampering in the first place. The platform's firmware-resilience protections cannot do their job on a node in this state. This is also the rare entry where the defect ships with the hardware rather than accruing over time, so a node that has never been patched since delivery is affected by construction.","attack_vector":"An attacker with privileged logical access to the host, or physical access to the server internals - so an insider, a technician during an RMA or rack move, or a tenant with root on a bare-metal node during their tenancy. Nothing is reachable over the network.","remediation":"Flash UEFI to U8E126I-2.20 or later, which closes Manufacturing Mode. As a system firmware update it applies on the next reboot, so it costs a drain and a maintenance window on a GPU node. The operational action beyond the flash: check delivery dates against the June 2021 - July 2023 window and verify Boot Guard/Manufacturing Mode state per node, because a firmware version check alone will not tell you whether a given unit shipped in this condition.","references":["https://support.lenovo.com/us/en/product_security/LEN-150020","https://nvd.nist.gov/vuln/detail/CVE-2024-23591"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-02-16"},{"id":"CVE-2025-21096","cve":"CVE-2025-21096","aliases":[],"title":"Intel TDX firmware: Improper buffer restrictions in TDX firmware reachable by a privileged host user for privilege","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel TDX firmware","year":"2025","cvss_score":1.9,"severity":"low","kev":false,"impact":"Improper buffer restrictions in TDX firmware reachable by a privileged host user for privilege escalation. Bundled with the other 2025 TDX firmware fixes.","attack_vector":"Privileged host user.","remediation":"Update the Intel TDX module. The TDX module is loaded by the SEAM loader at boot, so the practical rollout is: stage the new module, drain every trust domain off the node, and reboot. It is not a live-patchable component and running TDs cannot be migrated through it. After the update, every TD must re-attest because the TDX module SVN is part of the attestation report - so anything that pinned the old measurement will fail until you update your attestation policy too. No OEM BIOS release needed for the module itself, which makes this materially faster than a platform firmware update.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21096","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01312.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2025-08-12"},{"id":"CVE-2025-0029","cve":"CVE-2025-0029","aliases":[],"title":"AMD SEV-SNP - selective DMA write drops on host-induced faults: By inducing faults, a high-privileged local attacker","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV-SNP - selective DMA write drops on host-induced faults","year":"2025","cvss_score":1.8,"severity":"low","kev":false,"impact":"By inducing faults, a high-privileged local attacker can selectively drop a confidential guest's DMA writes, costing SEV-SNP guest memory integrity. Selectivity is what makes this more than noise: an attacker who can choose *which* writes vanish can corrupt a computation in a targeted way - dropping a gradient update, a checkpoint write, or a log entry - rather than just breaking the VM.","attack_vector":"Local, high-privileged host attacker, against a confidential guest doing DMA.","remediation":"Fixed in AMD SEV firmware / AGESA and reaches you as an OEM SBIOS package - AMD hands AGESA to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before shipping BIOS. **Budget one to six months of OEM lag**, longer on older platforms and sometimes never on end-of-support SKUs. Applying it means draining the host and doing a full power cycle. Because the fix moves the platform's reported SEV-SNP TCB version, you must also pull fresh VCEK certificates from AMD's Key Distribution Service and update any attestation policy your tenants pin - otherwise guests will start failing launch validation the moment the BIOS lands. Some SEV firmware can alternatively be staged from linux-firmware (amd/amd_sev_*.sbin) and committed via the ccp driver at boot, which is faster than waiting on BIOS - check whether your platform supports firmware hot-load before assuming the OEM is the only route. CVSS 1.8 badly undersells this for anyone whose product claim is 'the host operator cannot tamper with your workload'; rate it against your own trust story, not the score.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0029","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"],"published":"2026-02-10"},{"id":"CVE-2025-48509","cve":"CVE-2025-48509","aliases":[],"title":"AMD SEV firmware - missing checks around RMP initialization (AMD-SB-3023): Missing checks around RMP initialization","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV firmware - missing checks around RMP initialization (AMD-SB-3023)","year":"2025","cvss_score":1.8,"severity":"low","kev":false,"impact":"Missing checks around RMP initialization lead to I/O memory being misidentified, weakening the boundary the reverse-map table enforces. Lowest-severity member of the AMD-SB-3023 RMP batch; include it in the same firmware pass rather than tracking it separately.","attack_vector":"Local, privileged, during SNP platform initialization.","remediation":"Fixed in AMD PI/AGESA firmware and delivered only as an OEM SBIOS package - AMD ships the PI drop to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before releasing BIOS. **Budget one to six months of OEM lag**, and note that several CVEs in this batch are marked 'no fix planned' on Naples (EPYC 7001) - for those the only remediation is retiring the hardware. Applying it means cordon, drain and a full power cycle per node; there is no driver reload, no live patch and no VBIOS step. Because this touches the SEV-SNP trust boundary, the update moves the platform TCB version: refresh VCEK certificates from AMD's KDS and update tenant attestation policy, or confidential guest launches will fail immediately after the BIOS lands.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-48509","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3023.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"tags":["tenant-isolation"],"published":"2026-02-10"},{"id":"CVE-2026-15791","cve":"CVE-2026-15791","aliases":[],"title":"BuildKit: Crafted low-level API message deletes the contents of the host /tmp","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"BuildKit","year":"2026","cvss_score":1.8,"severity":"low","kev":false,"impact":"Crafted low-level API message deletes the contents of the host /tmp","attack_vector":"Anyone with build API access","remediation":"Upgrade BuildKit","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-15791"],"status":"curated","published":"2026-07-21"},{"id":"CVE-2023-20518","cve":"CVE-2023-20518","aliases":[],"title":"AMD Secure Processor - incomplete cleanup exposing the Master Encryption Key (AMD-SB-3003): Incomplete cleanup in the","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor - incomplete cleanup exposing the Master Encryption Key (AMD-SB-3003)","year":"2023","cvss_score":1.6,"severity":"low","kev":false,"impact":"Incomplete cleanup in the ASP exposes the platform Master Encryption Key. CVSS 1.6 is the lowest score in this database and the description is the most alarming sentence in it - the MEK is the key underneath the platform's memory encryption. Scores measure exploitability under a specific model; they do not measure what an attacker walks away with. If the MEK leaks, patching afterwards does not put it back, and the node's cryptographic identity is spent.","attack_vector":"Local, privileged, and per AMD's scoring hard to reach in practice - which is why the number is low.","remediation":"Fixed in AMD PI/AGESA firmware and delivered only as an OEM SBIOS package - AMD ships the PI drop to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before releasing BIOS. **Budget one to six months of OEM lag**, and note that several CVEs in this batch are marked 'no fix planned' on Naples (EPYC 7001) - for those the only remediation is retiring the hardware. Applying it means cordon, drain and a full power cycle per node; there is no driver reload, no live patch and no VBIOS step. Marked 'no fix planned' on some generations. Judge this on the asset at risk rather than the score: if you offer confidential computing, an unpatched-and-unpatchable MEK exposure path is something to know about when you write the SLA, not something to leave at the bottom of a CVSS-sorted queue.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-20518","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3003.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"tags":["tenant-isolation"],"published":"2024-08-13"},{"id":"NCVD-2026-012-qct-quanta-cloud-technology-serv","cve":null,"aliases":[],"title":"QCT (Quanta Cloud Technology) server security centre: QCT firmware is unmeasurable from public data despite","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"QCT (Quanta Cloud Technology) server security centre","year":"2026","cvss_score":0,"severity":"unscored","kev":false,"impact":"QCT firmware is unmeasurable from public data despite the appearance of a PSIRT. This is worse than having no page, because an operator doing vendor due diligence sees a security centre, ticks the box, and moves on - while in reality no QCT-originated BMC or BIOS vulnerability has ever been disclosed publicly, and the channel has been dormant for six years. Quanta is one of the largest ODM server manufacturers in the world and its boards are widely present in cloud and AI datacenter builds, so a substantial amount of deployed BMC firmware has no disclosure history whatsoever. An operator cannot distinguish 'this firmware has no known vulnerabilities' from 'nobody has ever looked, or looked and never told you'. A PSIRT page does exist at qct.io/Press-Releases/index/PR/Server/Security-Center and is readable, but every entry on it is an Intel advisory passed through - Intel-SA-00086, 00088, 00115, 00161, 00125/00131, 00233 - and the newest item dates to 2019. QCT has published no first-party BMC or BIOS advisory at all.","attack_vector":"Not an attack path - a disclosure gap that applies to any operator running QCT or Quanta-manufactured server and GPU chassis hardware.","remediation":"Nothing to flash and nothing to subscribe to. Practical steps: route firmware and security questions through your QCT account team in writing and keep the responses, since the public channel will not serve you; require a firmware support and vulnerability-notification commitment in the purchase agreement before the next order; and in the absence of vendor disclosure, apply the generic BMC controls that do not depend on knowing about specific CVEs - disable IPMI-over-LAN in favour of Redfish over TLS, use per-node unique BMC credentials, disable virtual media and SSH/SMASH on the BMC where unused, isolate the management VLAN, and baseline firmware hashes at turnup so change is at least detectable.","references":["https://www.qct.io/Press-Releases/index/PR/Server/Security-Center"],"status":"curated"},{"id":"NCVD-2026-013-supermicro-s-public-security-adv","cve":null,"aliases":[],"title":"Supermicro's public security advisory portal itself: An operator cannot programmatically track Supermicro firmware","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Supermicro's public security advisory portal itself","year":"2026","cvss_score":0,"severity":"unscored","kev":false,"impact":"An operator cannot programmatically track Supermicro firmware advisories. Supermicro publishes real, detailed BMC and BIOS advisories on a roughly quarterly cadence, but the pages are unreadable to any scanner, SBOM pipeline, or vulnerability-management tool that fetches them without a browser session. The practical result is that Supermicro firmware CVEs enter operator awareness late and by hand, via NVD or a vendor account manager, and that fleet-wide 'are we patched' questions cannot be answered automatically. On a Supermicro-heavy GPU fleet this is a measurement gap, not a vulnerability - but it is the reason the vulnerability entries above have vendor advisory links that will not resolve for your tooling. (supermicro.com/en/support/security_center and the dated security_BMC_IPMI_* / security_BIOS_* advisory pages). Every one of them returns HTTP 403 from ordinary automated clients, including the site root and deliberately bogus paths, which means it is a blanket WAF block rather than a missing page.","attack_vector":"Not an attack - a visibility failure. It affects anyone trying to automate firmware advisory ingestion for a Supermicro fleet from a datacenter or CI egress IP rather than a human browser.","remediation":"There is no fix an operator can apply to the vendor's WAF. What works: subscribe to Supermicro's security notification mailing list through your reseller or account team so advisories arrive by email rather than by scraping; mirror each advisory's contents into your own internal tracker when it lands, since you cannot re-fetch it later; and drive automated detection off NVD and the CVE Program's cvelistV5 records, which do carry the Supermicro CNA entries and are freely fetchable. Budget a human in the loop for every Supermicro advisory cycle.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36435","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2025/12xxx/CVE-2025-12006.json"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2026-014-tyan-mitac-computing-psirt","cve":null,"aliases":[],"title":"Tyan / MiTAC Computing PSIRT: For Tyan, this vendor's firmware is unmeasurable from public data","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Tyan / MiTAC Computing PSIRT","year":"2026","cvss_score":0,"severity":"unscored","kev":false,"impact":"For Tyan, this vendor's firmware is unmeasurable from public data - the only substantive Tyan BMC vulnerability in the public record, the S5552 TLS private key disclosure, was published by a third-party research lab rather than the vendor, and there is now no vendor channel at all through which fixed firmware or future advisories could be obtained. An operator with Tyan hardware in the fleet has a BMC with a known credential-grade disclosure bug and no path to a fix. For Advantech the situation is different but leads to the same operator conclusion: a working PSIRT exists, so subscribe to it, but do not expect it to cover server BMC or BIOS firmware, because it never has. Tyan has been absorbed into MiTAC Computing; www.tyan.com now serves a TLS certificate for *.mitaccomputing.com that does not match its own hostname, www.tyan.com.tw is a stub linking onward to MiTAC, and mitaccomputing.com returns HTTP 403 to automated clients. Advantech, checked alongside as an industrial-server vendor, does run a real and actively maintained PSIRT at advantech.com/en/security-advisory with ACIRT/AQIRT advisory IDs and an RSS feed - but every advisory on it covers IoT, wireless and software products, with no BMC or BIOS content.","attack_vector":"Not an attack path - a vendor-continuity and disclosure gap. It applies to any operator carrying Tyan-branded server boards, which persist in secondhand and budget capacity builds long after the brand's own support channel has gone.","remediation":"For Tyan hardware, assume no firmware fix will arrive. Approach MiTAC Computing directly through a sales channel if you need firmware, and otherwise treat these nodes as permanently unpatched: replace the BMC TLS certificate with one you control, isolate the management VLAN with an explicit allowlist, disable IPMI-over-LAN and virtual media, use per-node unique credentials, and plan the hardware out of the fleet on a defined timeline rather than indefinitely. For Advantech, subscribe to the ACIRT feed but keep server firmware tracking on a separate mechanism, since the feed does not cover it.","references":["https://www.advantech.com/en/security-advisory","https://www.nozominetworks.com/labs/vulnerability-advisories-cve-2023-2538/"],"status":"curated"},{"id":"NCVD-2026-015-wiwynn-celestica-ingrasys-foxcon","cve":null,"aliases":[],"title":"Wiwynn / Celestica / Ingrasys (Foxconn) / AIC BMC firmware: This vendor's firmware is unmeasurable from public data","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Wiwynn / Celestica / Ingrasys (Foxconn) / AIC BMC firmware","year":"2026","cvss_score":0,"severity":"unscored","kev":false,"impact":"This vendor's firmware is unmeasurable from public data, and that is the finding. An operator running Wiwynn, Celestica, Ingrasys or AIC hardware cannot answer basic questions: is there a known vulnerability in this BMC, has a fix shipped, what version am I supposed to be on. There is no feed to subscribe to and no advisory to correlate against a CVE. Since these builders ship BMC firmware derived from the same AMI MegaRAC and ASPEED lineage as everything else in this database, the realistic assumption is that they inherit the same vulnerability classes - unauthenticated REST handlers, weak firmware verification, credential exposure - without any of the disclosure that would let an operator act. A neocloud whose fleet is largely ODM whitebox is running an out-of-band management plane whose security posture it has no mechanism to assess. Four ODM whitebox and OCP chassis builders whose hardware carries a meaningful share of hyperscale and neocloud GPU capacity. Their corporate websites were fetched and read directly and none of them publishes a security advisory page, a PSIRT contact, or a CVE disclosure channel of any kind. Ingrasys serves HTTP 200 for every path including nonsense ones, and its /security and /psirt paths render the site's 404 message.","attack_vector":"Not an attack path - a disclosure gap. It applies to any operator whose GPU capacity sits on OCP or ODM whitebox chassis from builders who sell to hyperscalers under contract and have never built a public-facing security function.","remediation":"There is no patch, because there is no advisory. What an operator can actually do: make PSIRT existence a procurement requirement and get firmware-update commitments and a security contact written into the purchase contract, since these vendors will respond to a customer of size even without a public channel. Obtain firmware through the integrator or hyperscaler channel that sourced the hardware. Treat these BMCs as permanently unpatched and isolate them accordingly - dedicated management VLAN, no route from tenant networks, explicit management-host allowlist. Finally, measure independently: capture a firmware hash baseline at node turnup so you can at least detect change, since you will never be told about a vulnerability.","references":["https://www.wiwynn.com/","https://www.celestica.com/","https://www.ingrasys.com/","https://www.aicipc.com/"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2012-2934","cve":"CVE-2012-2934","aliases":[],"title":"Xen on older AMD CPUs - 64-bit PV guest processor erratum: Xen 4.0 and 4.1 running a 64-bit PV guest on older AMD CPUs","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Xen on older AMD CPUs - 64-bit PV guest processor erratum","year":"2012","cvss_score":null,"severity":"unscored","kev":false,"impact":"Xen 4.0 and 4.1 running a 64-bit PV guest on older AMD CPUs did not protect against a processor erratum, letting a guest OS user hang the host. Historical, but it is the earliest entry in a long pattern worth naming: AMD CPU errata that a hypervisor must actively work around, where forgetting the workaround hands guests a host-availability lever.","attack_vector":"From inside a 64-bit PV guest on affected legacy AMD silicon.","remediation":"Fixed in Xen (XSA-9). Hypervisor update plus reboot. Affected silicon is long retired; carried here for completeness of the AMD virtualisation history rather than as an action item.","references":["https://nvd.nist.gov/vuln/detail/CVE-2012-2934"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2012-12-03"},{"id":"CVE-2013-6885","cve":"CVE-2013-6885","aliases":["Erratum 793"],"title":"AMD 16h processor microcode - locked instructions vs write-combined memory: Interaction between locked instructions","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD 16h processor microcode - locked instructions vs write-combined memory","year":"2013","cvss_score":null,"severity":"unscored","kev":false,"impact":"Interaction between locked instructions and write-combined memory types hangs the system, reachable from an unprivileged local application. A tenant can wedge the whole machine with a small crafted program - no privilege needed, no recovery short of a hard reset. Included for completeness on legacy AMD hardware; it is the archetype of the 'any tenant can halt the node' class that keeps recurring in CPU errata.","attack_vector":"Local, unprivileged. Any process on the machine, including inside a container.","remediation":"Fixed by a BIOS/microcode update carrying the erratum 793 workaround. Affects AMD 16h family processors, long out of production - if you have any of these in a fleet, the realistic answer is that they are past firmware support and should be retired rather than patched.","references":["https://nvd.nist.gov/vuln/detail/CVE-2013-6885"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"published":"2013-11-29"},{"id":"CVE-2018-0021","cve":"CVE-2018-0021","aliases":[],"title":"Juniper Junos OS MACsec key configuration (CKN/CAK): If you configure a MACsec connectivity-association name or key","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Juniper Junos OS MACsec key configuration (CKN/CAK)","year":"2018","cvss_score":null,"severity":"unscored","kev":false,"impact":"If you configure a MACsec connectivity-association name or key shorter than its full length, Junos silently zero-fills the remainder. A 16-character passphrase you believed was a 256-bit key is 16 characters followed by a long run of zeros, and it falls to dictionary and brute-force attack. MACsec is what protects inter-site and inter-pod links carrying every tenant's traffic, so a recoverable CAK means an attacker with a tap decrypts the lot — and nothing in the config output tells you the key is weak.","attack_vector":"An attacker with a passive tap on the MACsec-protected link who recovers the key offline. No access to the devices is needed.","remediation":"Config change, not a patch: reconfigure every MACsec association with the full 64-digit CKN and full 32-digit CAK, generated from a CSPRNG. Rekeying a MACsec link drops it briefly, so do redundant links one at a time. Then audit every MACsec key in the fabric for length — this is the kind of defect that survives for years because the config looks fine.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-0021"],"status":"curated","tags":["tenant-isolation"],"published":"2018-04-11"},{"id":"CVE-2020-10255","cve":"CVE-2020-10255","aliases":["TRRespass"],"title":"DDR4 / LPDDR4 DRAM - Target Row Refresh mitigation: Many-sided Rowhammer defeats the in-DRAM Target Row Refresh","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"DDR4 / LPDDR4 DRAM - Target Row Refresh mitigation","year":"2020","cvss_score":null,"severity":"unscored","kev":false,"impact":"Many-sided Rowhammer defeats the in-DRAM Target Row Refresh mitigation that vendors marketed as making Rowhammer solved. On a GPU host node this is the substrate under your host memory: bit flips in page tables or in a hypervisor's structures are a classic path to cross-VM compromise, and on a bare-metal GPU rental it is a path from tenant code to host.","attack_vector":"Local code on the node with the ability to allocate and access memory at a controlled rate. Any tenant container or VM qualifies; no privileges needed.","remediation":"No universal fix. Practical controls, in order of value: use ECC DIMMs and actually monitor correctable-error rates (a Rowhammer campaign shows up as a correctable-error storm before it succeeds), enable the platform's refresh-rate and RFM settings in BIOS where offered, and prefer DDR5 modules with on-die ECC. Cost: a BIOS setting change is a drain and reboot; a DIMM refresh is a hardware refresh cycle. Effectively UNPATCHABLE in software.","references":["https://www.vusec.net/projects/trrespass/","https://download.vusec.net/papers/trrespass_sp20.pdf","https://nvd.nist.gov/vuln/detail/CVE-2020-10255"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"published":"2020-03-10"},{"id":"CVE-2021-0891","cve":"CVE-2021-0891","aliases":["PowerVR uninitialized heap disclosure"],"title":"Imagination PowerVR GPU driver - memory residue: An unprivileged application gets the GPU driver to hand back","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Imagination PowerVR GPU driver - memory residue","year":"2021","cvss_score":null,"severity":"unscored","kev":false,"impact":"An unprivileged application gets the GPU driver to hand back uninitialized heap memory, disclosing whatever the previous consumer left there. Same residual-memory class as LeftoverLocals and included here as evidence the class is a driver-design pattern rather than a one-off - if you evaluate a non-NVIDIA accelerator, ask the vendor specifically what their driver zeroes on allocation.","attack_vector":"An unprivileged local application with access to the GPU device node.","remediation":"Update to the fixed Imagination DDK / vendor driver. In the datacenter this matters as a due-diligence question rather than a fleet action: PowerVR is not a datacenter part. Cost: driver update and reload where it applies.","references":["https://source.android.com/security/bulletin/2022-08-01","https://nvd.nist.gov/vuln/detail/CVE-2021-0891"],"status":"curated","tags":["tenant-isolation"],"published":"2022-08-24"},{"cwe":["CWE-20"],"fleet":{"pain_class":"daemon-restart"},"id":"CVE-2021-37914","cve":"CVE-2021-37914","aliases":["GHSA-h563-xh25-x54q"],"title":"Argo Workflows (controller, expression template evaluation of input parameters): When EXPRESSION_TEMPLATES is on and","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo Workflows (controller, expression template evaluation of input parameters)","year":"2021","cvss_score":null,"severity":"unscored","kev":false,"impact":"When EXPRESSION_TEMPLATES is on and untrusted users can set workflow input parameters, the supplied value gets evaluated as an expression template. A tenant can disrupt or distort another workflow's execution by feeding in crafted parameters rather than plain data.","attack_vector":"Any user permitted to specify input parameters on a workflow run, on a controller with expression templates enabled.","remediation":"Upgrade the controller to 3.1.6 or later and restart. If you cannot upgrade immediately, disable EXPRESSION_TEMPLATES or stop accepting workflow input parameters from untrusted principals.","references":["https://github.com/argoproj/argo-workflows/issues/6441","https://nvd.nist.gov/vuln/detail/CVE-2021-37914"],"status":"curated","tags":["tenant-isolation"]},{"id":"CVE-2022-23645","cve":"CVE-2022-23645","aliases":[],"title":"swtpm (state blob header parsing): An invalid hdrsize in swtpm's saved state header causes an out-of-bounds access","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"swtpm (state blob header parsing)","year":"2022","cvss_score":null,"severity":"unscored","kev":false,"impact":"An invalid hdrsize in swtpm's saved state header causes an out-of-bounds access, crashing swtpm or preventing it from starting. The operationally sharp version: a VM whose vTPM will not start cannot unseal its disk key, so a corrupted or tampered state blob is a data-availability event, not just a crash. On a GPU cloud that migrates or restores VMs, a bad state file taken across a migration bricks the guest's boot.","attack_vector":"Whoever can write the swtpm state file - host-level access, a compromised migration path, or corruption in the storage holding VM state. Not reachable from inside a well-isolated guest.","remediation":"Package update to swtpm 0.5.3 / 0.6.2 / 0.7.1 or later on hypervisor hosts, then restart swtpm processes - package-level, no reboot of the host required. Worth pairing with an operational control: treat vTPM state blobs as data whose integrity you protect and back up, because losing one is equivalent to losing the guest's disk encryption key.","references":["https://github.com/stefanberger/swtpm/security/advisories/GHSA-2qgm-8xf4-3hqw","https://nvd.nist.gov/vuln/detail/CVE-2022-23645"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2022-02-18"},{"id":"CVE-2022-42335","cve":"CVE-2022-42335","aliases":["XSA-430"],"title":"Xen (shadow paging): x86 shadow paging arbitrary pointer dereference - host crash or worse","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen (shadow paging)","year":"2022","cvss_score":null,"severity":"unscored","kev":false,"impact":"x86 shadow paging arbitrary pointer dereference - host crash or worse","attack_vector":"Tenant VM guest","remediation":"Hypervisor patch + reboot/evacuation","references":["https://xenbits.xen.org/xsa/advisory-430.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2023-04-25"},{"id":"CVE-2022-50617","cve":"CVE-2022-50617","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu/powerplay/psm): A memory or reference-count leak","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu/powerplay/psm)","year":"2022","cvss_score":null,"severity":"unscored","kev":false,"impact":"A memory or reference-count leak in the amdgpu power management (SMU/powerplay). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amdgpu/powerplay/psm: Fix memory leak in power state init","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-50617","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-12-08"},{"id":"CVE-2022-50619","cve":"CVE-2022-50619","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A memory or reference-count leak in the amdkfd (KFD","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2022","cvss_score":null,"severity":"unscored","kev":false,"impact":"A memory or reference-count leak in the amdkfd (KFD compute driver, /dev/kfd). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amdkfd: Fix memory leak in kfd_mem_dmamap_userptr()","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-50619","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-12-08"},{"id":"CVE-2022-50718","cve":"CVE-2022-50718","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): A memory or reference-count leak in the amdgpu kernel driver core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2022","cvss_score":null,"severity":"unscored","kev":false,"impact":"A memory or reference-count leak in the amdgpu kernel driver core. Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amdgpu: fix pci device refcount leak","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-50718","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-12-24"},{"id":"CVE-2022-50760","cve":"CVE-2022-50760","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): A memory or reference-count leak","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2022","cvss_score":null,"severity":"unscored","kev":false,"impact":"A memory or reference-count leak in the amdgpu firmware, ACPI and IP-block initialisation. Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amdgpu: Fix PCI device refcount leak in amdgpu_atrm_get_bios()","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-50760","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-12-24"},{"id":"CVE-2022-50781","cve":"CVE-2022-50781","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (amdgpu/pm): An out-of-bounds access in the amdgpu power","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (amdgpu/pm)","year":"2022","cvss_score":null,"severity":"unscored","kev":false,"impact":"An out-of-bounds access in the amdgpu power management (SMU/powerplay) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: amdgpu/pm: prevent array underflow in vega20_odn_edit_dpm_table()","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-50781","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-12-24"},{"id":"CVE-2022-50844","cve":"CVE-2022-50844","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu): Missing or insufficient validation of user-supplied","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu)","year":"2022","cvss_score":null,"severity":"unscored","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu power management (SMU/powerplay). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdgpu: Fix type of second parameter in odn_edit_dpm_table() callback","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-50844","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-12-30"},{"id":"CVE-2023-40286","cve":"CVE-2023-40286","aliases":[],"title":"Supermicro BMC (IPMI web interface): Part of the same 2023 Supermicro BMC web-interface batch","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC (IPMI web interface)","year":"2023","cvss_score":null,"severity":"unscored","kev":false,"impact":"Part of the same 2023 Supermicro BMC web-interface batch - an injection flaw in the management UI that feeds the session-hijack-to-firmware-flash chain.","attack_vector":"Network reach to the BMC web interface; operator interaction for the injection to land.","remediation":"BMC firmware flash per board, bundled with the rest of the batch. Score not independently confirmed - treat it as equivalent to its siblings for prioritisation.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-40286"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2024-03-27"},{"id":"CVE-2023-52482","cve":"CVE-2023-52482","aliases":[],"title":"Linux x86/srso - SRSO mitigation missing for Hygon processors: The kernel's Speculative Return Stack Overflow","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux x86/srso - SRSO mitigation missing for Hygon processors","year":"2023","cvss_score":null,"severity":"unscored","kev":false,"impact":"The kernel's Speculative Return Stack Overflow (Inception) mitigation was not applied to Hygon processors, which are AMD Zen derivatives and carry the same defect. Any Hygon node in the fleet therefore ran with the SRSO mitigation silently inactive - a cross-privilege speculative disclosure channel that your vulnerability dashboard reported as mitigated.","attack_vector":"Local, cross-privilege speculative execution on Hygon silicon.","remediation":"Fixed in the Linux kernel by extending the SRSO mitigation to Hygon. Distro kernel update plus reboot; no firmware step. Verify afterwards by reading /sys/devices/system/cpu/vulnerabilities/spec_rstack_overflow on Hygon nodes rather than trusting the CPU vendor string to have been handled correctly - this bug existed precisely because it was not.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52482"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-02-29"},{"id":"CVE-2023-53703","cve":"CVE-2023-53703","aliases":[],"title":"Linux HID/amd_sfh - shift out of bounds: A shift operation in the AMD Sensor Fusion Hub driver exceeds the maximum","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux HID/amd_sfh - shift out of bounds","year":"2023","cvss_score":null,"severity":"unscored","kev":false,"impact":"A shift operation in the AMD Sensor Fusion Hub driver exceeds the maximum valid shift value, producing undefined behaviour in the kernel. Same story as the SFH use-after-free: low real exposure on a server, and a good reminder that unused autoloading drivers cost you something.","attack_vector":"Local, on hosts with amd_sfh loaded.","remediation":"Distro kernel update plus reboot, or blacklist the module on server images and carry neither the bug nor the next one.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53703"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-10-22"},{"id":"CVE-2023-53723","cve":"CVE-2023-53723","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): A race condition or locking defect in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu)","year":"2023","cvss_score":null,"severity":"unscored","kev":false,"impact":"A race condition or locking defect in the amdgpu RAS / GPU reset and recovery path. Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdgpu: disable sdma ecc irq only when sdma RAS is enabled in suspend","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53723","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-10-22"},{"id":"CVE-2023-53780","cve":"CVE-2023-53780","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2023","cvss_score":null,"severity":"unscored","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: fix FCLK pstate change underflow","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53780","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-12-09"},{"id":"CVE-2023-54144","cve":"CVE-2023-54144","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A race condition or locking defect in the amdkfd (KFD","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2023","cvss_score":null,"severity":"unscored","kev":false,"impact":"A race condition or locking defect in the amdkfd (KFD compute driver, /dev/kfd). Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdkfd: Fix kernel warning during topology setup","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-54144","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-12-24"},{"id":"CVE-2023-54150","cve":"CVE-2023-54150","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd): An out-of-bounds access in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd)","year":"2023","cvss_score":null,"severity":"unscored","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd: Fix an out of bounds error in BIOS parser","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-54150","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-12-24"},{"id":"CVE-2023-54261","cve":"CVE-2023-54261","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A NULL pointer dereference in the amdkfd (KFD compute","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2023","cvss_score":null,"severity":"unscored","kev":false,"impact":"A NULL pointer dereference in the amdkfd (KFD compute driver, /dev/kfd). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdkfd: Add missing gfx11 MQD manager callbacks","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-54261","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-12-30"},{"id":"CVE-2023-54296","cve":"CVE-2023-54296","aliases":[],"title":"Linux KVM/SVM - source vCPU selection in SEV-ES intra-host migration: KVM fetched source vCPUs from the wrong VM","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux KVM/SVM - source vCPU selection in SEV-ES intra-host migration","year":"2023","cvss_score":null,"severity":"unscored","kev":false,"impact":"KVM fetched source vCPUs from the wrong VM during SEV-ES intra-host migration. Operating on the wrong VM's vCPU structures during a confidential-VM migration is a cross-VM state confusion bug - exactly the failure class you do not want on the code path that moves encrypted guest state around.","attack_vector":"Through the KVM migration ioctl path, from the VMM.","remediation":"Fixed in the Linux kernel. Take the distro kernel update (RHEL/Rocky, Ubuntu, SLES) and reboot the host - no firmware, VBIOS or AGESA step. On a GPU fleet this is a cordon, drain and rolling reboot; plan it as normal kernel maintenance.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-54296"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-12-30"},{"id":"CVE-2024-29038","cve":"CVE-2024-29038","aliases":[],"title":"tpm2-tools (tpm2_checkquote TPM2_GENERATED magic validation): tpm2_checkquote does not verify that the structure","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"tpm2-tools (tpm2_checkquote TPM2_GENERATED magic validation)","year":"2024","cvss_score":null,"severity":"unscored","kev":false,"impact":"tpm2_checkquote does not verify that the structure it is checking was actually generated by a TPM, so a fabricated quote passes validation. The sibling of the tpm2-tss issue in the same disclosure, and it hits the command-line tool that most operators' attestation scripts actually shell out to. Result is the same: a machine with no TPM, or a machine whose TPM state is wrong, can present as attested.","attack_vector":"A malicious or compromised endpoint producing its own quote artefacts. The attacker is the thing claiming to be healthy.","remediation":"Package update to a fixed tpm2-tools wherever verification runs, then restart the verifying service and re-run attestation across the fleet - old results are not evidence. Package-level fix, no reboot. Check whether your attestation pipeline calls tpm2_checkquote, the tpm2-tss FAPI, or its own verifier, because all three had this same missing check and they patch through different channels.","references":["https://github.com/tpm2-software/tpm2-tools/security/advisories/GHSA-5495-c38w-gr6f","https://nvd.nist.gov/vuln/detail/CVE-2024-29038"],"status":"curated","fleet":{"pain_class":"hot-patch"},"published":"2024-06-28"},{"id":"CVE-2024-29039","cve":"CVE-2024-29039","aliases":[],"title":"tpm2-tools (tpm2_checkquote PCR selection handling): tpm2_checkquote does not validate the TPML_PCR_SELECTION","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"tpm2-tools (tpm2_checkquote PCR selection handling)","year":"2024","cvss_score":null,"severity":"unscored","kev":false,"impact":"tpm2_checkquote does not validate the TPML_PCR_SELECTION in the supplied PCR input file, so an attacker who controls that file makes digests map to the wrong PCR slots and banks. The verifier then reports a perfectly valid signature over a misread picture of the machine's state - a node running a tampered boot chain can be made to look like it booted the golden image. Rated critical by the maintainers. For anyone whose product promise is 'verified clean bare metal' or whose scheduler gates on attestation, this is the tool that silently makes the check meaningless.","attack_vector":"The attacker is the attested machine, or anything that can influence the PCR input file the verifier reads. If your verifier consumes quote artefacts uploaded by the node being checked - which is the common design - the node controls its own grade.","remediation":"Package update to a fixed tpm2-tools on every host that verifies quotes, plus a restart of the verifying service. No firmware, no reboot, so it is cheap - but the second half is not: every attestation result produced by the old tool proves nothing, so re-attest the fleet after updating. Structurally, the verifier should pin the expected PCR selection itself rather than accepting one supplied alongside the quote.","references":["https://github.com/tpm2-software/tpm2-tools/security/advisories/GHSA-8rjm-5f5f-h4q6","https://nvd.nist.gov/vuln/detail/CVE-2024-29039"],"status":"curated","published":"2024-06-28"},{"id":"CVE-2024-35907","cve":"CVE-2024-35907","aliases":["mlxbf_gige call request_irq() after NAPI initialized"],"title":"Linux kernel mlxbf_gige (BlueField out-of-band management NIC): NULL function-pointer dereference when the DPU's","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel mlxbf_gige (BlueField out-of-band management NIC)","year":"2024","cvss_score":null,"severity":"unscored","kev":false,"impact":"NULL function-pointer dereference when the DPU's oob_net0 interface comes up, reproduced on BlueField-3 under kdump. oob_net0 is the out-of-band management interface on the DPU - the path you use to reach a DPU whose in-band networking is broken. Losing it during a crash-dump or recovery boot means losing the remote hands you were counting on.","attack_vector":"Local on the DPU - triggered by interface bring-up during kdump or an unusual boot sequence, not by an external attacker.","remediation":"Upgrade the DPU Arm-side kernel to 6.9 or a stable backport (5.15.154, 6.1.85, 6.6.26, 6.8.5), delivered as a DOCA/BFOS package upgrade or BFB re-image with a DPU reset per node. Bundle with other DPU-side fixes into one maintenance pass.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-35907","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2024/CVE-2024-35907.json"],"status":"curated","published":"2024-05-19"},{"id":"CVE-2024-42142","cve":"CVE-2024-42142","aliases":["net/mlx5 E-switch: create ingress ACL when needed"],"title":"Linux kernel mlx5_core eswitch ingress ACL: The eswitch ingress ACL - the table that enforces per-VF ingress policy","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel mlx5_core eswitch ingress ACL","year":"2024","cvss_score":null,"severity":"unscored","kev":false,"impact":"The eswitch ingress ACL - the table that enforces per-VF ingress policy - is only created when vport metadata match or prio tag is on. Turn vport metadata match off via devlink and bring up an active-backup LAG, and the driver panics on a missing ingress ACL. Beyond the crash, this is a config in which the VF ingress enforcement structure is simply absent when the driver expects it.","attack_vector":"Requires an administrator to have set esw_port_metadata=false via devlink and to be running active-backup LAG. Not tenant-triggered, but it is a realistic operator configuration on bonded ConnectX hosts.","remediation":"Upgrade the host kernel to 6.10 or a stable backport (6.1.98, 6.6.39, 6.9.9). Rolling reboot. Immediate config-level mitigation: keep esw_port_metadata at its default (true) on LAG hosts - a devlink change, no reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-42142","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2024/CVE-2024-42142.json"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2024-07-30"},{"id":"CVE-2024-53241","cve":"CVE-2024-53241","aliases":["XSA-466"],"title":"Xen (x86 speculation): Xen hypercall page unsafe against speculative attacks - guest leaks hypervisor memory","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen (x86 speculation)","year":"2024","cvss_score":null,"severity":"unscored","kev":false,"impact":"Xen hypercall page unsafe against speculative attacks - guest leaks hypervisor memory","attack_vector":"Tenant VM guest","remediation":"Hypervisor patch + host reboot with evacuation","references":["https://xenbits.xen.org/xsa/advisory-466.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2024-12-24"},{"id":"CVE-2025-1713","cve":"CVE-2025-1713","aliases":["XSA-467"],"title":"Xen (VT-d passthrough): Deadlock potential with VT-d and legacy PCI device pass-through - host hang","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen (VT-d passthrough)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"Deadlock potential with VT-d and legacy PCI device pass-through - host hang","attack_vector":"Tenant VM guest with PCI passthrough","remediation":"Hypervisor patch + host reboot. Directly in the path of GPU passthrough, the default topology for a GPU cloud","references":["https://xenbits.xen.org/xsa/advisory-467.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-07-17"},{"id":"CVE-2025-21601","cve":"CVE-2025-21601","aliases":[],"title":"Juniper Junos OS (httpd / J-Web on QFX5120, EX, SRX, MX): Crafted HTTP requests to the web management process drive CPU","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Juniper Junos OS (httpd / J-Web on QFX5120, EX, SRX, MX)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"Crafted HTTP requests to the web management process drive CPU consumption up until the device stops responding. QFX5120 is a common data-center leaf, and the vulnerability is in the management daemon, so an attacker who reaches the management interface can take the switch's control plane out without credentials.","attack_vector":"Unauthenticated, remote — reachability to the device's web management service. Only exposed if J-Web / HTTP management is enabled, which many operators leave on for convenience.","remediation":"Junos upgrade (24.2R2 or later on QFX5120) plus reboot. The far cheaper immediate fix is a config change: disable J-Web entirely (`delete system services web-management`) and manage the fabric through NETCONF or the CLI. Most data-center operators should have this off regardless.","references":["https://supportportal.juniper.net/s/article/2025-04-Security-Bulletin-Junos-OS-SRX-and-EX-Series-MX240-MX480-MX960-QFX5120-Series-When-web-management-is-enabled-for-specific-services-an-attacker-may-cause-a-CPU-spike-by-sending-genuine-packets-to-the-device-CVE-2025-21601","https://nvd.nist.gov/vuln/detail/CVE-2025-21601"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["fabric-dos"],"published":"2025-04-09"},{"id":"CVE-2025-21602","cve":"CVE-2025-21602","aliases":[],"title":"Juniper Junos OS / Junos OS Evolved (rpd, BGP UPDATE): A crafted BGP UPDATE crashes the routing protocol daemon. In a","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Juniper Junos OS / Junos OS Evolved (rpd, BGP UPDATE)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"A crafted BGP UPDATE crashes the routing protocol daemon. In a BGP-underlay leaf/spine the rpd going down is the rack losing reachability; because BGP UPDATEs propagate, a single malicious or malformed announcement can ripple to every device that accepts it, which is how a one-switch bug becomes a fabric-wide outage.","attack_vector":"Unauthenticated attacker able to get a crafted UPDATE into the BGP mesh — either as a peer, or upstream of one, since the message is forwarded along.","remediation":"Junos/Junos Evolved upgrade plus reboot, staged so ECMP paths are never both down. Interim: BGP UPDATE filtering and strict inbound policy at the fabric edge, plus `bgp-error-tolerance` style handling where the release supports it — live config changes.","references":["https://supportportal.juniper.net/s/article/2025-01-Security-Bulletin-Junos-OS-and-Junos-OS-Evolved-Receipt-of-specially-crafted-BGP-update-packet-causes-RPD-crash-CVE-2025-21602","https://nvd.nist.gov/vuln/detail/CVE-2025-21602"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["fabric-dos"],"published":"2025-01-09"},{"id":"CVE-2025-2296","cve":"CVE-2025-2296","aliases":["GHSA-6pp6-cm5h-86g5"],"title":"EDK II OvmfPkg (X86QemuLoadImageLib, QemuLoadKernelImage direct-boot path): With Secure Boot on, a kernel","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"EDK II OvmfPkg (X86QemuLoadImageLib, QemuLoadKernelImage direct-boot path)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"With Secure Boot on, a kernel whose signature is not in the allowed database is correctly rejected by image verification - and then loaded anyway, because the code falls back to the legacy loader. The guest boots an unsigned kernel while reporting that Secure Boot is enforcing. Any tenant-isolation or attestation story that rests on guest Secure Boot in a QEMU/KVM GPU environment is simply not true on affected OVMF builds.","attack_vector":"Whoever controls the kernel image or the direct-boot command line for the VM - a tenant with access to their own instance's boot configuration, or an attacker who has compromised the image pipeline. Requires the VM to use QEMU direct kernel boot rather than a normal bootloader.","remediation":"Hypervisor-side firmware package update (edk2/OVMF), not a server BIOS flash - update the ovmf/edk2 build on your hosts and restart guests onto the new firmware; no host reboot needed. Config workaround: stop using QEMU direct kernel boot for guests that rely on Secure Boot, and boot through a signed shim/bootloader from a virtual disk instead.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-2296","https://github.com/tianocore/edk2/security/advisories/GHSA-6pp6-cm5h-86g5"],"status":"curated","published":"2025-12-09"},{"id":"CVE-2025-23374","cve":"CVE-2025-23374","aliases":[],"title":"Dell Enterprise SONiC (sensitive information in log files): Sensitive information is written into log files","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell Enterprise SONiC (sensitive information in log files)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"Sensitive information is written into log files on the switch. Switch logs are routinely shipped wholesale to a central syslog or observability stack that far more people can read than can log into the switch — so secrets written to the log leak to a much wider audience than the device's own access control implies.","attack_vector":"Anyone with read access to the switch's logs or to the log aggregation pipeline they are shipped to.","remediation":"Upgrade to Enterprise SONiC 4.4.1 or 4.2.3 or later — NOS image upgrade plus reboot. Also purge historical logs from your aggregator and rotate anything that appeared in them; that cleanup is the part people skip.","references":["https://www.dell.com/support/kbdoc/en-us/000340083/dsa-2025-275-security-update-for-dell-enterprise-sonic-distribution-vulnerabilities","https://nvd.nist.gov/vuln/detail/CVE-2025-23374"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-01-30"},{"id":"CVE-2025-2826","cve":"CVE-2025-2826","aliases":["Arista Security Advisory 0120"],"title":"Arista EOS (ingress ACL enforcement on ethernet/LAG): With IPv4 ingress, MAC ingress, or IPv6 standard ingress ACLs","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (ingress ACL enforcement on ethernet/LAG)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"With IPv4 ingress, MAC ingress, or IPv6 standard ingress ACLs applied to one or more ethernet or LAG interfaces, the policies may not be enforced at all. The third ACL-enforcement defect in the same family — worth treating Arista ingress ACLs as a control that needs periodic active verification rather than a set-and-forget boundary.","attack_vector":"Any traffic arriving on an affected interface. No attacker capability required.","remediation":"EOS upgrade plus reload on affected platforms. Because ACL enforcement is the thing that fails, the only trustworthy verification is sending traffic that should be dropped and confirming it is — do that as a standing test in your fabric CI, not just after this patch.","references":["https://www.arista.com/en/support/advisories-notices/security-advisory/21414-security-advisory-0120"],"status":"curated","tags":["tenant-isolation"],"published":"2025-05-27"},{"id":"CVE-2025-37866","cve":"CVE-2025-37866","aliases":[],"title":"Linux kernel mlxbf-bootctl (BlueField secure boot fuse state): The BlueField boot-control driver mishandles the sysfs","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel mlxbf-bootctl (BlueField secure boot fuse state)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"The BlueField boot-control driver mishandles the sysfs buffer when reporting secure-boot fuse state, producing a kernel warning. Low direct impact, but it sits on the interface an operator uses to verify that a DPU's secure boot fuses are actually burned - if you attest DPU secure-boot state by reading this sysfs node, the reporting path itself was not sound.","attack_vector":"Local on the DPU - triggered by reading the secure_boot_fuse_state sysfs attribute on BlueField.","remediation":"Upgrade the DPU Arm-side kernel via a DOCA/BFOS package upgrade or BFB re-image; requires a DPU reset per node. Low urgency on its own - fold it into the next scheduled BFB refresh rather than taking a dedicated window.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-37866","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2025/CVE-2025-37866.json"],"status":"curated","published":"2025-05-09"},{"id":"CVE-2025-40148","cve":"CVE-2025-40148","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add NULL pointer checks in dc_stream cursor attribute functions","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-40148","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-11-12"},{"id":"CVE-2025-40191","cve":"CVE-2025-40191","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A NULL pointer dereference in the amdkfd (KFD compute","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"A NULL pointer dereference in the amdkfd (KFD compute driver, /dev/kfd). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdkfd: Fix kfd process ref leaking when userptr unmapping","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-40191","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-11-12"},{"id":"CVE-2025-40288","cve":"CVE-2025-40288","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): Memory is handed to a consumer without being","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"Memory is handed to a consumer without being initialised or cleared in the amdgpu GEM/VM/command-submission ioctl surface. Whatever the previous owner left behind is readable - and on a GPU node the previous owner is very often a different tenant's job. This is the classic residual-data leak between workloads sharing a card: model weights, activations, keys or tokens from the prior tenant can surface in a fresh allocation. Upstream fix: drm/amdgpu: Fix NULL pointer dereference in VRAM logic for APU devices","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-40288","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-12-06"},{"id":"CVE-2025-40289","cve":"CVE-2025-40289","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): A correctness defect in the amdgpu RAS / GPU reset","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu: hide VRAM sysfs attributes on GPUs without VRAM","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-40289","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-12-06"},{"id":"CVE-2025-40300","cve":"CVE-2025-40300","aliases":["VMSCAPE"],"title":"AMD Zen 1-Zen 5 - branch predictor isolation between guest and userspace hypervisor (AMD-SB-7046): Insufficient","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"AMD Zen 1-Zen 5 - branch predictor isolation between guest and userspace hypervisor (AMD-SB-7046)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"Insufficient branch-predictor isolation between a guest VM and the **userspace** hypervisor process - QEMU - lets a malicious guest train the predictor and then steer speculation inside the VMM that manages it. The VMM has the guest's memory mapped, so this is a Spectre-class read of confidential guest state from a process the guest can influence. AMD rates it High and, unusually, it spans **Zen 1 through Zen 5** - there is no 'we are on new silicon' escape from this one.","attack_vector":"From inside a guest VM. Tenant-reachable, no host privilege required. Affects every AMD generation currently in datacenter service.","remediation":"Fixed in the Linux kernel by issuing a conditional IBPB after every VMexit before returning to userspace. Take the distro kernel update and reboot the host - no firmware, BIOS or microcode step, which makes it one of the cheaper fixes to deploy. **Expect a real performance cost**: the kernel commit itself notes the IBPB duplicates the context-switch IBPB and that workloads switching frequently between hypervisor and userspace absorb the most overhead. On a virtualised GPU fleet with heavy device emulation, benchmark before and after rather than assuming it is free.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-40300","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-7046.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-09-11"},{"id":"CVE-2025-40310","cve":"CVE-2025-40310","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (amd/amdkfd): A NULL pointer dereference in the amdkfd (KFD compute","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (amd/amdkfd)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"A NULL pointer dereference in the amdkfd (KFD compute driver, /dev/kfd). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: amd/amdkfd: resolve a race in amdgpu_amdkfd_device_fini_sw","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-40310","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-12-08"},{"id":"CVE-2025-40332","cve":"CVE-2025-40332","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A race condition or locking defect in the amdkfd (KFD","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"A race condition or locking defect in the amdkfd (KFD compute driver, /dev/kfd). Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdkfd: Fix mmap write lock not release","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-40332","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-12-09"},{"id":"CVE-2025-40335","cve":"CVE-2025-40335","aliases":[],"title":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu): Missing or insufficient validation of","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu user-mode queues (doorbell submission path). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdgpu: validate userq input args","attack_vector":"Local. Reachable by any process with a render node open that can create user-mode queues - the normal ROCm submission path, reachable from an unprivileged container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-40335","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-12-09"},{"id":"CVE-2025-40339","cve":"CVE-2025-40339","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A NULL pointer dereference in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"A NULL pointer dereference in the amdgpu GEM/VM/command-submission ioctl surface. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: fix nullptr err of vm_handle_moved","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-40339","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-12-09"},{"id":"CVE-2025-49133","cve":"CVE-2025-49133","aliases":[],"title":"libtpms (CryptHmacSign, vTPM): Out-of-bounds read when signKey and signScheme are mismatched, aborting the vTPM","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"libtpms (CryptHmacSign, vTPM)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"Out-of-bounds read when signKey and signScheme are mismatched, aborting the vTPM. libtpms is the TPM behind QEMU/KVM guests, so on a GPU cloud that rents VMs this is a guest reaching out and killing its own virtual TPM - which takes down measured boot and any disk unlock or key sealing that depended on it, and on some stacks takes the guest with it. The libtpms instance of the same defect as the TCG reference implementation issue, which is worth noting because they patch through completely different channels.","attack_vector":"A guest able to issue TPM commands to its vTPM - i.e. any tenant in a VM you provisioned with a virtual TPM. No escape required, no host access.","remediation":"libtpms package update on the hypervisor hosts plus a restart of affected guests' swtpm processes - so a rolling VM restart rather than a firmware flash, which is the cheap end of this database. Do not assume patching the host TPM stack covers your physical TPMs or your platform firmware TPM; those are separate code paths with separate fixes.","references":["https://github.com/stefanberger/libtpms/security/advisories/GHSA-25w5-6fjj-hf8g","https://nvd.nist.gov/vuln/detail/CVE-2025-49133"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2025-06-10"},{"id":"CVE-2025-52953","cve":"CVE-2025-52953","aliases":["JSA100059"],"title":"Juniper Junos OS / Junos OS Evolved (rpd BGP session handling): A genuine, valid BGP UPDATE message resets a live BGP","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Juniper Junos OS / Junos OS Evolved (rpd BGP session handling)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"A genuine, valid BGP UPDATE message resets a live BGP session between peers. Because the message is valid, no filter catches it — the fabric tears down and rebuilds sessions on demand from an adjacent attacker. Repeated, it keeps the underlay in permanent reconvergence, which for a synchronous training job is indistinguishable from an outage.","attack_vector":"Unauthenticated attacker with adjacent network access able to originate a BGP UPDATE into the fabric.","remediation":"Junos upgrade plus reboot, staged across ECMP pairs. There is no clean config workaround because the triggering message is legitimate; tighten which peers you accept sessions from and monitor for session-reset storms in the meantime.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-52953"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["fabric-dos"],"published":"2025-07-11"},{"id":"CVE-2025-52989","cve":"CVE-2025-52989","aliases":[],"title":"Juniper Junos OS / Junos OS Evolved (annotate configuration command): The `annotate` configuration command can be used","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Juniper Junos OS / Junos OS Evolved (annotate configuration command)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"The `annotate` configuration command can be used to escalate privileges. A user with limited configuration rights — the sort of account you give an automation system or a junior operator — gains more than their class allows on a device that may be a spine. Junos login classes are the mechanism most Juniper shops use for internal separation of duties, and this puts a hole in it.","attack_vector":"Authenticated low-privileged user with configuration access to the device.","remediation":"Junos upgrade plus reboot. Interim: restrict the `annotate` command in affected login classes via the allow-/deny-commands regex — a live config change that closes it without a maintenance window.","references":["https://supportportal.juniper.net/s/article/2025-07-Security-Bulletin-Junos-OS-and-Junos-OS-Evolved-Annotate-configuration-command-can-be-used-for-privilege-escalation-CVE-2025-52989"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-07-11"},{"id":"CVE-2025-54505","cve":"CVE-2025-54505","aliases":["XSA-488"],"title":"Xen / x86 CPU: Floating Point Divider State Sampling - transient-execution leak of FP divider state across domains","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen / x86 CPU","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"Floating Point Divider State Sampling - transient-execution leak of FP divider state across domains","attack_vector":"Tenant VM guest; any tenant process in a container","remediation":"Microcode + hypervisor/kernel mitigation + reboot; standing perf cost on FP-heavy workloads","references":["https://xenbits.xen.org/xsa/advisory-488.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"published":"2026-04-27"},{"id":"CVE-2025-58147","cve":"CVE-2025-58147","aliases":["XSA-475"],"title":"Xen (Viridian): Incorrect input sanitisation in Viridian (Hyper-V enlightenment) hypercalls","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen (Viridian)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"Incorrect input sanitisation in Viridian (Hyper-V enlightenment) hypercalls - guest attacks the hypervisor","attack_vector":"Tenant VM guest (Windows guests using Viridian)","remediation":"Hypervisor patch + reboot/evacuation","references":["https://xenbits.xen.org/xsa/advisory-475.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-10-31"},{"id":"CVE-2025-59957","cve":"CVE-2025-59957","aliases":[],"title":"Juniper Junos OS (QFX5000-Series, EX4600-Series): A physical-access path into affected QFX5000 and EX4600 switches","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Juniper Junos OS (QFX5000-Series, EX4600-Series)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"A physical-access path into affected QFX5000 and EX4600 switches. QFX5000 is a mainstream data-center leaf platform. Physical-access bugs matter more in this market than they used to: GPU capacity is increasingly resold, subleased and colocated, so 'someone was alone in the rack' is a routine event rather than an exotic threat model.","attack_vector":"Physical access to the switch. Realistic in colocation, shared cages, subleased capacity, and during rack-and-stack by contractors.","remediation":"Junos upgrade plus reboot on affected QFX5000/EX4600 platforms. Pair with physical controls — locked cabinets, console-port discipline, and tamper-evident seals — because a firmware patch does not address the underlying access.","references":["https://supportportal.juniper.net/s/article/2025-10-Security-Bulletin-Junos-OS"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-10-09"},{"id":"CVE-2025-59969","cve":"CVE-2025-59969","aliases":[],"title":"Juniper Junos OS Evolved (QFX5000 / PTX, multicast packet handling): Crafted multicast packets crash and restart","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Juniper Junos OS Evolved (QFX5000 / PTX, multicast packet handling)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"Crafted multicast packets crash and restart evo-aftmand / evo-pfemand — the forwarding-plane management daemons on Junos Evolved. Affects QFX5000-series and PTX, both of which appear in AI-cluster spine and super-spine roles. Multicast is reachable from any tenant that can source it, and the crash takes the forwarding plane with it.","attack_vector":"An attacker able to send crafted multicast packets into the fabric — any tenant workload on a VLAN the device serves.","remediation":"Junos Evolved upgrade plus reboot. Interim: rate-limit or filter unexpected multicast at the fabric edge with a control-plane policer — live config. If your cluster does not use multicast (most GPU fabrics do not, outside of some MPI transports), block it outright at the edge.","references":["https://supportportal.juniper.net/s/article/2026-04-Security-Bulletin-Junos-OS-Evolved-QFX5000-Series-and-PTX-Series-An-attacker-sending-crafted-multicast-packets-will-cause-evo-aftmand-evo-pfemand-to-crash-and-restart-CVE-2025-59969"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["fabric-dos"],"published":"2026-04-09"},{"id":"CVE-2025-68172","cve":"CVE-2025-68172","aliases":[],"title":"ASPEED crypto/ACRY accelerator driver (drivers/crypto/aspeed): The ACRY driver's probe error path and its remove path","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ASPEED crypto/ACRY accelerator driver (drivers/crypto/aspeed)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"The ACRY driver's probe error path and its remove path both hand-free a clock that the device-managed allocator will free again, giving a double free in the BMC kernel. It sits in the crypto accelerator the BMC uses for TLS and firmware signature work, which is the wrong place to have allocator corruption: a double free in a driver that touches signing and TLS paths is the kind of primitive that turns into controlled kernel memory reuse rather than just a crash. Realistic near-term impact is BMC kernel instability during driver load failures or module removal.","attack_vector":"BMC-local. Reached through driver probe failure or an explicit driver removal, so it needs root on the BMC or a boot-time condition that makes probe fail. Not host- or network-reachable directly.","remediation":"Kernel fix, backported into 6.6.117, 6.12.58 and 6.17.8 and later. In practice: BMC firmware flash per node, out-of-band, whenever your ODM rebases - which for a fix this recent will realistically be a full release cycle away. No config mitigation; the crypto driver is loaded because bmcweb's TLS wants it. Track it, do not run a special campaign for it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-68172","https://git.kernel.org/stable/c/e8407dfd267018f4647ffb061a9bd4a6d7ebacc6"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2025-12-16"},{"id":"CVE-2025-68180","cve":"CVE-2025-68180","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Fix NULL deref in debugfs odm_combine_segments","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-68180","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-12-16"},{"id":"CVE-2025-68190","cve":"CVE-2025-68190","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/atom): A NULL pointer dereference","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/atom)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"A NULL pointer dereference in the amdgpu firmware, ACPI and IP-block initialisation. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu/atom: Check kcalloc() for WS buffer in amdgpu_atom_execute_table_locked()","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-68190","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-12-16"},{"id":"CVE-2025-68196","cve":"CVE-2025-68196","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Cache streams targeting link when performing LT automation","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-68196","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-12-16"},{"id":"CVE-2025-68201","cve":"CVE-2025-68201","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): A correctness defect in the amdgpu kernel driver core reachable","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"A correctness defect in the amdgpu kernel driver core reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu: remove two invalid BUG_ON()s","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-68201","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2025-12-16"},{"id":"CVE-2025-68230","cve":"CVE-2025-68230","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdgpu): An out-of-bounds access in the amdkfd (KFD compute","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdgpu)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"An out-of-bounds access in the amdkfd (KFD compute driver, /dev/kfd) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: fix gpu page fault after hibernation on PF passthrough","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-68230","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2025-12-16"},{"id":"CVE-2025-68798","cve":"CVE-2025-68798","aliases":[],"title":"Linux perf/x86/amd - general protection fault from a NULL event on enable: A subtle race lets cpuc","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux perf/x86/amd - general protection fault from a NULL event on enable","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"A subtle race lets cpuc->events[idx] become NULL on AMD machines, and enabling the event then takes a general protection fault, panicking the host. Same family as the other AMD perf races: the monitoring stack crashes the machine it is monitoring.","attack_vector":"Local, through perf event scheduling. Reachable by observability agents and by tenants where perf is permitted.","remediation":"Distro kernel update plus reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-68798"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-01-13"},{"id":"CVE-2025-68816","cve":"CVE-2025-68816","aliases":["net/mlx5 fw_tracer validate format string parameters"],"title":"Linux kernel mlx5_core firmware tracer (diag/fw_tracer): The firmware tracer took format strings directly from device","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel mlx5_core firmware tracer (diag/fw_tracer)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"The firmware tracer took format strings directly from device firmware and passed them to kernel formatting with no validation. Malicious or compromised NIC/DPU firmware supplying %s, %p or %n reads arbitrary kernel memory. This is the clearest firmware-as-attacker case in the mlx5 stack, and it matters specifically for operators who accept hardware from third parties, run rented bare metal, or cannot fully attest NIC firmware provenance - the host kernel was trusting the device.","attack_vector":"Requires control of the NIC or DPU firmware image - a supply-chain or prior-tenant-persistence scenario on bare metal, or an attacker who already flashed the adapter. Not reachable from ordinary network traffic.","remediation":"Upgrade the host kernel to 6.19 or a stable backport (5.10.248, 5.15.198, 6.1.160, 6.6.120, 6.12.64, 6.18.3) - the fix restricts the tracer to integer and hex specifiers. Rolling reboot. Pair it with the real control: enforce signed firmware and re-flash adapters to a known-good version between bare-metal tenants.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-68816","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2025/CVE-2025-68816.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2026-01-13"},{"id":"CVE-2025-7026","cve":"CVE-2025-7026","aliases":["VU#746790"],"title":"Gigabyte UEFI firmware (SMM, unchecked RBX pointer): An attacker-controlled register is used as an unchecked pointer","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Gigabyte UEFI firmware (SMM, unchecked RBX pointer)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"An attacker-controlled register is used as an unchecked pointer inside System Management Mode, giving arbitrary write in ring -2. SMM sits beneath the hypervisor and beneath the OS kernel, so code that lands here can disable Secure Boot, tamper with firmware, and persist through a full disk wipe and OS reinstall. On a rented GPU node this is the canonical 'tenant leaves something behind for the next tenant' primitive, and no host-based tooling can detect it.","attack_vector":"Local privileged code on the host - kernel-level or a driver, so a tenant with root on bare metal, or an attacker who already got kernel execution. Not remote, but on bare-metal rental the precondition is exactly what you sell.","remediation":"UEFI/BIOS firmware update from Gigabyte, per board model, requiring a host reboot - which on GPU nodes means draining running training jobs. Gigabyte shipped fixed firmware; the practical problem is coverage, because affected models span consumer and server lines and not every SKU gets an image. There is no config workaround for an SMM callout. If you buy Gigabyte boards, make the firmware version part of your node-acceptance check, not a post-hoc audit.","references":["https://kb.cert.org/vuls/id/746790","https://nvd.nist.gov/vuln/detail/CVE-2025-7026"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2025-07-11"},{"id":"CVE-2025-7027","cve":"CVE-2025-7027","aliases":["VU#746790"],"title":"Gigabyte UEFI firmware (SMM, NVRAM double pointer dereference): An unvalidated NVRAM variable is dereferenced twice","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Gigabyte UEFI firmware (SMM, NVRAM double pointer dereference)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"An unvalidated NVRAM variable is dereferenced twice inside SMM, letting a local attacker write to SMRAM and gain ring -2 execution. NVRAM variables are writable from the OS on many platforms, which makes this a comparatively short path from host root to firmware-level persistence.","attack_vector":"Local privileged code on the host, able to set the UEFI variable. Any tenant with root on a bare-metal node qualifies.","remediation":"Gigabyte BIOS update per board model plus reboot. No config workaround - locking down NVRAM variable writes is not generally available to operators. Bundle with the other three SMM CVEs in the same advisory; they ship in the same image.","references":["https://kb.cert.org/vuls/id/746790","https://nvd.nist.gov/vuln/detail/CVE-2025-7027"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2025-07-11"},{"id":"CVE-2025-7028","cve":"CVE-2025-7028","aliases":["VU#746790"],"title":"Gigabyte UEFI firmware (SMM, unvalidated flash function pointers): Function pointer structures governing SPI flash","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Gigabyte UEFI firmware (SMM, unvalidated flash function pointers)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"Function pointer structures governing SPI flash operations are not validated, so an attacker in SMM context can redirect the routines that read, write and erase the platform firmware. This is the worst of the four: it hands the attacker the flash-write primitive directly, meaning a permanent bootkit rather than a runtime compromise.","attack_vector":"Local privileged code on the host.","remediation":"Gigabyte BIOS update per board plus reboot. Because the payoff is a flash write, assume any node you believe was compromised needs firmware re-flashed from a known-good image and its integrity independently verified - a BIOS update applied by a compromised system does not prove anything. For high-value nodes, external SPI verification is the only real assurance.","references":["https://kb.cert.org/vuls/id/746790","https://nvd.nist.gov/vuln/detail/CVE-2025-7028"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2025-07-11"},{"id":"CVE-2025-7029","cve":"CVE-2025-7029","aliases":["VU#746790"],"title":"Gigabyte UEFI firmware (SMM, OcHeader/OcData pointer control): Unchecked register use lets the attacker control","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Gigabyte UEFI firmware (SMM, OcHeader/OcData pointer control)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"Unchecked register use lets the attacker control the OcHeader and OcData pointers handled in SMM, producing arbitrary read and write at ring -2 - the same firmware-persistence outcome as the rest of the batch.","attack_vector":"Local privileged code on the host.","remediation":"Gigabyte BIOS update per board plus reboot; delivered in the same firmware image as CVE-2025-7026 through 7028, so treat all four as one change.","references":["https://kb.cert.org/vuls/id/746790","https://nvd.nist.gov/vuln/detail/CVE-2025-7029"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"published":"2025-07-11"},{"id":"CVE-2026-21444","cve":"CVE-2026-21444","aliases":[],"title":"libtpms (OpenSSL 3.x symmetric cipher IV handling): libtpms 0.10.0/0.10.1 built against OpenSSL 3.x returned","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"libtpms (OpenSSL 3.x symmetric cipher IV handling)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"libtpms 0.10.0/0.10.1 built against OpenSSL 3.x returned the initial IV instead of the last IV for certain symmetric ciphers, weakening every subsequent encrypt and decrypt step in the chain. Anything a tenant's vTPM encrypted - sealed keys, protected blobs - is protected less than the cryptography claims. Quiet failure mode: nothing errors, the data is just weaker than the threat model assumes, and you only find out when someone attacks it.","attack_vector":"No active attacker needed to introduce the weakness - it is present in every affected operation. Exploiting it requires an attacker who obtains the ciphertext, which for vTPM state means host-level access or a leaked VM state file.","remediation":"libtpms package update on hypervisor hosts and restart the swtpm processes. Because the weakness is in data already produced, rotate anything a vulnerable vTPM sealed rather than assuming the update is retroactive. Package-level, no firmware flash - but the re-sealing step is the part that takes planning on a fleet with long-lived guests.","references":["https://github.com/stefanberger/libtpms/security/advisories/GHSA-7jxr-4j3g-p34f","https://nvd.nist.gov/vuln/detail/CVE-2026-21444"],"status":"curated","published":"2026-01-02"},{"id":"CVE-2026-23034","cve":"CVE-2026-23034","aliases":[],"title":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu/userq): A memory or reference-count leak","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu/userq)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A memory or reference-count leak in the amdgpu user-mode queues (doorbell submission path). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amdgpu/userq: Fix fence reference leak on queue teardown v2","attack_vector":"Local. Reachable by any process with a render node open that can create user-mode queues - the normal ROCm submission path, reachable from an unprivileged container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-23034","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-01-31"},{"id":"CVE-2026-23051","cve":"CVE-2026-23051","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): A NULL pointer dereference in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A NULL pointer dereference in the amdgpu firmware, ACPI and IP-block initialisation. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: fix drm panic null pointer when driver not support atomic","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-23051","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-02-04"},{"id":"CVE-2026-23554","cve":"CVE-2026-23554","aliases":["XSA-480"],"title":"Xen (EPT): Use-after-free of EPT paging structures - HVM guest to host compromise","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen (EPT)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Use-after-free of EPT paging structures - HVM guest to host compromise","attack_vector":"Tenant VM guest (HVM)","remediation":"Hypervisor patch + host reboot with guest evacuation. Highest-severity recent Xen item for a multi-tenant HVM fleet","references":["https://xenbits.xen.org/xsa/advisory-480.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-03-23"},{"id":"CVE-2026-34878","cve":"CVE-2026-34878","aliases":["TFV-15","FIP ToC offset validation"],"title":"Arm Trusted Firmware-A BL1/BL2 boot stages on platforms that load firmware from a Firmware Image Package (FIP)","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arm Trusted Firmware-A BL1/BL2 boot stages on platforms that load firmware from a Firmware Image Package (FIP) container","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"FIP headers are not signed, so BL1/BL2 parse attacker-influenceable metadata before authentication happens. Bad offsets and lengths make the loader read from secure memory mapped in the EL3 translation regime and land that data in non-secure memory. That is a pre-authentication leak of secure-world contents at boot - keys, secure heap, whatever is mapped. For a bare-metal provider this is a tenant-handoff problem: a tenant who can write the boot FIP (firmware update path, A/B slot, provisioning share) gets to exfiltrate secure-world state on the next boot for the next tenant.","attack_vector":"An attacker who can modify or supply the FIP image the platform boots from - anyone with the firmware-update path, a writable boot partition, or control of the provisioning pipeline. On bare-metal GPU rental that is the previous tenant if you do not verify firmware between handoffs.","remediation":"Update TF-A past commits 48351 and 49485 (strict ToC bounds validation, overflow-safe arithmetic, short-read detection) and rebuild BL1/BL2 for the platform. That is an OEM-shipped firmware build, flashed to each node, reboot and drain. Independently: treat the FIP as tenant-writable until proven otherwise, and re-verify or re-flash boot firmware from a known-good image at every tenant handoff rather than relying on the parser being safe.","references":["https://trustedfirmware-a.readthedocs.io/en/latest/security_advisories/security-advisory-tfv-15.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2026-42487","cve":"CVE-2026-42487","aliases":["XSA-491"],"title":"Xen (x86 HVM): x86 HVM I/O port list traversal flaw","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen (x86 HVM)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"x86 HVM I/O port list traversal flaw","attack_vector":"Tenant VM guest (HVM)","remediation":"Hypervisor patch + reboot/evacuation","references":["https://xenbits.xen.org/xsa/advisory-491.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-06-18"},{"id":"CVE-2026-61810","cve":"CVE-2026-61810","aliases":["DMTF-2026-0003"],"title":"DMTF SPDM specification DSP0274 1.4 (FINISH transcript definition): A specification-level defect rather than","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"DMTF SPDM specification DSP0274 1.4 (FINISH transcript definition)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A specification-level defect rather than an implementation bug. SPDM 1.4 added OpaqueDataLength and OpaqueData fields to FINISH and FINISH_RSP, but the transcript definitions used to compute the FINISH signature and the verify-data HMACs were never updated to cover them. A literal implementation therefore authenticates a message while leaving part of that message unauthenticated - so an attacker in the path can alter the opaque fields without breaking the signature. This is exactly the class of flaw that undermines device attestation quietly: everything validates, and the validated thing is not what was sent.","attack_vector":"An attacker able to modify SPDM traffic between requester and responder - a PCIe interposer, a compromised switch or retimer in the path, or a malicious intermediary in a disaggregated fabric where SPDM crosses a network rather than a board trace.","remediation":"Requires a specification erratum plus updated implementations on both ends, so the fix arrives as vendor firmware updates once DMTF publishes and vendors rebase - expect a long tail and no OS-level patch. Nothing to configure. The operator action today is to know that SPDM 1.4 opaque data is not integrity-protected in affected implementations, and not to build any policy decision on values carried there.","references":["https://github.com/DMTF/libspdm/security/advisories/GHSA-chjj-xvqx-c8w4","https://nvd.nist.gov/vuln/detail/CVE-2026-61810"],"status":"curated"},{"id":"CVE-2026-62428","cve":"CVE-2026-62428","aliases":["XSA-500"],"title":"Xen (grant tables): Type confusion in grant-copy - guest corrupts hypervisor state","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen (grant tables)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Type confusion in grant-copy - guest corrupts hypervisor state","attack_vector":"Tenant VM guest","remediation":"Hypervisor patch + reboot/evacuation. Grant tables are on the hot path for every PV driver, so this is unavoidable exposure","references":["https://xenbits.xen.org/xsa/advisory-500.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-07-28"},{"id":"CVE-2026-62435","cve":"CVE-2026-62435","aliases":["XSA-501"],"title":"Xen (grant tables): Grant-table version change racing with other operations","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen (grant tables)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Grant-table version change racing with other operations","attack_vector":"Tenant VM guest","remediation":"Hypervisor patch + reboot/evacuation","references":["https://xenbits.xen.org/xsa/advisory-501.html"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-07-28"},{"id":"CVE-2026-63878","cve":"CVE-2026-63878","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): Missing or insufficient validation of","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu GEM/VM/command-submission ioctl surface. A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdgpu: check num_entries in GEM_OP GET_MAPPING_INFO","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63878","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-07-19"},{"id":"CVE-2026-63880","cve":"CVE-2026-63880","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): Missing or insufficient validation of","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu GEM/VM/command-submission ioctl surface. A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdgpu: fix lock leak on ENOMEM in AMDGPU_GEM_OP_GET_MAPPING_INFO","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63880","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-07-19"},{"id":"CVE-2026-63882","cve":"CVE-2026-63882","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A NULL pointer dereference in the amdkfd (KFD compute","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A NULL pointer dereference in the amdkfd (KFD compute driver, /dev/kfd). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdkfd: fix NULL pointer bug in svm_range_set_attr","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63882","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-07-19"},{"id":"CVE-2026-64229","cve":"CVE-2026-64229","aliases":[],"title":"Linux x86/mm - broadcast TLB flush with PCID disabled: Booting with nopcid clears the PCID feature but broadcast TLB","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux x86/mm - broadcast TLB flush with PCID disabled","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Booting with nopcid clears the PCID feature but broadcast TLB flushing stayed enabled, leaving TLB invalidation in an inconsistent configuration. Stale TLB entries are a memory-isolation problem: a translation that should have been invalidated but was not means one address space can still reach a mapping that was revoked.","attack_vector":"Local, on hosts booted with nopcid. Not attacker-selected unless the attacker controls boot parameters - but plenty of fleets set nopcid for debugging or for old mitigation workarounds and forget it.","remediation":"Fixed in the Linux kernel. Distro kernel update plus reboot. Also audit your boot parameters: nopcid is a performance and now correctness liability that is often left in place long after the reason for it is gone.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-64229"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-07-24"},{"id":"CVE-2026-64309","cve":"CVE-2026-64309","aliases":[],"title":"Linux crypto/ccp - SNP initialization on ioctl(SNP_COMMIT): The ccp driver initialised SNP from the SNP_COMMIT ioctl","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux crypto/ccp - SNP initialization on ioctl(SNP_COMMIT)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"The ccp driver initialised SNP from the SNP_COMMIT ioctl path, so a userspace process holding /dev/sev could drive SNP platform initialisation at a time the host was not expecting it - including while ordinary VMs are running. Triggering SNP platform state transitions underneath live guests is a route to destabilising the host and the confidential-computing state machine on it.","attack_vector":"Local, from a process with access to /dev/sev. On a well-run host that is the VMM or a management daemon, so the realistic path is a compromised control-plane component rather than a tenant.","remediation":"Fixed in the Linux kernel - KVM/x86 SEV code or the ccp/PSP driver. Take the distro kernel update (RHEL/Rocky, Ubuntu, SLES) and **reboot the host**; SEV/SNP hypervisor paths cannot be live-patched in any meaningful way, and SNP platform init/shutdown is not safe to cycle under running guests. Drain confidential-VM tenants, reboot, then re-admit. No firmware, VBIOS or AGESA step needed, which makes this one of the cheaper classes of SEV fix to roll out. Worth checking who actually has /dev/sev open on your hosts - the permissions on that node are the difference between 'root only' and 'any service account that got a bit too much'.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-64309"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-07-25"},{"id":"CVE-2026-68102","cve":"CVE-2026-68102","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): A memory or reference-count leak","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A memory or reference-count leak in the amdgpu firmware, ACPI and IP-block initialisation. Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amdgpu: fix aperture mapping leak","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68102","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-10"},{"id":"CVE-2026-68105","cve":"CVE-2026-68105","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): Missing or insufficient validation of","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu GEM/VM/command-submission ioctl surface. A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdgpu: Fix kernel panic during driver load failure","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68105","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-08-10"},{"id":"CVE-2026-68109","cve":"CVE-2026-68109","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/sdma7.1): A correctness defect in the amdgpu RAS /","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/sdma7.1)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/sdma7.1: replace BUG_ON() with WARN_ON()","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68109","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-10"},{"id":"CVE-2026-68110","cve":"CVE-2026-68110","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/sdma4.4.2): A correctness defect in the amdgpu RAS /","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/sdma4.4.2)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/sdma4.4.2: replace BUG_ON() with WARN_ON()","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68110","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-10"},{"id":"CVE-2026-68111","cve":"CVE-2026-68111","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/gfx9): A correctness defect in the amdgpu RAS / GPU","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/gfx9)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/gfx9: replace BUG_ON() with WARN_ON()","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68111","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-10"},{"id":"CVE-2026-68112","cve":"CVE-2026-68112","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/gfx9.4.3): A correctness defect in the amdgpu RAS /","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/gfx9.4.3)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/gfx9.4.3: replace BUG_ON() with WARN_ON()","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68112","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-10"},{"id":"CVE-2026-68113","cve":"CVE-2026-68113","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/gfx12): A correctness defect in the amdgpu RAS / GPU","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/gfx12)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/gfx12: replace BUG_ON() with WARN_ON()","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68113","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-10"},{"id":"CVE-2026-68114","cve":"CVE-2026-68114","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/gfx12.1): A correctness defect in the amdgpu RAS /","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/gfx12.1)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/gfx12.1: replace BUG_ON() with WARN_ON()","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68114","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-10"},{"id":"CVE-2026-68115","cve":"CVE-2026-68115","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/gfx10): A correctness defect in the amdgpu RAS / GPU","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/gfx10)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/gfx10: replace BUG_ON() with WARN_ON()","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68115","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-10"},{"id":"CVE-2026-68234","cve":"CVE-2026-68234","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A memory or reference-count leak","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A memory or reference-count leak in the amdgpu GEM/VM/command-submission ioctl surface. Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amdgpu: fix bo->pin leaking in amdgpu_bo_create_reserved","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68234","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-10"},{"id":"CVE-2026-68235","cve":"CVE-2026-68235","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: dce100: skip non-DP stream encoders for DP MST","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68235","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-10"},{"id":"CVE-2026-68237","cve":"CVE-2026-68237","aliases":[],"title":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu/userq): A race condition or locking defect","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu/userq)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A race condition or locking defect in the amdgpu user-mode queues (doorbell submission path). Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdgpu/userq: fix indefinite fence wait during GPU reset","attack_vector":"Local. Reachable by any process with a render node open that can create user-mode queues - the normal ROCm submission path, reachable from an unprivileged container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68237","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-10"},{"id":"CVE-2026-68238","cve":"CVE-2026-68238","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): A memory or reference-count leak","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A memory or reference-count leak in the amdgpu firmware, ACPI and IP-block initialisation. Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amdgpu: Release VFCT ACPI table reference","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68238","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-10"},{"id":"CVE-2026-68243","cve":"CVE-2026-68243","aliases":[],"title":"Linux i915 GPU kernel driver (context SSEU parameter): NULL dereference reachable by setting a context engine slot","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux i915 GPU kernel driver (context SSEU parameter)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"NULL dereference reachable by setting a context engine slot to an invalid engine class and then querying the SSEU context parameter. A tenant can panic the GPU node from an ordinary unprivileged ioctl sequence.","attack_vector":"Any local user or container with a DRM render node - i.e. any tenant that was scheduled a GPU. No privileged capability needed.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68243","https://git.kernel.org/stable/c/2b56757a9a7456825eb668fde92299e01c5e2721"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-10"},{"id":"CVE-2026-68246","cve":"CVE-2026-68246","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/gfx11): A correctness defect in the amdgpu RAS / GPU","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/gfx11)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/gfx11: replace BUG_ON() with WARN_ON()","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68246","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-10"},{"id":"CVE-2026-68249","cve":"CVE-2026-68249","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/sdma5.0): A correctness defect in the amdgpu RAS /","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/sdma5.0)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/sdma5.0: replace BUG_ON() with WARN_ON()","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68249","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-10"},{"id":"CVE-2026-68250","cve":"CVE-2026-68250","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/sdma5.2): A correctness defect in the amdgpu RAS /","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/sdma5.2)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/sdma5.2: replace BUG_ON() with WARN_ON()","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68250","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-10"},{"id":"CVE-2026-68251","cve":"CVE-2026-68251","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/sdma6.0): A correctness defect in the amdgpu RAS /","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/sdma6.0)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/sdma6.0: replace BUG_ON() with WARN_ON()","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68251","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-10"},{"id":"CVE-2026-68252","cve":"CVE-2026-68252","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/sdma7.0): A correctness defect in the amdgpu RAS /","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/sdma7.0)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/sdma7.0: replace BUG_ON() with WARN_ON()","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68252","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-10"},{"id":"CVE-2026-68256","cve":"CVE-2026-68256","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A memory or reference-count leak in the amdgpu display core","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A memory or reference-count leak in the amdgpu display core (DC/DM). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amd/display: detect_link_and_local_sink: DP alt mode timeout path leaks prev_sink reference","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68256","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-10"},{"id":"CVE-2026-68259","cve":"CVE-2026-68259","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A correctness defect in the amdkfd (KFD compute","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A correctness defect in the amdkfd (KFD compute driver, /dev/kfd) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdkfd: Check bounds in allocate_event_notification_slot","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68259","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-10"},{"id":"CVE-2026-68272","cve":"CVE-2026-68272","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): Missing or insufficient validation of","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu GEM/VM/command-submission ioctl surface. A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdgpu: validate CP_GFX_SHADOW chunk size in CS pass1","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68272","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-08-10"},{"id":"CVE-2026-68275","cve":"CVE-2026-68275","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A NULL pointer dereference in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A NULL pointer dereference in the amdgpu GEM/VM/command-submission ioctl surface. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: check amdgpu_vm_bo_find() result in GET_MAPPING_INFO","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68275","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-10"},{"id":"CVE-2026-68276","cve":"CVE-2026-68276","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/gfx): An out-of-bounds access in the amdgpu","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/gfx)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"An out-of-bounds access in the amdgpu firmware, ACPI and IP-block initialisation - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu/gfx: fix cleaner shader IB buffer overflow","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68276","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-10"},{"id":"CVE-2026-68347","cve":"CVE-2026-68347","aliases":[],"title":"Linux iommu/amd - IRQ-unsafe locking in guest domain allocation: An IRQ-unsafe lock taken during AMD IOMMU guest domain","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux iommu/amd - IRQ-unsafe locking in guest domain allocation","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"An IRQ-unsafe lock taken during AMD IOMMU guest domain allocation can deadlock the host. Guest domain allocation happens when you attach a device to a VM - so this fires on the GPU passthrough path, wedging the host at exactly the moment you are provisioning a tenant's accelerator.","attack_vector":"Local, on the IOMMU guest-domain allocation path - reachable by whatever attaches devices to guests, i.e. the VMM or the orchestrator.","remediation":"Fixed in the Linux kernel. Distro kernel update plus reboot; no firmware step.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68347"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-10"},{"id":"CVE-2026-68364","cve":"CVE-2026-68364","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A race condition or locking defect in the amdgpu display","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A race condition or locking defect in the amdgpu display core (DC/DM). Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amd/display: Fix ISM dc_lock deadlock during suspend","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68364","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-10"},{"id":"CVE-2026-68430","cve":"CVE-2026-68430","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/gfx8): A correctness defect in the amdgpu RAS / GPU","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/gfx8)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/gfx8: drop unecessary BUG_ON()","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68430","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-12"},{"id":"CVE-2026-68436","cve":"CVE-2026-68436","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: use kvzalloc to allocate struct dc","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68436","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-12"},{"id":"CVE-2026-72140","cve":"CVE-2026-72140","aliases":["i2c: mlxbf use-after-free in mlxbf_i2c_init_resource()"],"title":"Linux kernel i2c-mlxbf (BlueField DPU I2C controller): mlxbf_i2c_init_resource() frees a resource struct and then reads","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel i2c-mlxbf (BlueField DPU I2C controller)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"mlxbf_i2c_init_resource() frees a resource struct and then reads a field out of it to build the error code - a use-after-free on the DPU during I2C controller initialisation. Same subsystem as the earlier BlueField I2C stack overflow, which is the real signal: the DPU's platform glue has repeatedly shipped memory-safety bugs, and that is the layer your infrastructure services sit on.","attack_vector":"Local on the DPU, reached on the I2C init error path during driver probe.","remediation":"Upgrade the DPU Arm-side kernel to 7.2 or one of the wide stable backports (5.10.261 through 7.1.5), delivered as a DOCA/BFOS package upgrade or BFB re-image plus DPU reset. Low urgency alone - bundle into the next BFB refresh alongside the other BlueField platform fixes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-72140","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-72140.json"],"status":"curated","published":"2026-08-15"},{"id":"CVE-2026-72237","cve":"CVE-2026-72237","aliases":[],"title":"Linux perf/x86/amd/brs - kernel address leakage through Branch Sampling: A user-only branch stack collected via AMD","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux perf/x86/amd/brs - kernel address leakage through Branch Sampling","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A user-only branch stack collected via AMD Branch Sampling can contain branches that originated in the kernel, so unprivileged profiling leaks kernel addresses. That is a KASLR break handed to any tenant allowed to profile their own code - and on AI clusters, letting tenants profile their own GPU and CPU kernels is a feature people ask for. Combine it with any of the amdgpu memory-safety bugs in this database and you have a reliable local privilege escalation.","attack_vector":"Local, from a process permitted to use perf branch sampling. How reachable this is depends entirely on your perf_event_paranoid setting - if you relaxed it so tenants can profile, you granted this.","remediation":"Fixed in the Linux kernel by filtering kernel branches out of user-only branch stacks. Distro kernel update plus reboot; no firmware step. In the meantime, review perf_event_paranoid on GPU nodes: the value that makes tenant profiling work is the same value that exposes this.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-72237"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-08-15"},{"id":"CVE-2026-72325","cve":"CVE-2026-72325","aliases":[],"title":"Linux perf/x86/amd/core - Branch Sampling enabled from the SVM reload path: Branch Sampling and Last Branch Record","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux perf/x86/amd/core - Branch Sampling enabled from the SVM reload path","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Branch Sampling and Last Branch Record are mutually exclusive, and the KVM/SVM reload path could enable BRS anyway. Conflicting branch-tracing state on a virtualisation host means corrupted profiling data and, on the SVM path, incorrect state restored around guest entry - a correctness problem sitting on the boundary between host and guest execution.","attack_vector":"Local, on hosts running KVM guests with branch sampling in use.","remediation":"Distro kernel update plus reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-72325"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-15"},{"id":"CVE-2026-74353","cve":"CVE-2026-74353","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A correctness defect in the amdkfd (KFD compute","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A correctness defect in the amdkfd (KFD compute driver, /dev/kfd) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdkfd: always resume_all after suspend_all","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-74353","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"published":"2026-08-15"},{"id":"CVE-2026-74448","cve":"CVE-2026-74448","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): Memory is handed to a consumer without being","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Memory is handed to a consumer without being initialised or cleared in the amdkfd (KFD compute driver, /dev/kfd). Whatever the previous owner left behind is readable - and on a GPU node the previous owner is very often a different tenant's job. This is the classic residual-data leak between workloads sharing a card: model weights, activations, keys or tokens from the prior tenant can surface in a fresh allocation. Upstream fix: drm/amdkfd: fix QID bit leak in pqm_create_queue()","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-74448","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"},"tags":["tenant-isolation"],"published":"2026-08-15"},{"id":"NCVD-0000-001-aspeed-bmc-host-to-bmc-bridges-g","cve":null,"aliases":[],"title":"ASPEED BMC (host-to-BMC bridges generally): The ASPEED LPC/PCIe bridge architecture exists to let the host talk","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ASPEED BMC (host-to-BMC bridges generally)","year":"2019-2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"The ASPEED LPC/PCIe bridge architecture exists to let the host talk to the BMC; disabling it breaks legitimate in-band management (ipmitool, firmware update tooling). Operators frequently leave it on","attack_vector":"Local, host CPU","remediation":"Explicit per-fleet decision: lock the AHB bridges and lose in-band management tooling, or accept a host→BMC escalation path. There is no configuration that gives both","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-6260"],"status":"curated"},{"id":"NCVD-0000-002-ipmi-over-lan-as-a-protocol","cve":null,"aliases":[],"title":"IPMI over LAN as a protocol: IPMI has no transport confidentiality guarantees worth relying on, weak session handling","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"IPMI over LAN as a protocol","year":"2013-2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"IPMI has no transport confidentiality guarantees worth relying on, weak session handling, and per-vendor implementations that diverge from spec. Every BMC on the fleet speaks it by default","attack_vector":"Network, management VLAN","remediation":"Fleet policy decision to disable IPMI-over-LAN and force Redfish-only; costs a rewrite of provisioning/monitoring tooling and loses compatibility with older ODM chassis","references":["https://nvd.nist.gov/vuln/detail/CVE-2013-4786"],"status":"curated"},{"id":"NCVD-0000-003-internet-exposed-bmc","cve":null,"aliases":[],"title":"Internet-exposed BMC: Shodan-visible BMCs are a recurring finding at colo/neocloud buildouts","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Internet-exposed BMC","year":"2019-2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Shodan-visible BMCs are a recurring finding at colo/neocloud buildouts — a management interface reachable from the internet turns every BMC CVE above into a directly exploitable, pre-auth, below-OS foothold","attack_vector":"Network, from the public internet","remediation":"Continuous external attack-surface scanning of the management ranges plus an enforced out-of-band network design; the operational cost is that remote-hands and vendor support workflows often depend on the exposure","references":["https://www.shodan.io/search?query=ipmi"],"status":"curated"},{"id":"NCVD-0000-004-infiniband-subnet-manager-opensm","cve":null,"aliases":[],"title":"InfiniBand subnet manager (OpenSM / UFM): The IB subnet manager has unilateral authority over LID assignment, routing","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"InfiniBand subnet manager (OpenSM / UFM)","year":"2019-2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"The IB subnet manager has unilateral authority over LID assignment, routing and partitioning for the entire fabric, and the management datagram (MAD) path has historically had weak authentication. There is no per-tenant trust boundary in the SM","attack_vector":"Fabric-local","remediation":"Requires an architectural control: pinned partition keys (pkeys) per tenant, SM redundancy, and treating the SM host as tier-0 infrastructure. No patch exists because it is the protocol model","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0130"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-0000-005-rdma-roce","cve":null,"aliases":["ReDMArk"],"title":"RDMA / RoCE: RoCE and IB RDMA have no cryptographic authentication of the QP connection setup or of subsequent","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"RDMA / RoCE","year":"2021-2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"RoCE and IB RDMA have no cryptographic authentication of the QP connection setup or of subsequent RDMA reads/writes; an on-fabric attacker can inject and impersonate, and remote memory reads bypass the target CPU entirely. No CVE — it is the RDMA specification","attack_vector":"Fabric-local, tenant-to-tenant","remediation":"Requires fabric-level isolation (per-tenant pkeys / VXLAN-isolated RoCE domains) or PSP/IPsec offload on the NIC. Shared-fabric multi-tenancy without this is an unmitigated tenant-to-tenant read primitive","references":["https://www.usenix.org/conference/usenixsecurity21/presentation/rothenberger"],"status":"curated"},{"id":"NCVD-0000-006-nvme-of-over-rdma","cve":null,"aliases":["NeVerMore"],"title":"NVMe-oF over RDMA: NVMe-over-Fabrics inherits RDMA's lack of authentication","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVMe-oF over RDMA","year":"2022-2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"NVMe-over-Fabrics inherits RDMA's lack of authentication; storage targets are addressable by any host on the fabric, so a compromised tenant node can reach other tenants' namespaces. No CVE — design","attack_vector":"Fabric-local","remediation":"Enforce NVMe-oF host NQN allowlisting plus DH-HMAC-CHAP authentication and separate storage fabric; costs throughput and adds provisioning complexity","references":["https://arxiv.org/abs/2202.08080"],"status":"curated"},{"id":"NCVD-0000-007-facility-power-dcim-as-a-class","cve":null,"aliases":[],"title":"Facility power / DCIM as a class: PDUs, CRAC controllers, BMS and DCIM platforms run long-lived embedded firmware, sit","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Facility power / DCIM as a class","year":"2023-2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"PDUs, CRAC controllers, BMS and DCIM platforms run long-lived embedded firmware, sit on flat facility networks, and are frequently outside the compute operator's change-control. An attacker with PDU control has a physical-availability weapon against a GPU cluster","attack_vector":"Network, facility LAN","remediation":"Contractual and architectural: demand facility-network segmentation and a firmware SLA from the colo provider, and monitor the power plane independently. For a tenant in someone else's datacenter there is no patch you can apply yourself","references":["https://thehackernews.com/2023/08/multiple-flaws-in-cyberpower-and.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-0000-008-firmware-signing-key-compromise","cve":null,"aliases":[],"title":"Firmware signing-key compromise as a class: Firmware trust anchors (Boot Guard KM/BPM, UEFI PK/KEK, BMC image-signing","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Firmware signing-key compromise as a class","year":"2023-2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Firmware trust anchors (Boot Guard KM/BPM, UEFI PK/KEK, BMC image-signing keys) are held by ODMs with weaker security postures than the clouds that deploy their hardware, and several have been breached. A neocloud inherits the ODM's key hygiene","attack_vector":"Supply chain","remediation":"Requires diligence on ODM key custody at purchase time and an independent measured-boot / firmware-integrity baseline so a signed-but-malicious image is still detectable. Cannot be remediated after the fact","references":["https://www.binarly.io/blog/pkfail-untrusted-platform-keys-undermine-secure-boot-on-uefi-ecosystem"],"status":"curated"},{"id":"NCVD-0000-009-redfish-implementations-all-vend","cve":null,"aliases":[],"title":"Redfish implementations (all vendors): Redfish replaced IPMI but reintroduced the same class of flaws at the HTTP layer","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Redfish implementations (all vendors)","year":"2019-2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Redfish replaced IPMI but reintroduced the same class of flaws at the HTTP layer — CVE-2024-54085, CVE-2023-34330, CVE-2023-25191/25192, CVE-2018-15774 are all Redfish-surface bugs. The spec mandates no minimum implementation assurance, and every ODM ships its own stack","attack_vector":"Network / management VLAN","remediation":"Treat Redfish as an untrusted-by-default surface: mTLS or a management-plane proxy in front of every BMC, per-node unique credentials, and no direct operator access. This is architecture work, not patching","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-54085"],"status":"curated"},{"id":"NCVD-0000-010-kvm-over-ip-virtual-media","cve":null,"aliases":[],"title":"KVM-over-IP / virtual media: The BMC's virtual-media function can mount an arbitrary ISO as the host's boot device","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"KVM-over-IP / virtual media","year":"2019-2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"The BMC's virtual-media function can mount an arbitrary ISO as the host's boot device. Any BMC compromise therefore converts directly into arbitrary host boot, bypassing disk encryption and OS controls (CVE-2019-16649 is the concrete instance)","attack_vector":"Network / BMC","remediation":"Disable virtual media in the BMC baseline except during provisioning windows, and gate the KVM/vmedia ports at the management-network edge. Costs the remote-hands workflow that most operations teams rely on","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-16649"],"status":"curated"},{"id":"NCVD-0000-011-serial-console-servers-out-of-ba","cve":null,"aliases":[],"title":"Serial console servers / out-of-band access appliances: Console servers (Opengear, Lantronix, Digi and similar) hold","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Serial console servers / out-of-band access appliances","year":"2019-2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Console servers (Opengear, Lantronix, Digi and similar) hold plaintext-equivalent access to every device's console including switches and BMCs, run long-lived embedded Linux, and are the deliberate bypass path around all network segmentation","attack_vector":"Network / OOB LAN","remediation":"Bring the console-server fleet into the same patch and credential-rotation cadence as the compute; require per-device authentication rather than a shared console password. Frequently owned by the facility, not the operator","references":["https://nvd.nist.gov/vuln/detail/CVE-2011-3997"],"status":"curated"},{"id":"NCVD-0000-012-gpu-accelerator-firmware-vbios-g","cve":null,"aliases":[],"title":"GPU / accelerator firmware (VBIOS, GSP, NVSwitch): GPU-resident firmware sits below the host OS and is not covered","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU / accelerator firmware (VBIOS, GSP, NVSwitch)","year":"2023-2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"GPU-resident firmware sits below the host OS and is not covered by host EDR or reimaging. NVIDIA ships GPU firmware fixes through opaque bundled updates, so a neocloud cannot independently verify what changed or attest the running image between tenants","attack_vector":"Local, from a tenant with device access","remediation":"Enforce a firmware re-flash and attestation step at every tenant handoff rather than trusting a host reimage. NVIDIA's bundle model means the operator cannot audit the fix, only apply it","references":["https://www.nvidia.com/en-us/security/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"NCVD-0000-013-tenant-handoff-on-bare-metal","cve":null,"aliases":[],"title":"Tenant handoff on bare metal: Reimaging the host disk clears nothing in the BMC, UEFI/SPI flash, NIC/DPU firmware, GPU","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Tenant handoff on bare metal","year":"2019-2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Reimaging the host disk clears nothing in the BMC, UEFI/SPI flash, NIC/DPU firmware, GPU VBIOS or the ESP. Every below-OS CVE in this table is therefore also a cross-tenant persistence primitive on any bare-metal GPU offering","attack_vector":"Local, from a prior tenant","remediation":"Requires a full firmware re-provision and remote attestation between tenants — measured boot with verified PCRs, BMC reflash, NIC firmware verification. This is the single largest unpriced operational cost in bare-metal GPU rental","references":["https://www.binarly.io/reports/pkfail"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"NCVD-2011-001-ata-secure-erase-nvme-sanitize-f","cve":null,"aliases":["NIST SP 800-88","Reliably Erasing Data From Flash-Based Solid State Drives","sanitize verification gap"],"title":"ATA Secure Erase / NVMe Sanitize / Format NVM across SSD vendors","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ATA Secure Erase / NVMe Sanitize / Format NVM across SSD vendors - drives that report sanitization success while…","year":"2011","cvss_score":null,"severity":"unscored","kev":false,"impact":"CLASS ENTRY and the empirical foundation under every specific CVE in this file. UCSD researchers physically desoldered NAND from twelve SSDs after running the built-in sanitize commands and read the raw flash back. Findings, verbatim in substance: built-in commands are effective but manufacturers sometimes implement them INCORRECTLY; of eight drives claiming ATA SECURITY support only four executed ERASE UNIT reliably; two more silently erased just the first LBA unless the firmware had recently been reset; and one drive REPORTED SANITIZATION SUCCESS WHILE ALL DATA REMAINED INTACT - the filesystem was still mountable afterwards. They also found that overwriting the full visible address space twice is usually but not always sufficient, and that single-file sanitization techniques consistently FAIL on SSDs, because the FTL keeps copies at physical addresses the host cannot address. BREAKS TENANT HANDOFF generically: your 'secure erase between customers' step may silently do nothing, and the drive's success code is not evidence. The next tenant carves the previous tenant's checkpoints, datasets and cloud credentials out of blocks your wipe never reached.","attack_vector":"The next tenant on the reclaimed bare-metal host, reading unallocated or remapped blocks; or anyone who obtains the physical drive later and is willing to read the NAND directly, which is the only method that sees over-provisioned, retired and bad blocks at all.","remediation":"There is no patch - this is a verification and architecture problem. (1) Follow NIST SP 800-88 Rev 1 and pick the sanitization tier that matches the data: for anything that held tenant data, Purge (cryptographic erase or a verified device sanitize) at minimum, and Destroy for high-sensitivity media. (2) Stop trusting return codes. Sample-verify: after sanitize, read back a statistical sample of LBAs and confirm they are zeroed or random, and keep the evidence. This does NOT cover over-provisioned or retired blocks, which no host-side read can reach - be honest in your compliance story about that limit rather than claiming a guarantee you cannot make. (3) The only sanitization that is both fast and verifiable at fleet scale is throwing away a key YOU control: run LUKS/dm-crypt per tenant with the key in your KMS, so reclaim is a key-destruction event you can log and audit in milliseconds, and the drive's own erase behaviour stops mattering. (4) Qualify each SKU's sanitize implementation once, in a lab, before it enters the fleet - the researchers' own conclusion was that every implementation must be individually tested before it can be trusted. Verifying erase across a 10,000-drive fleet is a weeks-long operation that most operators never perform at all, which is exactly why a drive that lies about it survives undetected for years.","references":["https://www.usenix.org/legacy/events/fast11/tech/full_papers/Wei.pdf","https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-88r1.pdf","https://nvd.nist.gov/vuln/detail/CVE-2021-33082"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2011-002-ata-secure-erase-nvme-sanitize-f","cve":null,"aliases":["NIST SP 800-88","Reliably Erasing Data From Flash-Based Solid State Drives","sanitize verification gap"],"title":"ATA Secure Erase / NVMe Sanitize / Format NVM across SSD vendors","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ATA Secure Erase / NVMe Sanitize / Format NVM across SSD vendors - drives that report sanitization success while…","year":"2011","cvss_score":null,"severity":"unscored","kev":false,"impact":"CLASS ENTRY and the empirical foundation under every specific CVE in this file. UCSD researchers physically desoldered NAND from twelve SSDs after running the built-in sanitize commands and read the raw flash back. Findings, verbatim in substance: built-in commands are effective but manufacturers sometimes implement them INCORRECTLY; of eight drives claiming ATA SECURITY support only four executed ERASE UNIT reliably; two more silently erased just the first LBA unless the firmware had recently been reset; and one drive REPORTED SANITIZATION SUCCESS WHILE ALL DATA REMAINED INTACT - the filesystem was still mountable afterwards. They also found that overwriting the full visible address space twice is usually but not always sufficient, and that single-file sanitization techniques consistently FAIL on SSDs, because the FTL keeps copies at physical addresses the host cannot address. BREAKS TENANT HANDOFF generically: your 'secure erase between customers' step may silently do nothing, and the drive's success code is not evidence. The next tenant carves the previous tenant's checkpoints, datasets and cloud credentials out of blocks your wipe never reached.","attack_vector":"The next tenant on the reclaimed bare-metal host, reading unallocated or remapped blocks; or anyone who obtains the physical drive later and is willing to read the NAND directly, which is the only method that sees over-provisioned, retired and bad blocks at all.","remediation":"There is no patch - this is a verification and architecture problem. (1) Follow NIST SP 800-88 Rev 1 and pick the sanitization tier that matches the data: for anything that held tenant data, Purge (cryptographic erase or a verified device sanitize) at minimum, and Destroy for high-sensitivity media. (2) Stop trusting return codes. Sample-verify: after sanitize, read back a statistical sample of LBAs and confirm they are zeroed or random, and keep the evidence. This does NOT cover over-provisioned or retired blocks, which no host-side read can reach - be honest in your compliance story about that limit rather than claiming a guarantee you cannot make. (3) The only sanitization that is both fast and verifiable at fleet scale is throwing away a key YOU control: run LUKS/dm-crypt per tenant with the key in your KMS, so reclaim is a key-destruction event you can log and audit in milliseconds, and the drive's own erase behaviour stops mattering. (4) Qualify each SKU's sanitize implementation once, in a lab, before it enters the fleet - the researchers' own conclusion was that every implementation must be individually tested before it can be trusted. Verifying erase across a 10,000-drive fleet is a weeks-long operation that most operators never perform at all, which is exactly why a drive that lies about it survives undetected for years.","references":["https://www.usenix.org/legacy/events/fast11/tech/full_papers/Wei.pdf","https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-88r1.pdf","https://nvd.nist.gov/vuln/detail/CVE-2021-33082"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2015-001-hdd-and-ssd-controller-firmware","cve":null,"aliases":["Equation Group","nls_933w.dll","GrayFish","drive firmware implant"],"title":"HDD and SSD controller firmware as a persistence surface","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HDD and SSD controller firmware as a persistence surface - demonstrated against Seagate, Western Digital, Toshiba…","year":"2015","cvss_score":null,"severity":"unscored","kev":false,"impact":"CLASS ENTRY. Kaspersky documented a module that reprograms the microcontroller firmware of drives from at least five major vendors - described as the most powerful tool in that actor's arsenal. Once the implant is in controller firmware it runs below the operating system, survives reformatting and reimaging because those operations only rewrite the data area and never touch the microcontroller, and even reflashing does not reliably remove it: some firmware regions are not covered by an update, and the drive may report it is already on the latest version and refuse. TOP TENANT-HANDOFF RISK IN THIS CATEGORY. A bare-metal tenant with root has, by definition, the host-side access needed to attempt a firmware write over the standard command set. If it lands, your entire reclaim process - wipe, reimage, revalidate, re-rent - is theatre, because the malicious code is in a place none of those steps inspect. The implant sees every block the next tenant writes, and can lie about sanitize, lock state and firmware version to every tool you own.","attack_vector":"A tenant with root on the bare-metal host issuing firmware-write commands over the standard storage admin command set - no physical access needed. Also reachable by anyone in the supply chain or the RMA/decommission path who has the drive in hand. Detection is the hard part: a compromised controller is the thing reporting its own firmware version to you, so host-side attestation of drive firmware is self-referential and unreliable.","remediation":"Assume UNPATCHABLE and UNVERIFIABLE once suspected - you cannot trust a controller's self-report of its own integrity. Controls are preventative and procedural. (1) Deny tenants the ability to write drive firmware at all: block firmware-download/commit and vendor-specific pass-through commands at the hypervisor, IOMMU/VFIO policy or storage-controller layer, and do not hand tenants unfiltered raw block devices where the workload does not require it. (2) Prefer drives that enforce signed firmware, and confirm the vendor's signing story in procurement rather than assuming it. (3) Record expected firmware version per drive serial in a system the host cannot edit, and alert on any change - especially a version that goes backwards. (4) For high-sensitivity tenancies, retire local media at end of tenancy rather than recycling it into the pool - physical destruction is the only remediation with a guarantee attached, and it costs one drive versus an undetectable persistent foothold across every future customer on that node. (5) Push the risk off the local drive entirely: keep tenant data on network storage with operator-held encryption so that a compromised local controller sees only ciphertext and transient scratch.","references":["https://securelist.com/equation-the-death-star-of-malware-galaxy/68750/","https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-88r1.pdf"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2015-002-hdd-and-ssd-controller-firmware","cve":null,"aliases":["Equation Group","nls_933w.dll","GrayFish","drive firmware implant"],"title":"HDD and SSD controller firmware as a persistence surface","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HDD and SSD controller firmware as a persistence surface - demonstrated against Seagate, Western Digital, Toshiba…","year":"2015","cvss_score":null,"severity":"unscored","kev":false,"impact":"CLASS ENTRY. Kaspersky documented a module that reprograms the microcontroller firmware of drives from at least five major vendors - described as the most powerful tool in that actor's arsenal. Once the implant is in controller firmware it runs below the operating system, survives reformatting and reimaging because those operations only rewrite the data area and never touch the microcontroller, and even reflashing does not reliably remove it: some firmware regions are not covered by an update, and the drive may report it is already on the latest version and refuse. TOP TENANT-HANDOFF RISK IN THIS CATEGORY. A bare-metal tenant with root has, by definition, the host-side access needed to attempt a firmware write over the standard command set. If it lands, your entire reclaim process - wipe, reimage, revalidate, re-rent - is theatre, because the malicious code is in a place none of those steps inspect. The implant sees every block the next tenant writes, and can lie about sanitize, lock state and firmware version to every tool you own.","attack_vector":"A tenant with root on the bare-metal host issuing firmware-write commands over the standard storage admin command set - no physical access needed. Also reachable by anyone in the supply chain or the RMA/decommission path who has the drive in hand. Detection is the hard part: a compromised controller is the thing reporting its own firmware version to you, so host-side attestation of drive firmware is self-referential and unreliable.","remediation":"Assume UNPATCHABLE and UNVERIFIABLE once suspected - you cannot trust a controller's self-report of its own integrity. Controls are preventative and procedural. (1) Deny tenants the ability to write drive firmware at all: block firmware-download/commit and vendor-specific pass-through commands at the hypervisor, IOMMU/VFIO policy or storage-controller layer, and do not hand tenants unfiltered raw block devices where the workload does not require it. (2) Prefer drives that enforce signed firmware, and confirm the vendor's signing story in procurement rather than assuming it. (3) Record expected firmware version per drive serial in a system the host cannot edit, and alert on any change - especially a version that goes backwards. (4) For high-sensitivity tenancies, retire local media at end of tenancy rather than recycling it into the pool - physical destruction is the only remediation with a guarantee attached, and it costs one drive versus an undetectable persistent foothold across every future customer on that node. (5) Push the risk off the local drive entirely: keep tenant data on network storage with operator-held encryption so that a compromised local controller sees only ciphertext and transient scratch.","references":["https://securelist.com/equation-the-death-star-of-malware-galaxy/68750/","https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-88r1.pdf"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2018-001-amd-sev-memory-encryption-host-c","cve":null,"aliases":["SEVered"],"title":"AMD SEV memory encryption - host-controlled guest physical to host physical mapping: The original demonstration","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV memory encryption - host-controlled guest physical to host physical mapping","year":"2018","cvss_score":null,"severity":"unscored","kev":false,"impact":"The original demonstration that SEV does not protect against the host: the hypervisor remaps a guest's physical pages while the guest is servicing a network request, and reads the plaintext of the entire guest memory out of the responses the guest itself sends back. No CVE was ever assigned - AMD's position was that SEV was not designed to resist this - which is why it is easy to miss when auditing. It matters historically and commercially: it is the reason SEV-ES and then SEV-SNP exist, and it means any 'confidential computing' claim made on plain SEV hardware was never true against the operator.","attack_vector":"Malicious or compromised hypervisor against a SEV guest, using ordinary VMM control over second-level page tables plus any network service running in the guest. No exploit primitive needed beyond normal host capabilities.","remediation":"UNPATCHABLE on SEV. The architectural fix is SEV-SNP with the Reverse Map Table (Milan and later) and it is a hardware generation, not a firmware update. Practical guidance for an operator: do not market plain SEV or SEV-ES as protection from yourself, check that SNP is actually enabled rather than merely supported on the SKU, and make sure your attestation flow proves SNP is on rather than proving only that the CPU could do it.","references":["https://www.amd.com/en/developer/sev.html","https://arxiv.org/abs/1805.09604"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2018-002-ecc-ddr3-server-memory-on-intel","cve":null,"aliases":["ECCploit"],"title":"ECC DDR3 server memory on Intel Xeon (Haswell, Sandy Bridge) and AMD Opteron platforms","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"ECC DDR3 server memory on Intel Xeon (Haswell, Sandy Bridge) and AMD Opteron platforms; the technique generalises to…","year":"2018","cvss_score":null,"severity":"unscored","kev":false,"impact":"ECC is the answer most operators give when asked about Rowhammer, and ECCploit is why that answer is wrong. Correcting a flip takes measurably longer than a clean read, so the attacker gets a timing side channel that tells them exactly which bits they flipped - turning ECC from a defence into a feedback oracle for template building. With that feedback they place three flips in one word, which ECC neither corrects nor detects, producing silent corruption. For an AI datacenter this is the nastiest variant: the corruption is by construction invisible to the machine-check path, so your fleet health dashboard shows green while a tenant's memory is being rewritten.","attack_vector":"Unprivileged local code on an ECC server sharing DRAM with the victim. Attack time was about 32 minutes when corrections were directly observable and up to a week in noisy production-like conditions - slow, but a long-running batch tenant has that time.","remediation":"Do not treat ECC as a Rowhammer mitigation; treat it as error reporting. Make sure the reporting is actually wired up - EDAC or the equivalent collecting correctable-error counts per DIMM, exported to your monitoring, with alerting on rate rather than absolute count, since a burst of corrections on one rank is the strongest hammering signal you will get. Confirm firmware and OS handle uncorrectable errors by isolating rather than silently continuing. Retire DIMM SKUs that show elevated correctable-error rates. The structural fix is the same as every other entry here: do not share a memory controller between untrusted tenants.","references":["https://www.vusec.net/projects/eccploit/","https://download.vusec.net/papers/eccploit_sp19.pdf"],"status":"curated"},{"id":"NCVD-2018-004-ecc-ddr3-server-memory-on-intel","cve":null,"aliases":["ECCploit"],"title":"ECC DDR3 server memory on Intel Xeon (Haswell, Sandy Bridge) and AMD Opteron platforms","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"ECC DDR3 server memory on Intel Xeon (Haswell, Sandy Bridge) and AMD Opteron platforms; the technique generalises to…","year":"2018","cvss_score":null,"severity":"unscored","kev":false,"impact":"ECC is the answer most operators give when asked about Rowhammer, and ECCploit is why that answer is wrong. Correcting a flip takes measurably longer than a clean read, so the attacker gets a timing side channel that tells them exactly which bits they flipped - turning ECC from a defence into a feedback oracle for template building. With that feedback they place three flips in one word, which ECC neither corrects nor detects, producing silent corruption. For an AI datacenter this is the nastiest variant: the corruption is by construction invisible to the machine-check path, so your fleet health dashboard shows green while a tenant's memory is being rewritten.","attack_vector":"Unprivileged local code on an ECC server sharing DRAM with the victim. Attack time was about 32 minutes when corrections were directly observable and up to a week in noisy production-like conditions - slow, but a long-running batch tenant has that time.","remediation":"Do not treat ECC as a Rowhammer mitigation; treat it as error reporting. Make sure the reporting is actually wired up - EDAC or the equivalent collecting correctable-error counts per DIMM, exported to your monitoring, with alerting on rate rather than absolute count, since a burst of corrections on one rank is the strongest hammering signal you will get. Confirm firmware and OS handle uncorrectable errors by isolating rather than silently continuing. Retire DIMM SKUs that show elevated correctable-error rates. The structural fix is the same as every other entry here: do not share a memory controller between untrusted tenants.","references":["https://www.vusec.net/projects/eccploit/","https://download.vusec.net/papers/eccploit_sp19.pdf"],"status":"curated"},{"id":"NCVD-2018-004-microsoft-bitlocker-windows-edri","cve":null,"aliases":["ADV180028","BitLocker hardware encryption trusted by default","eDrive offload"],"title":"Microsoft BitLocker / Windows eDrive hardware-encryption offload on any TCG Opal or IEEE-1667 self-encrypting drive","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Microsoft BitLocker / Windows eDrive hardware-encryption offload on any TCG Opal or IEEE-1667 self-encrypting drive","year":"2018","cvss_score":null,"severity":"unscored","kev":false,"impact":"BitLocker's default was to hand encryption to the drive whenever the drive claimed to support it, and to then skip software encryption entirely. Every SED firmware weakness therefore became a full BitLocker bypass, silently, on machines whose operators believed they were encrypted. For a GPU operator this is the general lesson in its sharpest form: a self-attested hardware security claim was accepted with no verification, so the whole encryption posture of the fleet inherited the weakest drive firmware in it. BREAKS TENANT HANDOFF wherever Windows bare-metal nodes are re-let, and it fails silently - the management console reports the volume as encrypted and compliant the entire time.","attack_vector":"Anyone who obtains the physical drive from a Windows node that used hardware offload - next tenant, RMA path, decommission channel. The attacker exploits whatever SED firmware flaw the drive has; BitLocker simply removed the software layer that would have stopped them.","remediation":"Policy change, no firmware needed, but a full re-encrypt: set Group Policy 'Configure use of hardware-based encryption for fixed/operating system data drives' to Disabled, then fully decrypt and re-encrypt each volume - toggling the policy alone does NOT re-encrypt already-provisioned disks, which is the step operators most often miss and which leaves the fleet reporting compliant while still using drive crypto. Verify per host with 'manage-bde -status' and confirm the encryption method is a software AES-XTS value, not 'Hardware Encryption'. Budget a full re-encrypt window per node; on a large Windows bare-metal estate this is a rolling multi-week drain-and-re-image campaign, and there is no way to sample it - a node you did not re-encrypt is a node still relying on the drive.","references":["https://msrc.microsoft.com/update-guide/vulnerability/ADV180028","https://kb.cert.org/vuls/id/395981","https://www.ru.nl/en/research/research-news/radboud-university-researchers-discover-security-flaw-in-ssd-hard-drives"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"NCVD-2019-001-aspeed-ast2400-ast2500-ast2600-p","cve":null,"aliases":["P2A bridge default-on","iLPC2AHB","X-DMA","pantsdown attack surface"],"title":"ASPEED AST2400 / AST2500 / AST2600 (PCIe VGA P2A bridge, iLPC2AHB, X-DMA, SoC debug UART): The design-level problem","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ASPEED AST2400 / AST2500 / AST2600 (PCIe VGA P2A bridge, iLPC2AHB, X-DMA, SoC debug UART)","year":"2019","cvss_score":null,"severity":"unscored","kev":false,"impact":"The design-level problem behind pantsdown, and it outlives the patch. ASPEED silicon deliberately exposes several host-side windows into the BMC's own AHB address space: a PCIe VGA peer-to-peer (P2A) aperture, the LPC-to-AHB bridge reachable from the host's SuperIO, X-DMA, and a UART debug console. On a lot of shipped firmware these are left open because ODM tooling, vendor flashing utilities and iKVM features rely on them. A tenant that gets kernel or root on a bare-metal GPU node - or anything that reaches the host PCIe config space, which includes a rogue driver in a passthrough VM on some configurations - can write the BMC's RAM and flash directly. That is a host-to-BMC escalation with no BMC credentials and no management-network access, and it yields an implant on a processor that keeps running when the node is powered off and survives OS reimage entirely.","attack_vector":"Local to the host: code with kernel privilege on the server the BMC is attached to, or PCIe config-space access from a passthrough device. No network path to the BMC needed. On a bare-metal GPU rental fleet, this is any tenant who gets root on the box they rented.","remediation":"Not a single patch - a per-platform hardening audit. The BMC firmware must explicitly lock the P2A bridge, disable the iLPC2AHB path via the SuperIO/eSPI configuration, and clear the SoC debug UART enable at boot. OpenBMC upstream added kernel-side gating but ODM builds frequently re-enable pieces for their own flashing flows, so you cannot assume your image is safe because it is 'post-2019'. Verifying requires reading the BMC's own register state per platform, then a BMC firmware flash to fix - out-of-band, per node, ODM-rebase-lagged, and bricking risk if power is lost mid-write. Config-only partial mitigation: on a bare-metal fleet, wipe and re-flash BMC firmware between tenants rather than trusting the running image, and treat any node that has hosted an untrusted tenant as having a potentially dirty BMC.","references":["https://www.flamingspork.com/blog/2019/01/23/cve-2019-6260:-gaining-control-of-bmc-from-the-host-processor/","https://security.netapp.com/advisory/ntap-20190314-0001/","https://www.theregister.com/2019/01/24/bmc_pantsdown_bug/","https://support.lenovo.com/us/en/solutions/ps500241-aspeed-ast-series-bmc-vulnerability"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"NCVD-2019-001-pcie-address-translation-service","cve":null,"aliases":["Thunderclap","PCIe ATS translated-request trust","IOMMU bypass via Address Translation Services"],"title":"PCIe Address Translation Services on hosts using an IOMMU/SMMU for device isolation","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"PCIe Address Translation Services on hosts using an IOMMU/SMMU for device isolation - affects any DMA-capable…","year":"2019","cvss_score":null,"severity":"unscored","kev":false,"impact":"The IOMMU/SMMU is the entire basis for saying a passed-through GPU or NIC cannot read host or other-tenant memory. ATS undermines it by design: a device marked ATS-capable is trusted to have already translated an address, and the root complex forwards those requests without re-checking them. A malicious or compromised endpoint simply sets the translated bit and issues DMA at any physical address it likes. Thunderclap showed the broader problem too - even without ATS, real OS IOMMU policies map far more than the buffer in question, leaving windows onto adjacent kernel memory. In a GPU-passthrough fleet the consequence is a tenant with device control reading host memory and other tenants' data, which defeats the isolation model the whole product rests on.","attack_vector":"An attacker who controls a DMA-capable PCIe device: a tenant with GPU or NIC passthrough who can flash or exploit device firmware, an attacker with physical access to a slot or an external PCIe/Thunderbolt port, or a supply-chain-modified card. Not reachable from software alone on a well-configured host - the entry point is device control.","remediation":"There is no single patch; this is configuration you must actively verify, and most fleets have never checked. Concretely: disable ATS unless a workload genuinely needs it (`pci=noats` on Linux, or the equivalent BIOS switch), and never leave ATS enabled for a device assigned to a tenant. Confirm PCIe ACS is enabled on every upstream port and switch so peer-to-peer traffic between passed-through devices is forced up through the IOMMU rather than routed directly - many server BIOSes ship ACS off, and several 'GPU peer-to-peer performance' tuning guides tell you to turn it off, which silently deletes the isolation boundary between two tenants' GPUs on the same switch. Verify per-device IOMMU groups are not lumping unrelated functions together. Then decide deliberately whether the peer-to-peer bandwidth you gain by disabling ACS is worth the tenant-isolation guarantee you lose, and write that decision down.","references":["http://thunderclap.io/","http://thunderclap.io/thunderclap-paper-ndss2019.pdf","https://www.kernel.org/doc/html/latest/admin-guide/kernel-parameters.html"],"status":"curated"},{"id":"NCVD-2019-004-hpe-sas-ssds-20-models-incl-vo04","cve":null,"aliases":["HPE 32,768-hour SSD bug","HPE bulletin a00092491en_us","HPD8"],"title":"HPE SAS SSDs (20 models incl. VO0480JFDGT, VO0960JFDGU, VO1920JFDGV, VO3840JFDHA, MO0400JFFCF-MO3200JFFCL)","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HPE SAS SSDs (20 models incl. VO0480JFDGT, VO0960JFDGU, VO1920JFDGV, VO3840JFDHA, MO0400JFFCF-MO3200JFFCL) with…","year":"2019","cvss_score":null,"severity":"unscored","kev":false,"impact":"PHYSICAL / FLEET-WIDE AVAILABILITY EVENT. At exactly 32,768 power-on hours - about 3 years, 270 days - the drive fails permanently. HPE states that neither the drive nor the data on it can be recovered. The killer property is SYNCHRONY: drives bought and racked together cross the threshold together, so RAID redundancy provides no protection because you lose several members at once. One administrator reported eight drives failing within MINUTES of each other. For a GPU operator this is the whole-cluster failure mode nothing in your HA design accounts for: your redundancy assumes independent failures, and a firmware counter overflow makes them perfectly correlated. Not an attack - a latent time bomb sitting in the fleet with a known detonation date, which is exactly why it belongs in an operator vulnerability database.","attack_vector":"No attacker. The trigger is elapsed powered-on time, and every affected drive with a similar install date reaches it simultaneously. The exposure is determined entirely by your procurement and racking history - a bulk purchase deployed in one window is the worst case.","remediation":"Flash to firmware HPD8 or later BEFORE the threshold; after the drive fails there is no recovery and you restore from backup. HPE shipped HPD8 for the first eight models from 22 November 2019 and the remaining twelve in mid-December 2019. The urgent operational step is inventory, not patching: pull power-on hours for every SAS SSD in the fleet (HPE Smart Storage Administrator, or smartctl attribute 9) and sort by hours remaining, because your window is defined by the oldest drives. Across a large estate this is a rolling drain-and-flash campaign of weeks, and it must be sequenced by remaining hours rather than by rack - and critically you must STAGGER it, because flashing a whole batch on one day recreates the correlated-failure problem for the next latent bug. Standing control: alert on power-on-hours thresholds fleet-wide, and deliberately mix drive batches and vendors across redundancy groups so a single firmware defect cannot take out every member of a RAID set at once.","references":["https://www.bleepingcomputer.com/news/hardware/hp-warns-that-some-ssd-drives-will-fail-at-32-768-hours-of-use/","https://support.hpe.com/hpesc/public/docDisplay?docId=emr_na-a00092491en_us"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"NCVD-2019-005-hpe-sas-ssds-20-models-incl-vo04","cve":null,"aliases":["HPE 32,768-hour SSD bug","HPE bulletin a00092491en_us","HPD8"],"title":"HPE SAS SSDs (20 models incl. VO0480JFDGT, VO0960JFDGU, VO1920JFDGV, VO3840JFDHA, MO0400JFFCF-MO3200JFFCL)","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HPE SAS SSDs (20 models incl. VO0480JFDGT, VO0960JFDGU, VO1920JFDGV, VO3840JFDHA, MO0400JFFCF-MO3200JFFCL) with…","year":"2019","cvss_score":null,"severity":"unscored","kev":false,"impact":"PHYSICAL / FLEET-WIDE AVAILABILITY EVENT. At exactly 32,768 power-on hours - about 3 years, 270 days - the drive fails permanently. HPE states that neither the drive nor the data on it can be recovered. The killer property is SYNCHRONY: drives bought and racked together cross the threshold together, so RAID redundancy provides no protection because you lose several members at once. One administrator reported eight drives failing within MINUTES of each other. For a GPU operator this is the whole-cluster failure mode nothing in your HA design accounts for: your redundancy assumes independent failures, and a firmware counter overflow makes them perfectly correlated. Not an attack - a latent time bomb sitting in the fleet with a known detonation date, which is exactly why it belongs in an operator vulnerability database.","attack_vector":"No attacker. The trigger is elapsed powered-on time, and every affected drive with a similar install date reaches it simultaneously. The exposure is determined entirely by your procurement and racking history - a bulk purchase deployed in one window is the worst case.","remediation":"Flash to firmware HPD8 or later BEFORE the threshold; after the drive fails there is no recovery and you restore from backup. HPE shipped HPD8 for the first eight models from 22 November 2019 and the remaining twelve in mid-December 2019. The urgent operational step is inventory, not patching: pull power-on hours for every SAS SSD in the fleet (HPE Smart Storage Administrator, or smartctl attribute 9) and sort by hours remaining, because your window is defined by the oldest drives. Across a large estate this is a rolling drain-and-flash campaign of weeks, and it must be sequenced by remaining hours rather than by rack - and critically you must STAGGER it, because flashing a whole batch on one day recreates the correlated-failure problem for the next latent bug. Standing control: alert on power-on-hours thresholds fleet-wide, and deliberately mix drive batches and vendors across redundancy groups so a single firmware defect cannot take out every member of a RAID set at once.","references":["https://www.bleepingcomputer.com/news/hardware/hp-warns-that-some-ssd-drives-will-fail-at-32-768-hours-of-use/","https://support.hpe.com/hpesc/public/docDisplay?docId=emr_na-a00092491en_us"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"NCVD-2019-005-pcie-address-translation-service","cve":null,"aliases":["Thunderclap","PCIe ATS translated-request trust","IOMMU bypass via Address Translation Services"],"title":"PCIe Address Translation Services on hosts using an IOMMU/SMMU for device isolation","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"PCIe Address Translation Services on hosts using an IOMMU/SMMU for device isolation - affects any DMA-capable…","year":"2019","cvss_score":null,"severity":"unscored","kev":false,"impact":"The IOMMU/SMMU is the entire basis for saying a passed-through GPU or NIC cannot read host or other-tenant memory. ATS undermines it by design: a device marked ATS-capable is trusted to have already translated an address, and the root complex forwards those requests without re-checking them. A malicious or compromised endpoint simply sets the translated bit and issues DMA at any physical address it likes. Thunderclap showed the broader problem too - even without ATS, real OS IOMMU policies map far more than the buffer in question, leaving windows onto adjacent kernel memory. In a GPU-passthrough fleet the consequence is a tenant with device control reading host memory and other tenants' data, which defeats the isolation model the whole product rests on.","attack_vector":"An attacker who controls a DMA-capable PCIe device: a tenant with GPU or NIC passthrough who can flash or exploit device firmware, an attacker with physical access to a slot or an external PCIe/Thunderbolt port, or a supply-chain-modified card. Not reachable from software alone on a well-configured host - the entry point is device control.","remediation":"There is no single patch; this is configuration you must actively verify, and most fleets have never checked. Concretely: disable ATS unless a workload genuinely needs it (`pci=noats` on Linux, or the equivalent BIOS switch), and never leave ATS enabled for a device assigned to a tenant. Confirm PCIe ACS is enabled on every upstream port and switch so peer-to-peer traffic between passed-through devices is forced up through the IOMMU rather than routed directly - many server BIOSes ship ACS off, and several 'GPU peer-to-peer performance' tuning guides tell you to turn it off, which silently deletes the isolation boundary between two tenants' GPUs on the same switch. Verify per-device IOMMU groups are not lumping unrelated functions together. Then decide deliberately whether the peer-to-peer bandwidth you gain by disabling ACS is worth the tenant-isolation guarantee you lose, and write that decision down.","references":["http://thunderclap.io/","http://thunderclap.io/thunderclap-paper-ndss2019.pdf","https://www.kernel.org/doc/html/latest/admin-guide/kernel-parameters.html"],"status":"curated"},{"id":"NCVD-2019-007-hpe-sas-ssds-20-models-incl-vo04","cve":null,"aliases":["HPE 32,768-hour SSD bug","HPE bulletin a00092491en_us","HPD8"],"title":"HPE SAS SSDs (20 models incl. VO0480JFDGT, VO0960JFDGU, VO1920JFDGV, VO3840JFDHA, MO0400JFFCF-MO3200JFFCL)","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HPE SAS SSDs (20 models incl. VO0480JFDGT, VO0960JFDGU, VO1920JFDGV, VO3840JFDHA, MO0400JFFCF-MO3200JFFCL) with…","year":"2019","cvss_score":null,"severity":"unscored","kev":false,"impact":"PHYSICAL / FLEET-WIDE AVAILABILITY EVENT. At exactly 32,768 power-on hours - about 3 years, 270 days - the drive fails permanently. HPE states that neither the drive nor the data on it can be recovered. The killer property is SYNCHRONY: drives bought and racked together cross the threshold together, so RAID redundancy provides no protection because you lose several members at once. One administrator reported eight drives failing within MINUTES of each other. For a GPU operator this is the whole-cluster failure mode nothing in your HA design accounts for: your redundancy assumes independent failures, and a firmware counter overflow makes them perfectly correlated. Not an attack - a latent time bomb sitting in the fleet with a known detonation date, which is exactly why it belongs in an operator vulnerability database.","attack_vector":"No attacker. The trigger is elapsed powered-on time, and every affected drive with a similar install date reaches it simultaneously. The exposure is determined entirely by your procurement and racking history - a bulk purchase deployed in one window is the worst case.","remediation":"Flash to firmware HPD8 or later BEFORE the threshold; after the drive fails there is no recovery and you restore from backup. HPE shipped HPD8 for the first eight models from 22 November 2019 and the remaining twelve in mid-December 2019. The urgent operational step is inventory, not patching: pull power-on hours for every SAS SSD in the fleet (HPE Smart Storage Administrator, or smartctl attribute 9) and sort by hours remaining, because your window is defined by the oldest drives. Across a large estate this is a rolling drain-and-flash campaign of weeks, and it must be sequenced by remaining hours rather than by rack - and critically you must STAGGER it, because flashing a whole batch on one day recreates the correlated-failure problem for the next latent bug. Standing control: alert on power-on-hours thresholds fleet-wide, and deliberately mix drive batches and vendors across redundancy groups so a single firmware defect cannot take out every member of a RAID set at once.","references":["https://www.bleepingcomputer.com/news/hardware/hp-warns-that-some-ssd-drives-will-fail-at-32-768-hours-of-use/","https://support.hpe.com/hpesc/public/docDisplay?docId=emr_na-a00092491en_us"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"NCVD-2020-001-amd-zen-1-zen-zen-2-l1d-cache-wa","cve":null,"aliases":["Take A Way","Collide+Probe","Load+Reload"],"title":"AMD Zen 1 / Zen+ / Zen 2 - L1D cache way predictor: AMD's L1D way predictor hashes virtual addresses to predict which","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD Zen 1 / Zen+ / Zen 2 - L1D cache way predictor","year":"2020","cvss_score":null,"severity":"unscored","kev":false,"impact":"AMD's L1D way predictor hashes virtual addresses to predict which cache way holds a line. Collisions in that hash are observable, giving an attacker a channel to leak metadata about a victim's memory access pattern - the researchers used it to break KASLR, to recover an AES key from a table-based implementation, and to build a covert channel between processes. It works from JavaScript in a browser and across VMs, which is unusually broad reach for a microarchitectural channel.","attack_vector":"Local, unprivileged, co-resident with the victim on the same physical core. Affects Zen 1, Zen+ and Zen 2 (2017-2019 EPYC generations).","remediation":"**No CVE was assigned and AMD issued no microcode fix**, taking the position that existing side-channel guidance and secret-independent software already cover it. Treat this as unpatchable on affected silicon. Operator-side controls: do not co-schedule tenants on the same physical core, disable SMT on mixed-tenancy nodes, and prefer newer EPYC generations for workloads where cross-tenant leakage is in your threat model. Nothing here requires a reboot or firmware - it is a scheduling and fleet-composition decision.","references":["https://mlq.me/download/takeaway.pdf","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"tags":["tenant-isolation"]},{"id":"NCVD-2020-001-nvidia-multi-instance-gpu-mig-pa","cve":null,"aliases":["MIG side-channel caveat"],"title":"NVIDIA Multi-Instance GPU (MIG) partitioning: MIG gives each instance its own SM slice, L2 slice, memory slice and","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Multi-Instance GPU (MIG) partitioning","year":"2020","cvss_score":null,"severity":"unscored","kev":false,"impact":"MIG gives each instance its own SM slice, L2 slice, memory slice and memory-bandwidth allocation, and NVIDIA documents it as providing fault and performance isolation between instances. What it does NOT claim is protection against microarchitectural side channels, and it does not encrypt instance memory. Operators routinely sell MIG partitions as if they were independent GPUs. They are strong performance and fault isolation, not a cryptographic boundary - the shared GPU chip, shared power/thermal domain and shared memory controller remain observable.","attack_vector":"A tenant holding one MIG instance, observing shared resources used by a tenant on another instance of the same physical GPU. No exploit is needed for the telemetry channel; the frequency/power domain is shared by construction.","remediation":"Not a patchable defect - it is the documented scope of the feature. If your product promises isolation stronger than performance isolation, back MIG with NVIDIA Confidential Computing (Hopper and later, and note that CC and MIG have version-dependent compatibility constraints), or with whole-GPU allocation. Verify what you tell customers matches what MIG actually guarantees; also confirm MIG instances are actually destroyed and recreated between tenants rather than reused, because instance teardown is what triggers memory scrubbing.","references":["https://docs.nvidia.com/datacenter/tesla/mig-user-guide/","https://docs.nvidia.com/confidential-computing-deployment-guide/"],"status":"curated","tags":["tenant-isolation"]},{"id":"NCVD-2020-003-nvidia-multi-instance-gpu-mig-pa","cve":null,"aliases":["MIG side-channel caveat"],"title":"NVIDIA Multi-Instance GPU (MIG) partitioning: MIG gives each instance its own SM slice, L2 slice, memory slice and","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Multi-Instance GPU (MIG) partitioning","year":"2020","cvss_score":null,"severity":"unscored","kev":false,"impact":"MIG gives each instance its own SM slice, L2 slice, memory slice and memory-bandwidth allocation, and NVIDIA documents it as providing fault and performance isolation between instances. What it does NOT claim is protection against microarchitectural side channels, and it does not encrypt instance memory. Operators routinely sell MIG partitions as if they were independent GPUs. They are strong performance and fault isolation, not a cryptographic boundary - the shared GPU chip, shared power/thermal domain and shared memory controller remain observable.","attack_vector":"A tenant holding one MIG instance, observing shared resources used by a tenant on another instance of the same physical GPU. No exploit is needed for the telemetry channel; the frequency/power domain is shared by construction.","remediation":"Not a patchable defect - it is the documented scope of the feature. If your product promises isolation stronger than performance isolation, back MIG with NVIDIA Confidential Computing (Hopper and later, and note that CC and MIG have version-dependent compatibility constraints), or with whole-GPU allocation. Verify what you tell customers matches what MIG actually guarantees; also confirm MIG instances are actually destroyed and recreated between tenants rather than reused, because instance teardown is what triggers memory scrubbing.","references":["https://docs.nvidia.com/datacenter/tesla/mig-user-guide/","https://docs.nvidia.com/confidential-computing-deployment-guide/"],"status":"curated","tags":["tenant-isolation"]},{"id":"NCVD-2020-004-hpe-sas-ssds-ek0800jvypn-eo1600j","cve":null,"aliases":["HPE 40,000-hour SSD bug","HPE bulletin a00097382en_us","HPD7"],"title":"HPE SAS SSDs EK0800JVYPN, EO1600JVYPP, MK0800JVYPQ, MO1600JVYPR (800GB/1.6TB 12G SAS) with firmware prior to HPD7","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HPE SAS SSDs EK0800JVYPN, EO1600JVYPP, MK0800JVYPQ, MO1600JVYPR (800GB/1.6TB 12G SAS) with firmware prior to HPD7","year":"2020","cvss_score":null,"severity":"unscored","kev":false,"impact":"PHYSICAL / FLEET-WIDE AVAILABILITY EVENT, and the second one in four months - which is the real finding. At exactly 40,000 power-on hours, roughly 4.6 years, these drives fail permanently and HPE states that neither the data nor the drive can be recovered. Same correlated-failure property as the 32,768-hour bug: co-installed drives die together, so RAID fault tolerance is exceeded and the array goes with them. Affects HPE ProLiant, Synergy, Apollo 4200 and StoreEasy platforms - general-purpose server and storage hardware of exactly the kind that gets repurposed into GPU and AI build-outs. That two independent counter-overflow bombs surfaced in one vendor's SAS SSD line within months tells you this is a recurring firmware-engineering failure mode, not a freak event.","attack_vector":"No attacker. Elapsed powered-on hours, hitting every drive from the same deployment batch at the same moment. HPE projected the first failures would begin around October 2020.","remediation":"Flash to HPD7 or later before the threshold; HPE released it on 20 March 2020 with VMware ESXi, Windows and Linux packages. Post-failure there is no recovery - restore from backup. As with the 32,768-hour bug the first action is a fleet-wide power-on-hours audit (HPE Smart Storage Administrator or smartctl attribute 9) to find which drives are closest to the line, then a staggered rolling drain-and-flash sequenced by hours remaining rather than by rack. Make the standing controls permanent rather than treating this as a one-off: monitor power-on hours as a first-class fleet metric with alerting well ahead of any known threshold, subscribe to drive-vendor firmware bulletins as an operational feed, and deliberately mix procurement batches and vendors across redundancy groups so no single firmware defect can reach every member of an array simultaneously.","references":["https://www.bleepingcomputer.com/news/hardware/hpe-warns-of-new-bug-that-kills-ssd-drives-after-40-000-hours/","https://support.hpe.com/hpesc/public/docDisplay?docId=emr_na-a00097382en_us"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"NCVD-2021-001-amd-platform-secure-boot-psb-oem","cve":null,"aliases":["PSB left unfused","AMD Platform Secure Boot not enabled"],"title":"AMD Platform Secure Boot (PSB) OEM key fusing on EPYC server boards: PSB is the fuse-backed root of trust that makes","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Platform Secure Boot (PSB) OEM key fusing on EPYC server boards","year":"2021","cvss_score":null,"severity":"unscored","kev":false,"impact":"PSB is the fuse-backed root of trust that makes the PSP verify the OEM's BIOS signature before the x86 cores ever start. It is not enabled by default - the OEM has to burn their key hash into one-time-programmable fuses at manufacture, and a large fraction of shipped EPYC boards, including whitebox and ODM boards common in neocloud builds, arrive unfused. On an unfused board every firmware persistence flaw in this database becomes trivially exploitable: modified BIOS or PSP firmware boots without complaint, and a bare-metal tenant can leave something behind that owns every customer who follows. There is no CVE because it is a configuration state, not a bug, which is exactly why nobody checks it.","attack_vector":"Anyone who can write the SPI flash - a bare-metal tenant with ring 0, a supply-chain touch point between the factory and your rack, or an attacker chaining one of the SPI protection bypasses in this set. On an unfused board no signature check stands in the way.","remediation":"Verify PSB status at hardware intake, before the node ever enters the fleet - it is readable through the PSP mailbox and via open tooling (psb_status in the fwupd/amd-psb tooling family). Fusing is IRREVERSIBLE and is the OEM's action, not yours: you cannot fuse it yourself after the fact in most designs, and a wrong fuse bricks the board. So this is a procurement requirement - make 'PSB fused with the OEM key' a line item in your server RFP and reject boards that ship unfused. For hardware already deployed unfused, compensate with flash write protection, verified-boot measurement of the SPI image between tenancies, and not selling bare metal on those SKUs.","references":["https://www.amd.com/system/files/documents/amd-security-white-paper.pdf","https://github.com/fwupd/fwupd","https://www.amd.com/en/corporate/product-security"],"status":"curated"},{"id":"NCVD-2021-002-ddr4-dram-with-in-dram-trr-a-cou","cve":null,"aliases":["Half-Double"],"title":"DDR4 DRAM with in-DRAM TRR; a coupling effect that reaches rows at distance two rather than immediate neighbours","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"DDR4 DRAM with in-DRAM TRR; a coupling effect that reaches rows at distance two rather than immediate neighbours","year":"2021","cvss_score":null,"severity":"unscored","kev":false,"impact":"Google showed Rowhammer coupling is not confined to adjacent rows - hammering a row disturbs rows two away, and the mitigation logic that only watches immediate neighbours refreshes exactly the wrong rows. Operationally this means the TRR generation your DIMM vendor sold as a fix has a structural blind spot rather than a tuning problem, and mitigations designed around distance-one coupling need redesign. Same downstream consequences as any Rowhammer primitive: page-table corruption to host escape, or silent corruption of a co-tenant's data with no error signalled.","attack_vector":"Unprivileged local code sharing a memory controller with the victim. No CVE was assigned because this is a property of the DRAM, not a defect in a shippable product.","remediation":"Nothing you can install. The industry answer is Refresh Management (RFM) in the DDR5 spec plus revised in-DRAM tracking, which means new DIMMs and a memory-controller generation that drives RFM - capex on a refresh cycle, not a patch window. In the meantime: raise the refresh rate where BIOS allows it, and treat memory-controller sharing between untrusted tenants as a policy you have chosen to accept rather than a boundary you have.","references":["https://security.googleblog.com/2021/05/introducing-half-double-new-hammering.html","https://github.com/google/hammer-kit"],"status":"curated"},{"id":"NCVD-2021-002-discrete-tpm-lpc-spi-bus-unencry","cve":null,"aliases":["TPM bus sniffing","LPC/SPI interposer"],"title":"Discrete TPM (LPC / SPI bus, unencrypted sessions): A discrete TPM talks to the CPU over LPC or SPI in the clear unless","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Discrete TPM (LPC / SPI bus, unencrypted sessions)","year":"2021","cvss_score":null,"severity":"unscored","kev":false,"impact":"A discrete TPM talks to the CPU over LPC or SPI in the clear unless the software explicitly uses parameter-encrypted sessions, which most does not. An attacker who clips a logic analyser onto the bus - or onto the exposed pins of a socketed TPM header - reads the sealed key as it is released. The demonstrated case is recovering a BitLocker/LUKS volume key in minutes with a cheap probe, on a machine whose owner believed the disk was hardware-protected. For an operator, this is the flaw that turns a physically accessible node into a full data disclosure, and it leaves no trace in any log.","attack_vector":"Physical access to the motherboard for the duration of one boot. In practice: a colo cage neighbour, remote-hands staff, a decommissioning or RMA handler, or hardware intercepted in shipping. No credentials, no software exploit, no persistence needed.","remediation":"No patch exists - it is a property of the bus, not a bug. Mitigations are architectural: enable TPM parameter encryption / encrypted sessions in the software that unseals (recent Linux and Windows stacks support it, older ones do not), require a PIN or second factor so the TPM value alone is not sufficient to unlock, prefer fTPM where the bus is internal to the package, and use tamper-evident chassis with a documented seal check on every physical touch. Treat any node that left your custody as untrusted until re-provisioned.","references":["https://dolosgroup.io/blog/2021/7/9/from-stolen-laptop-to-inside-the-company-network","https://pulsesecurity.co.nz/articles/TPM-sniffing"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2021-005-intel-cpus-with-sgx-attacked-ove","cve":null,"aliases":["VoltPillager","hardware undervolting attack on SGX"],"title":"Intel CPUs with SGX, attacked over the SVID serial bus between the voltage regulator and the CPU package: Re-runs","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel CPUs with SGX, attacked over the SVID serial bus between the voltage regulator and the CPU package","year":"2021","cvss_score":null,"severity":"unscored","kev":false,"impact":"Re-runs the Plundervolt fault injection using a cheap external microcontroller wired to the motherboard's voltage-regulator bus, so it works on hosts that have already applied the Plundervolt microcode fix - that fix only disabled the software MSR, it did not stop voltage manipulation reaching the package. Recovers keys from enclaves and corrupts enclave computation. Intel's position is that physical attacks are outside the SGX threat model, so there is no patch and there will not be one. The operator consequence is concrete: if an attacker can put hands on the chassis, SGX and by extension the confidential-computing guarantee on that machine are void.","attack_vector":"Physical access to the motherboard, roughly 30 dollars of hardware, and a few minutes to attach to the SVID bus. Relevant threat models: colocation facilities where you do not control the cage, hardware in transit, decommissioned or RMA'd nodes, hosting partners, and anyone with datacenter floor access. Not reachable remotely.","remediation":"UNPATCHABLE by design - Intel classifies it as out of scope, so there is no microcode, BIOS or kernel fix to deploy. Mitigation is entirely physical and procedural: chassis intrusion detection wired into the BMC and actually alarmed, tamper-evident seals, controlled cage access with audited entry logs, and refusing to run confidential workloads on hardware whose physical custody you cannot vouch for. If you sell confidential compute, this is a contractual and facility-security question, not an engineering one - and it should shape which sites you are willing to place attested workloads in.","references":["https://www.usenix.org/conference/usenixsecurity21/presentation/chen-zitai","https://github.com/zt-chen/voltpillager"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2021-013-argo-workflows-argo-server-tls-k","cve":null,"aliases":["GHSA-6c73-2v8x-qpvm"],"title":"Argo Workflows (Argo Server, TLS keys baked into the container image): Argo Server's TLS private keys ship inside the","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo Workflows (Argo Server, TLS keys baked into the container image)","year":"2021","cvss_score":null,"severity":"unscored","kev":false,"impact":"Argo Server's TLS private keys ship inside the published container image, so anyone who can pull the image extracts them. With those keys plus network position an attacker decrypts Argo Server traffic or forges requests to it, which means impersonating the operators and tenants who drive workflow scheduling.","attack_vector":"An attacker on any network path to Argo Server. Affects Argo Server before 3.0 with --secure=true, or 3.0+ where --secure is left unspecified.","remediation":"Upgrade to 3.0.9 or 3.1.6 and restart Argo Server. Do not expose Argo Server directly - terminate TLS at a load balancer or ingress holding certificates you issued, and keep the pod reachable only from inside the cluster.","references":["https://github.com/argoproj/argo-workflows/security/advisories/GHSA-6c73-2v8x-qpvm"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-285"],"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2021-014-argo-workflows-argo-server-auth","cve":null,"aliases":["GHSA-prqf-xr2j-xf65"],"title":"Argo Workflows (Argo Server, --auth-mode=client on Kubernetes 1.19+ outside a pod): In this configuration the client's","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo Workflows (Argo Server, --auth-mode=client on Kubernetes 1.19+ outside a pod)","year":"2021","cvss_score":null,"severity":"unscored","kev":false,"impact":"In this configuration the client's supplied key is ignored and the server's own Kubernetes credentials are used instead. Every request silently runs at the server's privilege level, so a tenant with a narrow client identity gets whatever the Argo Server account can do across the cluster.","attack_vector":"Any client authenticating to an Argo Server that runs outside a Kubernetes pod (bare metal or VM) with --auth-mode=client on Kubernetes 1.19 or newer, where the server account is more privileged than the client.","remediation":"Upgrade to 3.0.9 or 3.1.6 and restart. There is no workaround for the affected versions - if you cannot upgrade, run Argo Server inside the cluster as a pod and scope its service account down to the minimum.","references":["https://github.com/argoproj/argo-workflows/security/advisories/GHSA-prqf-xr2j-xf65"],"status":"curated","tags":["tenant-isolation"]},{"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2021-015-argo-workflows-argo-server-defau","cve":null,"aliases":["GHSA-rc7p-gmvh-xfx2"],"title":"Argo Workflows (Argo Server default --auth-mode=server before 3.0): Before 3.0 the Argo Server defaulted to","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo Workflows (Argo Server default --auth-mode=server before 3.0)","year":"2021","cvss_score":null,"severity":"unscored","kev":false,"impact":"Before 3.0 the Argo Server defaulted to --auth-mode=server, meaning every request ran as the server's own service account. An exposed UI therefore lets anonymous internet users submit workflows that execute arbitrary containers on the cluster. This is the configuration behind the observed crypto-mining campaigns against Argo Workflows clusters.","attack_vector":"Any unauthenticated user who can reach the Argo Workflows UI, which in the reported incidents meant the open internet.","remediation":"Move to --auth-mode=client and pull the UI off the internet immediately, then upgrade Argo Server to 3.x or later. Because this has been actively exploited in the wild, audit the cluster for unexpected workflows and workload pods before assuming it is clean.","references":["https://github.com/argoproj/argo-workflows/security/advisories/GHSA-rc7p-gmvh-xfx2","https://www.intezer.com/blog/container-security/new-attacks-on-kubernetes-via-misconfigured-argo-workflows/"],"status":"curated","tags":["tenant-isolation"]},{"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2021-018-argo-workflows-argo-server-tls-k","cve":null,"aliases":["GHSA-6c73-2v8x-qpvm"],"title":"Argo Workflows (Argo Server, TLS keys baked into the container image): Argo Server's TLS private keys ship inside the","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo Workflows (Argo Server, TLS keys baked into the container image)","year":"2021","cvss_score":null,"severity":"unscored","kev":false,"impact":"Argo Server's TLS private keys ship inside the published container image, so anyone who can pull the image extracts them. With those keys plus network position an attacker decrypts Argo Server traffic or forges requests to it, which means impersonating the operators and tenants who drive workflow scheduling.","attack_vector":"An attacker on any network path to Argo Server. Affects Argo Server before 3.0 with --secure=true, or 3.0+ where --secure is left unspecified.","remediation":"Upgrade to 3.0.9 or 3.1.6 and restart Argo Server. Do not expose Argo Server directly - terminate TLS at a load balancer or ingress holding certificates you issued, and keep the pod reachable only from inside the cluster.","references":["https://github.com/argoproj/argo-workflows/security/advisories/GHSA-6c73-2v8x-qpvm"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-285"],"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2021-019-argo-workflows-argo-server-auth","cve":null,"aliases":["GHSA-prqf-xr2j-xf65"],"title":"Argo Workflows (Argo Server, --auth-mode=client on Kubernetes 1.19+ outside a pod): In this configuration the client's","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo Workflows (Argo Server, --auth-mode=client on Kubernetes 1.19+ outside a pod)","year":"2021","cvss_score":null,"severity":"unscored","kev":false,"impact":"In this configuration the client's supplied key is ignored and the server's own Kubernetes credentials are used instead. Every request silently runs at the server's privilege level, so a tenant with a narrow client identity gets whatever the Argo Server account can do across the cluster.","attack_vector":"Any client authenticating to an Argo Server that runs outside a Kubernetes pod (bare metal or VM) with --auth-mode=client on Kubernetes 1.19 or newer, where the server account is more privileged than the client.","remediation":"Upgrade to 3.0.9 or 3.1.6 and restart. There is no workaround for the affected versions - if you cannot upgrade, run Argo Server inside the cluster as a pod and scope its service account down to the minimum.","references":["https://github.com/argoproj/argo-workflows/security/advisories/GHSA-prqf-xr2j-xf65"],"status":"curated","tags":["tenant-isolation"]},{"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2021-020-argo-workflows-argo-server-defau","cve":null,"aliases":["GHSA-rc7p-gmvh-xfx2"],"title":"Argo Workflows (Argo Server default --auth-mode=server before 3.0): Before 3.0 the Argo Server defaulted to","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo Workflows (Argo Server default --auth-mode=server before 3.0)","year":"2021","cvss_score":null,"severity":"unscored","kev":false,"impact":"Before 3.0 the Argo Server defaulted to --auth-mode=server, meaning every request ran as the server's own service account. An exposed UI therefore lets anonymous internet users submit workflows that execute arbitrary containers on the cluster. This is the configuration behind the observed crypto-mining campaigns against Argo Workflows clusters.","attack_vector":"Any unauthenticated user who can reach the Argo Workflows UI, which in the reported incidents meant the open internet.","remediation":"Move to --auth-mode=client and pull the UI off the internet immediately, then upgrade Argo Server to 3.x or later. Because this has been actively exploited in the wild, audit the cluster for unexpected workflows and workload pods before assuming it is clean.","references":["https://github.com/argoproj/argo-workflows/security/advisories/GHSA-rc7p-gmvh-xfx2","https://www.intezer.com/blog/container-security/new-attacks-on-kubernetes-via-misconfigured-argo-workflows/"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-306","CWE-1188"],"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2022-004-beegfs-client-to-metadata-storag","cve":null,"aliases":["BeeGFS connAuthFile disabled","BeeGFS unauthenticated client trust"],"title":"BeeGFS (client-to-metadata/storage service authentication, connAuthFile): Class entry, not a single CVE. Before BeeGFS","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"BeeGFS (client-to-metadata/storage service authentication, connAuthFile)","year":"2022","cvss_score":null,"severity":"unscored","kev":false,"impact":"Class entry, not a single CVE. Before BeeGFS 7.2.8 / 7.3.2 (October 2022) a default install ran with no connection authentication at all, and any process on a node that mounts BeeGFS could talk to the metadata and storage services directly as a trusted peer. That is the mechanism behind CVE-2019-15897, and the same posture persists on any current install where the operator answered the new prompt by disabling authentication rather than configuring a connAuthFile. Because a parallel filesystem is commonly mounted by several clusters, this is also a lateral-movement path from one cluster into another. Two adjacent misconfigurations make it worse: a world-readable connAuthFile hands the shared secret to every local user, and mounting without nosuid lets a tenant plant a setuid binary on shared storage and pick it up as root elsewhere.","attack_vector":"Any user with a shell or a job on a node that mounts BeeGFS. No credential is required when authentication is disabled. Relying on connNetFilterFile.conf or connInterfacesFile.conf does not substitute - those restrict interfaces and networks, not identity.","remediation":"Run BeeGFS 7.2.8 / 7.3.2 or later so the install refuses to start without an explicit authentication decision, then actually configure a connAuthFile rather than disabling authentication. Set the connAuthFile mode so only the BeeGFS service account can read it, and mount every client nosuid. Applying this means restarting the BeeGFS services and remounting clients cluster-wide. Note that ThinkParQ does not publish a security-advisory index at beegfs.io - CVE-2019-15897 and this configuration guidance both come from the third-party researcher who reported it.","references":["https://www.hpcsec.com/2023/01/02/beegfs-security/","https://doc.beegfs.io/7.3.2/release_notes.html","https://www.hpcsec.com/2019/12/04/cve-2019-15897/"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-306","CWE-1188"],"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2022-005-beegfs-client-to-metadata-storag","cve":null,"aliases":["BeeGFS connAuthFile disabled","BeeGFS unauthenticated client trust"],"title":"BeeGFS (client-to-metadata/storage service authentication, connAuthFile): Class entry, not a single CVE. Before BeeGFS","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"BeeGFS (client-to-metadata/storage service authentication, connAuthFile)","year":"2022","cvss_score":null,"severity":"unscored","kev":false,"impact":"Class entry, not a single CVE. Before BeeGFS 7.2.8 / 7.3.2 (October 2022) a default install ran with no connection authentication at all, and any process on a node that mounts BeeGFS could talk to the metadata and storage services directly as a trusted peer. That is the mechanism behind CVE-2019-15897, and the same posture persists on any current install where the operator answered the new prompt by disabling authentication rather than configuring a connAuthFile. Because a parallel filesystem is commonly mounted by several clusters, this is also a lateral-movement path from one cluster into another. Two adjacent misconfigurations make it worse: a world-readable connAuthFile hands the shared secret to every local user, and mounting without nosuid lets a tenant plant a setuid binary on shared storage and pick it up as root elsewhere.","attack_vector":"Any user with a shell or a job on a node that mounts BeeGFS. No credential is required when authentication is disabled. Relying on connNetFilterFile.conf or connInterfacesFile.conf does not substitute - those restrict interfaces and networks, not identity.","remediation":"Run BeeGFS 7.2.8 / 7.3.2 or later so the install refuses to start without an explicit authentication decision, then actually configure a connAuthFile rather than disabling authentication. Set the connAuthFile mode so only the BeeGFS service account can read it, and mount every client nosuid. Applying this means restarting the BeeGFS services and remounting clients cluster-wide. Note that ThinkParQ does not publish a security-advisory index at beegfs.io - CVE-2019-15897 and this configuration guidance both come from the third-party researcher who reported it.","references":["https://www.hpcsec.com/2023/01/02/beegfs-security/","https://doc.beegfs.io/7.3.2/release_notes.html","https://www.hpcsec.com/2019/12/04/cve-2019-15897/"],"status":"curated","tags":["tenant-isolation"]},{"id":"NCVD-2023-001-ddr4-chips-from-all-three-major","cve":null,"aliases":["RowPress"],"title":"DDR4 chips from all three major DRAM manufacturers; worsens as process nodes shrink: A different read-disturbance","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"DDR4 chips from all three major DRAM manufacturers; worsens as process nodes shrink","year":"2023","cvss_score":null,"severity":"unscored","kev":false,"impact":"A different read-disturbance mechanism from Rowhammer: instead of repeatedly opening and closing a row, hold it open. That cuts the activation count needed for a flip by one to two orders of magnitude, and in the extreme a single activation of an adjacent row suffices. The operator consequence is that activation-counting mitigations - which is what essentially every deployed Rowhammer defence is - can be under threshold and still let flips through. Anything you were told about 'activations per refresh window' as a safety argument needs re-deriving.","attack_vector":"Unprivileged local code on a shared host, demonstrated on a real DDR4 system that already had Rowhammer protection enabled. Sharing a memory controller with the victim is the only requirement.","remediation":"Not patchable. Mitigations have to be redesigned to bound how long a row stays open, not just how often it is activated; the authors show existing Rowhammer defences can be adapted at low additional cost, but that lands in future memory controllers and DRAM, not in your installed base. For now the honest answer to a customer asking 'is our memory isolated from the tenant next door' is no, and the only lever you control is not putting them there.","references":["https://arxiv.org/abs/2306.17061","https://github.com/CMU-SAFARI/RowPress"],"status":"curated"},{"id":"NCVD-2023-001-gigabyte-uefi-firmware-oem-updat","cve":null,"aliases":["Gigabyte App Center backdoor"],"title":"Gigabyte UEFI firmware (OEM update-dropper in firmware): Gigabyte firmware shipped a UEFI module that writes a Windows","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Gigabyte UEFI firmware (OEM update-dropper in firmware)","year":"2023","cvss_score":null,"severity":"unscored","kev":false,"impact":"Gigabyte firmware shipped a UEFI module that writes a Windows executable to disk at every boot and has it fetch and run further code from Gigabyte-controlled URLs - one of them plain HTTP, with weak certificate validation on the others. It is not malware, it is the vendor's own updater, but it behaves exactly like a firmware implant: an unremovable, boot-persistent downloader with no user consent and no way to disable it from the OS. Anyone who can MITM that fetch, or who compromises the vendor's distribution point, gets code execution on every affected machine at every boot.","attack_vector":"A network attacker in path of the update fetch, or a supply-chain compromise of the vendor endpoint. No access to the machine needed.","remediation":"Gigabyte published firmware updates that fix the transport and validation; applying them is a per-board BIOS flash plus reboot. Where the platform allows it, disable the 'APP Center Download & Install' option in BIOS setup - a config-only mitigation you can push faster than a firmware campaign. The broader operator lesson: audit what your board vendor's firmware talks to on the network before a node ever carries tenant workload, and block outbound egress from the provisioning network by default.","references":["https://eclypsium.com/research/supply-chain-risk-from-gigabyte-app-center-backdoor/","https://kb.cert.org/vuls/id/287178"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"NCVD-2023-001-intel-processors-with-linear-add","cve":null,"aliases":["SLAM","Spectre based on Linear Address Masking"],"title":"Intel processors with Linear Address Masking (LAM): SLAM: Linear Address Masking, a feature intended to let software","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors with Linear Address Masking (LAM)","year":"2023","cvss_score":null,"severity":"unscored","kev":false,"impact":"SLAM: Linear Address Masking, a feature intended to let software store metadata in unused address bits, makes previously impractical Spectre gadgets exploitable by widening the set of usable pointer-chasing gadgets - so a performance and usability feature became a security regression. Notable as a case where the mitigation was to ship the hardware feature disabled by default in Linux.","attack_vector":"Local unprivileged code on a host with LAM enabled.","remediation":"Linux disables LAM by default in response; keep it disabled unless you have a specific requirement and have assessed the tradeoff. Kernel-level configuration - a kernel update and reboot, no microcode or BIOS. Verify LAM state on your nodes rather than assuming the default held through a kernel upgrade.","references":["https://www.vusec.net/projects/slam/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"NCVD-2023-001-msi-intel-boot-guard-oem-key-lea","cve":null,"aliases":["no CVE"],"title":"MSI / Intel Boot Guard OEM key leak: The Money Message ransomware dump exposed MSI's firmware image-signing private","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"MSI / Intel Boot Guard OEM key leak","year":"2023","cvss_score":null,"severity":"unscored","kev":false,"impact":"The Money Message ransomware dump exposed MSI's firmware image-signing private keys for 57 products and Intel Boot Guard KM/BPM private keys for 116 products, reportedly touching Intel, Lenovo and Supermicro platforms. An attacker can sign a firmware image that the hardware root of trust accepts — Boot Guard is effectively void on affected silicon and the implant survives any OS reinstall","attack_vector":"Supply chain / local flash","remediation":"There is no patch. Boot Guard keys are fused into the CPU at manufacture, so revocation is impossible on shipped hardware. The only response is to treat Boot Guard as non-authoritative on affected platforms and add an independent firmware-measurement/attestation layer","references":["https://www.helpnetsecurity.com/2023/05/08/msi-private-keys-leaked/"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2023-003-integrated-gpu-graphics-data-com","cve":null,"aliases":["GPU.zip","graphics data compression side channel"],"title":"Integrated GPU graphics data compression (Intel, AMD, Apple, Arm, Qualcomm, NVIDIA): GPUs apply data-dependent lossless","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Integrated GPU graphics data compression (Intel, AMD, Apple, Arm, Qualcomm, NVIDIA)","year":"2023","cvss_score":null,"severity":"unscored","kev":false,"impact":"GPUs apply data-dependent lossless compression to framebuffer traffic even when software never asked for it, and the resulting DRAM traffic pattern is measurable from a co-resident context. The published attack recovered pixel data cross-origin from a browser. For a datacenter operator the realistic worry is remote-desktop, cloud-gaming and VDI fleets where the rendered surface is a customer's screen. No CVE was assigned and the affected vendors declined to ship fixes.","attack_vector":"A co-resident attacker able to render and time - in the original work, a web page in another browser tab. On a shared render host, another tenant's session.","remediation":"UNPATCHABLE. Vendors treated the compression as working-as-designed and the browsers mitigated the specific web attack by restricting cross-origin iframe rendering. There is no GPU driver or firmware update. If you sell shared remote-desktop or cloud-gaming capacity, the only real control is not co-residing untrusted sessions on the same GPU.","references":["https://www.hertzbleed.com/gpu.zip/","https://arxiv.org/abs/2310.01187"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"tags":["tenant-isolation"]},{"id":"NCVD-2023-004-nvidia-confidential-computing-h1","cve":null,"aliases":["NVIDIA CC DevTools mode"],"title":"NVIDIA Confidential Computing (H100/H200/B100/B200/GB200) - CC-DevTools operating mode: NVIDIA GPU confidential","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Confidential Computing (H100/H200/B100/B200/GB200) - CC-DevTools operating mode","year":"2023","cvss_score":null,"severity":"unscored","kev":false,"impact":"NVIDIA GPU confidential computing has three modes - CC-Off, CC-On and CC-DevTools. DevTools mode exists so that profilers and debuggers keep working, and it does so by disabling the protections: the GPU still produces attestation material and still looks like a confidential GPU to tooling, but the confidentiality guarantee is not in force. This is a configuration state, not a bug, and it is the single most likely way an operator ships a 'confidential' GPU that is not one.","attack_vector":"No attacker skill required - the risk is that the mode is set wrong, or set right and then changed by an operator debugging a performance problem and never changed back. Anyone with host root can set it.","remediation":"Not patchable; it is a control you must enforce. Verify CC mode per GPU as a continuous check rather than a build-time one, and make the attestation policy reject DevTools mode explicitly rather than accepting any signed report. Changing CC mode requires a GPU reset and therefore a node drain, which is also why operators are tempted to leave a node in DevTools mode once they set it.","references":["https://docs.nvidia.com/confidential-computing-deployment-guide/","https://docs.nvidia.com/cc-deployment-guide-tee.pdf"],"status":"curated","fleet":{"pain_class":"node-drain"},"tags":["tenant-isolation"]},{"id":"NCVD-2023-005-integrated-gpu-graphics-data-com","cve":null,"aliases":["GPU.zip","graphics data compression side channel"],"title":"Integrated GPU graphics data compression (Intel, AMD, Apple, Arm, Qualcomm, NVIDIA): GPUs apply data-dependent lossless","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Integrated GPU graphics data compression (Intel, AMD, Apple, Arm, Qualcomm, NVIDIA)","year":"2023","cvss_score":null,"severity":"unscored","kev":false,"impact":"GPUs apply data-dependent lossless compression to framebuffer traffic even when software never asked for it, and the resulting DRAM traffic pattern is measurable from a co-resident context. The published attack recovered pixel data cross-origin from a browser. For a datacenter operator the realistic worry is remote-desktop, cloud-gaming and VDI fleets where the rendered surface is a customer's screen. No CVE was assigned and the affected vendors declined to ship fixes.","attack_vector":"A co-resident attacker able to render and time - in the original work, a web page in another browser tab. On a shared render host, another tenant's session.","remediation":"UNPATCHABLE. Vendors treated the compression as working-as-designed and the browsers mitigated the specific web attack by restricting cross-origin iframe rendering. There is no GPU driver or firmware update. If you sell shared remote-desktop or cloud-gaming capacity, the only real control is not co-residing untrusted sessions on the same GPU.","references":["https://www.hertzbleed.com/gpu.zip/","https://arxiv.org/abs/2310.01187"],"status":"curated","tags":["tenant-isolation"]},{"id":"NCVD-2023-007-nvidia-confidential-computing-h1","cve":null,"aliases":["NVIDIA CC DevTools mode"],"title":"NVIDIA Confidential Computing (H100/H200/B100/B200/GB200) - CC-DevTools operating mode: NVIDIA GPU confidential","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Confidential Computing (H100/H200/B100/B200/GB200) - CC-DevTools operating mode","year":"2023","cvss_score":null,"severity":"unscored","kev":false,"impact":"NVIDIA GPU confidential computing has three modes - CC-Off, CC-On and CC-DevTools. DevTools mode exists so that profilers and debuggers keep working, and it does so by disabling the protections: the GPU still produces attestation material and still looks like a confidential GPU to tooling, but the confidentiality guarantee is not in force. This is a configuration state, not a bug, and it is the single most likely way an operator ships a 'confidential' GPU that is not one.","attack_vector":"No attacker skill required - the risk is that the mode is set wrong, or set right and then changed by an operator debugging a performance problem and never changed back. Anyone with host root can set it.","remediation":"Not patchable; it is a control you must enforce. Verify CC mode per GPU as a continuous check rather than a build-time one, and make the attestation policy reject DevTools mode explicitly rather than accepting any signed report. Changing CC mode requires a GPU reset and therefore a node drain, which is also why operators are tempted to leave a node in DevTools mode once they set it.","references":["https://docs.nvidia.com/confidential-computing-deployment-guide/","https://docs.nvidia.com/cc-deployment-guide-tee.pdf"],"status":"curated","tags":["tenant-isolation"]},{"id":"NCVD-2024-001-amd-global-history-register-side","cve":null,"aliases":["Branch History Leak","AMD-SB-7026"],"title":"AMD - Global History Register side channel: A side channel through the branch predictor's Global History Register","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD - Global History Register side channel","year":"2024","cvss_score":null,"severity":"unscored","kev":false,"impact":"A side channel through the branch predictor's Global History Register, disclosed by Harbin Institute of Technology. Branch history is shared microarchitectural state, so an attacker observes the shape of a co-resident victim's control flow - which for an inference workload can reveal what model is running and what path a request took through it.","attack_vector":"Local, co-resident with the victim.","remediation":"**No CVE and no fix** - AMD's response is software best practices. As with the other predictor-state channels, the enforceable control is not co-scheduling mutually untrusted tenants on the same physical core. No patch, no reboot; this is a scheduling and fleet-composition decision.","references":["https://www.amd.com/en/resources/product-security/bulletin/amd-sb-7026.html","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"tags":["tenant-isolation"]},{"id":"NCVD-2024-001-intel-sgx-cache-side-channel-on","cve":null,"aliases":["TeeJam"],"title":"Intel SGX (cache side channel on sub-cacheline access): TeeJam: shows that SGX's cache-based side-channel resistance is","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SGX (cache side channel on sub-cacheline access)","year":"2024","cvss_score":null,"severity":"unscored","kev":false,"impact":"TeeJam: shows that SGX's cache-based side-channel resistance is weaker than assumed at sub-cacheline granularity, enabling practical key recovery against enclave implementations previously believed to be constant-time. Operationally this matters because 'we run it in an enclave' is often the entire argument for putting a key on a shared host.","attack_vector":"Local code on the same machine as the victim enclave, with the scheduling control a privileged host has.","remediation":"No single patch - mitigation lives in the enclave software (constant-time implementations hardened at sub-cacheline granularity) and in keeping the SGX SDK/PSW current. Operator action is to require enclave vendors to state which side-channel hardening they apply, and to keep microcode and PSW at current TCB so attestation reflects reality.","references":["https://www.intel.com/content/www/us/en/developer/topic-technology/software-security-guidance/overview.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"tags":["tenant-isolation"]},{"id":"NCVD-2024-001-platform-attestation-as-an-opera","cve":null,"aliases":[],"title":"Platform attestation as an operational control (fTPM vs discrete TPM trust): Design-level: on most GPU servers the TPM","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Platform attestation as an operational control (fTPM vs discrete TPM trust)","year":"2024","cvss_score":null,"severity":"unscored","kev":false,"impact":"Design-level: on most GPU servers the TPM that backs measured boot is a firmware TPM inside the CPU package or chipset, whose own integrity depends on the same firmware the attestation is supposed to be measuring. If the platform firmware or the management engine is compromised, the fTPM reports whatever that firmware tells it to, and a verifier cannot distinguish a clean node from a compromised one. Every entry in this database that yields SMM, BMC or CSME code execution collapses the attestation guarantee alongside it. Operators selling 'verified clean bare metal' or confidential GPU compute are usually asserting something their hardware cannot independently prove.","attack_vector":"Any attacker who reaches the firmware layer beneath the TPM - SMM code execution, BMC takeover, or a management-engine flaw. The attestation does not fail loudly; it keeps passing.","remediation":"No patch. Architectural: root attestation in a device that is independent of the firmware it measures (discrete TPM on its own bus, or a separate platform root-of-trust device such as an OCP Cerberus-style controller), pin expected measurements rather than accepting any well-formed quote, verify the freshness and provenance of quotes rather than just their signature, and pair attestation with an independent firmware-integrity scan. Document to customers what attestation does and does not prove rather than over-claiming it.","references":["https://www.opencompute.org/projects/security","https://trustedcomputinggroup.org/resource/tpm-library-specification/"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2024-002-amd-zen-2-zen-3-and-zen-4-platfo","cve":null,"aliases":["ZenHammer"],"title":"AMD Zen 2, Zen 3 and Zen 4 platforms with DDR4 (7/10 Zen 2 and 6/10 Zen 3 devices flipped) and DDR5 (1/10 devices)","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD Zen 2, Zen 3 and Zen 4 platforms with DDR4 (7/10 Zen 2 and 6/10 Zen 3 devices flipped) and DDR5 (1/10 devices)","year":"2024","cvss_score":null,"severity":"unscored","kev":false,"impact":"Removes the 'we run EPYC, Rowhammer research is all Intel' excuse. ETH Zurich reverse-engineered AMD's DRAM address mapping and got bit flips on Zen 2/3/4, then demonstrated the standard exploit chain - page-table manipulation, RSA key corruption and sudo privilege escalation. It is also the first public DDR5 bit flip on a commodity system. For a GPU cloud this matters because EPYC is the host CPU under most HGX and 8-way GPU nodes: the box holding your control plane, your tenant credentials and every tenant's data on its way to the GPUs is hammerable from a guest.","attack_vector":"Unprivileged local code on an AMD Zen 2/3/4 host sharing DRAM with the victim. A container tenant or VM on the CPU side of a GPU node is sufficient; no GPU access needed.","remediation":"No vendor patch. The DDR4 case is the same non-answer as TRRespass and Blacksmith: refresh-rate increase where BIOS exposes it, or do not co-tenant behind one memory controller. The DDR5 result is a single device out of ten, so DDR5 EPYC platforms are better but not clear. Before you argue you are safe, run the published ZenHammer fuzzer against a sample of your own DIMM SKUs - vendor and date code determine your exposure far more than the CPU does, and this is the cheapest ground truth available.","references":["https://comsec.ethz.ch/research/dram/zenhammer/","https://github.com/comsec-group/zenhammer"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2024-002-intel-processors-indirect-branch","cve":null,"aliases":["Indirector"],"title":"Intel processors (Indirect Branch Predictor structure): Indirector: reverse-engineering the Indirect Branch Predictor","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (Indirect Branch Predictor structure)","year":"2024","cvss_score":null,"severity":"unscored","kev":false,"impact":"Indirector: reverse-engineering the Indirect Branch Predictor and Branch Target Buffer on recent Intel parts yielded high-precision branch target injection that works despite existing eIBRS/IBPB deployment, and showed that IBPB does not clear as much predictor state as assumed. The operator takeaway is that the barrier primitive hypervisors use for tenant separation is weaker than the documentation implies.","attack_vector":"Local code on an affected processor; the work targets cross-process and cross-privilege leakage.","remediation":"No single CVE or microcode fix maps cleanly to this research. Mitigation is the existing toolkit applied more aggressively: keep microcode current, enable IBPB on context switch where your workload can absorb the cost, and treat co-tenancy of untrusted workloads on the same physical core as unsupported. Both cost throughput.","references":["https://indirector.cpusec.org/"],"status":"curated","fleet":{"pain_class":"microcode + reboot"},"tags":["tenant-isolation"]},{"id":"NCVD-2024-003-nvidia-remote-attestation-servic","cve":null,"aliases":["NRAS dependency","GPU attestation availability"],"title":"NVIDIA Remote Attestation Service (NRAS) / NVIDIA attestation SDK: Verifying an NVIDIA GPU attestation report","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Remote Attestation Service (NRAS) / NVIDIA attestation SDK","year":"2024","cvss_score":null,"severity":"unscored","kev":false,"impact":"Verifying an NVIDIA GPU attestation report by the default path means calling NVIDIA's hosted Remote Attestation Service and fetching Reference Integrity Manifests from NVIDIA infrastructure. That places a live third-party dependency inside the trust decision for every confidential workload you start: if NRAS is unreachable, your orchestration either blocks workload admission or fails open, and most implementations fail open because blocking looks like an outage. There is no CVE here - it is an architectural exposure that operators consistently discover during their first NRAS incident rather than during design.","attack_vector":"Not an attacker in the usual sense. The realistic events are an NRAS outage, a network egress policy that blocks it, or an air-gapped deployment where it was never reachable at all.","remediation":"Deploy the local verifier path rather than the remote one where your threat model allows: NVIDIA supports local attestation verification with cached RIMs and the device identity certificate chain, which removes the runtime dependency. Cost: you take on RIM caching and freshness management. Whichever you choose, test the failure mode deliberately - block NRAS in staging and confirm your admission controller denies rather than admits.","references":["https://docs.attestation.nvidia.com/","https://github.com/NVIDIA/nvtrust"],"status":"curated"},{"id":"NCVD-2024-006-intel-sgx-cache-side-channel-on","cve":null,"aliases":["TeeJam"],"title":"Intel SGX (cache side channel on sub-cacheline access): TeeJam: shows that SGX's cache-based side-channel resistance is","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SGX (cache side channel on sub-cacheline access)","year":"2024","cvss_score":null,"severity":"unscored","kev":false,"impact":"TeeJam: shows that SGX's cache-based side-channel resistance is weaker than assumed at sub-cacheline granularity, enabling practical key recovery against enclave implementations previously believed to be constant-time. Operationally this matters because 'we run it in an enclave' is often the entire argument for putting a key on a shared host.","attack_vector":"Local code on the same machine as the victim enclave, with the scheduling control a privileged host has.","remediation":"No single patch - mitigation lives in the enclave software (constant-time implementations hardened at sub-cacheline granularity) and in keeping the SGX SDK/PSW current. Operator action is to require enclave vendors to state which side-channel hardening they apply, and to keep microcode and PSW at current TCB so attestation reflects reality.","references":["https://www.intel.com/content/www/us/en/developer/topic-technology/software-security-guidance/overview.html"],"status":"curated","tags":["tenant-isolation"]},{"id":"NCVD-2024-007-intel-processors-indirect-branch","cve":null,"aliases":["Indirector"],"title":"Intel processors (Indirect Branch Predictor structure): Indirector: reverse-engineering the Indirect Branch Predictor","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (Indirect Branch Predictor structure)","year":"2024","cvss_score":null,"severity":"unscored","kev":false,"impact":"Indirector: reverse-engineering the Indirect Branch Predictor and Branch Target Buffer on recent Intel parts yielded high-precision branch target injection that works despite existing eIBRS/IBPB deployment, and showed that IBPB does not clear as much predictor state as assumed. The operator takeaway is that the barrier primitive hypervisors use for tenant separation is weaker than the documentation implies.","attack_vector":"Local code on an affected processor; the work targets cross-process and cross-privilege leakage.","remediation":"No single CVE or microcode fix maps cleanly to this research. Mitigation is the existing toolkit applied more aggressively: keep microcode current, enable IBPB on context switch where your workload can absorb the cost, and treat co-tenancy of untrusted workloads on the same physical core as unsupported. Both cost throughput.","references":["https://indirector.cpusec.org/"],"status":"curated","tags":["tenant-isolation"]},{"id":"NCVD-2025-001-amd-zen-3-zen-4-new-exploitation","cve":null,"aliases":["AMD-SB-7031"],"title":"AMD Zen 3 / Zen 4 - new exploitation method for SRSO (CVE-2023-20569): Google's security team demonstrated a new way to","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD Zen 3 / Zen 4 - new exploitation method for SRSO (CVE-2023-20569)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"Google's security team demonstrated a new way to exploit the existing Inception/SRSO defect on Zen 3 and Zen 4. No new CVE was assigned and AMD asserts the original mitigations still hold. Recorded here because it is the kind of update that never reaches a patch queue: nothing to install, but your risk assessment of an already-known issue should move if the exploitation bar just dropped.","attack_vector":"Local, cross-privilege speculative execution on Zen 3 and Zen 4.","remediation":"**No new action** if you already applied the original SRSO mitigations (microcode plus AGESA, MilanPI 1.0.0.C / GenoaPI 1.0.0.9 or later, plus the kernel's Safe RET). Verify that you did - read /sys/devices/system/cpu/vulnerabilities/spec_rstack_overflow across the fleet rather than assuming. Note the interaction with AMD-SB-7061 above: Safe RET, the mitigation you are relying on here, has its own open weakness.","references":["https://www.amd.com/en/resources/product-security/bulletin/amd-sb-7031.html","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"]},{"id":"NCVD-2025-001-aspeed-ast2600-ast2700-hardware","cve":null,"aliases":["ASPEED secure boot disabled by default","OpenBMC RoT not enabled"],"title":"ASPEED AST2600 / AST2700 hardware root of trust in OpenBMC builds: AST2600 has a fuse-backed secure boot that verifies","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ASPEED AST2600 / AST2700 hardware root of trust in OpenBMC builds","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"AST2600 has a fuse-backed secure boot that verifies the BMC firmware image against an RSA key, and AST2700 adds the OCP Caliptra-based scheme. Both ship off. The default OpenBMC build configuration leaves secure boot disabled because that is convenient for board bring-up, and plenty of production images never turn it on. Where it is off, anything that can write the BMC's SPI flash - a compromised BMC userspace process, a host-side AHB bridge write, or a supply-chain-tampered update package - installs code that persists across power cycles, OS reimage, and node redeployment to the next tenant, with nothing in the boot chain to reject it. On a GPU cluster this is the difference between a compromised node you can rebuild and a compromised node you have to physically retire.","attack_vector":"Any write path to BMC SPI flash: code execution on the BMC, a host-side AHB bridge, or a malicious/unsigned firmware update accepted by the update daemon. Not remotely reachable by itself - it is the amplifier that turns a one-shot BMC compromise into a permanent one.","remediation":"Enabling it is a one-way door: it means burning OTP fuses with your signing key on every node, which cannot be undone and cannot be done remotely. Practically this is an order-time decision with the ODM, not something an operator retrofits on a deployed fleet. For fleets already racked, audit whether secure boot is fused on each SKU, demand the answer in writing from the ODM, and where it is off, compensate with SPI flash content attestation (hash the image out-of-band and compare against a known-good) plus strict control over who can push BMC firmware. Note that a hardware root of trust also does not help if the signed image itself has a signature-verification bug - see the Supermicro image-parser entries.","references":["https://eclypsium.com/wp-content/uploads/OpenBMC-Security-in-Practice.pdf","https://developer.nvidia.com/blog/analyzing-baseboard-management-controllers-to-secure-data-center-infrastructure/","https://www.binarly.io/blog/old-but-gold-the-underestimated-potency-of-decades-old-attacks-on-bmc-security"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"NCVD-2025-001-intel-sgx-ddr4-memory-bus-physic","cve":null,"aliases":["WireTap"],"title":"Intel SGX / DDR4 memory bus (physical interposer): WireTap: a low-cost passive DDR4 interposer reads the memory bus of","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SGX / DDR4 memory bus (physical interposer)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"WireTap: a low-cost passive DDR4 interposer reads the memory bus of an SGX machine and recovers the platform's attestation key, letting the attacker forge quotes that Intel's own attestation service accepts. The consequence for an operator is that SGX remote attestation no longer proves anything about a machine an adversary has had physical access to - which includes colocation, transit, and any hardware that has left your custody. Notably the published work reports the attack against production DCAP attestation, not a lab-only configuration.","attack_vector":"Physical access to the machine long enough to install an interposer between the CPU and a DIMM. Not remote, but well within reach of anyone in the supply chain, a colo neighbour with cage access, or an insider in a datacenter you do not own.","remediation":"Unpatchable in the field on affected parts - deterministic memory encryption without integrity or freshness is a design property of SGX on these generations, not a bug with a patch. Operator response is procedural: treat physical custody as part of the SGX trust boundary, refuse to accept attestation from hardware outside your custody chain, and plan migration to platforms with stronger memory integrity. Watch for Intel TCB recovery advisories, but do not assume one will close this.","references":["https://wiretap.fail/"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"tags":["tenant-isolation"]},{"id":"NCVD-2025-002-amd-sev-es-sev-snp-transient-exe","cve":null,"aliases":["PowerHooK","AMD-SB-3032"],"title":"AMD SEV-ES / SEV-SNP - transient-execution-amplified power side channel: Graz researchers amplified the power side","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV-ES / SEV-SNP - transient-execution-amplified power side channel","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"Graz researchers amplified the power side channel with transient execution to extract AES key bytes from inside SEV-ES and SEV-SNP guests. The significance for an operator is the trend line: power telemetry keeps turning out to be a channel that memory encryption does not cover, and each iteration extracts more with less. If you sell confidential computing, the host's power meter is part of your attack surface.","attack_vector":"Requires a malicious hypervisor with access to power telemetry (RAPL) on a host running confidential guests.","remediation":"**No CVE and no microcode fix** - AMD's answer is to restrict or disable hypervisor RAPL access. That is a configuration change you can make today at no performance cost: ensure the host's energy interfaces are root-only and are not exposed to any process a tenant can influence, and do not pass power telemetry into guests. No reboot, no firmware. Combine with performance-determinism mode if you are also mitigating Collide+Power, accepting the throughput cost that carries.","references":["https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3032.html","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"hot-patch"},"tags":["tenant-isolation"]},{"id":"NCVD-2025-002-intel-sgx-and-amd-sev-snp-dram-i","cve":null,"aliases":["Battering RAM"],"title":"Intel SGX and AMD SEV-SNP / DRAM interposer (memory aliasing): Battering RAM: a cheap DRAM interposer that aliases","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SGX and AMD SEV-SNP / DRAM interposer (memory aliasing)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"Battering RAM: a cheap DRAM interposer that aliases physical addresses so the same ciphertext is served for two different locations, defeating the memory-encryption protections of both Intel SGX and AMD SEV-SNP and giving the attacker plaintext access to enclave and confidential-VM memory. Same operator conclusion as WireTap and it generalises across both vendors' confidential-compute stacks, so it is not a reason to switch silicon.","attack_vector":"Physical access to install an interposer on the DIMM path. Applies to any deployment where hardware is not continuously in your physical custody.","remediation":"No firmware or microcode fix on affected parts - the attack targets a design property of deterministic memory encryption. Treat confidential compute as protecting against a remote or software-privileged adversary, not a physical one, and write that into what you tell customers. Physical security, tamper-evident hardware and custody controls are the actual mitigation.","references":["https://batteringram.eu/"],"status":"curated","tags":["tenant-isolation"]},{"id":"NCVD-2025-003-amd-sev-snp-ciphertext-side-chan","cve":null,"aliases":["Relocate+Vote","Chosen Plaintext Oracle against SEV-SNP","AMD-SB-3021"],"title":"AMD SEV-SNP - ciphertext side channels amplified by hypervisor page movement: Two 2025 follow-ups to CipherLeaks","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV-SNP - ciphertext side channels amplified by hypervisor page movement","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"Two 2025 follow-ups to CipherLeaks - Toronto's Relocate+Vote and ETH Zurich's chosen-plaintext oracle - both combine ciphertext visibility with the hypervisor's ability to move or swap guest pages. Relocating a page changes which address the deterministic encryption is keyed to, giving the attacker the equivalent of a chosen-plaintext oracle against a confidential VM. AMD publishes this as informational with **no CVE**, which is the point worth noting: your vulnerability scanner will never mention it.","attack_vector":"Malicious hypervisor able to observe guest ciphertext and relocate guest pages.","remediation":"**No patch.** The available control is guest policy: SEV-SNP ABI 1.58 and later let a guest forbid hypervisor page move and swap, which removes the amplification. As the operator, support and document that policy bit so tenants can set it; as a tenant-facing claim, be honest that ciphertext visibility on Zen 3 and Zen 4 is architectural. The durable fix is Ciphertext Hiding on Zen 5 (Turin) - a hardware refresh, not a maintenance window.","references":["https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3021.html","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"tags":["tenant-isolation"]},{"id":"NCVD-2025-003-nvidia-gddr6-gpu-memory-a6000-cl","cve":null,"aliases":["GPUHammer"],"title":"NVIDIA GDDR6 GPU memory (A6000-class and similar discrete GPUs): The first demonstrated Rowhammer bit flips in GPU","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GDDR6 GPU memory (A6000-class and similar discrete GPUs)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"The first demonstrated Rowhammer bit flips in GPU memory. Researchers flipped bits in GDDR6 on an NVIDIA A6000 from a co-tenant GPU workload and used a single flip to degrade a victim's neural network accuracy from around 80% to near zero. The attack does not read the victim's data - it corrupts it - which makes it an integrity attack on a co-tenant's model or training run rather than a confidentiality one, and correspondingly hard to detect from inside the victim job.","attack_vector":"A co-tenant workload on the same physical GPU with the ability to allocate and hammer memory. No privileges, no driver bug. GPUs without ECC or with ECC disabled are the exposed population.","remediation":"NVIDIA's published response is to enable System-Level ECC, which is available and on by default on datacenter parts (H100, A100, and the Hopper/Blackwell line) but is off or absent on workstation-class parts. Verify with nvidia-smi -q -d ECC across the fleet and turn it on with nvidia-smi -e 1, which requires a GPU reset - so a node drain. Cost is real: ECC on these parts costs roughly 6-10% inference throughput and around 6% of usable VRAM. HBM3/HBM3e parts with on-die ECC are considered less exposed. There is no firmware patch that removes the underlying DRAM weakness.","references":["https://gpuhammer.com/","https://arxiv.org/abs/2507.08166"],"status":"curated","fleet":{"pain_class":"node-drain"},"tags":["tenant-isolation"]},{"id":"NCVD-2025-004-amd-sev-snp-rmp-entries-cached-i","cve":null,"aliases":["AMD-SB-3036"],"title":"AMD SEV-SNP - RMP entries cached in L1D/L2 leaking physical address bits: Reverse-map table entries cached in L1D and","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV-SNP - RMP entries cached in L1D/L2 leaking physical address bits","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"Reverse-map table entries cached in L1D and L2 leak up to six physical address bits to an **unprivileged** process. AMD and the researchers agree there is no immediate impact, and six bits is not a compromise on its own - but physical address bits are exactly the primitive that makes Rowhammer, cache-eviction-set construction and DMA targeting practical, so it is a building block for other people's attacks rather than an attack itself.","attack_vector":"Local, unprivileged - notably lower than the rest of the RMP family, which mostly needs hypervisor privilege.","remediation":"**No fix planned.** Nothing to install and nothing to reboot for. Treat it as a standing reminder that side-channel primitives accumulate: it lowers the cost of the next attack against your SNP hosts without ever appearing in a patch queue. Track AMD-SB-3036 in case AMD's assessment changes.","references":["https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3036.html","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"tags":["tenant-isolation"]},{"id":"NCVD-2025-005-amd-sev-snp-dimm-interposer-vari","cve":null,"aliases":["AMD-SB-3024"],"title":"AMD SEV-SNP - DIMM interposer variant of BadRAM (KU Leuven): A memory-bus interposer variant of the BadRAM aliasing","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV-SNP - DIMM interposer variant of BadRAM (KU Leuven)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"A memory-bus interposer variant of the BadRAM aliasing attack against SEV-SNP. AMD's response is **WONTFIX** - the attack is declared outside the SEV-SNP threat model because it requires physical interposition on the memory bus.","attack_vector":"Physical access with a memory-bus interposer.","remediation":"**No patch, and none coming** - AMD has scoped it out of the threat model. The operator control is physical: tamper-evident chassis handling, chain of custody, and not making confidential-computing claims that a tenant could reasonably read as covering an adversary with physical access to the DIMM slot. If a customer's threat model includes your own datacenter staff, this is a conversation to have explicitly rather than one to leave to the marketing page.","references":["https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3024.html","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"tags":["tenant-isolation"]},{"id":"NCVD-2025-006-amd-confidential-computing-ddr5","cve":null,"aliases":["AMD-SB-3040"],"title":"AMD confidential computing - DDR5 memory bus interposition against TEEs: Compromising trusted execution environments by","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD confidential computing - DDR5 memory bus interposition against TEEs","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"Compromising trusted execution environments by interposing on the DDR5 memory bus. Same family as the BadRAM interposer work and the same conclusion for an operator: SEV-SNP's guarantees are software- and firmware-scoped, and an adversary with hands on the hardware sits outside them.","attack_vector":"Physical access with DDR5 bus interposition hardware.","remediation":"**No patch** - physical attacks fall outside the declared SEV-SNP threat model. Mitigation is datacenter physical security, hardware chain of custody, and honest scoping of what confidential computing does and does not promise your tenants.","references":["https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3040.html","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"tags":["tenant-isolation"]},{"id":"NCVD-2025-007-amd-secure-processor-boot-rom-ph","cve":null,"aliases":["AMD-SB-7044"],"title":"AMD Secure Processor boot ROM - physical attacks bypassing secure boot: Physical attacks that bypass secure boot in the","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor boot ROM - physical attacks bypassing secure boot","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"Physical attacks that bypass secure boot in the ASP boot ROM. Boot ROM is mask-programmed silicon, so the flawed logic itself cannot be patched - only worked around by the firmware layered above it. An attacker who defeats ASP secure boot owns the platform's root of trust from power-on.","attack_vector":"Physical access to the platform.","remediation":"**Boot ROM is unpatchable by construction** - any mitigation is compensating logic in AGESA/PI firmware above it, delivered as an OEM BIOS package. The real controls are physical: chain of custody for hardware, tamper evidence, and treating any node returned from third-party hands as untrusted until its firmware is measured and reprovisioned. This is the argument for firmware attestation between bare-metal tenants rather than trusting a reimage.","references":["https://www.amd.com/en/resources/product-security/bulletin/amd-sb-7044.html","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"tags":["tenant-isolation"]},{"id":"NCVD-2025-008-nvidia-gpus-with-gddr6-and-no-on","cve":null,"aliases":["GPUHammer"],"title":"NVIDIA GPUs with GDDR6 and no on-die ECC - demonstrated on RTX A6000","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPUs with GDDR6 and no on-die ECC - demonstrated on RTX A6000. Not reproduced on A100 (HBM), H100 or RTX…","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"The first working Rowhammer against GPU memory, driven from ordinary user-level CUDA. It flips bits in a neighbouring tenant's GPU memory. The headline result is the one that should worry an AI datacenter: a single bit flip in the exponent of a floating-point weight took an ImageNet model from 80% to 0.1% accuracy. There is no crash, no XID, no ECC event on an ECC-off GPU - the victim's model simply stops working and they have no way to attribute it to you. In a fleet that time-slices or MIG-partitions GPUs between customers, this is a cross-tenant integrity attack against the actual product you sell, and it breaks tenant handoff: residual flips persist in the physical DRAM after the previous tenant leaves.","attack_vector":"A tenant running unprivileged CUDA on a GPU whose memory is shared with, or was previously allocated to, the victim - MIG partitions, time-sliced sharing, MPS, or simply the next tenant on a re-let card. No driver exploit and no privilege escalation required.","remediation":"Enable ECC on every GPU that does not have on-die ECC: `nvidia-smi -e 1`, then reset or reboot the GPU. It is not free - roughly a 10% inference slowdown on an A6000 and about 6.25% of memory capacity gone, which on a rental fleet is directly billable capacity you stop selling. Cards with on-die ECC (H100, GB200-class HBM3e, RTX 5090) are not affected in the demonstrated form. Practical fleet policy: refuse to share a physical GPU between untrusted tenants at all, scrub and reset GPU memory between tenants, and audit that ECC has not been disabled by a tenant with elevated access. Verify ECC state per GPU rather than assuming the fleet default held.","references":["https://gpuhammer.com/","https://arxiv.org/abs/2507.08166"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"NCVD-2025-011-nvidia-gpus-with-gddr6-and-no-on","cve":null,"aliases":["GPUHammer"],"title":"NVIDIA GPUs with GDDR6 and no on-die ECC - demonstrated on RTX A6000","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPUs with GDDR6 and no on-die ECC - demonstrated on RTX A6000. Not reproduced on A100 (HBM), H100 or RTX…","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"The first working Rowhammer against GPU memory, driven from ordinary user-level CUDA. It flips bits in a neighbouring tenant's GPU memory. The headline result is the one that should worry an AI datacenter: a single bit flip in the exponent of a floating-point weight took an ImageNet model from 80% to 0.1% accuracy. There is no crash, no XID, no ECC event on an ECC-off GPU - the victim's model simply stops working and they have no way to attribute it to you. In a fleet that time-slices or MIG-partitions GPUs between customers, this is a cross-tenant integrity attack against the actual product you sell, and it breaks tenant handoff: residual flips persist in the physical DRAM after the previous tenant leaves.","attack_vector":"A tenant running unprivileged CUDA on a GPU whose memory is shared with, or was previously allocated to, the victim - MIG partitions, time-sliced sharing, MPS, or simply the next tenant on a re-let card. No driver exploit and no privilege escalation required.","remediation":"Enable ECC on every GPU that does not have on-die ECC: `nvidia-smi -e 1`, then reset or reboot the GPU. It is not free - roughly a 10% inference slowdown on an A6000 and about 6.25% of memory capacity gone, which on a rental fleet is directly billable capacity you stop selling. Cards with on-die ECC (H100, GB200-class HBM3e, RTX 5090) are not affected in the demonstrated form. Practical fleet policy: refuse to share a physical GPU between untrusted tenants at all, scrub and reset GPU memory between tenants, and audit that ECC has not been disabled by a tenant with elevated access. Verify ECC state per GPU rather than assuming the fleet default held.","references":["https://gpuhammer.com/","https://arxiv.org/abs/2507.08166"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"NCVD-2025-012-intel-sgx-ddr4-memory-bus-physic","cve":null,"aliases":["WireTap"],"title":"Intel SGX / DDR4 memory bus (physical interposer): WireTap: a low-cost passive DDR4 interposer reads the memory bus of","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SGX / DDR4 memory bus (physical interposer)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"WireTap: a low-cost passive DDR4 interposer reads the memory bus of an SGX machine and recovers the platform's attestation key, letting the attacker forge quotes that Intel's own attestation service accepts. The consequence for an operator is that SGX remote attestation no longer proves anything about a machine an adversary has had physical access to - which includes colocation, transit, and any hardware that has left your custody. Notably the published work reports the attack against production DCAP attestation, not a lab-only configuration.","attack_vector":"Physical access to the machine long enough to install an interposer between the CPU and a DIMM. Not remote, but well within reach of anyone in the supply chain, a colo neighbour with cage access, or an insider in a datacenter you do not own.","remediation":"Unpatchable in the field on affected parts - deterministic memory encryption without integrity or freshness is a design property of SGX on these generations, not a bug with a patch. Operator response is procedural: treat physical custody as part of the SGX trust boundary, refuse to accept attestation from hardware outside your custody chain, and plan migration to platforms with stronger memory integrity. Watch for Intel TCB recovery advisories, but do not assume one will close this.","references":["https://wiretap.fail/"],"status":"curated","tags":["tenant-isolation"]},{"id":"NCVD-2025-013-intel-sgx-and-amd-sev-snp-dram-i","cve":null,"aliases":["Battering RAM"],"title":"Intel SGX and AMD SEV-SNP / DRAM interposer (memory aliasing): Battering RAM: a cheap DRAM interposer that aliases","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SGX and AMD SEV-SNP / DRAM interposer (memory aliasing)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"Battering RAM: a cheap DRAM interposer that aliases physical addresses so the same ciphertext is served for two different locations, defeating the memory-encryption protections of both Intel SGX and AMD SEV-SNP and giving the attacker plaintext access to enclave and confidential-VM memory. Same operator conclusion as WireTap and it generalises across both vendors' confidential-compute stacks, so it is not a reason to switch silicon.","attack_vector":"Physical access to install an interposer on the DIMM path. Applies to any deployment where hardware is not continuously in your physical custody.","remediation":"No firmware or microcode fix on affected parts - the attack targets a design property of deterministic memory encryption. Treat confidential compute as protecting against a remote or software-privileged adversary, not a physical one, and write that into what you tell customers. Physical security, tamper-evident hardware and custody controls are the actual mitigation.","references":["https://batteringram.eu/"],"status":"curated","tags":["tenant-isolation"]},{"id":"NCVD-2025-014-nvidia-gddr6-gpu-memory-a6000-cl","cve":null,"aliases":["GPUHammer"],"title":"NVIDIA GDDR6 GPU memory (A6000-class and similar discrete GPUs): The first demonstrated Rowhammer bit flips in GPU","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GDDR6 GPU memory (A6000-class and similar discrete GPUs)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"The first demonstrated Rowhammer bit flips in GPU memory. Researchers flipped bits in GDDR6 on an NVIDIA A6000 from a co-tenant GPU workload and used a single flip to degrade a victim's neural network accuracy from around 80% to near zero. The attack does not read the victim's data - it corrupts it - which makes it an integrity attack on a co-tenant's model or training run rather than a confidentiality one, and correspondingly hard to detect from inside the victim job.","attack_vector":"A co-tenant workload on the same physical GPU with the ability to allocate and hammer memory. No privileges, no driver bug. GPUs without ECC or with ECC disabled are the exposed population.","remediation":"NVIDIA's published response is to enable System-Level ECC, which is available and on by default on datacenter parts (H100, A100, and the Hopper/Blackwell line) but is off or absent on workstation-class parts. Verify with nvidia-smi -q -d ECC across the fleet and turn it on with nvidia-smi -e 1, which requires a GPU reset - so a node drain. Cost is real: ECC on these parts costs roughly 6-10% inference throughput and around 6% of usable VRAM. HBM3/HBM3e parts with on-die ECC are considered less exposed. There is no firmware patch that removes the underlying DRAM weakness.","references":["https://gpuhammer.com/","https://arxiv.org/abs/2507.08166"],"status":"curated","tags":["tenant-isolation"]},{"id":"NCVD-2026-001-dmtf-libspdm-get-measurement-ext","cve":null,"aliases":["DMTF-2026-0002"],"title":"DMTF libspdm (GET_MEASUREMENT_EXTENSION_LOG offset/length wrap): Wrapping addition of the Offset and Length fields","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"DMTF libspdm (GET_MEASUREMENT_EXTENSION_LOG offset/length wrap)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Wrapping addition of the Offset and Length fields in GET_MEASUREMENT_EXTENSION_LOG lets a requester read memory outside the measurement log from a libspdm responder. The attacker is the host or fabric peer asking a device for its measurements, and what it gets back is device firmware memory - which is where attestation keys and secrets live. No CVE was assigned, so this will not appear in NVD, OSV or any scanner your fleet runs.","attack_vector":"Any SPDM requester that can reach an affected responder with MEL_CAP and CHUNK_CAP set - in practice the host talking to its own accelerators or NICs, so host-side root on a bare-metal node is enough.","remediation":"libspdm update embedded in device or platform firmware, plus - importantly - integrator hygiene: the bug only bites when the integrator's libspdm_copy_mem() has its assertions compiled out, which is a build-configuration decision your device vendor made and you cannot see. There is no config mitigation and no way to detect affected devices from the outside. Ask vendors directly for their libspdm version and build flags as part of hardware acceptance; that question is the only real control here.","references":["https://github.com/DMTF/libspdm/security/advisories/GHSA-m4wc-xmvg-369f"],"status":"curated"},{"id":"NCVD-2026-001-linux-safe-ret-srso-mitigation-o","cve":null,"aliases":["Safe RET Interrupt Vulnerability","AMD-SB-7061"],"title":"Linux Safe RET SRSO mitigation on AMD Zen 1-Zen 4 - interrupt-induced weakening: An attacker executing code on the","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux Safe RET SRSO mitigation on AMD Zen 1-Zen 4 - interrupt-induced weakening","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"An attacker executing code on the machine injects an interrupt at a precise moment to disrupt **Safe RET**, which is the default SRSO/Inception mitigation on Linux. Disrupting it weakens the mitigation and can leak information across privilege boundaries. The reason to care operationally is that your tooling will report SRSO as mitigated while the mitigation is being defeated - demonstrated on Zen 1 and Zen 2, suspected on Zen 3 and Zen 4.","attack_vector":"Local, requires code execution on the system and precise interrupt timing - so reachable from a tenant workload, not from the network.","remediation":"**No fix published as of August 2026** - this is the live, open gap in the SRSO story. AMD's assessment attributes it to the Linux Safe RET implementation rather than to silicon, so the eventual fix is expected to be a kernel change rather than microcode or BIOS. Until then, treat SRSO as partially mitigated on Zen 1 through Zen 4 rather than closed: for workloads where cross-privilege speculative leakage is genuinely in your threat model, dedicated nodes rather than shared ones is the only control that holds. Track AMD-SB-7061 for the fix.","references":["https://www.amd.com/en/resources/product-security/bulletin/amd-sb-7061.html","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"tags":["tenant-isolation"]},{"id":"NCVD-2026-002-amd-sev-firmware-arbitrary-code","cve":null,"aliases":["AMD-SB-3033"],"title":"AMD SEV firmware - arbitrary code execution on the AMD Security Processor (physical): An academic disclosure achieving","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV firmware - arbitrary code execution on the AMD Security Processor (physical)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"An academic disclosure achieving arbitrary code execution on the AMD Security Processor itself. Code execution in the ASP means control of SEV key management and attestation for every confidential guest on the node. AMD scopes it out as requiring physical access, but shipped defence-in-depth firmware anyway - which is a reasonable signal about how seriously to take it.","attack_vector":"Physical access to the platform.","remediation":"AMD shipped **defence-in-depth** PI updates - MilanPI 1.0.0.J and GenoaPI 1.0.0.H (both December 2025) - so there is something to deploy despite the WONTFIX-adjacent framing. Delivered as an OEM SBIOS package with the usual lag and a power cycle. The mitigation state is **tenant-verifiable via Platform Info Bit 5** in the attestation report, which is worth advertising to confidential-computing customers: they can check you applied it rather than taking your word.","references":["https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3033.html","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"},"tags":["tenant-isolation"]},{"id":"NCVD-2026-002-dmtf-libspdm-cryptlib-mbedtls-cs","cve":null,"aliases":["DMTF-2026-0001"],"title":"DMTF libspdm (cryptlib_mbedtls CSR generation, stack overflow): An over-long Common Name in a GET_CSR request writes","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"DMTF libspdm (cryptlib_mbedtls CSR generation, stack overflow)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"An over-long Common Name in a GET_CSR request writes past a stack array inside the responder, corrupting stack data and opening the door to code execution in the device's root-of-trust firmware. The responder here is the thing that is supposed to prove the device is trustworthy, so a compromise at this point does not just break one device - it makes that device's attestation statements attacker-authored. No CVE assigned, so scanners will not flag it.","attack_vector":"Any SPDM requester able to send GET_CSR to a responder that supports CSR_CAP and builds on the mbedTLS crypt backend. On a server that is the host, so local root on a bare-metal node reaches it.","remediation":"libspdm update, delivered as device or platform firmware - per-node flash with vendor rebase lag, no package path, no config toggle. Where a device exposes CSR generation you do not use, ask the vendor whether CSR_CAP can be disabled in their build; turning off an unused capability is the only lever an operator has short of the firmware update.","references":["https://github.com/DMTF/libspdm/security/advisories/GHSA-j54w-759w-xj3m"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"NCVD-2026-003-amd-rep-string-execution-unit-sc","cve":null,"aliases":["AMD-SB-7069"],"title":"AMD - REP-string execution unit scheduler contention side channel: A newer variant of the SQUIP scheduler-contention","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"AMD - REP-string execution unit scheduler contention side channel","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A newer variant of the SQUIP scheduler-contention channel, this time reached through REP-string instructions. Same shape as SQUIP - contention on execution-unit scheduler queues observable from a co-resident thread - and the same operator consequence: cross-tenant leakage that no software isolation boundary sees.","attack_vector":"Local, requires co-residency with the victim on the same physical core.","remediation":"**No fix.** AMD points at existing best practices - constant-time, secret-independent code - which you cannot impose on a tenant's workload. The controls that actually work are yours: do not co-schedule distinct trust domains on sibling SMT threads, or disable SMT on mixed-tenancy nodes. Both are scheduling decisions; disabling SMT needs a reboot, whole-core allocation policy needs a kubelet restart and a drain. On a GPU fleet where the CPU is rarely the bottleneck, the throughput cost is smaller than it looks.","references":["https://www.amd.com/en/resources/product-security/bulletin/amd-sb-7069.html","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"},"tags":["tenant-isolation"]},{"id":"NCVD-2026-005-das-u-boot-fit-image-signature-v","cve":null,"aliases":["BRLY-2026-038","BRLY-2026-041"],"title":"Das U-Boot (FIT image signature verification): Binarly disclosed a cluster of flaws in U-Boot's FIT image handling","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Das U-Boot (FIT image signature verification)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Binarly disclosed a cluster of flaws in U-Boot's FIT image handling - stack buffer underflow and NULL dereference during signature verification, unbounded recursion in FIT validation, and an SPL out-of-bounds write while loading a FIT image. U-Boot is the first-stage bootloader on essentially every ASPEED-based BMC, so these sit below the BMC firmware itself: an attacker who lands here owns the root of trust for the management controller and nothing running on the host can see it.","attack_vector":"Requires the ability to present a crafted FIT image to the bootloader - in BMC terms, an attacker who already achieved a firmware write, or a malicious/compromised firmware update image. It is the persistence layer of a BMC compromise rather than the initial entry.","remediation":"No CVE IDs assigned as of disclosure. Fix arrives as a U-Boot update embedded in a full BMC firmware image from your board vendor, which means an out-of-band per-node BMC flash and an ODM rebase lag measured in months. There is no config mitigation - the compensating control is signed-update enforcement plus strict isolation of the management network.","references":["https://www.binarly.io/advisories"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"NCVD-2026-006-uefi-secure-boot-microsoft-2011","cve":null,"aliases":["Secure Boot 2026 certificate expiry"],"title":"UEFI Secure Boot (Microsoft 2011 CA/KEK expiry): Not an exploitable flaw but a fleet-wide trust-anchor deadline","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"UEFI Secure Boot (Microsoft 2011 CA/KEK expiry)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Not an exploitable flaw but a fleet-wide trust-anchor deadline. The 2011-era Microsoft UEFI CA and KEK certificates baked into most server firmware reach end of validity in 2026. Nodes that never receive updated certificates stop receiving valid Secure Boot revocation updates and, depending on firmware behaviour, may refuse to boot newly signed bootloaders - meaning Secure Boot either silently stops being enforceable or turns into an outage. For a GPU operator this hits the exact control the bare-metal tenant-handoff story depends on.","attack_vector":"No attacker required. The risk is that stale certificates leave you unable to revoke a future vulnerable bootloader, so every bypass in this cluster becomes permanent on affected nodes.","remediation":"Inventory Secure Boot certificate validity per node now, before the deadline. Updated certificates arrive by OEM firmware/BIOS update or OS vendor channel and require a reboot; older boards may never get them, in which case plan hardware refresh or accept that those nodes have no working revocation path. Treat certificate expiry as a scheduled fleet program, not an incident.","references":["https://techcommunity.microsoft.com/blog/windows-itpro-blog/act-now-secure-boot-certificates-expire-in-june-2026/4426856"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"NCVD-2026-013-community-open-source-sonic-soni","cve":null,"aliases":["community SONiC","sonic-net","open-source network OS blind spot"],"title":"Community / open-source SONiC (sonic-net): Community SONiC — the open-source NOS that a growing share of cost-optimised","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Community / open-source SONiC (sonic-net)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Community SONiC — the open-source NOS that a growing share of cost-optimised GPU-cluster fabrics run on — has essentially no published CVE history. A CPE-based NVD query returns zero records, and the project's own process routes reports privately to its security committee. This is not evidence that SONiC is secure; it is evidence that you have no vulnerability feed for the operating system running your leaf/spine. Every downstream commercial distribution that has been examined (Dell Enterprise SONiC) has produced multiple 9.x-severity findings including authentication bypass, command injection, hard-coded and default credentials — which is what you would expect the shared upstream to look like too. Operators running community SONiC are patching blind.","attack_vector":"Not a specific vulnerability. The exposure is process-level: an operator cannot subscribe to a feed that tells them when their switch OS needs patching, so known-vulnerable images stay in production indefinitely.","remediation":"No patch to apply. Practical controls: pin to a distribution that issues advisories (Dell, Edgecore or a vendor-supported build) rather than self-built community images; track the sonic-buildimage git history for security-relevant commits since there is no advisory feed; scan the SONiC container images for known-vulnerable component versions, because most of the real risk is the Debian base and the bundled daemons (lldpd, FRR, redis, the SDK) rather than SONiC-specific code; and keep switch management interfaces on an isolated OOB network on the assumption that you will not learn about the next flaw in time.","references":["https://github.com/sonic-net/SONiC/security","https://www.dell.com/support/kbdoc/en-us/000245655/dsa-2024-449-security-update-for-dell-enterprise-sonic-distribution-vulnerabilities"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2026-014-weka-data-platform-and-vast-data","cve":null,"aliases":["WEKA","VAST Data","AI-storage advisory blind spot"],"title":"WEKA Data Platform and VAST Data (published-advisory coverage): Neither WEKA nor VAST Data","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"WEKA Data Platform and VAST Data (published-advisory coverage)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Neither WEKA nor VAST Data — two of the most commonly deployed storage platforms under large AI training clusters — has a meaningful public CVE record. NVD keyword searches return nothing attributable to either product. As with community SONiC, this is an absence of disclosure rather than an absence of vulnerabilities: both are complex distributed systems with kernel-level clients, RDMA data paths and multi-tenant namespace separation, and comparable platforms (Lustre, Spectrum Scale, Ceph, BeeGFS) all have significant published histories including cross-tenant access-control failures. An operator running WEKA or VAST has no external feed telling them when to patch the layer that holds every tenant's training data.","attack_vector":"Not a specific vulnerability. The exposure is that these platforms' security posture is visible only to the vendor, so an operator's patch decisions depend entirely on vendor-issued release notes and on asking directly.","remediation":"No patch. Make it contractual and operational: require the vendor to notify you of security-relevant fixes in release notes and to state a disclosure policy; ask for their most recent third-party penetration-test summary at renewal; keep the storage cluster's management plane on an isolated network; and verify multi-tenant namespace separation yourself with an actual cross-tenant read test rather than trusting the product claim. Track the kernel-client packages these platforms install, since those inherit Linux CVEs on your normal patch cycle even when the platform itself publishes nothing.","references":["https://nvd.nist.gov/vuln/search/results?query=weka","https://nvd.nist.gov/vuln/search/results?query=vast+data"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2026-018-rack-pdu-and-ups-management-esta","cve":null,"aliases":["shared PDU/UPS credentials","commissioning credential reuse"],"title":"Rack PDU and UPS management estates as a class (all vendors): PHYSICAL, and the most common real-world finding","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Rack PDU and UPS management estates as a class (all vendors)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"PHYSICAL, and the most common real-world finding in this whole layer. PDU and UPS management interfaces are almost universally commissioned with a single shared credential set, applied by whoever racked the gear, never rotated, and documented in a spreadsheet or a runbook. There is no per-device identity, no rotation, and no logging that would show misuse. Anyone who obtains that one credential - a departing contractor, a leaked runbook, one of the many device-side credential-disclosure CVEs listed above - can switch outlets across the entire hall. No CVE will ever be assigned to this, and it is more likely to be used against you than any of the memory-corruption bugs in this file.","attack_vector":"Anyone who reaches the PDU/UPS management network with the shared credential. On many builds that network is the same one the BMCs sit on, which is the same one a tenant with host root can sometimes see.","remediation":"Not a patch. Per-device credentials issued from a secret manager, PDU management interfaces on a VLAN unreachable from any tenant-facing network or from the BMC network, outlet switching disabled on PDUs that do not need it, and authentication logs from power devices shipped somewhere you actually read. Doing this across an existing hall is a few engineer-weeks and touches every rack, which is why it does not get done.","references":["https://www.cisa.gov/news-events/cybersecurity-advisories/aa24-060a"],"status":"curated"},{"id":"NCVD-2026-019-leased-colocation-facility-infra","cve":null,"aliases":["facility VLAN ownership gap","colo landlord equipment"],"title":"Leased colocation facility infrastructure (power, cooling, access control) as a class: Most GPU operators lease space","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Leased colocation facility infrastructure (power, cooling, access control) as a class","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Most GPU operators lease space rather than own buildings, which means the UPS, switchgear, PDU upstream feeds, CRAH units and access control that determine whether their GPUs stay powered and cooled are the landlord's equipment, on the landlord's network, patched on the landlord's schedule - which is frequently never, because facility gear is treated as a fixed asset rather than as software. Every CVE in this file is therefore, for many operators, a vulnerability they carry the consequences of and have no authority to fix. An availability event here takes out a hall regardless of how well the compute plane is run.","attack_vector":"Whoever can reach the landlord's facility network - which typically includes the landlord's own vendors, remote-monitoring contractors, and any building-management remote access path the operator has never seen.","remediation":"Contractual, not technical. Require in the colocation agreement: a current inventory of facility control equipment with firmware versions, evidence of a patch cadence, segmentation of the facility network from any operator-reachable network, notification of remote-access paths granted to third parties, and the right to audit. Then verify rather than trust. Where a landlord will not commit, price the risk into the site decision - this belongs in site selection, not in the security backlog.","references":["https://www.cisa.gov/topics/critical-infrastructure-security-and-resilience"],"status":"curated","tags":["physical-impact"]},{"id":"NCVD-2026-023-nvme-admin-command-set-firmware","cve":null,"aliases":["NVMe Firmware Image Download","NVMe Firmware Commit","firmware downgrade attack","unsigned drive firmware"],"title":"NVMe admin command set - Firmware Image Download (opcode 11h) and Firmware Commit (opcode 10h) reachable from the host","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVMe admin command set - Firmware Image Download (opcode 11h) and Firmware Commit (opcode 10h) reachable from the…","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"CLASS ENTRY, and the single highest-leverage tenant-handoff risk in this whole category. The NVMe specification puts firmware download and firmware commit in the standard admin command set, meaning a host with root can push a firmware image to the drive using nothing more exotic than nvme-cli. Whether that is safe depends entirely on the drive enforcing signature verification and anti-rollback - and that enforcement is uneven across vendors and generations, is not something the host can independently verify, and has been shown breakable even where present (the Kioxia CM6/PM6/PM7 JTAG work bypasses the RSA signature check outright). Where verification is weak or absent, a departing tenant flashes a modified image and owns the drive controller permanently: it outlives your wipe, your reimage and your reprovision, sees everything the next tenant writes, and lies to every host-side tool about its own version, its lock state and whether a sanitize succeeded. Even where signature checking holds, downgrade to an older SIGNED image with known vulnerabilities is a live path unless the drive enforces anti-rollback - and the RPMB replay flaw (CVE-2020-13799) undermines exactly that anti-rollback state across eMMC, UFS and all NVMe versions.","attack_vector":"A tenant with root on the bare-metal host, issuing standard NVMe admin commands to a locally attached drive. No physical access, no exotic tooling, no exploit needed where the drive does not enforce signing - the command path is a documented, supported feature. Also relevant in SR-IOV and DPU/computational-storage designs, where whether the admin queue and Security Send/Receive are properly filtered from a tenant-controlled function is a per-platform question most operators have never actually tested.","remediation":"Policy and platform configuration; there is no patch because the command path is by design. (1) Block it at the platform: filter Firmware Image Download and Firmware Commit - and vendor-specific and Security Send/Receive pass-through - so tenant-controlled hosts and virtual functions cannot reach them. Do not expose raw NVMe admin queues to tenants unless the product genuinely requires it. (2) TEST it rather than assuming: on a representative node, try to flash a drive from a tenant-equivalent shell and confirm you are refused. Most operators have never run this check and will be surprised by the result on at least one SKU. (3) Make signed firmware and enforced anti-rollback a written procurement requirement, and get the vendor to state it per SKU. (4) Inventory expected firmware version per drive serial in a store the host cannot write to, and alert on any change or any version that decreases. (5) For sensitive tenancies, retire local media at end of tenancy instead of recycling it. Assume that verifying firmware integrity across a 10,000-drive fleet is not achievable with host-side tooling - the controller is the thing answering your questions - so the control has to be preventing the write and controlling the media's lifecycle, not detecting the implant afterwards.","references":["https://www.nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0c-2022.10.04-Ratified.pdf","https://github.com/google/security-research/security/advisories/GHSA-3hh8-94j4-62rh","https://www.kb.cert.org/vuls/id/231329","https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-88r1.pdf"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2026-024-bacnet-bacnet-ip-as-a-protocol-f","cve":null,"aliases":[],"title":"BACnet / BACnet IP as a protocol (facility control plane): BACnet has no authentication, no integrity protection and no","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"BACnet / BACnet IP as a protocol (facility control plane)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"BACnet has no authentication, no integrity protection and no encryption at the network layer. Any device that can put a BACnet frame on the wire can issue a WriteProperty to any object on any controller that will accept it - fan speed, damper position, chilled-water setpoint, occupancy schedule, alarm enable. There is no credential to steal because none exists, and there is no log entry that distinguishes a legitimate command from a forged one. This is the single highest-leverage weakness in the whole facility stack for an AI datacenter: an attacker who reaches the BACnet segment does not need to exploit anything, they simply operate the building. Against a hall of 40-140 kW GPU racks, writing setpoints or zeroing fan commands crosses accelerator thermal-shutdown thresholds in minutes, taking down every in-flight training job and stressing hardware through repeated thermal cycles. It also breaks tenant handoff in a shared building: BACnet gives no way to scope one tenant's control authority away from another's equipment, so a compromised neighbour on the same building segment can command your cooling.","attack_vector":"Any host on the BACnet/IP segment, unauthenticated, using off-the-shelf tooling (YABE, Wireshark's BACnet dissector, the open-source BACnet stack utilities). Reachability is everything: the segment typically includes mechanical rooms, IDF closets, the fire and lighting integrators' gear, the landlord's building network, and every controls contractor's laptop that has ever been plugged in. BACnet/IP uses UDP 47808 and relies on broadcast, so it also crosses VLANs wherever a BBMD (BACnet Broadcast Management Device) has been configured to bridge them - operators routinely do not know where their BBMDs are. BACnet MS/TP behind a BACnet router is reachable from IP through that router.","remediation":"Unpatchable by design; the protocol will never authenticate. Three real options, in order of what most operators can actually do. First, segmentation: BACnet on a dedicated VLAN with no route to tenant, corporate or internet networks, an inventory of every BBMD, and switch-level port security or 802.1X on ports serving mechanical spaces. Second, physical security: locked mechanical rooms and control panels, because RS-485 field bus access is a wirecutter away. Third, the actual protocol fix - BACnet Secure Connect (BACnet/SC), which adds TLS and certificate-based device identity; it is supported by newer controller generations and is a controller-replacement project, so treat it as a capital line item for any new build and a multi-year migration for an existing one. For a leased colo the honest answer is that you cannot fix this yourself: it is the landlord's control network. Put it in the contract - require BACnet segment isolation, require disclosure of BBMD placement, require that no tenant network can route to it, and require evidence rather than assurance.","references":["https://www.cisa.gov/news-events/ics-advisories/icsa-26-078-08","https://nvd.nist.gov/vuln/detail/CVE-2026-32666","https://www.cisa.gov/resources-tools/resources/secure-design-alert-security-design-improvements-scada-devices"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2026-024-nvme-of-fabric-authentication-as","cve":null,"aliases":[],"title":"NVMe-oF fabric authentication as deployed - host NQN allowlisting on Linux nvmet, SPDK and most storage appliances","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"NVMe-oF fabric authentication as deployed - host NQN allowlisting on Linux nvmet, SPDK and most storage appliances","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"There is no CVE for this because it is the specification working as designed, and it is the single largest tenant-isolation gap in NVMe-oF as it is actually deployed. Outside of in-band DH-HMAC-CHAP (NVMe TP-8006, Linux 6.0+) and NVMe/TCP TLS (Linux 6.7+), the only thing binding a namespace to a tenant is the host NQN the initiator asserts in its Connect command. The NQN is a self-declared string. Any host that can reach the target's transport port can claim another tenant's NQN and be granted that tenant's namespaces - full read and write of another customer's dataset or checkpoint volume, with no exploit, no memory corruption and nothing anomalous in the target logs. Most neocloud NVMe-oF deployments run exactly this way because auth and TLS are off by default and cost throughput.","attack_vector":"Any host with IP or RDMA reachability to the target's transport port (TCP 4420 or the RDMA CM port) - a tenant bare-metal node, a compromised BMC bridged onto the storage VLAN, or anything that lands on the storage network. The attacker needs to learn or guess the victim's host NQN, which is usually derived from a predictable pattern or readable from orchestration metadata.","remediation":"Not patchable - this is a configuration and architecture decision. Enable DH-HMAC-CHAP in-band authentication on every subsystem with per-host keys, provisioned by the same system that provisions the namespace, and enable NVMe/TCP TLS where the kernel and target support it. Both cost CPU and some latency, which is why they are skipped; measure it rather than assuming. Independently: put the storage fabric on its own VLAN/VRF with per-initiator ACLs so an unknown host cannot open a connection at all, and treat host NQNs as secrets rather than as inventory labels. Note that turning on DH-HMAC-CHAP exposes you to the nvmet-auth parsing bugs elsewhere in this database, so patch the target kernel in the same change.","references":["https://nvmexpress.org/specifications/","https://docs.kernel.org/admin-guide/nvme-multipath.html"],"status":"curated"},{"id":"NCVD-2026-025-landlord-owned-facility-control","cve":null,"aliases":[],"title":"Landlord-owned facility control network in a leased colo or wholesale hall (governance gap): Almost every neocloud","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Landlord-owned facility control network in a leased colo or wholesale hall (governance gap)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Almost every neocloud and bare-metal GPU provider runs in space it does not own, which means the entire cooling and BMS layer - the thing that can kill a whole hall of accelerators in minutes - is operated by a third party the tenant has no technical visibility into and no ability to patch. The tenant carries the full financial consequence of a thermal event (lost training runs, missed SLAs, damaged accelerators, customer churn) while holding none of the controls. In practice the tenant does not know which BMS vendor is deployed, what firmware the controllers run, whether the BACnet segment is isolated from the building's office network, whether the mechanical contractor has a permanent remote-access tunnel in, or who else has a badge to the mechanical rooms. In multi-tenant halls it is worse: cooling is shared infrastructure, so a compromise driven through a neighbouring tenant's contractor lands on your racks. This also directly breaks tenant handoff - when a hall is re-let, nothing in the facility control layer is reset, re-flashed or credential-rotated between occupants.","attack_vector":"Not a single vector - a structural one. The realistic entry points are the mechanical contractor's remote-support path, the landlord's corporate network being flat with the building VLAN, an unmonitored BBMD bridging BACnet across zones, a shared BMS supervisor serving all tenants, and physical access to mechanical rooms and control panels by staff and contractors who are outside the tenant's security program entirely.","remediation":"You cannot patch this - it is the landlord's equipment. The honest remediation is contractual and it needs to be in the lease or the master services agreement, not a security questionnaire answered once. Ask for and get in writing: the BMS/BAS vendor and product versions serving your halls; a network diagram showing the facility VLAN and every route off it; confirmation that no tenant or corporate network can route to the BACnet/Modbus segments; the list of remote-access paths into building controls and who holds them; a commitment to notify you within a defined window when CISA publishes an ICS advisory affecting deployed equipment, with a patch SLA; the right to have a third party validate the segmentation; and physical access logs for mechanical rooms. Also negotiate what you actually need operationally - independent temperature telemetry you own (your own sensors on your own network, not the landlord's BMS feed), so you can detect a thermal excursion without trusting a system you cannot audit, and a documented, tested time-to-thermal-shutdown figure for your rack density so you know how many minutes of margin you are buying.","references":["https://www.cisa.gov/news-events/ics-advisories","https://www.cisa.gov/resources-tools/resources/layering-network-security-through-segmentation"],"status":"curated"},{"id":"NCVD-2026-025-raid-hba-controller-firmware-upd","cve":null,"aliases":[],"title":"RAID/HBA controller firmware update path as a class","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"RAID/HBA controller firmware update path as a class - Broadcom MegaRAID and LSI 9400/9500/9600 HBAs, Microchip…","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"On most of these SKUs the controller firmware can be written from the running host OS by a root process using the vendor CLI (StorCLI, arcconf, perccli, ssacli) over the normal driver ioctl path. There is no host-verifiable attestation of what firmware the controller is actually running - the version string you read back is reported by the same firmware you are trying to verify. The controller is a PCIe device with DMA to host memory that sits below the operating system and below the boot chain, and its firmware is not covered by UEFI Secure Boot or by any measured-boot chain that a normal reimage re-establishes. For a bare-metal GPU rental, that means a tenant with root - which every bare-metal tenant has - can write persistent code beneath the next tenant's operating system, and a wipe-and-reimage handoff does not remove it. This is the concrete mechanism behind 'bare-metal tenant handoff is not a reimage', on a component operators rarely inventory at all.","attack_vector":"Local root on the bare-metal host: the legitimate tenant during their rental, or anyone who achieved root through any other path. No physical access, no BMC access, and no reboot required to stage the flash on most controllers.","remediation":"Not patchable and largely unmitigated on current SKUs. What operators can actually do: (1) make controller firmware version and checksum part of the handoff checklist and reflash from a vendor-signed image between tenants rather than trusting the reported version - budget the reboot into the OEM update utility and the node drain, and expect the OEM package to trail Broadcom/Microchip by months; (2) prefer platforms where the controller participates in a platform root of trust that the BMC can attest, and make that a procurement requirement rather than a hope; (3) blacklist or restrict the management ioctl path from tenant workloads where the workload does not need it; (4) accept and price the residual risk for SKUs where the firmware cannot be independently verified, and keep those nodes out of the pool you offer for security-sensitive tenants.","references":["https://www.broadcom.com/support/resources/product-security-center","https://www.microchip.com/en-us/solutions/embedded-security/how-to-report-potential-product-security-vulnerabilities"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"NCVD-2026-026-leased-colocation-facility-infra","cve":null,"aliases":["facility VLAN ownership gap","colo landlord equipment"],"title":"Leased colocation facility infrastructure (power, cooling, access control) as a class: Most GPU operators lease space","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Leased colocation facility infrastructure (power, cooling, access control) as a class","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Most GPU operators lease space rather than own buildings, which means the UPS, switchgear, PDU upstream feeds, CRAH units and access control that determine whether their GPUs stay powered and cooled are the landlord's equipment, on the landlord's network, patched on the landlord's schedule - which is frequently never, because facility gear is treated as a fixed asset rather than as software. Every CVE in this file is therefore, for many operators, a vulnerability they carry the consequences of and have no authority to fix. An availability event here takes out a hall regardless of how well the compute plane is run.","attack_vector":"Whoever can reach the landlord's facility network - which typically includes the landlord's own vendors, remote-monitoring contractors, and any building-management remote access path the operator has never seen.","remediation":"Contractual, not technical. Require in the colocation agreement: a current inventory of facility control equipment with firmware versions, evidence of a patch cadence, segmentation of the facility network from any operator-reachable network, notification of remote-access paths granted to third parties, and the right to audit. Then verify rather than trust. Where a landlord will not commit, price the risk into the site decision - this belongs in site selection, not in the security backlog.","references":["https://www.cisa.gov/topics/critical-infrastructure-security-and-resilience"],"status":"curated","tags":["physical-impact"]},{"id":"NCVD-2026-026-ses-scsi-enclosure-services-encl","cve":null,"aliases":[],"title":"SES (SCSI Enclosure Services) enclosure management on shared SAS JBODs and expanders: SES is how a host controls","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"SES (SCSI Enclosure Services) enclosure management on shared SAS JBODs and expanders","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"SES is how a host controls a drive enclosure - slot power, locate LEDs, fan and power-supply state - and it has no authentication or authorization of any kind. In a shared JBOD or a multi-initiator SAS topology, every attached host can issue SES commands to the enclosure, including commands affecting slots that belong to another host's drives. An operator running a dense shared-enclosure build therefore has a path where one tenant's node can power-cycle or spin down drives serving a different tenant, or induce enough enclosure-level disruption to fail arrays, with nothing more exotic than a standard SCSI command through /dev/sg. The kernel-side ses driver has also had its own out-of-bounds bugs (CVE-2023-53675, CVE-2023-53431) that make malformed enclosure descriptors a host-crash primitive, which matters when the enclosure firmware is the thing supplying those descriptors.","attack_vector":"Local root on any host attached to the shared SAS fabric, issuing SES commands through the generic SCSI interface. Requires no privileged position beyond being one of the initiators the enclosure already trusts - which is the design.","remediation":"Not patchable at the protocol level; SES has no notion of an authenticated initiator. Mitigations are topological: use SAS zoning on the expander so each host only sees its own drive groups and the enclosure services it needs, avoid multi-tenant shared JBODs entirely for bare-metal rentals, and keep enclosure firmware current for the expander-side parsing bugs. Patch the host kernel for the ses driver out-of-bounds issues so a misbehaving or malicious enclosure cannot crash the initiator. Treat SAS zoning configuration as tenant-isolation configuration and audit it the way you would audit FC zoning.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53675","https://nvd.nist.gov/vuln/detail/CVE-2023-53431","https://www.t10.org/"],"status":"curated"},{"id":"NCVD-2026-027-modbus-tcp-as-an-unauthenticated","cve":null,"aliases":[],"title":"Modbus TCP as an unauthenticated control channel on facility gear: Modbus TCP has no authentication, no authorization","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Modbus TCP as an unauthenticated control channel on facility gear","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Modbus TCP has no authentication, no authorization and no integrity checking. A write-single-register from any host on the segment is indistinguishable from a legitimate one. In a datacenter Modbus is the lingua franca for exactly the equipment you least want touched: CDU and liquid-cooling controllers, chiller and CRAH controllers, generator and transfer-switch controllers, power meters, and the gateways that expose all of the above to the BMS and DCIM. Where a register is writable, an attacker can command an ATS to transfer, a chiller to stop, or a CDU pump to change speed. Where it is read-only, they can still feed false telemetry to the systems that make automated decisions from it. The physical consequence in a GPU hall is immediate: coolant flow or air handling commanded away from 40-140 kW racks means thermal shutdown of the fleet in minutes and hardware stress from the cycle. Modbus DoS is equally cheap - malformed frames crash many embedded Modbus stacks outright, as the Socomec DIRIS Digiware cluster shows.","attack_vector":"Any host that can reach TCP 502 on the device. No credentials. Many facility devices additionally expose Modbus RTU tunnelled over TCP on non-standard ports, which operators forget to inventory. The gear is on the facility VLAN, generally managed by the landlord or the mechanical contractor rather than by the datacenter operator, and frequently reachable from the DCIM collector - so a compromised monitoring server is a direct control path. Internet-exposed Modbus on port 502 remains a standing Shodan finding for building and industrial gear.","remediation":"Unpatchable by design. Controls in order of effectiveness: put Modbus devices on a dedicated segment with a default-deny policy that permits TCP 502 only from the specific poller address; where the device supports it, disable Modbus writes entirely and run read-only, which many facility integrations do not actually need; where a Modbus-to-BACnet or Modbus-to-SNMP gateway exists, terminate Modbus at the gateway and never route it further; and monitor for write function codes (5, 6, 15, 16, 22, 23) on a segment that should only ever see reads - that is a cheap, high-signal detection most operators do not have. For leased space, the equipment and the segment belong to the landlord, so this becomes a contractual item: demand to know which facility devices speak Modbus, whether writes are enabled, and what filters exist in front of port 502.","references":["https://www.cisa.gov/news-events/cybersecurity-advisories/aa22-103a","https://nvd.nist.gov/vuln/detail/CVE-2024-48882","https://www.cisa.gov/resources-tools/resources/secure-design-alert-security-design-improvements-scada-devices"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2026-028-nvme-admin-command-set-firmware","cve":null,"aliases":["NVMe Firmware Image Download","NVMe Firmware Commit","firmware downgrade attack","unsigned drive firmware"],"title":"NVMe admin command set - Firmware Image Download (opcode 11h) and Firmware Commit (opcode 10h) reachable from the host","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVMe admin command set - Firmware Image Download (opcode 11h) and Firmware Commit (opcode 10h) reachable from the…","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"CLASS ENTRY, and the single highest-leverage tenant-handoff risk in this whole category. The NVMe specification puts firmware download and firmware commit in the standard admin command set, meaning a host with root can push a firmware image to the drive using nothing more exotic than nvme-cli. Whether that is safe depends entirely on the drive enforcing signature verification and anti-rollback - and that enforcement is uneven across vendors and generations, is not something the host can independently verify, and has been shown breakable even where present (the Kioxia CM6/PM6/PM7 JTAG work bypasses the RSA signature check outright). Where verification is weak or absent, a departing tenant flashes a modified image and owns the drive controller permanently: it outlives your wipe, your reimage and your reprovision, sees everything the next tenant writes, and lies to every host-side tool about its own version, its lock state and whether a sanitize succeeded. Even where signature checking holds, downgrade to an older SIGNED image with known vulnerabilities is a live path unless the drive enforces anti-rollback - and the RPMB replay flaw (CVE-2020-13799) undermines exactly that anti-rollback state across eMMC, UFS and all NVMe versions.","attack_vector":"A tenant with root on the bare-metal host, issuing standard NVMe admin commands to a locally attached drive. No physical access, no exotic tooling, no exploit needed where the drive does not enforce signing - the command path is a documented, supported feature. Also relevant in SR-IOV and DPU/computational-storage designs, where whether the admin queue and Security Send/Receive are properly filtered from a tenant-controlled function is a per-platform question most operators have never actually tested.","remediation":"Policy and platform configuration; there is no patch because the command path is by design. (1) Block it at the platform: filter Firmware Image Download and Firmware Commit - and vendor-specific and Security Send/Receive pass-through - so tenant-controlled hosts and virtual functions cannot reach them. Do not expose raw NVMe admin queues to tenants unless the product genuinely requires it. (2) TEST it rather than assuming: on a representative node, try to flash a drive from a tenant-equivalent shell and confirm you are refused. Most operators have never run this check and will be surprised by the result on at least one SKU. (3) Make signed firmware and enforced anti-rollback a written procurement requirement, and get the vendor to state it per SKU. (4) Inventory expected firmware version per drive serial in a store the host cannot write to, and alert on any change or any version that decreases. (5) For sensitive tenancies, retire local media at end of tenancy instead of recycling it. Assume that verifying firmware integrity across a 10,000-drive fleet is not achievable with host-side tooling - the controller is the thing answering your questions - so the control has to be preventing the write and controlling the media's lifecycle, not detecting the implant afterwards.","references":["https://www.nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0c-2022.10.04-Ratified.pdf","https://github.com/google/security-research/security/advisories/GHSA-3hh8-94j4-62rh","https://www.kb.cert.org/vuls/id/231329","https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-88r1.pdf"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2026-030-raid-hba-controller-firmware-upd","cve":null,"aliases":[],"title":"RAID/HBA controller firmware update path as a class","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"RAID/HBA controller firmware update path as a class - Broadcom MegaRAID and LSI 9400/9500/9600 HBAs, Microchip…","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"On most of these SKUs the controller firmware can be written from the running host OS by a root process using the vendor CLI (StorCLI, arcconf, perccli, ssacli) over the normal driver ioctl path. There is no host-verifiable attestation of what firmware the controller is actually running - the version string you read back is reported by the same firmware you are trying to verify. The controller is a PCIe device with DMA to host memory that sits below the operating system and below the boot chain, and its firmware is not covered by UEFI Secure Boot or by any measured-boot chain that a normal reimage re-establishes. For a bare-metal GPU rental, that means a tenant with root - which every bare-metal tenant has - can write persistent code beneath the next tenant's operating system, and a wipe-and-reimage handoff does not remove it. This is the concrete mechanism behind 'bare-metal tenant handoff is not a reimage', on a component operators rarely inventory at all.","attack_vector":"Local root on the bare-metal host: the legitimate tenant during their rental, or anyone who achieved root through any other path. No physical access, no BMC access, and no reboot required to stage the flash on most controllers.","remediation":"Not patchable and largely unmitigated on current SKUs. What operators can actually do: (1) make controller firmware version and checksum part of the handoff checklist and reflash from a vendor-signed image between tenants rather than trusting the reported version - budget the reboot into the OEM update utility and the node drain, and expect the OEM package to trail Broadcom/Microchip by months; (2) prefer platforms where the controller participates in a platform root of trust that the BMC can attest, and make that a procurement requirement rather than a hope; (3) blacklist or restrict the management ioctl path from tenant workloads where the workload does not need it; (4) accept and price the residual risk for SKUs where the firmware cannot be independently verified, and keep those nodes out of the pool you offer for security-sensitive tenants.","references":["https://www.broadcom.com/support/resources/product-security-center","https://www.microchip.com/en-us/solutions/embedded-security/how-to-report-potential-product-security-vulnerabilities"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"NCVD-2026-033-wiegand-reader-to-controller-wir","cve":null,"aliases":[],"title":"Wiegand reader-to-controller wiring and legacy 125 kHz proximity / MIFARE Classic credentials: Two structural","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Wiegand reader-to-controller wiring and legacy 125 kHz proximity / MIFARE Classic credentials","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Two structural weaknesses in nearly every badge system installed before roughly the last decade, and plenty installed since. First, the Wiegand protocol between the card reader and the door controller is unauthenticated, unencrypted plaintext over a few wires - anyone who can reach the back of the reader, which is on the unsecured side of the door, can splice in a small logger to capture every credential presented, or inject a previously captured credential to open the door. Second, 125 kHz proximity cards and MIFARE Classic credentials are cloneable in seconds with a sub-hundred-dollar reader from a pocket's distance, so an attacker who stands near an employee in a coffee shop can walk into the hall an hour later. Neither of these produces an anomalous event: the badge system records a valid credential at a valid door at a plausible time. What follows from being inside the cage is the whole point - drives with model weights and customer data, server console ports, an unlocked out-of-band switch, and the ability to plant a hardware implant that survives everything your software security program looks at. For a multi-tenant operator this also defeats the cage boundary you sell to customers, and it defeats it in a way that leaves no forensic trace.","attack_vector":"Physical presence. For Wiegand tapping: brief access to the reader housing, which is mounted outside the secured area by definition and usually held on with a security screw. For card cloning: proximity to any credential holder, or access to a credential left in a desk. No network access is required for either, which is exactly why network-centric security programs miss them. Note that many datacenter cages are protected by a single badge reader with no second factor, and that contractor and landlord staff badges frequently open more doors than the tenant realises.","remediation":"Not patchable - these are design properties of the installed hardware. The real fixes are hardware replacements and they cost money: move reader-to-controller communication from Wiegand to OSDP v2 with Secure Channel (encrypted and mutually authenticated), which requires readers and controllers that support it and a rewiring pass; and migrate credentials from 125 kHz prox and MIFARE Classic to a cryptographic credential (DESFire EV2/EV3 with a properly managed site key, or mobile credentials) which requires new cards and new readers. In the meantime: add a second independent factor at the hall and cage doors (PIN pad or biometric, on a separate system from the badge reader), fit tamper switches on reader housings and alarm on them, put a camera covering every cage door with retention long enough to be useful, and run a periodic reconciliation of who actually holds a badge that opens your hall - including landlord and contractor staff. For leased space, cage-door credential technology is a lease negotiation item; ask what credential format is in use and treat '125 kHz prox' as a finding.","references":["https://www.cisa.gov/news-events/ics-advisories/icsa-23-061-03","https://nvd.nist.gov/vuln/detail/CVE-2022-40633","https://www.securityindustry.org/industry-standards/open-supervised-device-protocol/"],"status":"curated"},{"id":"NCVD-2026-035-leased-colocation-facility-infra","cve":null,"aliases":["facility VLAN ownership gap","colo landlord equipment"],"title":"Leased colocation facility infrastructure (power, cooling, access control) as a class: Most GPU operators lease space","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Leased colocation facility infrastructure (power, cooling, access control) as a class","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Most GPU operators lease space rather than own buildings, which means the UPS, switchgear, PDU upstream feeds, CRAH units and access control that determine whether their GPUs stay powered and cooled are the landlord's equipment, on the landlord's network, patched on the landlord's schedule - which is frequently never, because facility gear is treated as a fixed asset rather than as software. Every CVE in this file is therefore, for many operators, a vulnerability they carry the consequences of and have no authority to fix. An availability event here takes out a hall regardless of how well the compute plane is run.","attack_vector":"Whoever can reach the landlord's facility network - which typically includes the landlord's own vendors, remote-monitoring contractors, and any building-management remote access path the operator has never seen.","remediation":"Contractual, not technical. Require in the colocation agreement: a current inventory of facility control equipment with firmware versions, evidence of a patch cadence, segmentation of the facility network from any operator-reachable network, notification of remote-access paths granted to third parties, and the right to audit. Then verify rather than trust. Where a landlord will not commit, price the risk into the site decision - this belongs in site selection, not in the security backlog.","references":["https://www.cisa.gov/topics/critical-infrastructure-security-and-resilience"],"status":"curated","tags":["physical-impact"]},{"cwe":["CWE-287"],"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2026-039-htcondor-access-point-daemons-co","cve":null,"aliases":["HTCONDOR-2026-0001"],"title":"HTCondor (Access Point daemons, condor identity): A user with WRITE authorization on an Access Point - i.e. anyone who","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HTCondor (Access Point daemons, condor identity)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A user with WRITE authorization on an Access Point - i.e. anyone who can submit a job - can impersonate the condor service account to that AP's daemons. That grants ADMINISTRATOR-level commands (hold and remove anyone's jobs, submit as other users, shut daemons down), the ability to edit any attribute of any job ClassAd, and the ability to mint an IDToken for the condor identity. Because IDToken signing keys are usually shared across a pool, that token is then usable against other machines, so a single submit-capable tenant escalates to pool-wide control.","attack_vector":"Any user with WRITE authorization to an Access Point, from any host. The HTCondor team rates the effort as medium - custom tooling is required, but no privileged position is.","remediation":"Upgrade the Access Point to HTCondor 24.0.22, 24.12.22, 25.0.12 or 25.11.1 and restart its daemons. Because a forged condor IDToken survives the patch, rotate the IDToken signing keys across every machine that shares them with the affected AP and reissue tokens. Fix date was 2026-07-21; no CVE ID had been published for this advisory as of 2026-08-20.","references":["https://htcondor.org/security/vulnerabilities/HTCONDOR-2026-0001.html","https://htcondor.org/security/vulnerabilities/"],"status":"curated","tags":["tenant-isolation"]},{"id":"NCVD-2026-039-nvme-admin-command-set-firmware","cve":null,"aliases":["NVMe Firmware Image Download","NVMe Firmware Commit","firmware downgrade attack","unsigned drive firmware"],"title":"NVMe admin command set - Firmware Image Download (opcode 11h) and Firmware Commit (opcode 10h) reachable from the host","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVMe admin command set - Firmware Image Download (opcode 11h) and Firmware Commit (opcode 10h) reachable from the…","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"CLASS ENTRY, and the single highest-leverage tenant-handoff risk in this whole category. The NVMe specification puts firmware download and firmware commit in the standard admin command set, meaning a host with root can push a firmware image to the drive using nothing more exotic than nvme-cli. Whether that is safe depends entirely on the drive enforcing signature verification and anti-rollback - and that enforcement is uneven across vendors and generations, is not something the host can independently verify, and has been shown breakable even where present (the Kioxia CM6/PM6/PM7 JTAG work bypasses the RSA signature check outright). Where verification is weak or absent, a departing tenant flashes a modified image and owns the drive controller permanently: it outlives your wipe, your reimage and your reprovision, sees everything the next tenant writes, and lies to every host-side tool about its own version, its lock state and whether a sanitize succeeded. Even where signature checking holds, downgrade to an older SIGNED image with known vulnerabilities is a live path unless the drive enforces anti-rollback - and the RPMB replay flaw (CVE-2020-13799) undermines exactly that anti-rollback state across eMMC, UFS and all NVMe versions.","attack_vector":"A tenant with root on the bare-metal host, issuing standard NVMe admin commands to a locally attached drive. No physical access, no exotic tooling, no exploit needed where the drive does not enforce signing - the command path is a documented, supported feature. Also relevant in SR-IOV and DPU/computational-storage designs, where whether the admin queue and Security Send/Receive are properly filtered from a tenant-controlled function is a per-platform question most operators have never actually tested.","remediation":"Policy and platform configuration; there is no patch because the command path is by design. (1) Block it at the platform: filter Firmware Image Download and Firmware Commit - and vendor-specific and Security Send/Receive pass-through - so tenant-controlled hosts and virtual functions cannot reach them. Do not expose raw NVMe admin queues to tenants unless the product genuinely requires it. (2) TEST it rather than assuming: on a representative node, try to flash a drive from a tenant-equivalent shell and confirm you are refused. Most operators have never run this check and will be surprised by the result on at least one SKU. (3) Make signed firmware and enforced anti-rollback a written procurement requirement, and get the vendor to state it per SKU. (4) Inventory expected firmware version per drive serial in a store the host cannot write to, and alert on any change or any version that decreases. (5) For sensitive tenancies, retire local media at end of tenancy instead of recycling it. Assume that verifying firmware integrity across a 10,000-drive fleet is not achievable with host-side tooling - the controller is the thing answering your questions - so the control has to be preventing the write and controlling the media's lifecycle, not detecting the implant afterwards.","references":["https://www.nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0c-2022.10.04-Ratified.pdf","https://github.com/google/security-research/security/advisories/GHSA-3hh8-94j4-62rh","https://www.kb.cert.org/vuls/id/231329","https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-88r1.pdf"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"cwe":["CWE-918","CWE-444"],"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2026-040-kubeflow-pipelines-frontend-serv","cve":null,"aliases":["GHSA-gqww-5pj5-8fq7","CVE-2026-54745 (reserved, no MITRE record as of 2026-08-20)"],"title":"Kubeflow Pipelines (frontend server, /_proxy/ route in proxy-middleware.ts): The Kubeflow Pipelines frontend forwards","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Kubeflow Pipelines (frontend server, /_proxy/ route in proxy-middleware.ts)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"The Kubeflow Pipelines frontend forwards arbitrary user-supplied URLs through /_proxy/ with no allowlist, no IP filter and no auth gate, and the route is not covered by the auth middleware - so it stays open even with ENABLE_AUTHZ=true, the posture documented as multi-user secure. Method, body and headers pass through verbatim, giving an unauthenticated caller reach into link-local metadata (169.254.0.0/16), loopback, RFC1918 and any *.svc.cluster.local service, plus an HTTP smuggling primitive against them. On a GPU cluster the first stop is node IAM credentials.","attack_vector":"Anyone who can reach the Kubeflow Pipelines frontend, including via a crafted Referer header on unrelated paths. No credentials, and multi-user mode does not help.","remediation":"Track the kubeflow/pipelines advisory and upgrade the frontend image once a fixed tag ships. Until then, block the four /_proxy/ prefixes (/apis/v1beta1/_proxy/, /apis/v2beta1/_proxy/, and their /pipeline-prefixed forms) at the ingress, and rotate node/pod cloud credentials if the frontend was reachable from untrusted networks.","references":["https://github.com/kubeflow/pipelines/security/advisories/GHSA-gqww-5pj5-8fq7"],"status":"curated","tags":["tenant-isolation"]},{"id":"NCVD-2026-041-raid-hba-controller-firmware-upd","cve":null,"aliases":[],"title":"RAID/HBA controller firmware update path as a class","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"RAID/HBA controller firmware update path as a class - Broadcom MegaRAID and LSI 9400/9500/9600 HBAs, Microchip…","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"On most of these SKUs the controller firmware can be written from the running host OS by a root process using the vendor CLI (StorCLI, arcconf, perccli, ssacli) over the normal driver ioctl path. There is no host-verifiable attestation of what firmware the controller is actually running - the version string you read back is reported by the same firmware you are trying to verify. The controller is a PCIe device with DMA to host memory that sits below the operating system and below the boot chain, and its firmware is not covered by UEFI Secure Boot or by any measured-boot chain that a normal reimage re-establishes. For a bare-metal GPU rental, that means a tenant with root - which every bare-metal tenant has - can write persistent code beneath the next tenant's operating system, and a wipe-and-reimage handoff does not remove it. This is the concrete mechanism behind 'bare-metal tenant handoff is not a reimage', on a component operators rarely inventory at all.","attack_vector":"Local root on the bare-metal host: the legitimate tenant during their rental, or anyone who achieved root through any other path. No physical access, no BMC access, and no reboot required to stage the flash on most controllers.","remediation":"Not patchable and largely unmitigated on current SKUs. What operators can actually do: (1) make controller firmware version and checksum part of the handoff checklist and reflash from a vendor-signed image between tenants rather than trusting the reported version - budget the reboot into the OEM update utility and the node drain, and expect the OEM package to trail Broadcom/Microchip by months; (2) prefer platforms where the controller participates in a platform root of trust that the BMC can attest, and make that a procurement requirement rather than a hope; (3) blacklist or restrict the management ioctl path from tenant workloads where the workload does not need it; (4) accept and price the residual risk for SKUs where the firmware cannot be independently verified, and keep those nodes out of the pool you offer for security-sensitive tenants.","references":["https://www.broadcom.com/support/resources/product-security-center","https://www.microchip.com/en-us/solutions/embedded-security/how-to-report-potential-product-security-vulnerabilities"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"NCVD-2026-044-nvme-admin-command-set-firmware","cve":null,"aliases":["NVMe Firmware Image Download","NVMe Firmware Commit","firmware downgrade attack","unsigned drive firmware"],"title":"NVMe admin command set - Firmware Image Download (opcode 11h) and Firmware Commit (opcode 10h) reachable from the host","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVMe admin command set - Firmware Image Download (opcode 11h) and Firmware Commit (opcode 10h) reachable from the","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"CLASS ENTRY, and the single highest-leverage tenant-handoff risk in this whole category. The NVMe specification puts firmware download and firmware commit in the standard admin command set, meaning a host with root can push a firmware image to the drive using nothing more exotic than nvme-cli. Whether that is safe depends entirely on the drive enforcing signature verification and anti-rollback - and that enforcement is uneven across vendors and generations, is not something the host can independently verify, and has been shown breakable even where present (the Kioxia CM6/PM6/PM7 JTAG work bypasses the RSA signature check outright). Where verification is weak or absent, a departing tenant flashes a modified image and owns the drive controller permanently: it outlives your wipe, your reimage and your reprovision, sees everything the next tenant writes, and lies to every host-side tool about its own version, its lock state and whether a sanitize succeeded. Even where signature checking holds, downgrade to an older SIGNED image with known vulnerabilities is a live path unless the drive enforces anti-rollback - and the RPMB replay flaw (CVE-2020-13799) undermines exactly that anti-rollback state across eMMC, UFS and all NVMe versions.","attack_vector":"A tenant with root on the bare-metal host, issuing standard NVMe admin commands to a locally attached drive. No physical access, no exotic tooling, no exploit needed where the drive does not enforce signing - the command path is a documented, supported feature. Also relevant in SR-IOV and DPU/computational-storage designs, where whether the admin queue and Security Send/Receive are properly filtered from a tenant-controlled function is a per-platform question most operators have never actually tested.","remediation":"Policy and platform configuration; there is no patch because the command path is by design. (1) Block it at the platform: filter Firmware Image Download and Firmware Commit - and vendor-specific and Security Send/Receive pass-through - so tenant-controlled hosts and virtual functions cannot reach them. Do not expose raw NVMe admin queues to tenants unless the product genuinely requires it. (2) TEST it rather than assuming: on a representative node, try to flash a drive from a tenant-equivalent shell and confirm you are refused. Most operators have never run this check and will be surprised by the result on at least one SKU. (3) Make signed firmware and enforced anti-rollback a written procurement requirement, and get the vendor to state it per SKU. (4) Inventory expected firmware version per drive serial in a store the host cannot write to, and alert on any change or any version that decreases. (5) For sensitive tenancies, retire local media at end of tenancy instead of recycling it. Assume that verifying firmware integrity across a 10,000-drive fleet is not achievable with host-side tooling - the controller is the thing answering your questions - so the control has to be preventing the write and controlling the media's lifecycle, not detecting the implant afterwards.","references":["https://www.nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0c-2022.10.04-Ratified.pdf","https://github.com/google/security-research/security/advisories/GHSA-3hh8-94j4-62rh","https://www.kb.cert.org/vuls/id/231329","https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-88r1.pdf"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2026-046-raid-hba-controller-firmware-upd","cve":null,"aliases":[],"title":"RAID/HBA controller firmware update path as a class - Broadcom MegaRAID and LSI 9400/9500/9600 HBAs, Microchip Adaptec","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"RAID/HBA controller firmware update path as a class - Broadcom MegaRAID and LSI 9400/9500/9600 HBAs, Microchip","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"On most of these SKUs the controller firmware can be written from the running host OS by a root process using the vendor CLI (StorCLI, arcconf, perccli, ssacli) over the normal driver ioctl path. There is no host-verifiable attestation of what firmware the controller is actually running - the version string you read back is reported by the same firmware you are trying to verify. The controller is a PCIe device with DMA to host memory that sits below the operating system and below the boot chain, and its firmware is not covered by UEFI Secure Boot or by any measured-boot chain that a normal reimage re-establishes. For a bare-metal GPU rental, that means a tenant with root - which every bare-metal tenant has - can write persistent code beneath the next tenant's operating system, and a wipe-and-reimage handoff does not remove it. This is the concrete mechanism behind 'bare-metal tenant handoff is not a reimage', on a component operators rarely inventory at all.","attack_vector":"Local root on the bare-metal host: the legitimate tenant during their rental, or anyone who achieved root through any other path. No physical access, no BMC access, and no reboot required to stage the flash on most controllers.","remediation":"Not patchable and largely unmitigated on current SKUs. What operators can actually do: (1) make controller firmware version and checksum part of the handoff checklist and reflash from a vendor-signed image between tenants rather than trusting the reported version - budget the reboot into the OEM update utility and the node drain, and expect the OEM package to trail Broadcom/Microchip by months; (2) prefer platforms where the controller participates in a platform root of trust that the BMC can attest, and make that a procurement requirement rather than a hope; (3) blacklist or restrict the management ioctl path from tenant workloads where the workload does not need it; (4) accept and price the residual risk for SKUs where the firmware cannot be independently verified, and keep those nodes out of the pool you offer for security-sensitive tenants.","references":["https://www.broadcom.com/support/resources/product-security-center","https://www.microchip.com/en-us/solutions/embedded-security/how-to-report-potential-product-security-vulnerabilities"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"cwe":["CWE-287"],"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2026-059-htcondor-access-point-daemons-co","cve":null,"aliases":["HTCONDOR-2026-0001"],"title":"HTCondor (Access Point daemons, condor identity): A user with WRITE authorization on an Access Point - i.e. anyone who","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HTCondor (Access Point daemons, condor identity)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A user with WRITE authorization on an Access Point - i.e. anyone who can submit a job - can impersonate the condor service account to that AP's daemons. That grants ADMINISTRATOR-level commands (hold and remove anyone's jobs, submit as other users, shut daemons down), the ability to edit any attribute of any job ClassAd, and the ability to mint an IDToken for the condor identity. Because IDToken signing keys are usually shared across a pool, that token is then usable against other machines, so a single submit-capable tenant escalates to pool-wide control.","attack_vector":"Any user with WRITE authorization to an Access Point, from any host. The HTCondor team rates the effort as medium - custom tooling is required, but no privileged position is.","remediation":"Upgrade the Access Point to HTCondor 24.0.22, 24.12.22, 25.0.12 or 25.11.1 and restart its daemons. Because a forged condor IDToken survives the patch, rotate the IDToken signing keys across every machine that shares them with the affected AP and reissue tokens. Fix date was 2026-07-21; no CVE ID had been published for this advisory as of 2026-08-20.","references":["https://htcondor.org/security/vulnerabilities/HTCONDOR-2026-0001.html","https://htcondor.org/security/vulnerabilities/"],"status":"curated","tags":["tenant-isolation"]},{"cwe":["CWE-918","CWE-444"],"fleet":{"pain_class":"daemon-restart"},"id":"NCVD-2026-060-kubeflow-pipelines-frontend-serv","cve":null,"aliases":["GHSA-gqww-5pj5-8fq7","CVE-2026-54745 (reserved, no MITRE record as of 2026-08-20)"],"title":"Kubeflow Pipelines (frontend server, /_proxy/ route in proxy-middleware.ts): The Kubeflow Pipelines frontend forwards","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Kubeflow Pipelines (frontend server, /_proxy/ route in proxy-middleware.ts)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"The Kubeflow Pipelines frontend forwards arbitrary user-supplied URLs through /_proxy/ with no allowlist, no IP filter and no auth gate, and the route is not covered by the auth middleware - so it stays open even with ENABLE_AUTHZ=true, the posture documented as multi-user secure. Method, body and headers pass through verbatim, giving an unauthenticated caller reach into link-local metadata (169.254.0.0/16), loopback, RFC1918 and any *.svc.cluster.local service, plus an HTTP smuggling primitive against them. On a GPU cluster the first stop is node IAM credentials.","attack_vector":"Anyone who can reach the Kubeflow Pipelines frontend, including via a crafted Referer header on unrelated paths. No credentials, and multi-user mode does not help.","remediation":"Track the kubeflow/pipelines advisory and upgrade the frontend image once a fixed tag ships. Until then, block the four /_proxy/ prefixes (/apis/v1beta1/_proxy/, /apis/v2beta1/_proxy/, and their /pipeline-prefixed forms) at the ingress, and rotate node/pod cloud credentials if the frontend was reachable from untrusted networks.","references":["https://github.com/kubeflow/pipelines/security/advisories/GHSA-gqww-5pj5-8fq7"],"status":"curated","tags":["tenant-isolation"]},{"id":"NCVD-2026-064-nvme-admin-command-set-firmware","cve":null,"aliases":["NVMe Firmware Image Download","NVMe Firmware Commit","firmware downgrade attack","unsigned drive firmware"],"title":"NVMe admin command set - Firmware Image Download (opcode 11h) and Firmware Commit (opcode 10h) reachable from the host","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVMe admin command set - Firmware Image Download (opcode 11h) and Firmware Commit (opcode 10h) reachable from the","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"CLASS ENTRY, and the single highest-leverage tenant-handoff risk in this whole category. The NVMe specification puts firmware download and firmware commit in the standard admin command set, meaning a host with root can push a firmware image to the drive using nothing more exotic than nvme-cli. Whether that is safe depends entirely on the drive enforcing signature verification and anti-rollback - and that enforcement is uneven across vendors and generations, is not something the host can independently verify, and has been shown breakable even where present (the Kioxia CM6/PM6/PM7 JTAG work bypasses the RSA signature check outright). Where verification is weak or absent, a departing tenant flashes a modified image and owns the drive controller permanently: it outlives your wipe, your reimage and your reprovision, sees everything the next tenant writes, and lies to every host-side tool about its own version, its lock state and whether a sanitize succeeded. Even where signature checking holds, downgrade to an older SIGNED image with known vulnerabilities is a live path unless the drive enforces anti-rollback - and the RPMB replay flaw (CVE-2020-13799) undermines exactly that anti-rollback state across eMMC, UFS and all NVMe versions.","attack_vector":"A tenant with root on the bare-metal host, issuing standard NVMe admin commands to a locally attached drive. No physical access, no exotic tooling, no exploit needed where the drive does not enforce signing - the command path is a documented, supported feature. Also relevant in SR-IOV and DPU/computational-storage designs, where whether the admin queue and Security Send/Receive are properly filtered from a tenant-controlled function is a per-platform question most operators have never actually tested.","remediation":"Policy and platform configuration; there is no patch because the command path is by design. (1) Block it at the platform: filter Firmware Image Download and Firmware Commit - and vendor-specific and Security Send/Receive pass-through - so tenant-controlled hosts and virtual functions cannot reach them. Do not expose raw NVMe admin queues to tenants unless the product genuinely requires it. (2) TEST it rather than assuming: on a representative node, try to flash a drive from a tenant-equivalent shell and confirm you are refused. Most operators have never run this check and will be surprised by the result on at least one SKU. (3) Make signed firmware and enforced anti-rollback a written procurement requirement, and get the vendor to state it per SKU. (4) Inventory expected firmware version per drive serial in a store the host cannot write to, and alert on any change or any version that decreases. (5) For sensitive tenancies, retire local media at end of tenancy instead of recycling it. Assume that verifying firmware integrity across a 10,000-drive fleet is not achievable with host-side tooling - the controller is the thing answering your questions - so the control has to be preventing the write and controlling the media's lifecycle, not detecting the implant afterwards.","references":["https://www.nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0c-2022.10.04-Ratified.pdf","https://github.com/google/security-research/security/advisories/GHSA-3hh8-94j4-62rh","https://www.kb.cert.org/vuls/id/231329","https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-88r1.pdf"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2026-066-raid-hba-controller-firmware-upd","cve":null,"aliases":[],"title":"RAID/HBA controller firmware update path as a class - Broadcom MegaRAID and LSI 9400/9500/9600 HBAs, Microchip Adaptec","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"RAID/HBA controller firmware update path as a class - Broadcom MegaRAID and LSI 9400/9500/9600 HBAs, Microchip","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"On most of these SKUs the controller firmware can be written from the running host OS by a root process using the vendor CLI (StorCLI, arcconf, perccli, ssacli) over the normal driver ioctl path. There is no host-verifiable attestation of what firmware the controller is actually running - the version string you read back is reported by the same firmware you are trying to verify. The controller is a PCIe device with DMA to host memory that sits below the operating system and below the boot chain, and its firmware is not covered by UEFI Secure Boot or by any measured-boot chain that a normal reimage re-establishes. For a bare-metal GPU rental, that means a tenant with root - which every bare-metal tenant has - can write persistent code beneath the next tenant's operating system, and a wipe-and-reimage handoff does not remove it. This is the concrete mechanism behind 'bare-metal tenant handoff is not a reimage', on a component operators rarely inventory at all.","attack_vector":"Local root on the bare-metal host: the legitimate tenant during their rental, or anyone who achieved root through any other path. No physical access, no BMC access, and no reboot required to stage the flash on most controllers.","remediation":"Not patchable and largely unmitigated on current SKUs. What operators can actually do: (1) make controller firmware version and checksum part of the handoff checklist and reflash from a vendor-signed image between tenants rather than trusting the reported version - budget the reboot into the OEM update utility and the node drain, and expect the OEM package to trail Broadcom/Microchip by months; (2) prefer platforms where the controller participates in a platform root of trust that the BMC can attest, and make that a procurement requirement rather than a hope; (3) blacklist or restrict the management ioctl path from tenant workloads where the workload does not need it; (4) accept and price the residual risk for SKUs where the firmware cannot be independently verified, and keep those nodes out of the pool you offer for security-sensitive tenants.","references":["https://www.broadcom.com/support/resources/product-security-center","https://www.microchip.com/en-us/solutions/embedded-security/how-to-report-potential-product-security-vulnerabilities"],"status":"curated","fleet":{"pain_class":"firmware-flash"}}]}