NVIDIA DCGM - nv-hostengine: A network-reachable caller drives nv-hostengine into an unhandled error condition
Impact
A network-reachable caller drives nv-hostengine into an unhandled error condition, reaching limited code execution and privilege escalation. DCGM runs as root on every GPU node and holds the fleet's telemetry, so it is a high-value target sitting on an open port.
Who can reach it
Network, with low privileges. nv-hostengine listens on TCP 5555 by default and many operators leave it bound beyond localhost so a central collector can scrape it - that binding is the exposure.
What to do
Update DCGM per bulletin 5328 and restart nv-hostengine. Cost: restarting the host engine briefly interrupts telemetry but does not touch running GPU jobs - no drain needed. While you are there, bind nv-hostengine to localhost and scrape via a local exporter instead of exposing 5555.
References
Related entries
- NVIDIA DCGM - nv-hostengine: A heap-based buffer overflow reachable through the bound socket gives denial of serviceCVE-2023-0208 · NVIDIA DCGM - nv-hostengineHigh
- vGPU Manager: Improper permission managementCVE-2024-0085 · vGPU ManagerMedium
- NVIDIA NeMo: SaveRestoreConnector extracts .tar archives unsafely, so a crafted archive writes files outsideCVE-2024-0129 · NVIDIA NeMoMedium
- ConnectX: Access-control flaw in NIC firmwareCVE-2025-23262 · ConnectXMedium
- TensorRT-LLM: Code exec via insecure deserializationCVE-2026-24142 · TensorRT-LLMMedium
- TensorRT-LLM: Unsafe external file loadingCVE-2026-24226 · TensorRT-LLMMedium
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.