NVIDIA Resiliency Extension: A race condition in the checkpointing core reaches information disclosure, data tampering
Impact
A race condition in the checkpointing core reaches information disclosure, data tampering and privilege escalation. Checkpoints are the training run's crown jewels - a tampering primitive here means silently poisoned model state that survives every restart.
Who can reach it
Local, low privileges. An account on a node participating in the checkpoint write path.
What to do
Update the Resiliency Extension per bulletin 5746 and rebuild training images. Cost: image rebuild and job restart. Verify checkpoint integrity out-of-band if you suspect exposure - a corrupted checkpoint does not announce itself.
References
Related entries
- NVIDIA Resiliency Extension: Predictable log-file names in the log-aggregation path let an attacker pre-createCVE-2025-33225 · NVIDIA Resiliency ExtensionHigh
- NeMo Framework: Arbitrary Python code exec via code injectionCVE-2025-33236 · NeMo FrameworkHigh
- Megatron-Bridge: code injection via malicious input in the data merging and data shuffling tutorialsCVE-2025-33239 · Megatron-BridgeHigh
- NeMo Framework: Unsafe deserialization of untrusted model/config data allows arbitrary code executionCVE-2025-33241 · NeMo FrameworkHigh
- NeMo Framework: Local command injection via unsanitized input passed to shell executionCVE-2025-33246 · NeMo FrameworkHigh
- Megatron-LM: Local privesc / RCE via unsafe pickle deserializationCVE-2025-33247 · Megatron-LMHigh
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.