Database/Control plane, storage & DevOps
ECC DDR3 server memory on Intel Xeon (Haswell, Sandy Bridge) and AMD Opteron platforms
Impact
ECC is the answer most operators give when asked about Rowhammer, and ECCploit is why that answer is wrong. Correcting a flip takes measurably longer than a clean read, so the attacker gets a timing side channel that tells them exactly which bits they flipped - turning ECC from a defence into a feedback oracle for template building. With that feedback they place three flips in one word, which ECC neither corrects nor detects, producing silent corruption. For an AI datacenter this is the nastiest variant: the corruption is by construction invisible to the machine-check path, so your fleet health dashboard shows green while a tenant's memory is being rewritten.
Who can reach it
Unprivileged local code on an ECC server sharing DRAM with the victim. Attack time was about 32 minutes when corrections were directly observable and up to a week in noisy production-like conditions - slow, but a long-running batch tenant has that time.
What to do
Do not treat ECC as a Rowhammer mitigation; treat it as error reporting. Make sure the reporting is actually wired up - EDAC or the equivalent collecting correctable-error counts per DIMM, exported to your monitoring, with alerting on rate rather than absolute count, since a burst of corrections on one rank is the strongest hammering signal you will get. Confirm firmware and OS handle uncorrectable errors by isolating rather than silently continuing. Retire DIMM SKUs that show elevated correctable-error rates. The structural fix is the same as every other entry here: do not share a memory controller between untrusted tenants.
References
Related entries
- ECC DDR3 server memory on Intel Xeon (Haswell, Sandy Bridge) and AMD Opteron platformsNCVD-2018-002-ecc-ddr3-server-memory-on-intel · ECC DDR3 server memory on Intel Xeon (Haswell, Sandy Bridge) and AMD Opteron platforms; the technique generalises to…Unscored
- PCIe Address Translation Services on hosts using an IOMMU/SMMU for device isolationNCVD-2019-001-pcie-address-translation-service · PCIe Address Translation Services on hosts using an IOMMU/SMMU for device isolation - affects any DMA-capable…Unscored
- PCIe Address Translation Services on hosts using an IOMMU/SMMU for device isolationNCVD-2019-005-pcie-address-translation-service · PCIe Address Translation Services on hosts using an IOMMU/SMMU for device isolation - affects any DMA-capable…Unscored
- AMD Zen 1 / Zen+ / Zen 2 - L1D cache way predictor: AMD's L1D way predictor hashes virtual addresses to predict whichNCVD-2020-001-amd-zen-1-zen-zen-2-l1d-cache-wa · AMD Zen 1 / Zen+ / Zen 2 - L1D cache way predictorUnscored
- DDR4 DRAM with in-DRAM TRR; a coupling effect that reaches rows at distance two rather than immediate neighboursNCVD-2021-002-ddr4-dram-with-in-dram-trr-a-cou · DDR4 DRAM with in-DRAM TRR; a coupling effect that reaches rows at distance two rather than immediate neighboursUnscored
- BeeGFS (client-to-metadata/storage service authentication, connAuthFile): Class entry, not a single CVE. Before BeeGFSNCVD-2022-004-beegfs-client-to-metadata-storag · BeeGFS (client-to-metadata/storage service authentication, connAuthFile)Unscored
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.