Database/Kernel, userspace & hypervisor
Linux kernel x86/mm: pmd_modify() drops the dirty bit, losing written data on PMD-mapped THP
Impact
pmd_modify() masked out the hardware dirty bit, so pmd_mksaveddirty() had nothing left to carry into _PAGE_SAVED_DIRTY. Any pmd_modify() on a writable, dirty PMD - an mprotect(), or NUMA hinting going through do_huge_pmd_numa_page() - silently clears the dirty state. With MADV_FREE on a PMD-mapped anonymous THP, reclaim then finds a lazyfree folio with no dirty bit anywhere and discards data that was rewritten after the madvise; subsequent reads return fresh zero pages. PMD-mapped file THPs lose writeback of rewritten data the same way. This is silent user-space data loss, not a privilege boundary break, and it needs no attacker: the report notes production data lost by users of the Polars analytics library under the right mix of huge pages, MADV_FREE and reclaim pressure. On a GPU fleet the exposure is data-processing and training jobs under memcg pressure producing wrong results with no error anywhere.
Who can reach it
No attacker required - any local workload using MADV_FREE on PMD-mapped transparent huge pages, plus mprotect() or NUMA balancing, under memory pressure. Exposure is highest on nodes with THP enabled, memcg limits and NUMA balancing on, which describes most GPU hosts.
What to do
Update to a stable kernel that keeps _PAGE_DIRTY in pmd_modify()'s preserved mask and reboot each node; there is no runtime mitigation short of disabling transparent huge pages or avoiding MADV_FREE in affected workloads. The record lists stable commits only, no fixed release numbers, so map them onto your vendor kernel. Treat results produced by affected jobs as suspect.
References
Related entries
- Linux kernel powerpc/eeh: recursive locking hangs the EEH handler during PCI error recoveryCVE-2026-97948 · Linux kernel powerpc/eeh (recursive pci_rescan_remove_lock in eeh_rmv_device)Unscored
- Linux LIO iSCSI target: LUN_RESET on a WRITE_PENDING command deadlocks the target worker threadCVE-2026-97951 · Linux kernel SCSI target iSCSI frontend (aborted WRITE_PENDING dataout handling)Unscored
- Linux kernel vhost-vdpa: failed eventfd install leaves an ERR_PTR reachable by the config callbackCVE-2026-97993 · Linux kernel vhost-vdpa (ERR_PTR installed in v->config_ctx by VHOST_VDPA_SET_CONFIG_CALL)Unscored
- Linux kernel vhost-vdpa: queue size is not checked against the device maximum, giving an out-of-bounds descriptor readCVE-2026-97994 · Linux kernel vhost-vdpa (VHOST_SET_VRING_NUM validation)Unscored
- Linux kernel BPF verifier (bpf_loop nr_loops argument type): bpf_loop() declared nr_loops as ARG_ANYTHING, so aCVE-2026-98007 · Linux kernel BPF verifier (bpf_loop nr_loops argument type)Unscored
- Linux kernel BPF: bpf_btf_find_by_name_kind() can sleep in softirq context and install an fd into the interrupted taskCVE-2026-98046 · Linux kernel BPF helper bpf_btf_find_by_name_kind() (missing sleepable annotation)Unscored
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.