GPU VulnDB

Database/Firmware, BMC & network fabric

Linux kernel hfi1: PIO credit-return mmap installs a bad pfn and oopses the node

UnscoredCVE-2026-98216Firmware, BMC & network fabriccurated

Impact

hfi1_file_mmap()'s PIO_CRED case adds a byte offset to a struct credit_return *, so the computed address lands 256 KiB or 512 KiB past a 10240-byte allocation whenever the hardware send context index reaches 64 or 128. remap_pfn_range() then installs a frame above MAXPHYADDR and the first user read takes a "Bad pagetable" oops, taking the node and every job on it down. Because low send-context indices give offset zero, it is intermittent - whether a job trips it depends on which context it gets. Correcting only the arithmetic is not enough: dma_mmap_coherent() selects the page with vm_pgoff, which hfi1 sets to 0, so user space receives the wrong credit-return page and send PIO stalls forever. On a shared HPC node with Omni-Path, that is an unprivileged crash or hang reachable by any tenant job that opens a PSM2 endpoint.

Who can reach it

Local user with access to the hfi1 character device on a node with Omni-Path 100 fabric - any tenant or batch job allowed to run MPI/PSM2. No authentication beyond device access, no remote path.

What to do

Pick up a stable kernel containing the fix (commits linked in the record), then drain jobs and reboot each Omni-Path node; a kernel driver fix cannot be applied to a running node. No vendor advisory with a product version is present in the record. As an interim workaround the record shows PSM2_SDMA=2 (send PIO disabled) completing normally while the PIO path hangs - the crash itself happens at endpoint open, so that is not a reliable mitigation.

References

Related entries

All Firmware, BMC & network fabric entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.