Linux drm/amdkfd: integer underflow in the EOP ring size log calculation
Impact
The low 6 bits of cp_hqd_eop_control hold the base-2 log of the EOP ring size, and the calculation could in theory underflow if the ring buffer size were very small, writing a wrong size into a hardware queue descriptor. Upstream states explicitly that in practice the ring buffer size cannot be less than 4096, so the underflow is not reachable in any configuration the driver actually produces - this is defensive hardening of the AMD compute queue setup path, not a usable attack on a GPU node. The same defect was fixed in two places and AMD/upstream assigned two ids: this entry covers both CVE-2026-98176 (the ffs() form) and CVE-2026-98177 (the order_base_2() form) - one flaw class, one component, one remediation. Carry it as hygiene in your kernel currency tracking rather than as a reason to open a window.
Who can reach it
Local only, and no demonstrated path: it would require a KFD queue created with an EOP ring buffer smaller than the driver permits. Holding /dev/kfd is a prerequisite, but upstream does not describe a reachable trigger.
What to do
Fixed in stable; take the kernel carrying commits 82e1b8299c6d / c883d0a132d4 (98176) and 287a34c4d712 / 8ee521b8b189 (98177). Both are kernel-side, so they land with a kernel update and a node reboot - no separate action for the second id.
Also covers 1 CVE
The vendor assigned a separate id to each affected code path. They share this advisory, this score and this fix, so they are one entry here.
References
Related entries
- Linux drm/amdgpu: NULL dereference during GPU reset when KFD init failed after probeCVE-2026-98178 · Linux kernel drm/amdgpu (amdgpu_amdkfd_clear_kfd_mapping on GPU reset)Unscored
- Linux kernel amdgpu: register BAR mapping leaks on every driver unload or hot-unplugCVE-2026-98179 · Linux kernel amdgpu (register BAR iounmap on device removal)Unscored
- amdgpu nbio_v7_9: NULL dereference in hard IRQ when a RAS interrupt arrives before RAS late_initCVE-2026-98270 · Linux kernel amdgpu nbio_v7_9 (RAS controller interrupt handler)Unscored
- GPU / accelerator firmware (VBIOS, GSP, NVSwitch): GPU-resident firmware sits below the host OS and is not coveredNCVD-0000-012-gpu-accelerator-firmware-vbios-g · GPU / accelerator firmware (VBIOS, GSP, NVSwitch)Unscored
- NVIDIA Multi-Instance GPU (MIG) partitioning: MIG gives each instance its own SM slice, L2 slice, memory slice andNCVD-2020-001-nvidia-multi-instance-gpu-mig-pa · NVIDIA Multi-Instance GPU (MIG) partitioningUnscored
- NVIDIA Multi-Instance GPU (MIG) partitioning: MIG gives each instance its own SM slice, L2 slice, memory slice andNCVD-2020-003-nvidia-multi-instance-gpu-mig-pa · NVIDIA Multi-Instance GPU (MIG) partitioningUnscored
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.