Linux kernel accel/qaic: unbounded response message walk in resp_worker() reads past the slab allocation
Impact
The QAIC driver's response worker walks a message received from the accelerator device without the bounds checks that decode_message() performs, because an earlier fix for DBC deactivation open-coded the walk. A malformed wire message from the card can push the walk past the end of the slab allocation, spin in an endless loop, or panic the host kernel. On a node with Cloud AI 100 cards this is a host-kernel fault reached from the device side, so the blast radius is the whole node and every tenant workload scheduled on it, not just the process holding the accelerator. Exposure depends on whether anything other than trusted firmware can shape the device's responses - a tenant that can load its own firmware or workload image onto the card is the realistic path. Nodes without QAIC accelerators do not load this driver at all.
Who can reach it
Requires a QAIC (Cloud AI 100) accelerator present and the driver loaded. The malformed input comes from the device, so the attacker must be able to influence what the card sends back - firmware or on-card workload control. No unauthenticated network path.
What to do
Take the stable-kernel fix that makes resp_worker() call decode_message() instead of duplicating the walk; it is in the linked stable commits across the maintained branches. Rolling it out means installing the patched kernel and rebooting each accelerator node, or at minimum unbinding and reloading the qaic module with the card idle - neither is cheap on a node holding running jobs, so drain first. No vendor advisory with fixed version strings is present in the record.
References
Related entries
- GPU / accelerator firmware (VBIOS, GSP, NVSwitch): GPU-resident firmware sits below the host OS and is not coveredNCVD-0000-012-gpu-accelerator-firmware-vbios-g · GPU / accelerator firmware (VBIOS, GSP, NVSwitch)Unscored
- NVIDIA Multi-Instance GPU (MIG) partitioning: MIG gives each instance its own SM slice, L2 slice, memory slice andNCVD-2020-001-nvidia-multi-instance-gpu-mig-pa · NVIDIA Multi-Instance GPU (MIG) partitioningUnscored
- NVIDIA Multi-Instance GPU (MIG) partitioning: MIG gives each instance its own SM slice, L2 slice, memory slice andNCVD-2020-003-nvidia-multi-instance-gpu-mig-pa · NVIDIA Multi-Instance GPU (MIG) partitioningUnscored
- Integrated GPU graphics data compression (Intel, AMD, Apple, Arm, Qualcomm, NVIDIA): GPUs apply data-dependent losslessNCVD-2023-003-integrated-gpu-graphics-data-com · Integrated GPU graphics data compression (Intel, AMD, Apple, Arm, Qualcomm, NVIDIA)Unscored
- NVIDIA Confidential Computing (H100/H200/B100/B200/GB200) - CC-DevTools operating mode: NVIDIA GPU confidentialNCVD-2023-004-nvidia-confidential-computing-h1 · NVIDIA Confidential Computing (H100/H200/B100/B200/GB200) - CC-DevTools operating modeUnscored
- Integrated GPU graphics data compression (Intel, AMD, Apple, Arm, Qualcomm, NVIDIA): GPUs apply data-dependent losslessNCVD-2023-005-integrated-gpu-graphics-data-com · Integrated GPU graphics data compression (Intel, AMD, Apple, Arm, Qualcomm, NVIDIA)Unscored
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.