NVIDIA TensorRT-LLM: Concurrent requests race inside TensorRT-LLM and reach data tampering with a changed scope. On a
Impact
Concurrent requests race inside TensorRT-LLM and reach data tampering with a changed scope. On a shared LLM serving tier, a race between concurrent requests means one tenant's request can affect another's - the mechanism by which response bleed-through happens.
Who can reach it
Network, low privileges. Any client able to issue concurrent requests to the serving endpoint, which is every client.
What to do
Upgrade TensorRT-LLM per bulletin 5805 and roll the serving deployment. Cost: rolling restart. Until patched, the compensating control is reducing concurrency or dedicating an engine per tenant, both of which cost throughput.
References
Related entries
- TensorRT-LLM: Code exec via insecure deserialization on model loadCVE-2026-24220 · TensorRT-LLMMedium
- TensorRT-LLM: Missing authentication in configuration processingCVE-2026-24259 · TensorRT-LLMMedium
- NVIDIA GPU driver for Linux: race condition in the kernel mode layer leads to an out-of-bounds writeCVE-2026-47565 · NVIDIA GPU Display Driver for Linux (kernel mode layer race condition)Medium
- NVIDIA GPU driver: use-after-free in the kernel module, scored as requiring physical accessCVE-2026-47586 · NVIDIA GPU Display Driver kernel module (use-after-free)Medium
- NVIDIA vGPU Manager (vGPU plugin): Time-of-check to time-of-use on a shared resource between guest and host plugin. ACVE-2020-5969 · NVIDIA vGPU Manager (vGPU plugin)Medium
- NVIDIA vGPU Manager (vGPU plugin): The plugin keeps using a resource it validated after the guest has changed it - aCVE-2021-1061 · NVIDIA vGPU Manager (vGPU plugin)Medium
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.