GPU VulnDB

Database/AI/ML frameworks & serving

vLLM: crafted request to the Gemma4 unified parser crashes the inference server

CVSS 5.5CVE-2026-103241AI/ML frameworks & servingcurated

Impact

A remote caller with no authentication can send input that the Gemma4 unified parser mishandles, taking the vLLM server process down. On a GPU fleet that means the model replica stops serving and its GPUs sit idle until the process is restarted; in a shared cluster an attacker who can reach any public or tenant-facing vLLM endpoint can keep knocking replicas over and effectively deny other tenants the accelerators they are paying for. The record notes an exploit has been published, so this is cheap to trigger. Availability only - the record claims no confidentiality or integrity impact.

Who can reach it

Anyone who can send a request to the vLLM HTTP endpoint. No authentication required per the record's CVSS vector (PR:N), so any network path to the serving port - a tenant pod, an ingress, a misplaced LoadBalancer - is enough.

What to do

Upgrade vLLM to 0.29.1rc0 or later (fix commit 3439bad37e68ba9755a46f4f6b44a4aeaf1f60a9) and restart each serving process; replicas can be rolled one at a time behind the router, so no node drain or reboot is needed. Until then, keep vLLM endpoints off untrusted networks and front them with authentication and request rate limits. The record gives no configuration-only mitigation.

References

Related entries

All AI/ML frameworks & serving entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.