GPU VulnDB

Database/AI/ML frameworks & serving

vLLM: unauthenticated /tokenize caller sets video max_frames and fps, exhausting frontend memory

CVSS 5.3CVE-2026-105758AI/ML frameworks & servingcurated

Impact

The Qwen2-VL and Qwen3-VL video backends accept media_io_kwargs.video.max_frames and media_io_kwargs.video.fps from the request with no server-side ceiling. A caller hitting /tokenize can make the sampler decode every frame it selects from attacker-supplied video, consuming frontend memory out of proportion to the request size and potentially killing the API process before scheduling or admission control ever sees the work. That is a crash of the shared serving endpoint on a multi-tenant GPU node: the model must reload and every in-flight request on that replica is lost, while the GPUs sit idle through the restart. The Rust frontend is unaffected because it rejects media_io_kwargs outright.

Who can reach it

Anyone who can reach the Python frontend's /tokenize endpoint; the advisory states the caller need not be authenticated. Only deployments serving a Qwen2-VL or Qwen3-VL video model through the Python frontend are exposed.

What to do

Upgrade to vLLM 0.30.0 and restart the serving processes; all versions from 0.24.0 up to 0.30.0 are affected. This is a frontend-only change - roll replicas behind the load balancer, no node drain needed. If upgrading has to wait, strip media_io_kwargs at the gateway or run the Rust frontend, which already rejects the field.

References

Related entries

All AI/ML frameworks & serving entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.