Database/AI/ML frameworks & serving
vLLM: unauthenticated /tokenize caller sets video max_frames and fps, exhausting frontend memory
Impact
The Qwen2-VL and Qwen3-VL video backends accept media_io_kwargs.video.max_frames and media_io_kwargs.video.fps from the request with no server-side ceiling. A caller hitting /tokenize can make the sampler decode every frame it selects from attacker-supplied video, consuming frontend memory out of proportion to the request size and potentially killing the API process before scheduling or admission control ever sees the work. That is a crash of the shared serving endpoint on a multi-tenant GPU node: the model must reload and every in-flight request on that replica is lost, while the GPUs sit idle through the restart. The Rust frontend is unaffected because it rejects media_io_kwargs outright.
Who can reach it
Anyone who can reach the Python frontend's /tokenize endpoint; the advisory states the caller need not be authenticated. Only deployments serving a Qwen2-VL or Qwen3-VL video model through the Python frontend are exposed.
What to do
Upgrade to vLLM 0.30.0 and restart the serving processes; all versions from 0.24.0 up to 0.30.0 are affected. This is a frontend-only change - roll replicas behind the load balancer, no node drain needed. If upgrading has to wait, strip media_io_kwargs at the gateway or run the Rust frontend, which already rejects the field.
References
Related entries
- vLLM: GLMGA video backend builds an attacker-sized frame-index list, starving the shared media loaderCVE-2026-105760 · vLLM GLMGA video backend (pre-decode frame-index list from request media_io_kwargs)Medium
- BentoML OpenLLM 0.6.30 (async_run_command in src/openllm/common.py): A model repository directory name flows unescapedCVE-2026-15035 · BentoML OpenLLM 0.6.30 (async_run_command in src/openllm/common.py)Medium
- JupyterHub: unauthenticated logins write unbounded usernames to the log, exhausting storageCVE-2026-54338 · JupyterHub form-based login authenticators (failed-login logging)Medium
- vLLM: malformed JSON to the OpenAI-compatible endpoints returns server paths and versionsCVE-2026-73555 · vLLM OpenAI-compatible API server (validation_exception_handler, sanitize_message)Medium
- vLLM: attacker-supplied structured-output regex pins a CPU core and stalls the engine pathCVE-2026-73556 · vLLM structured outputs, lm-format-enforcer backend (structured_outputs.regex)Medium
- vLLM: integer overflow in the activation CUDA kernel leaks another batched request's outputCVE-2026-73558 · vLLM CUDA activation kernels (act_and_mul_kernel, activation_kernels.cu)Medium
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.