Database/AI/ML frameworks & serving
BentoML (bundled Gradio app, multipart boundary handling): Appending a long run of characters to a multipart boundary
Impact
Appending a long run of characters to a multipart boundary makes the server chew through each one, burning CPU until the model endpoint stops answering. An unauthenticated request takes a served model offline, and on a GPU node the pod keeps holding its device allocation while it is useless.
Who can reach it
Any unauthenticated client that can reach the BentoML serving endpoint. Single crafted HTTP request, no session.
What to do
Upgrade BentoML past 1.4.5 and restart the serving pods. Put a request-size and rate limit in front of the model endpoint, and set CPU limits on the serving container so the abuse cannot spread to co-located pods.
References
Related entries
- Ollama (GGUF import): Crafted GGUF causes DoS on model createCVE-2025-0312 · Ollama (GGUF import)High
- NVIDIA Triton (Python backend): Information disclosure from the Python backendCVE-2025-23320 · NVIDIA Triton (Python backend)High
- vLLM (weight loading): `hf_model_weights_iterator` uses `torch.load` without `weights_only`CVE-2025-24357 · vLLM (weight loading)High
- vLLM (ZeroMQ): DoS and data exposure over ZeroMQCVE-2025-30202 · vLLM (ZeroMQ)High
- vLLM (HTTP GET): Single HTTP GET crashes the serverCVE-2025-48956 · vLLM (HTTP GET)High
- PyTorch (`torch.linalg.lu`): DoS on slice operationCVE-2025-55551 · PyTorch (`torch.linalg.lu`)High
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.