GPU VulnDB

Database/AI/ML frameworks & serving

llama.cpp llama-server: use-after-free and double free in the chat tool-call parser reachable from /completion

CVSS 9.2CVE-2026-107183AI/ML frameworks & servingcurated

Impact

An unauthenticated caller that can reach llama-server's HTTP API can submit a chat_parser in a POST /completion request that emits a tool-id after a tool-close tag, leaving current_tool dangling. The result is a use-after-free plus double free in the serving process: reliably a crash of the inference endpoint, and per the report a shapeable heap write primitive, so code execution in the server process is in play. llama-server usually runs with the GPUs mapped into it and often without its own authentication, trusting a gateway in front; a crash drops the model off the node and a heap write lands in a process that already holds /dev/nvidia* handles and whatever model weights and keys the deployment gave it. Versions before build b11393 are affected.

Who can reach it

Anyone who can send an HTTP POST to llama-server's /completion endpoint - no authentication required. In practice that is any tenant, pod or user inside the network boundary where the inference port is exposed, including through a proxy that forwards request fields verbatim.

What to do

Upgrade llama.cpp to build b11393 or later (fix commit dbe4c3ed42343f0a7ba0fd7e808ffeaca404d29f, PR 29942) and restart each llama-server process; no node reboot or driver work is involved, so this is a rolling daemon restart per replica. Until then, keep the endpoint off any untrusted network and reject or strip client-supplied chat_parser/grammar fields at the gateway.

References

Related entries

All AI/ML frameworks & serving entries

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.