
Cloud Vulnerability DB
A community-led vulnerabilities database
CVE-2026-0599 is an Uncontrolled Resource Consumption (DoS) vulnerability in HuggingFace text-generation-inference (TGI) that allows unauthenticated remote attackers to exhaust server resources by exploiting unbounded external image fetching during input validation in VLM (Vision Language Model) mode. Affected versions are all releases up to and including 3.3.6; the issue is resolved in version 3.3.7. It carries a CVSS v3.0 base score of 7.5 (High) (Feedly, EUVD). The vulnerability was disclosed on February 2, 2026, and was reported via the huntr bug bounty platform (huntr).
The root cause is CWE-400 (Uncontrolled Resource Consumption): when TGI operates in VLM mode, the router's input validation logic scans incoming prompts for Markdown-formatted image links using a regex () and then performs a synchronous (blocking) HTTP GET request to fetch each linked image. Prior to the fix, the entire response body was read into memory and cloned before decoding — with no size cap — meaning an attacker could point the server at an arbitrarily large or slow-draining HTTP resource (Feedly, GitHub Commit). Critically, this fetch occurs during validation and is triggered even if the request is subsequently rejected for exceeding token limits, so no valid model interaction is required. The fix introduces a configurable max_image_fetch_size parameter (default 1 GiB) that checks the Content-Length header and enforces a hard read limit via a take() reader, returning an ImageTooLarge error if exceeded (GitHub Commit).
Successful exploitation causes resource exhaustion on the TGI host, including network bandwidth saturation, memory inflation, and CPU overutilization, potentially crashing the host machine entirely. Because the default TGI deployment configuration lacks both authentication and memory usage limits, any network-reachable instance is vulnerable without any credentials. There is no confidentiality or integrity impact; the vulnerability is purely an availability (DoS) risk affecting the inference service and potentially the underlying host (Feedly, EUVD).
No public exploit code or active in-the-wild exploitation has been reported as of the available data. The EPSS score is approximately 0.0027 (0.27%), indicating a low current probability of exploitation in the wild (Feedly). The vulnerability is not listed in the CISA Known Exploited Vulnerabilities (KEV) catalog. However, the attack requires no authentication, no user interaction, and low complexity — any unauthenticated attacker with network access to a TGI VLM endpoint can trigger it. The vulnerability is detectable by Qualys scanner (detection ID 5007355) (Feedly).
/generate or /v1/chat/completions.{"inputs": " Describe this image.", "parameters": {}}./generate endpoint without any authentication headers. The router's validation logic will immediately begin a blocking HTTP GET to the attacker's URL./generate or /v1/chat/completions containing Markdown image syntax ( in the inputs field; requests originating from a single or small set of source IPs in rapid succession.text-generation-router); high CPU utilization on the inference host without a corresponding increase in legitimate inference requests; OOM (Out of Memory) killer events in system logs (dmesg, /var/log/syslog) targeting the TGI process.Upgrade text-generation-inference to version 3.3.7 or later, which introduces the max_image_fetch_size parameter (defaulting to 1 GiB) to enforce a hard limit on image fetch size during validation (GitHub Commit). As an immediate workaround for those unable to upgrade, place TGI behind an authenticated reverse proxy (e.g., nginx with HTTP Basic Auth) to prevent unauthenticated access, and configure network-level egress filtering to restrict outbound HTTP requests from the TGI host. Additionally, set OS-level memory limits (e.g., via cgroups or Docker --memory flags) on the TGI container or process to bound the blast radius of any exploitation attempt (Feedly).
The vulnerability was reported through the huntr bug bounty platform and assigned by @huntr_ai (huntr). A brief mention appeared on Mastodon via @thehackerwire shortly after disclosure. Coverage was picked up by security aggregators including Vulners, CIRCL, VulDB, and INCIBE-CERT, as well as a technical write-up published at infinitsec.net. No major vendor statements or significant researcher controversy have been noted beyond standard advisory publication.
Source: This report was generated using AI
Free Vulnerability Assessment
Evaluate your cloud security practices across 9 security domains to benchmark your risk level and identify gaps in your defenses.
Get a personalized demo
"Best User Experience I have ever seen, provides full visibility to cloud workloads."
"Wiz provides a single pane of glass to see what is going on in our cloud environments."
"We know that if Wiz identifies something as critical, it actually is."