CVE-2026-105758:
NixOS 취약성 분석 및 완화
개요
CVE-2026-105758 is a denial-of-service vulnerability in vLLM, an inference and serving engine for large language models, caused by missing server-side resource limits on video sampling parameters in the Qwen2-VL and Qwen3-VL backends. Affecting vLLM versions >= 0.24.0 and < 0.30.0, the flaw allows unauthenticated remote attackers to exhaust API server memory by submitting arbitrarily large media_io_kwargs.video.max_frames and fps values to the /tokenize endpoint. The vulnerability was reported by Eva Crystal / 0xiviel (XSource Security), privately disclosed on September 23, 2026, and published to the GitHub Advisory Database on October 5, 2026. It carries a CVSS v3.1 base score of 5.3 (Medium) (Github Advisory, Feedly).
기술적 세부 사항
The root cause is CWE-770 (Allocation of Resources Without Limits or Throttling): the Qwen2VLVideoBackend and Qwen3VLVideoBackend classes in vllm/multimodal/video.py read max_frames and fps directly from the request-supplied media_io_kwargs dictionary without enforcing any server-side ceiling. Unlike other backends (e.g., GLMGAVideoBackend, which caps at _MAX_FRAMES=640 and _MAX_FPS=30), the Qwen samplers ignore the num_frames key entirely and use max_frames as the sole upper bound — which an attacker can set to an arbitrarily large value (e.g., 1e9). The vulnerable code path is reachable via unauthenticated HTTP POST to /tokenize and /invocations, as these endpoints are excluded from AuthenticationMiddleware even when --api-key is configured. An attacker can also explicitly select the qwen3_vl video backend on any deployment via the video_backend request field, making the attack applicable beyond Qwen-specific deployments (Github Advisory, vLLM PR #56729).
영향
Successful exploitation causes memory exhaustion of the vLLM API server process, potentially terminating it before scheduling or admission control can intervene, which affects all tenants sharing the instance. Measured impact shows that 74 extra bytes of JSON in a 1.43 MiB request body drove server peak RSS from 2,271 MiB to 13,629 MiB; at larger video sizes, the process can be OOM-killed entirely. There is no confidentiality or integrity impact — the vulnerability is purely an availability (denial-of-service) issue. The --limit-mm-per-prompt flag does not mitigate this, as it bounds media items rather than frames within an item (Github Advisory).
악용 가능성
No public proof-of-concept exploit code is known to exist, and there is no evidence of in-the-wild exploitation as of the advisory publication date (Feedly). The EPSS score is 0.0, and the vulnerability is not listed in the CISA Known Exploited Vulnerabilities (KEV) catalog. However, the attack requires no authentication, no special privileges, and no user interaction, making it trivially exploitable against any exposed vLLM instance serving Qwen2-VL or Qwen3-VL models with the Python frontend.
착취 단계
- Reconnaissance: Identify internet-facing vLLM API servers (versions 0.24.0–0.29.x) using tools like Shodan or Censys, looking for services exposing the vLLM OpenAI-compatible API (default port 8000). Confirm the deployment serves a Qwen2-VL or Qwen3-VL model, or plan to specify
video_backend: qwen3_vlexplicitly. - Prepare a high-frame-count video payload: Craft or obtain a video file with a large number of frames (e.g., a long, high-FPS video). Encode it as a
data:URI or host it at an attacker-controlled URL to bypass theVLLM_MAX_MEDIA_DOWNLOAD_SIZE_MBlimit (which does not apply todata:URIs). - Craft the malicious request: Construct an HTTP POST request to
/tokenize(or/invocations) with a JSON body that includes the video payload and setsmedia_io_kwargs.video.max_framesto a very large value (e.g.,1000000000) andmedia_io_kwargs.video.fpsto a high value (e.g.,1000000). Optionally includevideo_backend: qwen3_vlto force the vulnerable sampler on non-Qwen deployments. - Submit the request: Send the unauthenticated POST request to the target server. The frontend will decode every frame selected by the attacker-controlled parameters before any scheduling or admission control runs.
- Achieve denial of service: The server's memory consumption spikes dramatically (measured: 2,271 MiB → 13,629 MiB with 74 extra bytes of JSON), potentially triggering an OOM kill of the API process and denying service to all users of the instance (Github Advisory, vLLM PR #56729).
타협의 징후
- Network: Unauthenticated HTTP POST requests to
/tokenizeor/invocationsendpoints containingmedia_io_kwargsfields with abnormally largemax_frames(e.g., > 768) orfps(e.g., > 30) values; large request bodies (>1 MiB) to these endpoints from unexpected sources. - Logs: vLLM access logs showing repeated or large POST requests to
/tokenizeor/invocationswithmedia_io_kwargsin the body; Python OOM kill messages in system logs (e.g.,Out of memory: Killed process <pid> (python)). - Process: Sudden spike in RSS/memory usage of the vLLM Python process followed by process termination; unexpected restart of the vLLM API server process.
- File System: No specific file artifacts expected, as the attack is entirely in-memory and transient.
완화 및 해결 방법
Upgrade vLLM to version 0.30.0 or later, which applies class-level _MAX_FRAMES=768 and _MAX_FPS=30 caps to both Qwen2VLVideoBackend and Qwen3VLVideoBackend via commit ea723c8 (PR #56729), preventing request-level values from exceeding these ceilings (vLLM Release v0.30.0, vLLM PR #56729). If immediate patching is not possible, restrict network access to the /tokenize and /invocations endpoints using a reverse proxy or firewall rules, and implement external rate limiting and resource quotas on API requests containing video data. Note that configuring --api-key alone does not protect /tokenize or /invocations, as these endpoints are excluded from authentication middleware (Github Advisory).
커뮤니티 반응
The vulnerability was reported by Eva Crystal / 0xiviel (XSource Security) and coordinated by jperezdealgaba, who also authored the fix. The vLLM maintainer Isotr0py noted in the PR review that a more robust long-term solution would be making fps and max_frames static (non-overridable) kwargs, leaving the class-level cap as an interim measure pending further refactoring (vLLM PR #56729). Red Hat tracked the issue via Bugzilla and the advisory was picked up by GitLab's advisory database shortly after publication (Github Advisory).
추가 자료
근원: 이 보고서는 AI를 사용하여 생성되었습니다.
관련 NixOS 취약점:
무료 취약성 평가
클라우드 보안 태세를 벤치마킹합니다
9개의 보안 도메인에서 클라우드 보안 관행을 평가하여 위험 수준을 벤치마킹하고 방어의 허점을 식별합니다.
추가 Wiz 리소스
맞춤형 데모 받기
맞춤형 데모 신청하기
"내가 본 최고의 사용자 경험은 클라우드 워크로드에 대한 완전한 가시성을 제공합니다."
"Wiz는 클라우드 환경에서 무슨 일이 일어나고 있는지 볼 수 있는 단일 창을 제공합니다."
"우리는 Wiz가 무언가를 중요한 것으로 식별하면 실제로 중요하다는 것을 알고 있습니다."