CVE-2026-105922:
NixOS 취약성 분석 및 완화
개요
CVE-2026-105922 is a denial-of-service vulnerability in vllm-project vLLM affecting all versions up to and including 0.31.0. The flaw resides in the get_token_bin_counts_and_mask function within vllm/model_executor/layers/utils.py (the Penalty Handler component), and can be triggered remotely by an authenticated user. It was disclosed on October 6, 2026, with a public proof-of-concept exploit available at the time of disclosure. It carries a CVSS v3.1 base score of 4.3 (Medium) and a CVSS v4.0 base score of 2.1 (Low) (GitHub Advisory, Feedly).
기술적 세부 사항
The root cause is classified as CWE-404 (Improper Resource Shutdown or Release) and CWE-248 (Uncaught Exception). When a vLLM server is started with --enable-prompt-embeds, the InputBatch.add_request method never writes prompt token IDs for prompt_embeds requests (only setting is_token_ids=False), leaving those rows in token_ids_cpu uninitialized or containing stale data from previously reused request slots. The _make_prompt_token_ids_cpu_tensor() function subsequently copies these uninitialized positions, and get_token_bin_counts_and_mask passes them to a CUDA scatter_add_ operation without consulting is_token_ids. When a batch contains prompt_embeds requests of mixed lengths — specifically when the longest prompt is an embedding request — the stale token IDs used as scatter indices can fall outside the valid range [0, vocab_size], triggering a CUDA device-side assertion (scatter gather kernel index out of bounds) that kills the EngineCore process (GitHub Issue #57719, PoC Gist).
영향
Successful exploitation causes the vLLM EngineCore process to crash with a fatal CUDA error, rendering the entire inference server unresponsive — all in-flight and subsequent requests return HTTP 500 EngineDeadError and the server must be manually restarted. The impact is limited to availability; there is no confidentiality or integrity impact. Any client with API access to a server running with --enable-prompt-embeds can trigger this condition with a single burst of concurrent requests, making it a practical denial-of-service against shared or multi-tenant LLM serving deployments (GitHub Issue #57719, GitHub Advisory).
악용 가능성
A public proof-of-concept exploit script (repro_prompt_embeds_penalties_engine_crash.py) was published by the original reporter at the time of disclosure and is referenced in the GitHub advisory (PoC Gist). The EPSS score is approximately 0.303% (21st percentile), indicating a low but non-negligible probability of exploitation in the wild. No active in-the-wild exploitation or threat actor attribution has been reported, and the vulnerability is not listed in the CISA KEV catalog. Exploitation requires low privileges (authenticated API access) but no special configuration beyond the server being started with --enable-prompt-embeds (GitHub Advisory, Feedly).
착취 단계
- Identify a vulnerable target: Locate a vLLM server (version ≤ 0.31.0) started with the
--enable-prompt-embedsflag and accessible over the network. Obtain valid API credentials (low-privilege access is sufficient). - Obtain model embedding weights: Retrieve the target model's
embed_tokens.weighttensor from the Hugging Face cache or model files usingsafetensors. - Craft malicious requests: Construct at least two groups of
/v1/completionsrequests usingprompt_embeds(base64-encodedtorch.saveof embedding rows) with different prompt lengths — e.g., 900 tokens and 300 tokens — ensuring the longer group usesprompt_embeds. Include any penalty parameter (repetition_penalty,presence_penalty, orfrequency_penalty) in each request body. - Send concurrent burst: Dispatch ~16 concurrent HTTP POST requests to
/v1/completionswith the crafted bodies, mixing the two prompt lengths so the longest prompt in the batch is aprompt_embedsrequest. - Engine crash: The mixed-length batch causes
get_token_bin_counts_and_maskto scatter stale/uninitialized token IDs from reused request slots, triggering a CUDA out-of-bounds assertion that kills EngineCore. All subsequent requests return HTTP 500EngineDeadErroruntil the server is restarted (GitHub Issue #57719, PoC Gist).
타협의 징후
- Logs: CUDA device-side assertion errors in server logs:
ScatterGatherKernel.cu:203: Assertion 'idx_dim >= 0 && idx_dim < index_size && "scatter gather kernel index out of bounds"' failed; log entries showingEngineCore encountered a fatal errorandvllm.v1.engine.exceptions.EngineDeadError. - Network: Burst of concurrent POST requests to
/v1/completionswithprompt_embedsfields of varying sizes and penalty parameters (repetition_penalty,presence_penalty,frequency_penalty) set; subsequent requests returning HTTP 500 responses. - Process: vLLM EngineCore worker process terminating unexpectedly;
/healthendpoint returning non-200 responses after a burst ofprompt_embedsrequests;torch.AcceleratorError: CUDA error: device-side assert triggeredin process output (GitHub Issue #57719).
완화 및 해결 방법
The primary remediation is to upgrade vLLM to a version beyond 0.31.0, where a bugfix (commit 37a52fa) addressing the is_token_ids-unaware scatter in the penalty bin-count builder has been merged (GitHub Issue #57719, GitHub Advisory). As an immediate workaround, disable the --enable-prompt-embeds server flag if it is not required for your use case, as the vulnerability only manifests when this feature is active. Additionally, implement network access controls to restrict API access to trusted clients only, and consider rate-limiting concurrent requests to reduce the attack surface until a patched version is deployed.
커뮤니티 반응
The vulnerability was originally reported by researcher Yunzez via a GitHub issue on September 19, 2026, with a detailed root cause analysis and a self-contained reproducer script (GitHub Issue #57719). A vLLM maintainer (vadiklyutiy) initially could not reproduce the issue on a nightly build but confirmed the underlying code defect was present; the reporter subsequently confirmed reproduction on released v0.27.1 and v0.30. A fix was committed by Cross2pro on September 27, 2026, and the CVE was formally published on October 6, 2026. No significant broader media coverage or social media discussion has been observed beyond the GitHub issue thread.
추가 자료
근원: 이 보고서는 AI를 사용하여 생성되었습니다.
관련 NixOS 취약점:
무료 취약성 평가
클라우드 보안 태세를 벤치마킹합니다
9개의 보안 도메인에서 클라우드 보안 관행을 평가하여 위험 수준을 벤치마킹하고 방어의 허점을 식별합니다.
추가 Wiz 리소스
맞춤형 데모 받기
맞춤형 데모 신청하기
"내가 본 최고의 사용자 경험은 클라우드 워크로드에 대한 완전한 가시성을 제공합니다."
"Wiz는 클라우드 환경에서 무슨 일이 일어나고 있는지 볼 수 있는 단일 창을 제공합니다."
"우리는 Wiz가 무언가를 중요한 것으로 식별하면 실제로 중요하다는 것을 알고 있습니다."