CVE-2026-28350: 
Python vulnerability analysis and mitigation

Overview

CVE-2026-28350 is a <base> tag injection vulnerability in the lxml_html_clean (lxml-html-clean) Python library that allows attackers to hijack the resolution of all relative URLs on a sanitized HTML page. The flaw affects all versions up to and including 0.4.3, with version 0.4.4 containing the fix. It was published on March 2, 2026 via a GitHub Security Advisory and assigned a CVSS v3.1 base score of 6.1 (Moderate) (Github Advisory, GHSA Advisory).

Technical details

The root cause is that the <base> HTML tag was not included in the page_structure kill set of the Cleaner class (CWE-116: Improper Encoding or Escaping of Output). While the default Cleaner configuration with page_structure=True removes <html>, <head>, and <title> tags, it did not specifically handle <base>, allowing it to pass through unsanitized. Although the HTML specification requires <base> to reside inside <head>, browsers accept and process <base> tags placed anywhere in the document, meaning an attacker can inject a <base href="https://evil.com"> tag into user-supplied HTML that is then sanitized and rendered by the application. The fix (commit 9c5612c) adds base to the kill_tags set whenever head is being removed (GHSA Advisory, Patch Commit).

Impact

Successful exploitation enables three attack vectors against users of applications that render the sanitized HTML output: (1) Phishing & Credential Theft — relative navigation links (e.g., <a href="/login">) and form actions (e.g., <form action="/auth">) are silently redirected to an attacker-controlled domain; (2) Stored XSS Escalation — if the application loads JavaScript via relative paths (e.g., <script src="assets/app.js">), the browser fetches the script from the attacker's server, enabling full stored cross-site scripting and account compromise; (3) UI Defacement — relative image and stylesheet references are loaded from the attacker's server, enabling visual redressing. Availability is not impacted, but confidentiality and integrity are both affected at a low-to-moderate level (Github Advisory).

Exploitability

A proof-of-concept is publicly available in the GitHub Security Advisory itself, demonstrating that calling clean_html('<base href="https://evil.com">Account') preserves the malicious <base> tag in the output. There is no evidence of in-the-wild exploitation at this time. The EPSS score is approximately 0.016% (0.028% per Feedly), placing it in the 4th percentile for exploitation likelihood. The vulnerability is not listed in the CISA Known Exploited Vulnerabilities (KEV) catalog. Exploitation requires user interaction (a victim must visit or view the page containing the injected tag) (Github Advisory, GHSA Advisory).

Exploitation steps

  1. Identify a target application: Find a web application that uses lxml_html_clean (versions ≤ 0.4.3) to sanitize user-supplied HTML before rendering it to other users (e.g., a comment system, rich-text editor, or email preview feature).
  2. Craft the malicious payload: Construct an HTML snippet containing a <base> tag pointing to an attacker-controlled domain, such as: <base href="https://attacker.com/"><a href="/login">Click here to log in</a>.
  3. Inject the payload: Submit the crafted HTML through the application's user input mechanism (e.g., a comment field, profile bio, or message body) that is processed by clean_html() or the Cleaner class.
  4. Verify bypass: Confirm that the sanitized output still contains the <base> tag (the library does not strip it in vulnerable versions), meaning the injection survived sanitization.
  5. Trigger victim interaction: Wait for or socially engineer a victim user to view the page containing the injected content. The browser will resolve all relative URLs against the attacker's domain.
  6. Harvest credentials or execute scripts: Depending on the application's structure — if it uses relative paths for login forms or JavaScript assets — the attacker can capture submitted credentials via a phishing page on their server, or serve malicious JavaScript to achieve stored XSS (GHSA Advisory).

Indicators of compromise

  • Logs: Application logs showing user-submitted HTML content containing <base tags or href attributes pointing to external domains in fields that accept rich text or HTML input.
  • Network: Outbound browser requests from users' sessions to unexpected external domains for resources that should be served locally (e.g., JavaScript files, CSS, images, or form POST submissions going to non-application domains).
  • Application Output: Rendered HTML pages containing a <base href="..."> element with an external domain in the href attribute, particularly when the application uses lxml_html_clean for sanitization.
  • File System / Dependencies: Presence of lxml_html_clean or lxml-html-clean version 0.4.3 or earlier in the Python environment (pip show lxml-html-clean).

Mitigation and workarounds

The primary remediation is to upgrade lxml-html-clean to version 0.4.4 or later, which adds <base> to the kill tag set whenever <head> is removed, preventing the injection entirely (Patch Commit, Github Advisory). If immediate patching is not possible, implement a strict Content Security Policy (CSP) header — particularly base-uri 'self' or base-uri 'none' — to prevent browsers from honoring injected <base> tags. Additionally, validate and strip <base> tags from all user-supplied HTML as a defense-in-depth measure. Packages for Fedora and openSUSE have also been updated (openSUSE Advisory).

Community reactions

Red Hat tracked the vulnerability via Bugzilla (bug #2444963) and published a CVE advisory page, indicating it affects packages in their ecosystem (Red Hat CVE). The openSUSE security team issued security announcements for updated python311-lxml-html-clean packages. Tenable released a Nessus detection plugin (plugin ID 301400) for the vulnerability. No significant social media controversy or notable researcher commentary beyond the advisory itself has been observed.

Additional resources

Linux Distribution fix status

Fix availability across major Linux distributions and their releases.

Debian

Fixed

bookworm

lxml: 4.9.2-1+deb12u1

Fixed

sid

lxml-html-clean: 0.4.4-1

Fixed

trixie

lxml-html-clean: 0.4.4-1~deb13u1

Fixed

Ubuntu

Unknown

devel

lxml-html-clean

Unknown

noble

lxml-html-clean

Unknown

noble (esm-apps)

lxml-html-clean

Unknown

resolute

lxml-html-clean

Unknown

resolute (esm-apps)

lxml-html-clean

Unknown

Source: This report was generated using AI

Related Python vulnerabilities:

CVE ID

Severity

Score

Technologies

Component name

CISA KEV exploit

Has fix

Published date

GHSA-v2f8-6655-7grjCRITICAL10
  • Python logoPython
  • vibe-trading-ai
NoYesOct 02, 2026
CVE-2026-105782HIGH7.5
  • Python logoPython
  • scrapy
NoYesOct 06, 2026
GHSA-v853-p72q-4cfwHIGH7.5
  • Python logoPython
  • quart
NoYesOct 05, 2026
CVE-2026-105751MEDIUM6.9
  • Python logoPython
  • docling
NoYesOct 05, 2026
CVE-2026-105750MEDIUM5.9
  • Python logoPython
  • docling
NoYesOct 05, 2026

Free Vulnerability Assessment

Benchmark your Cloud Security Posture

Evaluate your cloud security practices across 9 security domains to benchmark your risk level and identify gaps in your defenses.

Request assessment

Get a personalized demo

Ready to see Wiz in action?

"Best User Experience I have ever seen, provides full visibility to cloud workloads."
David EstlickCISO
"Wiz provides a single pane of glass to see what is going on in our cloud environments."
Adam FletcherChief Security Officer
"We know that if Wiz identifies something as critical, it actually is."
Greg PoniatowskiHead of Threat and Vulnerability Management