CVE-2026-4229: 
Python vulnerability analysis and mitigation

Overview

CVE-2026-4229 is a SQL injection vulnerability in the remove_training_data function of vanna-ai/vanna, affecting versions up to and including 2.0.2. The flaw resides in src/vanna/legacy/google/bigquery_vector.py, where the id argument is interpolated directly into a raw SQL DELETE statement without sanitization or parameterization. Disclosed on March 16, 2026, the vulnerability was discovered via static code audit of Vanna v2.0.2 (commit 365d061) and reported by researcher YLChen-007; the vendor did not respond to early disclosure. It carries a CVSS v3.1 base score of 7.3 (High) and a CVSS v4.0 base score of 5.5 (Medium) (Feedly, GitHub Gist).

Technical details

The root cause is CWE-89 (SQL Injection) — the remove_training_data method in BigQuery_VectorStore constructs a DELETE query using a Python f-string: query = f"DELETE FROM \{self.table_id}` WHERE id = '{id}'", directly embedding the user-supplied idwithout escaping or parameterization ([GitHub Gist](https://gist.github.com/YLChen-007/b4f326eaecc29b192cf93dc5d6bc0623)). The attack vector is network-accessible via HTTP POST to/api/v0/remove_training_data, where the id field is extracted from the JSON request body with no type, format, or content validation. Critically, the default authentication backend (NoAuth) unconditionally returns Trueforis_logged_in(), meaning no credentials are required to reach the vulnerable endpoint. All other vector store backends (pgvector, Oracle, ChromaDB, OpenSearch) use parameterized queries or safe APIs — only the BigQuery backend is affected. A proof-of-concept payload such as ' OR '1'='1` causes the DELETE to match all rows in the training data table (GitHub Gist).

Impact

Successful exploitation allows an unauthenticated remote attacker to delete all rows in the BigQuery training data table with a single crafted request, causing irreversible loss of curated SQL examples, DDL schemas, and documentation used to train the AI assistant (GitHub Gist). Attackers can also perform surgical deletion of specific training data categories to silently degrade AI output quality. Depending on BigQuery error handling and the permissions of the connected service account, error-based or blind SQL injection techniques may enable data exfiltration from the training dataset or other tables within the same BigQuery dataset, impacting confidentiality in addition to integrity and availability (Feedly).

Exploitability

A public proof-of-concept exploit was published by researcher YLChen-007 on GitHub Gist on February 26, 2026, prior to CVE assignment (GitHub Gist). The CVSS v4.0 exploit maturity is rated PROOF_OF_CONCEPT, and no authentication is required under the default NoAuth configuration, making exploitation trivial for any network-reachable attacker (Feedly). The EPSS score is approximately 0.028% (0.000280), indicating low but non-zero probability of active exploitation in the near term. No threat actor attribution, exploit kit integration, or CISA KEV catalog listing has been reported as of the time of this report.

Exploitation steps

  1. Reconnaissance: Identify internet-facing Vanna Flask API instances (default port 8084) using Shodan, Censys, or similar tools. Confirm the BigQuery vector store backend is in use by probing the /api/v0/get_training_data endpoint.
  2. Verify unauthenticated access: Send a GET request to /api/v0/get_training_data. If the default NoAuth backend is active, the server responds with training data without requiring credentials.
  3. Craft the injection payload: Prepare a JSON POST body with a malicious id value, e.g., {"id": "' OR '1'='1"}.
  4. Send the exploit request: Issue POST /api/v0/remove_training_data with the crafted payload and Content-Type: application/json. The server constructs: DELETE FROM \project.dataset.training_data` WHERE id = '' OR '1'='1'`.
  5. Achieve mass deletion: BigQuery executes the injected SQL, deleting all rows in the training data table. Confirm by re-querying /api/v0/get_training_data and observing an empty result set.
  6. Optional — data exfiltration: Use error-based or blind SQL injection techniques (e.g., time-based inference via BigQuery-specific functions) to enumerate and extract data from the training dataset or adjacent tables accessible to the BigQuery service account (GitHub Gist).

Indicators of compromise

  • Network: Unexpected HTTP POST requests to /api/v0/remove_training_data from external or unknown IP addresses; repeated requests with unusual or non-UUID id field values (e.g., containing single quotes, OR, =, or SQL keywords).
  • Logs: Flask/web server access logs showing POST requests to /api/v0/remove_training_data with JSON bodies containing SQL metacharacters; BigQuery audit logs recording DELETE queries with WHERE clauses matching all rows (e.g., WHERE id = '' OR '1'='1').
  • BigQuery Audit: Sudden large-scale DELETE operations on the training data table in BigQuery's Data Access audit logs, especially from the Vanna service account at unexpected times.
  • Application Behavior: Vanna AI assistant returning degraded or empty responses due to missing training data; /api/v0/get_training_data returning an empty dataset after previously returning populated results (GitHub Gist).

Mitigation and workarounds

No vendor patch has been released as of the disclosure date, as the vendor did not respond to the researcher's disclosure (Feedly). Immediate mitigations include: (1) replacing the default NoAuth backend with a proper authentication implementation to restrict access to the /api/v0/remove_training_data endpoint; (2) patching bigquery_vector.py to use BigQuery parameterized queries (e.g., self.conn.query(query, job_config=bigquery.QueryJobConfig(query_parameters=[bigquery.ScalarQueryParameter('id', 'STRING', id)]))) instead of f-string interpolation; (3) applying network-level controls (firewall rules, VPN) to restrict access to the Vanna Flask API to trusted hosts only; and (4) validating the id parameter against a strict UUID format before passing it to the database layer (GitHub Gist).

Community reactions

The vulnerability was disclosed publicly via a GitHub Gist by researcher YLChen-007 on February 26, 2026, with a detailed static code audit and data flow trace (GitHub Gist). The researcher noted that the vendor was contacted early but did not respond in any way, which is reflected in the CVE description (Feedly). Coverage has appeared on threat intelligence aggregators including VulDB, ENISA EUVD, and radar.offseq.com, but no significant broader media coverage or notable community debate has been observed.

Additional resources


Source: This report was generated using AI

Related Python vulnerabilities:

CVE ID

Severity

Score

Technologies

Component name

CISA KEV exploit

Has fix

Published date

GHSA-v2f8-6655-7grjCRITICAL10
  • Python logoPython
  • vibe-trading-ai
NoYesOct 02, 2026
CVE-2026-105782HIGH7.5
  • Python logoPython
  • scrapy
NoYesOct 06, 2026
GHSA-v853-p72q-4cfwHIGH7.5
  • Python logoPython
  • quart
NoYesOct 05, 2026
CVE-2026-105751MEDIUM6.9
  • Python logoPython
  • docling
NoYesOct 05, 2026
CVE-2026-105750MEDIUM5.9
  • Python logoPython
  • docling
NoYesOct 05, 2026

Free Vulnerability Assessment

Benchmark your Cloud Security Posture

Evaluate your cloud security practices across 9 security domains to benchmark your risk level and identify gaps in your defenses.

Request assessment

Get a personalized demo

Ready to see Wiz in action?

"Best User Experience I have ever seen, provides full visibility to cloud workloads."
David EstlickCISO
"Wiz provides a single pane of glass to see what is going on in our cloud environments."
Adam FletcherChief Security Officer
"We know that if Wiz identifies something as critical, it actually is."
Greg PoniatowskiHead of Threat and Vulnerability Management