Marque uma demonstração personalizada
"A melhor experiência do usuário que eu já vi, fornece visibilidade total para cargas de trabalho na nuvem."
"A Wiz fornece um único painel de vidro para ver o que está acontecendo em nossos ambientes de nuvem."
"Sabemos que se a Wiz identifica algo como crítico, na verdade é."
Wiz Research | Technical Report
Security Engineers & Red Teamers looking to choose the right AI agent-model pairing for vulnerability discovery, penetration testing, and exploit development workflows.
Security Leaders & CISOs who need to assess the real-world offensive capabilities of AI agents to inform risk posture, tool investments, and responsible deployment decisions.
AI & ML Engineers building or evaluating agentic security tools who want to understand how agent architecture and model selection jointly determine performance on complex, multi-step tasks.
Security Researchers benchmarking frontier models and seeking a reproducible, deterministic evaluation framework that spans the full offensive lifecycle.
Full benchmark results across 25 agent-model combinations: 4 agents (Claude Code, Gemini CLI, OpenCode, Codex) × 8 models (Claude Opus 4.6, Opus 4.5, Sonnet 4.5, Haiku 4.5, GPT-5.2, Gemini 3 Pro, Gemini 3 Flash, Grok 4), with scores and runtimes for every pairing.
257 real-world challenges across five offensive security categories: Zero-Day Discovery, CVE Detection (176 challenges across Python, Go, and Java), API Security, Web Security (PHP CTF-style exploits), and Cloud Security (AWS, Azure, GCP, Kubernetes).
Detailed methodology and anti-cheating measures: deterministic scoring (no LLM-as-a-judge), pass@3 evaluation, network-isolated containers, dynamic validation, and session-specific flags.
Analysis of agent vs. model effects: how native-provider advantage, agent architecture, and domain-specific tooling independently shape offensive performance.
Category-by-category breakdowns: scoring methods, vulnerability types covered, and comparative charts showing where each combination leads or falls short.