パーソナライズされたデモを見る
"私が今まで見た中で最高のユーザーエクスペリエンスは、クラウドワークロードを完全に可視化します。"
"Wiz を使えば、クラウド環境で何が起こっているかを 1 つの画面で確認することができます"
"Wizが何かを重要視した場合、それは実際に重要であることを私たちは知っています。"
Wiz Research | Technical Report
Security Engineers & Red Teamers looking to choose the right AI agent-model pairing for vulnerability discovery, penetration testing, and exploit development workflows.
Security Leaders & CISOs who need to assess the real-world offensive capabilities of AI agents to inform risk posture, tool investments, and responsible deployment decisions.
AI & ML Engineers building or evaluating agentic security tools who want to understand how agent architecture and model selection jointly determine performance on complex, multi-step tasks.
Security Researchers benchmarking frontier models and seeking a reproducible, deterministic evaluation framework that spans the full offensive lifecycle.
Full benchmark results across 25 agent-model combinations: 4 agents (Claude Code, Gemini CLI, OpenCode, Codex) × 8 models (Claude Opus 4.6, Opus 4.5, Sonnet 4.5, Haiku 4.5, GPT-5.2, Gemini 3 Pro, Gemini 3 Flash, Grok 4), with scores and runtimes for every pairing.
257 real-world challenges across five offensive security categories: Zero-Day Discovery, CVE Detection (176 challenges across Python, Go, and Java), API Security, Web Security (PHP CTF-style exploits), and Cloud Security (AWS, Azure, GCP, Kubernetes).
Detailed methodology and anti-cheating measures: deterministic scoring (no LLM-as-a-judge), pass@3 evaluation, network-isolated containers, dynamic validation, and session-specific flags.
Analysis of agent vs. model effects: how native-provider advantage, agent architecture, and domain-specific tooling independently shape offensive performance.
Category-by-category breakdowns: scoring methods, vulnerability types covered, and comparative charts showing where each combination leads or falls short.