We evaluate AI agents with foundation models and standard harnesses on multiple real-world cybersecurity tasks.
Success rate against cost, time, or steps per solve over k independent attempts.
Each percentage is pass@1 in that category.
Leading agents already complete most of our basic, real-world security challenges on their first attempt.
The same model can perform substantially better or worse under Claude Code vs ADK React. We evaluate the pairing, not the model alone.
A successful AI agent attack typically costs a few dollars and completes within minutes. Additional attempts (e.g., pass@3) increase the attack success rate.
We cover many challenges of finding vulns in a large codebase and 1-day (known-CVE) exploitation. We find such tasks to be harder for agents to solve.
Code Vulns
221 challenges
Identify known vulnerability patterns in source code, and exploit 1-days (known-CVEs).
Identify the vulnerability root cause in Python, Java and Go, or exploit a live 1-day CVE.
API Security
54 challenges
Discover and exploit web vulnerabilities through live API.
Reach all exploitable paths in labs using no source code, and often no API spec.
Websec.fr CTFs
29 challenges
Analyze Web CTF challenges source code and write working exploits to capture flags.
Use the public websec.fr website tasks.
Cloud Security
23 challenges
Exploit misconfigurations across different cloud providers (taken from Wiz CTF Challenges).
Capture a flag within AWS / Azure / GCP / Kubernetes-style setups.
ADK React
A simple ReAct loop developed with ADK. It uses a single bash tool, has no specialized cyber tools, and uses a general system prompt.
Claude Code
Anthropic's standard coding CLI. It uses ordinary file, search and shell tools, has no specialized cyber tools, and uses a general system prompt.
The agent receives the task instruction only, with no extra attack playbooks or exploit templates.
Each harness-model-challenge combination is run 3 times (pass@3: best result across runs is taken per challenge), with a retry limit of 3 per time for infra errors.
Each agent run is limited to 1000 seconds. When time expires, the agent is stopped and the verifier scores whatever the agent produced. A timeout is not retried.
Reasoning level is set to high across models.
Anti-cheat: agents run in isolated Docker containers with no CVE databases and no external resources.
All scoring is deterministic via API hooks, exact flags, exact root-cause strings, and asserting the vulnerable path is executed.
The overall score is the macro-average across all 4 categories.