Cyber Model Arena

We evaluate AI agents with foundation models and standard harnesses on multiple real-world cybersecurity tasks.

Model Quality vs. Efficiency

Success rate against cost, time, or steps per solve over k independent attempts.

Dataset
X per Solve
Harness
Pass@k

Leaderboard

Each percentage is pass@1 in that category.

#
Agent
Overall @ 1
Code Vulns
API Sec
Web Sec
Cloud Sec
Efficiency / pass
1
ADK React
74.9%57%85.2%90.8%66.7%
5 min$1.6258 steps
2
Claude Code
71.2%55.2%84%81.6%63.8%
12.2 min$2.6319 steps
3
Claude Code
68.5%50.1%84.6%78.2%60.9%
5.7 min$1.5258 steps
4
ADK React
66.9%41.8%78.4%85.1%62.3%
15.3 min$7.4612 steps
5
Claude Code
65.4%44%79%74.7%63.8%
11.9 min$2.7715 steps
6
Claude Code
64.4%38.9%78.4%79.3%60.9%
22.4 min$3.4847 steps
7
Claude Code
64.2%33.8%82.7%79.3%60.9%
27.2 min$8.2957 steps
8
ADK React
61%25.8%80.2%81.6%56.5%
30.6 min$2.3139 steps
9
ADK React
60.9%35.3%77.2%75.9%55.1%
15.7 min$4.1612 steps
10
Claude Code
60.5%37.7%73.5%81.6%49.3%
22.5 min$8.9947 steps
11
ADK React
59.7%17%80.2%85.1%56.5%
38.7 min$0.9537 steps
12
Claude Code
59.3%36.3%76.5%67.8%56.5%
14.1 min$1.5816 steps
13
Claude Code
59.1%26.1%82.1%74.7%53.6%
27.4 min$6.4738 steps
14
Claude Code
59.1%39.4%75.9%69%52.2%
16 min$9.3240 steps
15
ADK React
58.7%31.4%81.5%60.9%60.9%
26.3 min$2.3719 steps
16
Claude Code
58%32.9%82.7%65.5%50.7%
28 min$3.1333 steps
17
Claude Code
54.4%37.9%72.8%62.1%44.9%
12.1 min$0.6720 steps
18
ADK React
53.6%17.3%74.7%75.9%46.4%
41.6 min$2.3546 steps
19
ADK React
52.8%27.5%77.2%52.9%53.6%
31.9 min$1.6149 steps
20
ADK React
52.5%36.3%76.5%44.8%52.2%
19.2 min$1.5317 steps
21
ADK React
52.2%27.5%72.2%59.8%49.3%
21.7 min$1.0920 steps
22
Claude Code
51%26.5%75.9%60.9%40.6%
28.6 min$1.8921 steps
23
Claude Code
50.5%32%68.5%50.6%50.7%
17.7 min$2.3932 steps
24
ADK React
49.1%21.7%75.3%51.7%47.8%
40.1 min$1.4028 steps
25
ADK React
46.6%23.4%56.2%67.8%39.1%
43 min$2.3916 steps
26
ADK React
46.5%19%71%48.3%47.8%
46 min$1.7559 steps
27
Claude Code
45.5%9.7%71.6%64.4%36.2%
51.2 min$2.8515 steps
28
ADK React
42%14.3%77.2%29.9%46.4%
42.5 min$1.6926 steps

What can defenders learn?

Agents can already perform standard cyber tasks

Leading agents already complete most of our basic, real-world security challenges on their first attempt.

The harness matters

The same model can perform substantially better or worse under Claude Code vs ADK React. We evaluate the pairing, not the model alone.

Attacks are inexpensive and take minutes

A successful AI agent attack typically costs a few dollars and completes within minutes. Additional attempts (e.g., pass@3) increase the attack success rate.

Code Vulns are hard

We cover many challenges of finding vulns in a large codebase and 1-day (known-CVE) exploitation. We find such tasks to be harder for agents to solve.

Datasets

Code Vulns

221 challenges

Identify known vulnerability patterns in source code, and exploit 1-days (known-CVEs).

Identify the vulnerability root cause in Python, Java and Go, or exploit a live 1-day CVE.

API Security

54 challenges

Discover and exploit web vulnerabilities through live API.

Reach all exploitable paths in labs using no source code, and often no API spec.

Websec.fr CTFs

29 challenges

Analyze Web CTF challenges source code and write working exploits to capture flags.

Use the public websec.fr website tasks.

Cloud Security

23 challenges

Exploit misconfigurations across different cloud providers (taken from Wiz CTF Challenges).

Capture a flag within AWS / Azure / GCP / Kubernetes-style setups.

Harnesses

ADK React

A simple ReAct loop developed with ADK. It uses a single bash tool, has no specialized cyber tools, and uses a general system prompt.

Claude Code

Anthropic's standard coding CLI. It uses ordinary file, search and shell tools, has no specialized cyber tools, and uses a general system prompt.

Methodology

  • The agent receives the task instruction only, with no extra attack playbooks or exploit templates.

  • Each harness-model-challenge combination is run 3 times (pass@3: best result across runs is taken per challenge), with a retry limit of 3 per time for infra errors.

  • Each agent run is limited to 1000 seconds. When time expires, the agent is stopped and the verifier scores whatever the agent produced. A timeout is not retried.

  • Reasoning level is set to high across models.

  • Anti-cheat: agents run in isolated Docker containers with no CVE databases and no external resources.

  • All scoring is deterministic via API hooks, exact flags, exact root-cause strings, and asserting the vulnerable path is executed.

  • The overall score is the macro-average across all 4 categories.