HTB-Challenger Benchmark
Last updated: 2026-09-06 09:32:44 UTC
The HTB-Challenger Benchmark evaluates how well large language models can find and exploit security vulnerabilities. It tests models in a controlled environment against a fixed set of 16 recent Hack The Box challenges of varying difficulty, then measures and compares their performance. To understand what these results do - and do not - say about a model, see What Does the HTB-Challenger Benchmark Actually Measure?.
For details about challenge selection, the test harness, run limits, and scoring, visit the HTB-Challenger Benchmark Methodology page.
This is an ongoing project, and new models are added as they are tested.
Cost vs. Benchmark Score
Each point represents one model; models closer to the upper-left achieve a higher benchmark score at a lower median cost per challenge.
Benchmark metrics
All metrics except Score are reported per challenge and calculated as medians to reduce the impact of extreme values.
Click a model name in the table to view more details about its benchmark results.
| Model | Score | Cost | Steps | Duration | Tokens |
|---|---|---|---|---|---|
GPT-5.6 Sol |
87.2% | $0.29 | 8.5 | 00:02:05 | 0.09M |
GLM 5.3 Flash |
81.0% | $0.01 | 24.5 | 00:10:52 | 0.47M |
Grok 4.6 |
76.1% | $0.23 | 14.0 | 00:03:38 | 0.24M |
GPT-5.6 Terra Pro |
75.1% | $0.91 | 11.0 | 00:08:22 | 0.66M |
GPT-5.6 Luna + GPT-5.6 Sol |
74.5% | $0.07 | 20.0 | 00:06:02 | 0.61M |
Muse Spark 1.3 |
72.9% | $0.38 | 23.0 | 00:08:05 | 0.63M |
Kimi K3 |
70.7% | $0.46 | 22.0 | 00:07:34 | 0.35M |
GPT-5.6 Terra |
64.6% | $0.24 | 16.5 | 00:03:16 | 0.24M |
Qwen3.8 Max |
62.7% | $2.00 | 29.0 | 00:16:02 | 0.96M |
GPT-5.6 Luna Pro |
51.3% | $0.10 | 21.5 | 00:08:55 | 1.57M |
Muse Spark 1.2 |
50.4% | $0.48 | 32.0 | 00:05:36 | 0.98M |
GLM 5.3 |
47.2% | $0.49 | 16.5 | 00:11:52 | 0.27M |
DeepSeek V4 Pro 0813 |
38.8% | $0.20 | 12.5 | 00:07:17 | 0.25M |
DeepSeek V4 Pro 0423 |
36.2% | $0.35 | 62.5 | 00:06:19 | 1.94M |
Hy3 |
34.8% | $0.08 | 18.5 | 00:15:21 | 0.42M |
GPT-5.6 Luna |
31.6% | $0.02 | 17.5 | 00:04:38 | 0.36M |
MiniMax M3 |
30.5% | $0.16 | 26.5 | 00:07:45 | 0.43M |
DeepSeek V4 Flash 0731 |
8.2% | $0.00 | 14.5 | 00:02:29 | 0.10M |
Honorable mentions
I also attempted to run the benchmark with the following frontier models from Anthropic, OpenAI and Google:
- Claude Opus 4.8
- Claude Opus 5
- Claude Fable 5
- GPT-6 Astra (even with active Daybreak - Trusted Access for Cyber program)
- Gemini 3.7 Flash
Unfortunately, each of these models eventually refused to continue testing, returning messages such as: “This content was flagged for possible cybersecurity risk.” As a result, these models are not included in the benchmark results.