Last updated: 2026-09-06 09:32:44 UTC

The HTB-Challenger Benchmark evaluates how well large language models can find and exploit security vulnerabilities. It tests models in a controlled environment against a fixed set of 16 recent Hack The Box challenges of varying difficulty, then measures and compares their performance. To understand what these results do - and do not - say about a model, see What Does the HTB-Challenger Benchmark Actually Measure?.

For details about challenge selection, the test harness, run limits, and scoring, visit the HTB-Challenger Benchmark Methodology page.

This is an ongoing project, and new models are added as they are tested.

Cost vs. Benchmark Score

Each point represents one model; models closer to the upper-left achieve a higher benchmark score at a lower median cost per challenge.

Cost vs. Benchmark Score

Benchmark metrics

All metrics except Score are reported per challenge and calculated as medians to reduce the impact of extreme values.

Click a model name in the table to view more details about its benchmark results.

Model Score Cost Steps Duration Tokens
openai icon GPT-5.6 Sol 87.2% $0.29 8.5 00:02:05 0.09M
z-ai icon GLM 5.3 Flash 81.0% $0.01 24.5 00:10:52 0.47M
x-ai icon Grok 4.6 76.1% $0.23 14.0 00:03:38 0.24M
openai icon GPT-5.6 Terra Pro 75.1% $0.91 11.0 00:08:22 0.66M
openai icon GPT-5.6 Luna + GPT-5.6 Sol 74.5% $0.07 20.0 00:06:02 0.61M
meta icon Muse Spark 1.3 72.9% $0.38 23.0 00:08:05 0.63M
moonshotai icon Kimi K3 70.7% $0.46 22.0 00:07:34 0.35M
openai icon GPT-5.6 Terra 64.6% $0.24 16.5 00:03:16 0.24M
qwen icon Qwen3.8 Max 62.7% $2.00 29.0 00:16:02 0.96M
openai icon GPT-5.6 Luna Pro 51.3% $0.10 21.5 00:08:55 1.57M
meta icon Muse Spark 1.2 50.4% $0.48 32.0 00:05:36 0.98M
z-ai icon GLM 5.3 47.2% $0.49 16.5 00:11:52 0.27M
deepseek icon DeepSeek V4 Pro 0813 38.8% $0.20 12.5 00:07:17 0.25M
deepseek icon DeepSeek V4 Pro 0423 36.2% $0.35 62.5 00:06:19 1.94M
tencent icon Hy3 34.8% $0.08 18.5 00:15:21 0.42M
openai icon GPT-5.6 Luna 31.6% $0.02 17.5 00:04:38 0.36M
minimax icon MiniMax M3 30.5% $0.16 26.5 00:07:45 0.43M
deepseek icon DeepSeek V4 Flash 0731 8.2% $0.00 14.5 00:02:29 0.10M

Honorable mentions

I also attempted to run the benchmark with the following frontier models from Anthropic, OpenAI and Google:

Unfortunately, each of these models eventually refused to continue testing, returning messages such as: “This content was flagged for possible cybersecurity risk.” As a result, these models are not included in the benchmark results.