Pwnbench: cybersecurity benchmarking for frontier AI models
Capability scores depend on what the model is pointed at
Capture-the-flag challenges test whether a model can solve a puzzle that was built to be solved. Production infrastructure is different, because it carries misconfigurations nobody documented, identities with more permissions than they need, pipelines that trust the wrong inputs, and defenses that sometimes work. Pwnbench builds environments with those properties on purpose, so a capability score reflects what a model can do against the systems it would meet outside the lab.
Building targets that behave like real infrastructure is what we do
Pwned Labs creates custom, reproducible scenarios that stand alone or interweave the way a real organization runs them, and our cyber ranges already run realistic multi-cloud attack scenarios for red teams. Pwnbench applies the same engineering to measuring what a model can do against them.
Deploy, attack, score
Offensive capability across the full attack chain
Built for the teams who sign off on a release
Pwnbench is built for frontier model providers, AI labs running dangerous capability evaluations, and the safety and evaluation teams who need evidence about offensive cybersecurity capability they can stand behind. It produces that evidence before a model ships and again after it changes, so a release decision rests on a measurement rather than an estimate.