Pwnbench: cybersecurity benchmarking for frontier AI models

Pwnbench by Pwned Labs gives frontier model providers a way to measure offensive cybersecurity capability against infrastructure that behaves like production. It deploys custom environments across multi-cloud, Kubernetes, CI/CD, Active Directory and hybrid on demand, gives a model an objective inside them, and records how far it gets.
aws
Azure
Google Cloud
Kubernetes
Active Directory
CI/CD
Hybrid
Cloud Kubernetes CI/CD Active Directory Hybrid Objective DEPLOYATTACKSCORE

Capability scores depend on what the model is pointed at

Capture-the-flag challenges test whether a model can solve a puzzle that was built to be solved. Production infrastructure is different, because it carries misconfigurations nobody documented, identities with more permissions than they need, pipelines that trust the wrong inputs, and defenses that sometimes work. Pwnbench builds environments with those properties on purpose, so a capability score reflects what a model can do against the systems it would meet outside the lab.

40,000+
practitioners already train on Pwned Labs environments
Multi-cloud to hybrid
scenarios across cloud, Kubernetes, CI/CD, Active Directory and hybrid
CrowdStrike and Centene
used by their employees

Building targets that behave like real infrastructure is what we do

Pwned Labs creates custom, reproducible scenarios that stand alone or interweave the way a real organization runs them, and our cyber ranges already run realistic multi-cloud attack scenarios for red teams. Pwnbench applies the same engineering to measuring what a model can do against them.

Multi-cloudKubernetesCI/CDActive DirectoryHybrid

Deploy, attack, score

1Deploy
A fresh, isolated environment is provisioned for each run across multi-cloud, Kubernetes, CI/CD, Active Directory and hybrid targets, from a known and reproducible state.
2Attack
The model receives access and an objective, then works through reconnaissance, exploitation, privilege escalation and lateral movement against real services, and every action is captured.
3Score
Each run records what the model attempted, achieved and where it stalled, scored against the objectives and compared across models, versions and prompting setups.
Bespoke evaluations, not a fixed benchmark
Pwnbench is not a single fixed test. We align a run to the evaluation framework you already use, support more than one across your program, and build fully bespoke scenarios and objectives for your models, so a measurement matches the threat model you care about and changes as your models do.
Run capable models safely
Each run happens in an isolated environment provisioned fresh for the evaluation, so a capable model can be tested against live infrastructure without touching anything real. Every environment starts from a known state, so a result can be repeated and audited rather than taken on trust.

Offensive capability across the full attack chain

Reconnaissance and enumeration
Exploitation of misconfigurations and vulnerable services
Privilege escalation and lateral movement
Persistence and access to sensitive data
Behavior under partial information
Handling of dead ends, rate limits and defensive noise

Built for the teams who sign off on a release

Pwnbench is built for frontier model providers, AI labs running dangerous capability evaluations, and the safety and evaluation teams who need evidence about offensive cybersecurity capability they can stand behind. It produces that evidence before a model ships and again after it changes, so a release decision rests on a measurement rather than an estimate.

Pwnbench is in early access
Request access and we will set up a call to talk through targets, objectives and reporting for your models.