PwnRange: cloud attack ranges for evaluating and training AI models

We spawn a fresh, unique environment for every run across AWS, Azure, GCP, Kubernetes and CI/CD, give the model an objective, record every action, and score how far it gets.
aws
Azure
Google Cloud
Kubernetes
Active Directory
CI/CD
Hybrid
Cloud Kubernetes CI/CD Active Directory Hybrid Objective DEPLOYATTACKSCORE


Ranges rather than a benchmark

 
Public cyber benchmarks are fixed task sets, and fixed sets get memorized. The challenges are in training data, scores have climbed to the ceiling, and a high score no longer says much about what a model can do inside a real environment. PwnRange generates a new instance every time, with randomized names, addresses, credentials, flag placement and decoy resources, so a result reflects capability rather than recall.

What we cover

 
Our ranges run on live control planes rather than emulations. Scenario families cover AWS IAM privilege escalation and role chaining, Entra ID and hybrid identity paths from a web foothold to on-premises Active Directory, GCP service account and federation abuse, Kubernetes RBAC, secrets and escape chains, and CI/CD and supply chain pivots through pipelines, runners and artifact registries. AI-system surfaces such as MCP servers, agent tool credentials and prompt injection into infrastructure can be composed into any family.
AWSAzureGCPKubernetesCI/CD

How it works

 
1Deploy
Each run gets an isolated environment provisioned from our infrastructure as code into a dedicated account, subscription or project, with no shared state between runs.
2Attack
The model connects through an Inspect-compatible agent bridge or an SSH entry point on an attack host, with the objective and the scope stated in the prompt. We can also deliver scenarios in OpenEnv format or into accounts you supply.
3Score
Progress is scored against milestones and hidden flags verified at generation time. Every command, tool call and cloud API call is captured, so a run produces a full trajectory and out-of-band telemetry rather than a single number.

Built for the standards labs now require

 
Ranges run with no internet access and with the model API as the only permitted egress. Network policy is deny by default and verified before each run, real-time monitoring can stop a run on out-of-scope traffic, no real credentials or real organization names exist inside a range, and transcripts and network logs are retained for audit.

Evaluation and training from one generator

 
For evaluation, a scenario family keeps a fixed attack graph and a difficulty label while every instance differs, so results compare across models and over time. For training, families vary in topology and can be produced at volume with verifiable rewards. Results can be reported through PwnBench, our scenario suite with difficulty labels and human expert baselines.

Who it is for

 
Model developers evaluating cyber capability under their safety frameworks, government and third-party evaluators who need realistic enterprise and cloud ranges, and security vendors building agents on frontier models who need somewhere real to test them.

Working with us

 
Early access starts with a scoped pilot of scenario families built to your requirements, delivered hosted or into your own accounts. Inference runs on your model and your keys.