AI Penetration Testing: Methodology, Tools and Scope

Pwned Labs
  • October 8, 2026

AI penetration testing is the authorized security assessment of an AI-powered application to find and exploit weaknesses before an attacker does. In practice it targets the system built around a large language model: how untrusted text reaches it, what tools and data it can touch, what it discloses, and what happens to its output. This guide covers what an AI penetration test is, how it differs from traditional testing, the taxonomy we use to map the attack surface, the methodology a professional follows, the tooling involved, and how the same work informs defense.

Key Takeaway: An AI penetration test is a scoped, authorized engagement that maps where untrusted input enters an LLM system, what the model is trusted to do, and how its output is handled, then exploits the gaps. The skills transfer directly from web and API testing, and you do not need a machine learning background to start.


What is AI penetration testing?

AI penetration testing, sometimes called AI red teaming, is a controlled, authorized assessment that simulates how an attacker would compromise an AI application and the systems around it. It goes beyond running a model through a benchmark. A real test looks at the whole application: the prompts and rules that steer the model, the retrieval pipeline that feeds it context, the tools and agents it can invoke, and the downstream systems that consume its output. The goal is to prove concrete impact, such as exfiltrating data, coercing an unauthorized action, or turning model output into code execution, so a finding reads as a demonstrated attack chain and not as a note that a model can be made to say something it should not.


How AI penetration testing differs from traditional pentesting

The mindset is familiar: enumerate the surface, find where untrusted input meets trusted logic, and abuse the gap. What is new is the medium. Instead of crafting a payload for a parser, you are steering a model with natural language, and instead of a single deterministic exploit, you sweep a catalog of techniques and record what works, because model behavior is probabilistic and the same prompt can succeed one time in five. The trust boundary also moves. System instructions, developer rules, retrieved documents, and the user message all arrive as one undifferentiated stream of text, so an attacker who controls any part of that stream can try to influence the model. The base skills transfer directly, and comfort with HTTP, reading JSON, an intercepting proxy, and a general security foundation are enough to begin.


The Pwned Labs AI attack taxonomy

Generic technique lists are hard to test against, so we group AI attacks into seven families, ordered by where they act on the pipeline, from the input the model reads through to the actions it takes. A family names the weakness we test for, and the techniques under it are what we actually send, so one attempt can cross two families. Modality, whether text, voice, or multimodal, cuts across all seven. Each family maps to the OWASP Top 10 for LLM Applications and to MITRE ATLAS, so findings line up with the frameworks a client already tracks.

Family What it attacks Example techniques OWASP
01 Instruction The prompt itself is hijacked Direct injection and jailbreak, obfuscation and smuggling, multi-turn escalation LLM01
02 Context Poison what the model is fed Indirect and data-borne injection, RAG and knowledge-base poisoning, embedding and vector abuse LLM01, LLM08
03 Model Attack the model itself Intent and NLU evasion, model and training extraction, supply-chain poisoning LLM03, LLM04
04 Disclosure Make it reveal what it knows System-prompt leakage, retrieved-context emission, verbose error and infrastructure leak LLM02, LLM07
05 Output Weaponize what comes out Improper output handling, rendered exfiltration sinks, downstream injection LLM05
06 Agency Abuse what it can do Tool and function abuse, confused-deputy escalation, business-logic bypass LLM06
07 Platform (cross-cutting) The control plane around the model: authorization, tenancy, session, quota Broken authorization on AI APIs, cross-tenant and cross-user reads, unbounded consumption LLM10 and API-side


The seven families in depth

01 Instruction

The prompt itself is the attack. Direct injection types an instruction that overrides the developer's intent, and the catalog here is large: instruction override, refusal suppression, persona role-play, skeleton key and policy-puppetry framings that present the instruction as configuration, many-shot and crescendo multi-turn escalation, and adversarial suffixes. Obfuscation hides the payload from input filters with base64 and hex encoding, unicode homoglyphs, zero-width and tag-block smuggling, and bidirectional-text overrides. The reason any of it works is instruction precedence: the order in which the model trusts system, developer, and user text is trained, not enforced, so nothing rejects an instruction because of the layer it arrived in. A model that surrenders its own rules on the first attempt usually surrenders more under pressure. This family maps to OWASP LLM01.

02 Context

Here the attacker poisons what the model is fed instead of talking to it directly. Instructions are planted where the model will later read them: a poisoned knowledge-base article, hidden text in a document such as an HTML comment, white-on-white, one-pixel, or metadata, invisible unicode in content, an uploaded file, conversation history, or the output of another tool. They fire when the model retrieves them during a normal request, which is what makes zero-click, retrieval-triggered exfiltration possible. EchoLeak (CVE-2025-32711, CVSS 9.3), disclosed in June 2025, was this against Microsoft 365 Copilot, a crafted email that exfiltrated data with no interaction. The family also covers cross-customer vector leakage, where weak isolation lets one tenant's query return another tenant's chunks. See our explainer on indirect prompt injection. This family maps to OWASP LLM01 and LLM08.

03 Model

This family goes at the model and its supply chain: adversarial misclassification and character-level perturbation against intent and classifier logic, confidence manipulation and fallback abuse, model and training-data extraction through repeated queries, and poisoning a model, adapter, or dataset pulled from a public hub. Much of it is hard to confirm from outside the model, so testing leans on provenance and on probing suspected behavior. This family maps to OWASP LLM03 and LLM04.

04 Disclosure

The goal is to make the model reveal what it holds. Verbatim prompt extraction pulls the system prompt directly, and indirect extraction gets it by asking the model to summarize, translate, or encode its own instructions. The same family covers secret and credential leakage, PII echo and cross-session recall, training and config-artifact extraction, and infrastructure details leaked through verbose errors. Any secret found in a system prompt is both a finding and a design flaw, since the prompt should be treated as readable by anyone. This family maps to OWASP LLM02 and LLM07.

05 Output

Model output is attacker-influenced data, and any sink that renders or executes it without validation inherits the risk. The catalog includes stored and reflected cross-site scripting through rendered output, markdown-image data exfiltration, server-side request forgery through link unfurling and auto-rendering, downstream SQL and command injection, server-side template injection, and formula injection in exported CSVs. The markdown-image path was first demonstrated by Johann Rehberger against Google Bard in 2023 and has recurred across assistants since. This family maps to OWASP LLM05.

06 Agency

When the model can call tools, a successful manipulation becomes a real action. Testing starts with tool enumeration, then parameter injection and argument tampering, server-side request forgery through a tool to a collector or the cloud instance metadata service, confused-deputy escalation, goal hijacking and action chaining, and human-in-the-loop bypass. ForcedLeak (CVSS 9.4, 2025) chained an indirect injection in Salesforce Agentforce, through a Web-to-Lead field, into CRM data exfiltration. Agentic systems add tool-description poisoning and cross-server instruction override: in MCP tool poisoning, named by Invariant Labs in April 2025, malicious instructions hidden in a tool's name or description, text the model reads but the user never sees, can coerce an agent into reading local secrets or rerouting a trusted tool. The impact of any of this is the tool's capability multiplied by the identity it runs as. This family maps to OWASP LLM06.

07 Platform

The seventh family is cross-cutting: the control plane around the model, its authorization, tenancy, runtime, and quotas. On the authorization side it covers vertical RBAC and horizontal IDOR failures, cross-service token scope, tokens leaked in URLs, and cross-customer data bleed. On the infrastructure side, serverless sandbox escape, runtime egress and DNS exfiltration, secrets-store access from an over-scoped execution role, and MCP connector poisoning. It also covers supply-chain weaknesses such as feedback-loop poisoning and weak model or connector provenance, and consumption abuse from token amplification and recursive prompts through to denial of wallet. These are the classic application-security failures that do not disappear because a model sits in the middle, and they are often the fastest path to impact. This family maps to OWASP LLM10 and the API layer.


The AI penetration testing methodology

A repeatable methodology keeps an assessment thorough and defensible. A typical engagement moves through these phases.

  1. Scoping and authorization. Agree the target systems, data, and rules of engagement in writing, and confirm the client owns or controls what is in scope. Third-party AI services and shared model providers sit outside that authorization unless the provider has agreed in writing.
  2. Surface mapping. Enumerate every place untrusted text enters the model, the tools and functions it can call, the data sources it retrieves from, the external servers it loads, and the systems that consume its output. The map, not the model, drives the test.
  3. Instruction and context testing. Work direct injection first to gauge instruction separation, then indirect injection through documents, tickets, emails, and retrieved web content. This phase usually opens up the rest of the engagement.
  4. Disclosure testing. Probe what the model reveals about its context, its sources, other tenants, and its own system prompt, and treat any secret found in the prompt as both a finding and a design flaw.
  5. Retrieval and data poisoning. Seed or edit content the model retrieves to test whether attacker-controlled text is trusted during a normal request, and check tenant isolation across the vector store.
  6. Agency and tool abuse. Coerce sensitive tool calls, shape their arguments, and test whether tool and retrieval responses are trusted as instructions, which is where excessive agency turns into real impact.
  7. Output handling. Treat model output as untrusted input to every downstream sink, testing for cross-site scripting, server-side request forgery, SQL injection, and command execution.
  8. Reporting. Document each attack chain, its demonstrated impact, and prioritized, practical remediation that an engineering team can act on.


How an attack chains together

The families matter most in combination. A customer support assistant answers from a knowledge base built out of incoming tickets and can call a tool to look up order and account records. An attacker files a support ticket whose body carries an instruction aimed at the model, not the staff member reading it. When someone later asks the assistant to summarize open tickets, the knowledge base returns the poisoned ticket, the model treats the planted instruction as trusted context, and it calls the lookup tool with an attacker-chosen account identifier. The tool returns another customer's record, and the model places that data into a reply or an outbound link. That single chain crosses the Context, Agency, and Output families with no direct access to the model at all, which is why we test the path and not each risk in isolation.


The assessor's toolkit

Tooling supports the methodology, it does not replace it, and the kit spans four jobs. Offensive frameworks automate parts of the technique sweep: garak, NVIDIA's LLM vulnerability scanner, PyRIT, Microsoft's Python Risk Identification Toolkit, and promptfoo, DeepTeam, Giskard, and FuzzyAI. Agentic and MCP scanners test the tool layer: Agentic Radar, Invariant's MCP-Scan, Cisco's MCP Scanner, and Snyk Agent Scan. The classic web stack still does much of the work in the AI lane: an intercepting proxy such as Burp Suite or Caido, curl for a reproducible single request, the browser and its developer tools for whatever the deployed chat UI actually sends, and jq, ffuf, and httpx for everything around it. For out-of-band proof, Burp Collaborator or ngrok in front of a controlled redirector confirms when a model reaches a destination we own. Attacker models themselves earn a place in the kit, since they generate payload variants faster than a human can type them. The value is in interpreting results and chaining findings into demonstrated impact, which is work no scanner does on its own.


Defending against AI application attacks

The same testing informs the defense, and the most useful finding a report can carry is that blast radius is a design property, set by each tool's scope and the identity its runtime holds, both chosen before any attacker shows up. The working assumption is that the model will be tricked, so the effort goes into what a tricked model can actually reach. Scope tools to least privilege and give the runtime the narrowest identity that does the job. Apply default-deny egress from the tool runtime, which collapses most exfiltration without ever detecting the injection. Keep a clear separation between instructions and data, treat every retrieved input as untrusted, and validate and encode model output at every sink. The detection signals are observable without catching the prompt itself: an egress fetch the agent would never make, data carried inside an image or link URL, a tool call that does not match the user's request, and retrieval that crosses a tenant boundary. Capturing the full prompt, tool-call, and retrieval trace is what lets a client reproduce the chain and regression-test the fix.


How to learn AI penetration testing

AI penetration testing is learned by doing it against real, authorized targets. If you are new to the area, start with the concepts and the beginner path in our introduction to AI hacking, then practice against dedicated targets in the PromptStorm cyber range. To go end to end across all seven families, from direct and indirect prompt injection through retrieval poisoning, tool abuse, and agentic compromise across CI/CD pipelines and MCP tool servers, the AI Systems Attack and Defense bootcamp leads to the AISRTP certification, with every technique run hands-on in a live environment and the defensive view taught alongside the offensive one. For the wider curriculum, see our AI security training overview. Only ever test systems you own or are explicitly authorized to assess.


Frequently asked questions

What is AI penetration testing?

It is an authorized security assessment of an AI-powered application that maps where untrusted input reaches the model, what it is trusted to do with tools and data, and how its output is handled, then exploits the gaps to prove real impact.

Is AI penetration testing the same as AI red teaming?

The terms overlap and are often used interchangeably. In common usage, an AI penetration test is a scoped assessment of a specific application against a defined methodology, while AI red teaming can describe a broader, objective-driven exercise that includes model-level safety testing. Both exploit the application around the model to demonstrate impact.

Do I need a machine learning background to do AI penetration testing?

No. AI penetration testing targets the application around the model. Web and API testing skills, comfort with HTTP and JSON, an intercepting proxy, and a general security foundation are enough to start.

How is AI penetration testing different from a model benchmark or eval?

An eval measures model quality or safety in isolation. A penetration test assesses the whole application in context, chaining prompt injection, retrieval poisoning, tool abuse, and output handling into concrete, demonstrated impact.

What tools are used for AI penetration testing?

An intercepting proxy such as Burp Suite or OWASP ZAP, a catalog of prompt injection and jailbreak techniques, and frameworks such as garak and PyRIT, with evaluation harnesses like promptfoo for repeatable runs. Tooling supports the methodology and does not replace it.

Is AI penetration testing legal?

Only when it is authorized. Test systems you own or have explicit written permission to assess, and use dedicated practice environments like Pwned Labs labs to build the skills safely and legally.

Related Articles

OWASP LLM Top 10 (2025): Risks and How to Test Them

October 8, 2026
The OWASP Top 10 for LLM Applications is the reference list of the most critical security risks in applications built...

Cloud Penetration Testing: A Beginner's Guide

July 17, 2026
Cloud penetration testing means safely simulating real attacks against a cloud environment to find the weaknesses...

GCP Penetration Testing Training

July 17, 2026
Google Cloud Platform is the least-covered major cloud provider in offensive security training, which creates both a...