The OWASP Top 10 for LLM Applications is the reference list of the most critical security risks in applications built on large language models. It is published by the OWASP GenAI Security Project, and the current edition is the 2025 list, released in November 2024. This guide explains all ten risks in plain terms, shows how several of them have already been exploited in the wild, gives a test and a mitigation for each, and shows how they combine in real incidents.
What is the OWASP Top 10 for LLM Applications?
The OWASP Top 10 for LLM Applications is a community-driven, consensus list of the ten most important security risks facing software that uses large language models, from chat assistants and retrieval systems to autonomous agents. It is maintained by the OWASP GenAI Security Project and mirrors the format of the long-established OWASP Top 10 for web applications. The first edition appeared in 2023; the current 2025 edition was published in November 2024 and reorganized the list to reflect how production LLM systems are actually attacked, including retrieval pipelines and agents. Each entry has a stable identifier, LLM01 through LLM10, so teams can reference a risk unambiguously in findings and standards.
The OWASP LLM Top 10 (2025)
Here are the ten risks in the current list, each with what it means, a concrete example, how it is tested, and how it is mitigated.
Prompt Injection (LLM01)
Crafted input causes the model to follow attacker instructions instead of the developer's. In direct injection the attacker types the instruction; in indirect injection it is planted in content the model later retrieves, such as a document, email, or web page. It is the top risk because it underpins most real-world LLM attacks. See our explainer on indirect prompt injection. A real example is EchoLeak (CVE-2025-32711), a zero-click indirect injection in Microsoft 365 Copilot that exfiltrated data from a crafted email with no user action. Testing delivers instructions through every untrusted path, the user message and any content the model later retrieves, and checks whether they override the system prompt or trigger an action. Mitigation is to treat all retrieved and user content as untrusted, keep a clear separation between instructions and data, and constrain what the model can do when it acts on either.
Sensitive Information Disclosure (LLM02)
The model reveals data it should not, such as personal information, credentials, proprietary content, or another tenant's data, either through its training, its context window, or connected sources. In practice this is secrets or API keys echoed back from the context window, or another customer's records surfaced through weak tenant isolation. Testing probes for the system prompt, credentials, and cross-tenant data with targeted queries, and checks what the model will repeat from its connected sources. Mitigation is least-privilege context, strict tenant isolation, and keeping secrets out of the prompt and out of the model's reach entirely.
Supply Chain (LLM03)
Weaknesses enter through third-party components: poisoned or backdoored pretrained models, tampered fine-tuning adapters, malicious datasets, or vulnerable libraries and plugins pulled into the pipeline. A model pulled from a public hub can ship with a backdoor, and a typosquatted package in the serving pipeline is the same classic dependency risk applied to AI. Testing inventories every model, dataset, adapter, and library in the stack and verifies provenance, signatures, and hashes. Mitigation is a tracked inventory of those components with verified provenance, signatures, and pinned versions.
Data and Model Poisoning (LLM04)
An attacker manipulates training, fine-tuning, or embedding data to introduce backdoors, bias, or hidden triggers that change the model's behavior under specific conditions. An attacker who can influence a fine-tuning set or a feedback loop can plant a trigger that behaves normally until a specific phrase appears. It is hard to confirm from outside the model, so testing focuses on data provenance and on probing suspected triggers, since proving a model clean from the outside is not feasible. Mitigation is controlling who can write to training and fine-tuning data, validating that data, and monitoring for anomalous behavior after deployment.
Improper Output Handling (LLM05)
Model output is passed to downstream systems without validation, so generated text becomes an injection payload, leading to cross-site scripting, SQL injection, server-side request forgery, or command execution in whatever consumes it. Rendered as HTML the output becomes cross-site scripting; passed to a shell or a database it becomes command or SQL injection. Testing treats model output as untrusted input and pushes injection payloads through the model into whatever renders or executes them. Mitigation is to validate and encode model output at every sink, the same way any untrusted input is handled.
Excessive Agency (LLM06)
The model or agent is granted too much functionality, permission, or autonomy, so a successful manipulation turns into a real action: calling a sensitive tool, moving data, or changing a record the user never requested. An assistant wired to tools that can email, delete records, or move money turns a single successful injection into real-world impact. Testing enumerates the tools and permissions the agent holds, then attempts to coerce the highest-impact tool calls through crafted input. Mitigation is least-privilege tooling, human confirmation for high-impact actions, and scoping each tool so a coerced call cannot reach sensitive state.
System Prompt Leakage (LLM07)
The system prompt is exposed or over-relied upon. If it contains secrets, rules, or hidden logic, leaking it hands an attacker the blueprint, and treating it as a security control at all is itself the weakness. Attackers extract the system prompt to read hidden rules, filters, or embedded credentials, then design around them. Testing uses known extraction prompts and, more importantly, checks whether anything sensitive, secrets or access decisions, was ever placed in the prompt to begin with. Mitigation is to assume the system prompt will leak and to place no secret, credential, or access decision inside it.
Vector and Embedding Weaknesses (LLM08)
Flaws in the retrieval-augmented generation stack: poisoning of indexed content, weak tenant isolation in the vector store, or embedding-inversion and cross-context leakage that expose data the retriever should have kept separate. Poisoning indexed content plants attacker text that a normal query later retrieves and trusts, and weak isolation lets one tenant's query pull another tenant's chunks. Testing seeds content into the knowledge base to check whether it influences answers, and probes retrieval across tenant boundaries. Mitigation is access control and provenance on indexed content, strict tenant partitioning in the vector store, and treating retrieved text as data.
Misinformation (LLM09)
The model produces false or misleading output, including hallucinated facts and unsafe code suggestions, that users and downstream systems then trust and act on. Hallucinated package names are a live supply-chain vector: attackers register the invented name so a developer who trusts the suggestion installs their code. Testing probes high-stakes flows for confident but wrong output and checks what downstream systems do with it unverified. Mitigation is to ground answers in verified sources, surface uncertainty, and keep a human in the loop wherever output drives a consequential action.
Unbounded Consumption (LLM10)
Uncontrolled resource use, from denial-of-service and runaway inference cost, sometimes called denial of wallet, to model extraction where repeated queries reconstruct proprietary behavior. This ranges from denial of wallet, where expensive queries run up the inference bill, to model extraction, where repeated queries reconstruct proprietary behavior. Testing measures rate and cost controls and attempts high-volume or deliberately expensive requests. Mitigation is rate limiting, quotas and cost caps, and monitoring for the query patterns that signal extraction or abuse.
How the risks combine
These categories rarely appear alone. In a real engagement, prompt injection (LLM01) is the entry technique that leads to the others: injected instructions reach sensitive data (LLM02), coerce a tool into a real action (LLM06), poison what a later query retrieves (LLM08), or produce output that attacks a downstream system (LLM05). The list is best read as a set of links in a chain, which is why a test maps how untrusted text enters, what the model is trusted to do with it, and where its output goes, instead of scoring each risk in isolation. The disclosed incidents below each map onto two or more categories at once. For a pipeline-oriented view of the same ground, we group these risks into seven attack families by where they act, from the input the model reads to the actions it takes, in our AI penetration testing guide.
What changed in the 2025 update
The 2025 edition sharpened the list around how LLM systems are built today. Three areas are effectively new or substantially expanded: System Prompt Leakage (LLM07) was added after a run of real cases where secrets and logic were placed in the system prompt and then extracted; Vector and Embedding Weaknesses (LLM08) was added to cover the retrieval-augmented generation and vector-store stack that most enterprise assistants now depend on; and Unbounded Consumption (LLM10) broadened the earlier "model denial of service" entry to include inference-cost abuse and model extraction. Prompt injection remained at number one, and the framing across the list shifted further toward the application and its integrations and away from the model in isolation.
The OWASP LLM Top 10 in real incidents
These categories are not theoretical. Each of the following is a disclosed incident that maps directly onto the list.
- EchoLeak (CVE-2025-32711, CVSS 9.3, June 2025): a zero-click indirect prompt injection in Microsoft 365 Copilot that could exfiltrate data with no user interaction. It combines LLM01 (prompt injection) and LLM02 (sensitive information disclosure). Patched by Microsoft.
- ForcedLeak (CVSS 9.4, 2025): indirect prompt injection in Salesforce Agentforce, delivered through a Web-to-Lead field, led to CRM data exfiltration. It combines LLM01 with LLM06 (excessive agency). Reported by Noma and patched by Salesforce.
- MCP tool poisoning (Invariant Labs, April 2025): malicious instructions hidden in a tool's name, description, or input schema, read by the model but unseen by the user, could coerce an agent into reading local secrets or rerouting a trusted tool. It combines LLM01 with LLM03 (supply chain) and LLM06.
- Slack AI (August 2024): data exfiltration from private channels via indirect prompt injection over retrieved messages, catalogued as MITRE ATLAS case study AML.CS0035. It combines LLM01 with LLM08 (vector and embedding weaknesses).
- Google Bard (2023): Johann Rehberger showed indirect injection exfiltrating chat history through a rendered markdown image, combining LLM01 with LLM05 (improper output handling). Fixed by Google, and the technique has recurred across assistants since.
How to test for the OWASP LLM Top 10
Testing an LLM application against this list is an authorized security assessment, not a one-off scan. The core method is to map where untrusted text enters the system, what the model is trusted to do with it, and what happens to its output. In practice that means probing direct and indirect prompt injection first, because it opens up most of the rest; checking what the model discloses about its context, sources, and system prompt; poisoning or tampering with retrieved content to test the vector stack; and coercing tool calls to measure excessive agency. Output handling is tested by treating model responses as untrusted input to whatever renders or executes them. Only ever test systems you own or are explicitly authorized to assess.
Practice AI security hands-on
The fastest way to understand these risks is to exploit and then defend them in a live environment. The Pwned Labs AI Systems Attack and Defense bootcamp works the application-layer risks on this list hands-on: direct and indirect prompt injection (LLM01), sensitive information and system-prompt disclosure (LLM02, LLM07), retrieval and embedding poisoning (LLM08), tool abuse and excessive agency (LLM06), and improper output handling (LLM05), through to full agentic compromise across CI/CD pipelines and MCP tool servers. You use the same tooling a tester uses, including garak, PyRIT and promptfoo, and validate the skill in the AISRTP certification. You can also drill individual techniques in the PromptStorm cyber range, or start with our AI security training overview.
Frequently asked questions
What is the OWASP Top 10 for LLM Applications?
It is a consensus list, maintained by the OWASP GenAI Security Project, of the ten most critical security risks in applications built on large language models. The current 2025 edition runs from LLM01 Prompt Injection to LLM10 Unbounded Consumption.
What is the number one risk in the OWASP LLM Top 10?
Prompt injection (LLM01). Crafted input makes the model follow attacker instructions instead of the developer's, and in its indirect form it underpins most disclosed LLM incidents, including EchoLeak and the Slack AI data exfiltration.
When was the OWASP LLM Top 10 last updated?
The current edition is the 2025 list, published in November 2024 by the OWASP GenAI Security Project. It added System Prompt Leakage, Vector and Embedding Weaknesses, and Unbounded Consumption relative to the earlier edition.
How is the OWASP LLM Top 10 different from the OWASP Top 10 for web apps?
It follows the same format and numbering idea but covers risks unique to language-model applications, such as prompt injection, retrieval poisoning, and excessive agency, instead of classic web vulnerabilities. Both can apply to the same product.
How do you test an application against the OWASP LLM Top 10?
Through an authorized assessment that maps untrusted input, model permissions, and output handling, then probes prompt injection, information disclosure, retrieval poisoning, and tool abuse. Dedicated practice environments like Pwned Labs labs let you learn the techniques safely and legally.