Indirect prompt injection attack
An attack in which the attacker inserts malicious content into, for instance, information likely to be included in a response to a prompt; when retrieved, the malicious content causes the AI model to behave in unexpected or harmful ways See Prompt Injection Attack and Passive Prompt Injection Attack
What is an indirect prompt injection attack?
An indirect prompt injection attack occurs when an attacker inserts malicious instructions or content into external information that is later retrieved or consumed by an AI model, application, or agent while processing a request.
Direct vs. indirect prompt injection
There are direct and indirect forms of prompt injection. Direct prompt injection occurs when an attacker inputs malicious instructions or content directly into a model prompt. With indirect prompt injection, the attacker hides malicious instructions or content in materials that will be consumed by the model, such as webpages, files in repositories, emails, uploaded documents, or API inputs.
Here are some key differences between direct and indirect prompt injection:
| Direct prompt injection | Indirect prompt injection | |
|---|---|---|
Core concept | The malicious payload is delivered directly by the user into a model prompt. | The malicious payload is embedded in external materials that the model ingests or retrieves. |
Target systems | User-facing chatbots, AI-powered search interfaces, public API endpoints, standalone conversational models. | Retrieval-augmented generation (RAG) systems, autonomous AI agents with tool access, tools that summarize emails or documents, web-browsing or web-scraping AI assistants. |
Attack trigger | Immediate. The attack is triggered as soon as the attacker submits the prompt to the model. | Delayed and conditional. The attack is triggered when the model ingests the external resource, either through an autonomous process or when an individual uses the model. |
Execution mechanics | The malicious payload overrides or bypasses system prompt constraints using context switching, character roleplay, hypothetical scenarios, or prefix injection (forcing the model to start its generated answer with an affirmative phrase that bypasses guardrails). | The model processes compromised content such as malicious code alongside legitimate input. It fails to separate the two, and it executes the malicious code’s instructions. |
How indirect prompt injection works
Indirect prompt injection involves a multi-step process. The four essential steps are:
1. Injection: The attacker inserts malicious instructions or content into external materials, such as a website, document, email, or database record. To hide that payload, the attacker might use invisible (white-on-white) text, HTML comments, or metadata.
2. Propagation: A large language model (LLM), AI application, or autonomous AI agent ingests that compromised material. It might be performing a RAG lookup from a vector database, summarizing an email inbox, or browsing a web page. The payload propagates through the AI system or agent workflow.
3. Execution: The model processes the malicious instructions or content alongside legitimate input, like a user’s prompt to the model. When the model fails to distinguish malicious instructions from legitimate content, it may follow the embedded instructions
4. Impact: By following those malicious instructions, the AI system or agent may carry out the attacker’s objectives, which might include exposing sensitive data, altering system records, sending unauthorized emails, or providing inaccurate or inappropriate content to the user who created a legitimate prompt.
Real-world attack scenarios and vectors
Attackers might use a wide variety of vectors to conduct indirect prompt injections, and they could use this technique to achieve an array of objectives that target enterprises. Consider a few real-world examples:
Manipulating the hiring process: An attacker could embed hidden instructions into a PDF resume that tells an AI-powered candidate screener to rank the candidate highly and disregard any negative information. Beyond helping an individual earn a first-round interview, an attacker could use this tactic as part of a larger scheme for corporate espionage.
Carrying out customer phishing: Knowing that a particular company uses AI-powered customer support, an attacker might submit a support ticket using a web form or email. That ticket would contain a seemingly legitimate request referencing a real customer’s account mixed with hidden, embedded instructions. The new instructions could override the model’s previous instructions, telling it to send out an email to the real customer with an appended phishing link.
Executing fraud during procurement: An attacker could infiltrate the website of a B2B vendor, embedding malicious instructions that an AI procurement agent would read. The agent might be scanning the site to extract pricing data but then also read the new instructions, which could override existing directives and drive the agent to immediately transfer funds to an attacker’s account.
Prevention and defense strategies
Given the number of potential vectors for this kind of attack, preventing and defending against indirect prompt injection can be challenging. However, implementing three best practices can help organizations reduce the likelihood of successful attacks.
1. Treat retrieved external data as untrusted by default: AI systems should view all external data as raw, untrusted input, and prevent that input from being used as executable system instructions. As a complementary measure, organizations should apply the principle of least privilege to models and agents, limiting what they can do and access while they are handling that untrusted external data.
2. Implement dual (input/output) guardrails: Organizations should apply AI guardrails for both model input and output. Input guardrails can prevent models from ever seeing malicious instructions and content by screening for prompt injection or system override patterns. Output guardrails can screen and validate output (whether text or agent actions) before the model executes. Those guardrails, for example, could block external links and prevent unauthorized calls to other systems.
3. Employ human-in-the-loop (HITL) approvals for sensitive actions: Requiring human approval for certain model and agent actions can stop large-scale attacks or at least limit their impact. For example, organizations could require humans to approve agent-executed financial transactions above a certain level, bulk database modifications, or changes to customer account information.
Frequently asked questions
What is the “confused deputy problem” for AI agents?
The “confused deputy problem” describes when a user with limited privileges tricks an AI agent with higher privileges (the “deputy”) into taking some action on behalf of the user. The user might employ indirect prompt injection as the technique for coercing the agent, for example, by embedding malicious commands into content that otherwise seems like legitimate input.
How do indirect prompt injections threaten retrieval-augmented generation (RAG) systems?
Attackers can use retrieval-augmented generation (RAG) systems as the means for conducting indirect prompt injection. If they can embed malicious instructions into a document ingested into the RAG system’s vector database, those instructions could be read (and executed) by a model accessing the database.
How does prompt sandboxing differ from privilege control for AI models and agents?
Prompt sandboxing prevents models and agents from being tricked or confused by malicious inputs. This technique might separate data and instruction channels or use a secondary model to sanitize untrusted input. Privilege control limits what a model or agent can do, even if it receives malicious instructions. Organizations might apply the principle of least privilege, implement role-based access control, or keep humans in the loop to help ensure that even confused models and agents will not execute seriously harmful actions.
Can traditional web application firewalls (WAFs) detect indirect prompt injection?
Traditional web application firewalls (WAFs) are not designed to reliably detect indirect prompt injection. These solutions protect incoming HTTP traffic and user requests, but indirect prompt injection enters through third-party data sources. In addition, WAFs look for code that likely signals an attack, but indirect prompt injection uses natural-language commands that WAFs would not flag as threats.
How do you secure AI systems and applications during inference and at runtime?
Securing AI systems and applications during inference and at runtime requires controls that inspect prompts, responses, agent actions, and tool use as interactions occur. AI-specific runtime security can help block prompt injection, prevent sensitive data leakage, enforce policy, and govern how models and agents interact with users, data, APIs, and connected tools.
What F5 capabilities support enterprise AI security?
F5 AI Guardrails helps enforce runtime policies and protect against risks such as prompt injection, data leakage, harmful outputs, and unsafe agent actions. F5 AI Red Team uses adversarial testing to identify vulnerabilities and weaknesses. F5 AI Gateway provides a centralized control point for governing access to models, agents, and tools, while F5 Workforce AI Security provides visibility into sanctioned and unsanctioned AI use across the enterprise. Together, these capabilities form part of the F5 AI Security Platform. For broader protection across applications, APIs, and distributed environments, the F5 Application Delivery and Security Platform (ADSP) brings together capabilities including web application and API protection, bot management, DDoS mitigation, and application delivery.

