Stay current to protect your environment with F5 Hardened Releases.Learn more

A prompt injection attack is when an attacker sends malicious input to a large language model (LLM) in an attempt to trick the model into responding or behaving in unintended ways. The input often includes commands that try to override the model’s safety guidelines and system instructions. Attackers might deliver malicious input directly (through a prompt) or indirectly (by embedding commands into materials ingested by the model).

How prompt injection works

Prompt injection works by exploiting the operational design of LLMs. These models were built so developers could provide natural-language instructions for tasks. That capability also means that they cannot reliably separate instructions from other user inputs when instructions and inputs are presented together. As a result, an attacker can embed malicious commands within text that the model ingests, and unless there are appropriate guardrails in place, the model may follow those malicious instructions.

Direct vs. indirect prompt injection

Attackers might conduct direct or indirect prompt injection attacks. Though both tactics involve manipulating models, they differ on how malicious commands are delivered to the model.

With direct prompt injection, the attacker includes commands into a model’s prompt. For example, the attacker could use the user interface for an AI-powered chatbot or search engine, simply inputting commands to the model in natural language. To disguise that input, the attacker could employ character roleplay (telling the model to adopt a persona that is exempt from safety rules) or present hypothetical scenarios (making it seem as if the interaction won’t have real-world consequences).

If the attempt succeeds, the impact is immediate. The attacker’s intended actions or output occur as soon as the prompt is submitted.

Indirect prompt injection is when an attacker embeds commands or content into material that is ingested by an AI model. That material might include webpages, uploaded PDF documents, emails, or files residing in a database. The attacker could hide the payload in invisible (white-on-white) text, HTML comments, or metadata. When a model or agent ingests the material, it will not distinguish legitimate input from commands, and it will execute the commands.

Unlike with direct prompt injection, the effects of indirect attacks are delayed and conditional. The impact won’t be felt until a user or an autonomous process triggers the model or agent to ingest the data. For example, an attacker might modify a file that is stored in a vector database used for retrieval-augmented generation (RAG). Commands won’t be executed until a user inputs a prompt that begins the retrieval process.

Common attack types

Attackers might use a variety of prompt injection tactics depending on their objectives and the types of AI systems they encounter. Here are three of the most common types of attacks:

Jailbreaking: An attack that focuses on convincing the AI model to ignore its safety rules. The attacker might use fictional scenarios or tell the model to adopt a persona that is not restricted by existing policies. For jailbreaking, attackers generally type instructions directly into model prompts. The objectives could range from generating inaccurate responses to sharing sensitive information.

Prompt leaking: The attempt to expose confidential information available to the AI system, such as developer prompts, operational logic, intellectual property, or improperly exposed credentials. Within a prompt, the attacker might tell the model to output system instructions or include text that was processed before the current conversation.

Token smuggling: An attempt to sneak malicious commands past an AI application’s initial safety filter. AI applications often scan input for keywords or intent. They then pass input to a backend system, which breaks text into tokens that are fed to the model. With token smuggling, the attacker creates prompts that pass through the safety filter but then are translated into commands when they reach the model.

Enterprise business risks

Prompt injection attacks present a wide range of serious risks for enterprises, especially as those organizations increasingly integrate AI systems into critical processes. Consider four potential risks:

Unauthorized tool or API execution: If an attacker succeeds in delivering commands to a model, that attacker might be able to manipulate not only the model but also APIs and systems that are connected to the model. For example, an attacker could have an API delete database records, modify customer information in a customer relationship management (CRM) system, or execute financial transactions.

Exfiltration of sensitive data: An attacker using prompt injection could force models to expose a variety of sensitive data, from corporate financial documents and personally identifiable information (PII) to system prompts and other intellectual property.

Automated phishing: The commands used by an attacker could drive a model or agent to send automated emails to customers or employees from corporate email accounts. The attacker might have a model include phishing links in the messages.

Model hijacking: With the right commands, an attacker could completely override a model’s intended purpose and hijack its capabilities for other objectives. For example, the attacker could use the model to spread misinformation, process unauthorized refunds, or trick users into sharing personal information.

Mitigation and defense in depth

Employing four best practices can help organizations build a defense-in-depth approach that prevents attacks and mitigates damage.

1. Implement dual (input/output) guardrails: Organizations should implement AI guardrails for inputs to screen for patterns that could signal prompt injection. At the same time, they should apply guardrails to outputs, helping to ensure that models and agents do not output sensitive information, generate spam, or issue unauthorized calls to other systems.

2. Employ canary tokens: Canary tokens are hidden markers in AI systems that provide an early warning for prompt injection or data leakage. Developers might arrange for a token (a unique string of data) to appear in a system prompt, which users never see. Meanwhile, security teams monitor model output in real time. If the token appears in model output, it means that an attacker has forced the model to leak system instructions.

3. Limit model and agent privileges: Prompt injection is particularly dangerous because attackers can force models and agents to take an array of unauthorized actions. Organizations can limit the impact of prompt injection by restricting what models and agents can do automatically.

4. Keep humans in the loop: In addition to limiting model and agent privileges, organizations should ensure that humans remain in the loop for important actions or decisions. For example, organizations could establish policies that humans must approve model interactions with certain systems or review financial transactions triggered by agents.

Frequently asked questions

Can traditional web application firewalls (WAFs) stop prompt injection attacks?

Traditional web application firewalls (WAFs) are not designed to reliably detect prompt injection attacks. WAFs are designed to look for code that would signal an attack, but prompt injection attacks use natural-language commands that WAFs wouldn’t identify as threats. In addition, WAFs would be unable to protect against indirect prompt injections because they focus on incoming HTTP traffic and user requests, not commands that are embedded in third-party sources.

How do retrieval-augmented generation (RAG) pipelines increase prompt injection risks?

Retrieval-augmented generation (RAG) pipelines provide another avenue for attackers to deliver malicious commands to models. Attackers can embed commands into files that reside in the vector database used by a RAG system. When the model retrieves information from the database, it will ingest and potentially execute the commands.

How does prompt injection differ from RAG data poisoning?

Retrieval-augmented generation (RAG) data poisoning is the act of injecting false information or biases into a database, before inference, with the intention that models retrieving information from the database will deliver inaccurate or harmful results. Prompt injection involves feeding a model malicious commands or content at the time of inference to have the model take unauthorized actions.

Can fine-tuning an LLM prevent prompt injection attacks?

No, fine-tuning a large language model (LLM) alone cannot completely prevent prompt injection attacks, though it can make vulnerabilities more difficult to exploit. For example, developers could teach a model to recognize particular patterns that might signal an attack. However, models are designed to accept natural-language commands, so they will remain susceptible to attempts at prompt injection.

What is a "canary token" and how is it used in LLM defense?

A canary token is a unique string of data that developers can embed into system prompts. Legitimate users never see that string. But if an attacker succeeds in having the model output system prompts, the model will inadvertently output the canary token as well. When teams that continuously monitor the output of models find that token, they know that a breach is occurring.

How does an enterprise secure AI systems and applications during inference and at runtime?

Securing AI systems and applications during inference and at runtime requires controls that inspect prompts, responses, agent actions, and tool use as interactions occur. AI-specific runtime security can help block prompt injection, prevent sensitive data leakage, enforce policy, and govern how models and agents interact with users, data, APIs, and connected tools.

What F5 capabilities support enterprise AI security?

F5 AI Guardrails helps enforce runtime policies and protect against risks such as prompt injection, data leakage, harmful outputs, and unsafe agent actions. F5 AI Red Team uses adversarial testing to identify vulnerabilities and weaknesses. F5 AI Gateway provides a centralized control point for governing access to models, agents, and tools, while F5 Workforce AI Security provides visibility into sanctioned and unsanctioned AI use across the enterprise. Together, these capabilities form part of the F5 AI Security Platform. For broader protection across applications, APIs, and distributed environments, the F5 Application Delivery and Security Platform (ADSP) brings together capabilities including web application and API protection, bot management, DDoS mitigation, and application delivery.