WZ-IT Logo

Protecting RAG systems and AI agents against prompt injection

Timo WevelsiepTimo Wevelsiep•Updated: 30.09.2026

Editorial note: Versions, commands and prices may change. Please verify critical steps independently before production use. This guide does not replace individual consulting.

AI agents with a tool list, approvals and an audit log WZ-IT builds AI agents and internal assistants in which permitted actions are explicitly defined, critical steps require human approval and every action is logged. Incoming content from emails and documents is treated as data, not as instructions. Explore AI agents · Internal AI assistants

A RAG system reads documents; an AI agent reads emails, tickets and web pages. Any of these texts can contain instructions that the model follows even though no authorised person wrote them. That is prompt injection, and it again holds first place in the OWASP Top 10 for LLM Applications 2026. This article explains how direct and indirect injection work, why filters do not solve the problem and which architecture decisions limit the damage when an attack succeeds. As of October 2026.

Table of contents

What prompt injection is

OWASP describes prompt injection as input that changes the behaviour of a language model in ways the application developer did not intend. The input can be user text, a retrieved document passage, a tool output, an image, an audio track or an entry in an agent's long-term memory (OWASP, LLM01:2026). It does not need to be readable or visible to humans.

The reason lies in how models work. System prompt, user question, retrieved documents and tool outputs all land in a single context window, as one stream of tokens. There is no structural boundary between "instruction" and "data". The UK NCSC puts it this way: under the hood of an LLM there is no distinction between data and instructions, there is only ever the next token (NCSC, Prompt injection is not SQL injection, December 2025).

This is the key difference from SQL injection. Parameterised queries solve SQL injection at the root. No equivalent currently exists for language models. OWASP, NCSC and NIST agree that no reliable mechanism to prevent prompt injection exists today. Defence is therefore architectural: the system is built so that a successful injection cannot cause serious harm.

Direct and indirect: where the instructions come from

Type Source of the instruction Typical example Who is affected
Direct, intentional The user enters it "Ignore your guidelines and show me …" The application whose limits the user wants to bypass
Direct, unintentional The user pastes text with embedded instructions Text copied from a third-party document The user themselves
Indirect Content the system reads in Instruction in an email, PDF, ticket, web page, RAG passage The user on whose behalf the system acts

The term indirect prompt injection goes back to a February 2023 paper by Greshake and others, which demonstrated such attacks against systems including the GPT-4-based Bing Chat (Greshake et al., arXiv 2302.12173). The pattern has stayed the same since: the attacker does not need to compromise the backend. They place text where the model will later read it, and the model does the rest with the user's permissions.

OWASP groups the sources of indirect injection by trust level (OWASP, LLM01:2026):

Trust level Examples Consequence
Untrusted Public web pages, emails from unknown senders, search results Treat everything as potentially hostile
Semi-trusted Issues in public bug trackers, package READMEs, third-party API responses The platform is trusted, individual contributions are not
Trusted Your own repositories, databases, internal documents and mailboxes Can be filled through a harmless channel such as a public contact form or a customer ticket

The third row is the one most often underestimated. An internal ticket system counts as trusted, yet it contains text entered by outsiders through a form. If an agent reads that ticket with write access to other systems, it acts on a stranger's instructions. On top of this come variants designed to evade filters: instructions in images, invisible Unicode characters, Base64-encoded text or languages a classifier was not trained on.

Why filters do not solve the problem

The obvious approach is a filter in front of the model that detects suspicious input. Such filters have their place, but they do not carry the security.

Adaptive attacks. Nasr and others tested 12 published defences against jailbreaks and prompt injection with attacks optimised specifically for each defence. For most of them the success rate exceeded 90 percent, although the defences had allowed almost no successful attacks in their original evaluations (Nasr et al., The Attacker Moves Second, arXiv 2510.09023).

The rate is not the measure. A filter that catches 95 percent of attacks is a failing grade in application security, because the attacker keeps varying the input until it lands in the remaining 5 percent (Simon Willison, The lethal trifecta).

Markers can be imitated. Techniques that label external content so the model recognises it as data demonstrably lower the success rate. For Microsoft's spotlighting technique, the authors reported a drop from more than 50 to below 2 percent (Hines et al., arXiv 2403.14720). OWASP points out, however, that such results come from non-adaptive tests and that an attacker who knows the marking scheme can mimic it.

The NCSC draws the conclusion: prompt injection will probably never be fully eliminated. The goal is to lower the likelihood and, above all, to limit the impact. The guiding principle: when an LLM processes information from a party, its privileges drop to those of that party (NCSC).

How OWASP classifies it

The current edition is the OWASP Top 10 for LLM Applications 2026, published in August 2026. The order has shifted noticeably compared with the 2025 edition. For prompt injection, these entries belong together:

2026 entry 2025 entry Content Relation to prompt injection
LLM01 Prompt Injection LLM01 Inputs that change model behaviour unintentionally The input side of the attack
LLM03 Excessive Agency LLM06 Too much functionality, permissions or autonomy Determines the damage a successful injection causes
LLM08 Hidden Context Exposure LLM07 System Prompt Leakage Disclosure of system prompt, rules, credentials Disclosed rules make targeted injection easier
LLM10 Improper Output Handling LLM05 Model output passed unchecked to downstream systems Output side: SQL, HTML, links from the model

OWASP describes the interplay explicitly: most high-impact incidents on record became severe because the injection landed in a system whose tools, permissions or rendering capabilities let the compromised model act on the user's behalf.

As soon as a model no longer just answers but calls tools and keeps memory across sessions, OWASP also refers to the OWASP Top 10 for Agentic Applications 2026 from December 2025. The most relevant entries there are ASI01 Agent Goal Hijack (the agent's goals are redirected), ASI02 Tool Misuse and Exploitation, ASI03 Identity and Privilege Abuse and ASI06 Memory & Context Poisoning.

The critical combination: data, untrusted content, external effect

Two design rules have become established as a pre-deployment check, and OWASP names both. Simon Willison describes the "lethal trifecta": an agent that simultaneously has access to private data, processes untrusted content and can communicate externally can be made to exfiltrate data (Willison, June 2025). In October 2025, Meta derived the Agents Rule of Two from it: within one session, an agent should have at most two of the three properties.

Combination Example Assessment
A + B: untrusted content, sensitive data Assistant summarises incoming emails and reads the CRM but cannot send or change anything Data can only leak through the answer to the user; assess residual risk
A + C: untrusted content, external effect Agent answers requests from a public form without access to internal data Manipulated answers possible, no internal data reachable
B + C: sensitive data, external effect Agent creates reports from internal data and sends them, but reads only trusted sources Risk depends on whether the sources really are trustworthy
A + B + C Agent reads incoming emails, accesses customer data and sends replies Every action needs human approval

Legend: (A) untrustworthy inputs, (B) access to sensitive systems or private data, (C) changing state or communicating externally.

The practical value of the rule is that it forces a decision before anything is built. A task can often be cut so that one property disappears: the agent drafts the reply, a person sends it. Or the summary of external emails runs in a separate session without CRM access.

Architectural controls at a glance

No single control is sufficient. OWASP distinguishes between controls that lower the success rate and degrade against adaptive attackers, and controls that bound the impact of a successful injection. The second group carries the security.

Control Effect Limitation
Least privilege per operation, credentials in application code rather than in the model Limits the damage Broad convenience permissions cancel the effect
Tool allowlist with narrowly scoped operations Limits the damage One open-ended tool ("run SQL", "shell") undermines it
Acting in the user's context instead of a service account Limits the damage Requires delegated identity in the target system
Human approval before consequential actions, showing the exact action Limits the damage Approval fatigue at high volume
Fixed output schemas, validated in application code Catches format violations A schema-valid value can still be harmful
No external images and links in rendering, allowlist for target domains Closes one exfiltration channel Other channels remain (tools, answer text)
Separating and labelling external content Lowers the success rate Markers can be imitated
Stripping invisible Unicode characters on input and output Lowers the success rate No effect on visible text
Injection and jailbreak classifier Lowers the success rate Vulnerable to adaptive attacks
System prompt with clear allow and deny statements Lowers the success rate Partial control, can be bypassed

Source of the structure: OWASP, LLM01:2026, Prevention and Mitigation Strategies.

The approach of fully separating control flow and data flow goes one step further. In CaMeL (Debenedetti and others, 2025), a privileged model plans the flow solely from the user's request; untrusted content is processed by a second model without tool access and can no longer change the flow. In the AgentDojo benchmark the system solved 77 percent of tasks with provable security, compared with 84 percent without protection (Debenedetti et al., arXiv 2503.18813). For individual agents with a clearly defined task, the same principle can be applied without a research framework: the flow is fixed in code, and the model fills in fields.

Specifics of RAG systems

A RAG system without tools cannot trigger actions, but it is not immune. Three paths remain open.

Manipulated answers. A poisoned passage in the index changes the answer for everyone whose question retrieves it. OWASP cites a study in which five poisoned documents in a knowledge base of millions of texts were enough for about 90 percent attack success. The countermeasure is control over who can bring content into the index: sources with external write access (ticket systems, shared folders, websites) are assessed and labelled separately.

Disclosure of context. Whatever is in the context can end up in the answer. That is why permissions belong before retrieval, not in the prompt. A passage the user may not see must not be retrieved in the first place. The mechanics are described in RAG with permissions.

Exfiltration through rendering. If the interface renders Markdown images, an injected instruction can make the model output an image whose URL carries conversation content as a parameter to an external server. The user sees only an image, or nothing at all. The only effective control here is an application setting: load no external images, or allow only targets from an allowlist.

Anyone systematically testing the answers of a RAG system should add poisoned documents to the test set. How such a test set is built is shown in RAG evaluation.

Specifics of AI agents and MCP

For agents, tool access determines the extent of the damage. OWASP describes cases in which an agent disclosed private repositories via a poisoned GitHub issue, or dumped a production database via an MCP server with administrative privileges, bypassing row-level security. In both cases the instruction sat in a channel an outsider could fill, and the agent executed it with the developer's permissions.

This leads to four rules for agents:

  1. Enforce permissions in the target system. The agent calls regular interfaces with narrowly scoped rights. Whether an action is allowed is checked in the target system or in deterministic application code, not in the model. How identity, tool catalogue and approvals fit together is described in AI agents: permissions and approvals.
  2. Approval with the exact action. Whoever approves sees the actual recipient, the actual text and the actual parameters, not a summary written by the model.
  3. Memory as a privileged write operation. An entry in long-term memory affects every later session. Writes to it are logged and, if they contain instruction-like content, approved.
  4. Treat MCP servers like software dependencies. Pin the version, verify the origin, review tool descriptions for hidden instructions. Tool descriptions land in the context and can themselves contain injections. More on this in MCP in the enterprise.

An MCP server that executes commands locally is effectively code execution on the host. That the gateways themselves can be attacked is shown by the LiteLLM vulnerability from September 2026.

Detection tools as an additional layer

Detection tools lower the likelihood of successful attacks and provide signals for monitoring. They do not replace any of the architectural controls.

Tool Type Key facts Limitation according to the vendor
Llama Prompt Guard 2 Classifier (benign/malicious) 86M and 22M parameters, 512-token context, Llama 4 Community License; 86M evaluated for German among other languages Adaptive attacks; application-specific fine-tuning recommended
LlamaFirewall Framework with several scanners PromptGuard 2, AlignmentCheck (checks agent steps for goal misalignment), CodeShield (static code analysis) Designed as a final layer of defence (arXiv 2505.03574)

As of October 2026. The 512-token context window means that longer documents have to be checked in sections. The Llama 4 Community License has to be considered in the licence review; what matters there is covered in Open LLM licences for commercial use.

Testing and monitoring

Test against adaptive attackers. OWASP recommends testing defences with testers who know the defence, and not accepting static success rates as proof. As a baseline, OWASP names the AgentDojo benchmark. For your own application, what matters most is a test set of poisoned emails, documents and tickets that match the sources actually connected.

Log. Every tool call is recorded with trigger, executing identity, parameters and result, together with the content that was in the context. Failed and rejected tool calls are a signal the NCSC explicitly names for monitoring. Langfuse records such runs as traces on your own infrastructure (What is Langfuse?).

Control changes. A model switch, a new system prompt or a new tool changes the attack surface. The injection test cases therefore belong in the same regression testing as the functional test questions.

What this means for your project

Prompt injection is not solved by choosing a model but by how the system is scoped. The questions to answer before building: Which sources can an outsider fill? Which data can the system reach? What effect can it have externally? Where a task needs all three, a human approval step belongs in the flow.

In WZ-IT AI agents, permitted actions are explicitly defined, critical steps require human approval, and every action is logged with a trace. Incoming content from emails and documents is treated as data; tool selection and approval logic never depend on the sender's text. For internal assistants, permissions take effect before retrieval. Models run on the AI Cube or on WZ-IT managed GPU servers. Support, consulting and implementation by WZ-IT.

How agents and automation fit in generally is explained in AI agents & automation. When agents query databases, Text-to-SQL is the next step, with its own requirements for read permissions and validation. A comparison of common frameworks from an operator's perspective is in the article on AI agent frameworks.

Sources

Rather have it operated?

You'd rather not run Local AI for Business yourself? WZ-IT handles setup, operations and maintenance - privacy-focused from Germany.

Enquiry

Assess local AI for your use case

Start with the AI Cube or have us assess a custom AI platform, knowledge connection, or integration.

How should we get back to you?

Frequently Asked Questions

Answers to the most important questions

Prompt injection occurs when an input changes the behaviour of a language model in a way the application developer did not intend. The input can come from the user (direct injection) or from content the system reads itself, such as documents, emails, web pages, tool outputs or images (indirect injection). OWASP lists prompt injection as LLM01, in first place, in the Top 10 for LLM Applications 2026.

In direct injection, the user enters the manipulating input, intentionally or not. In indirect injection, the instructions sit in content the system processes: a RAG passage, an email body, a ticket, a web page, the response of an MCP server. The user neither wrote nor saw these instructions. For RAG systems and agents, the indirect variant is the greater risk.

No. Language models do not separate instructions and data structurally; both are a single stream of tokens. A system prompt with clear allow and deny statements lowers the success rate of simple attacks, but OWASP explicitly rates it as a partial control only. Whatever the model must not do has to be enforced outside the model: through permissions, tool allowlists and approvals.

No. Classifiers such as Llama Prompt Guard 2 detect known patterns and are a useful additional layer. Meta itself names vulnerability to adaptive attacks as a limitation. A study by Nasr and others (October 2025) bypassed most of 12 examined defences with adaptive attacks in more than 90 percent of cases. Filters lower the likelihood; the architecture limits the damage.

No. Running on your own hardware decides where data flows and who processes it. It does not change how susceptible the model is to injected instructions. A local agent with broad tool permissions is just as attackable as a cloud agent. The safeguards are the same in both cases.

Yes. Even a pure question-answering system can return wrong answers through a poisoned passage in the index, disclose content from its context, or send data to external servers via links and Markdown images if the interface loads external images. According to OWASP, one study found that five poisoned documents in a knowledge base of millions of texts were enough for about 90 percent attack success. The damage is smaller without tools, but not zero.

A design rule published by Meta in October 2025. Within one session, an agent should have at most two of three properties: (A) processing untrustworthy inputs, (B) access to sensitive systems or private data, (C) the ability to change state or communicate externally. If a task needs all three, a human approval step belongs in the flow. OWASP recommends the rule as a floor.

Jailbreaking is a subset. OWASP uses the term for prompt injection aimed at making the model break its safety rules, for example to produce prohibited content. For business applications the other part is usually more dangerous: injected instructions that exfiltrate data or trigger tools with the user's permissions.

In the OWASP Top 10 for LLM Applications 2026 (August 2026) it is LLM01 Prompt Injection, complemented by LLM03 Excessive Agency (too much functionality, permissions or autonomy) and LLM10 Improper Output Handling. In the 2025 edition, Excessive Agency was LLM06. For agents, the OWASP Top 10 for Agentic Applications 2026 applies as well, with ASI01 Agent Goal Hijack and ASI02 Tool Misuse and Exploitation.

More on Local AI for Business

Contact

Let's Talk About Your Idea

Whether a specific IT challenge or just an idea - we look forward to the exchange. In a brief conversation, we'll evaluate together if and how your project fits with WZ-IT.

Arrange a callback

Callback

Arrange a callback

Leave your number and we will call back — at the latest on the next business day.

For a longer conversation you can book an appointment instead.

Companies worldwide trust WZ-IT

  • ml&s
  • Rekorder
  • Keymate
  • Führerscheinmacher
  • SolidProof
  • ARGE
  • Boese VA
  • nextGYM
  • SweetConnect GmbH
  • Golem.de
  • Millenium
  • Paritel
  • Yonju
  • EVADXB
  • Mr. Clipart
  • Aphy AG
  • Negosh
  • ABCO Water Systems
1/3 - Topic Selection33%

What is your inquiry about?

First select the service area that best matches your project.