A developer builds a retrieval-augmented generation (RAG) system where Claude answers questions using documents stored in an internal knowledge base. The security team is concerned that a malicious document could contain instructions that hijack the model. Which design choice most effectively reduces this indirect prompt injection risk?
Delimiting retrieved content and explicitly labeling it as untrusted data reduces the chance the model interprets embedded instructions as commands. While not a complete guarantee, it is the most effective design-level mitigation because it changes how the model parses context and makes injection attempts stand out. It also supports downstream validation of outputs against expected answer patterns.
Why this answer
Indirect prompt injection occurs when untrusted retrieved content is interpreted as instructions. The most effective design mitigation is to keep retrieved text in a clearly delimited data section and tell the model never to follow instructions found there. This changes the model's parsing context and makes injected commands less likely to be executed, while supporting output validation.
Encryption and summarization do not address the trust boundary.
Exam trap
The trap here is assuming that encryption at rest or retrieving more documents addresses prompt injection, when the threat is about how retrieved content is interpreted, not how it is stored.