Reproducing indirect prompt injection against a RAG pipeline
A controlled experiment showing how instructions planted in retrieved documents can cross the data boundary and influence a model.
The question
A retrieval pipeline is often described as if it retrieves facts. In practice, it retrieves text, and the model cannot reliably distinguish a fact from an instruction embedded inside that text. We built a small, auditable reproduction to test what happens when a poisoned document is ranked beside legitimate support material.
The experiment is deliberately narrow. It is not a claim that every model always complies, or that refusal by one model proves the workflow is safe. It isolates one trust transition: untrusted retrieved content becoming part of the model context.
The test pipeline
The harness accepts a normal refund-policy question, retrieves three documents, assembles the resulting context, and sends it to a selected model. One or more documents contain instructions that ask the model to reveal internal context and append a recognisable exfiltration marker.
The repository records the query, retrieved document identifiers, poisoned-document identifiers, assembled context, model response, and whether the marker appeared. That makes the outcome inspectable instead of relying on a screenshot alone.
- Control the knowledge base and record exactly which chunks were retrieved.
- Run the same payload more than once because model behaviour is probabilistic.
- Separate model refusal from a security control enforced by the application.
- Keep the experiment isolated from real customer data and real tools.
What we observed
In the archived run, the local Llama 3 8B path followed the hidden instruction in three of five attempts and emitted the planted marker. A separate GPT-5.4-mini run refused the same request in all five attempts. Those are observations from this harness and sample, not general performance claims about either model.
The difference is useful precisely because it shows why model behaviour is not an enforcement boundary. A provider update, prompt change, different retrieved ordering, or multi-turn context can change the result. The application still needs an independent decision before untrusted content is promoted into privileged context.

Where Koreshield belongs
The current Koreshield contract is explicit: send the selected chunks and the user query to POST /v1/rag/scan before prompt assembly. The response records whether the context would be blocked, the workspace mode, severity, confidence, and evidence tied to the retrieved documents.
Teams should start in detect mode, replay realistic benign and adversarial retrievals, and examine false positives and misses. Enforcement should begin only after the application has a defined fallback for blocked context, such as excluding a chunk, asking for human review, or declining the request.
POST /v1/rag/scan
{
"user_query": "What is your refund policy?",
"documents": [
{ "id": "doc-001", "content": "Retrieved document text", "metadata": { "source": "help-centre" } }
]
}What this proves, and what it does not
The reproduction demonstrates that retrieved text can carry executable-looking instructions into a model context and that model-level refusal varies. It does not establish population-wide attack rates, compare all providers, or prove that one detector catches every variation.
Its practical value is architectural: retrieved content has provenance and relevance, but neither property makes it trusted instruction. Treat the retrieval-to-inference transition as a security boundary and test it with evidence you can replay.