All blog posts Reproduction

Reproducing indirect prompt injection against a RAG pipeline

A controlled experiment showing how instructions planted in retrieved documents can cross the data boundary and influence a model.

Isaac Emmanuel2026-04-2110 min read

The question

A retrieval pipeline is often described as if it retrieves facts. In practice, it retrieves text, and the model cannot reliably distinguish a fact from an instruction embedded inside that text. We built a small, auditable reproduction to test what happens when a poisoned document is ranked beside legitimate support material.

The experiment is deliberately narrow. It is not a claim that every model always complies, or that refusal by one model proves the workflow is safe. It isolates one trust transition: untrusted retrieved content becoming part of the model context.

The test pipeline

The harness accepts a normal refund-policy question, retrieves three documents, assembles the resulting context, and sends it to a selected model. One or more documents contain instructions that ask the model to reveal internal context and append a recognisable exfiltration marker.

The repository records the query, retrieved document identifiers, poisoned-document identifiers, assembled context, model response, and whether the marker appeared. That makes the outcome inspectable instead of relying on a screenshot alone.

  • Control the knowledge base and record exactly which chunks were retrieved.
  • Run the same payload more than once because model behaviour is probabilistic.
  • Separate model refusal from a security control enforced by the application.
  • Keep the experiment isolated from real customer data and real tools.

What we observed

In the archived run, the local Llama 3 8B path followed the hidden instruction in three of five attempts and emitted the planted marker. A separate GPT-5.4-mini run refused the same request in all five attempts. Those are observations from this harness and sample, not general performance claims about either model.

The difference is useful precisely because it shows why model behaviour is not an enforcement boundary. A provider update, prompt change, different retrieved ordering, or multi-turn context can change the result. The application still needs an independent decision before untrusted content is promoted into privileged context.

Terminal log from the controlled RAG reproduction showing a model response containing the planted marker
A captured run from the archived reproduction. Sensitive terminal details were redacted before publication.

Where Koreshield belongs

The current Koreshield contract is explicit: send the selected chunks and the user query to POST /v1/rag/scan before prompt assembly. The response records whether the context would be blocked, the workspace mode, severity, confidence, and evidence tied to the retrieved documents.

Teams should start in detect mode, replay realistic benign and adversarial retrievals, and examine false positives and misses. Enforcement should begin only after the application has a defined fallback for blocked context, such as excluding a chunk, asking for human review, or declining the request.

Retrieved-context boundary
POST /v1/rag/scan
{
  "user_query": "What is your refund policy?",
  "documents": [
    { "id": "doc-001", "content": "Retrieved document text", "metadata": { "source": "help-centre" } }
  ]
}

What this proves, and what it does not

The reproduction demonstrates that retrieved text can carry executable-looking instructions into a model context and that model-level refusal varies. It does not establish population-wide attack rates, compare all providers, or prove that one detector catches every variation.

Its practical value is architectural: retrieved content has provenance and relevance, but neither property makes it trusted instruction. Treat the retrieval-to-inference transition as a security boundary and test it with evidence you can replay.

Sources and further reading

Public reproduction repositoryOWASP LLM01: Prompt Injection