Researchers show that hidden instructions in ordinary documents can steer language model outputs, even in sensitive workflows such as CV screening, and that detection does not reliably remove the effect. Treat all external content as untrusted data, separate it from instructions, and let inferences lead to actions only after human approval.
In a peer-reviewed study published on 23 June 2026, Milani, Franzoni and Florindi describe how malicious instructions hidden in seemingly innocuous documents change a language model's behaviour without the user entering a hostile prompt themselves. The researchers tested 27 CV documents, six consumer-facing language models and four attack scenarios, and found that all systems tested were susceptible to the attacks studied.
What exactly did the researchers establish about indirect prompt injection in documents?
They define indirect prompt injection as instructions embedded in external content that the model is supposed to analyse, such as an uploaded CV. The model does not merely process that text as data to be assessed, but sometimes follows the command hidden within it. Our reading: in a CV screening workflow such an injection can distort the comparison between candidates, because the hidden command can shift the weighting of one candidate against another.
The core of their contribution is a distinction that is often conflated in practice: a model can detect an injection and still remain partly influenced by that injection. Detection and neutralisation are therefore not the same. The researchers point out that simple keyword filtering falls short, because inserted instructions can be semantically plausible or obfuscated. Prompt injection is also known as a leading vulnerability at the OWASP Foundation, the non-profit that maintains the widely used list of top vulnerabilities for LLM applications; the study ties in with that broader security context.
Why does detecting an injection offer no protection against its effect?
Because detection only indicates that something suspicious is present in the content, not that the outcome remains unaffected by it. A warning in the output does not rule out that the hidden instruction has already shifted the judgement. This is precisely where organisations can misjudge their security.
A preprint by Zhu and colleagues dated 4 April 2026 extends this picture to agentic, multi-step tool-calling environments, in which injected content can influence an agent's behaviour and its tool use. Across nine models and four injection vectors they report high hijack rates and find that widely used surface-level defences — prompt warnings, keyword filters, paraphrasing, spotlighting and LLM-as-a-judge (a second model that assesses the output) — often offer limited protection. Their analysis places detection before a tool call above after-the-fact checking. This is a preprint and not yet peer-reviewed; we read it as supporting evidence, not as established fact.
OpenAI's official safety guidance, updated on 7 October 2026, describes prompt injection along the same lines: untrusted text or data that attempts to override an AI system's instructions, with possible consequences such as data leakage, unintended tool calls and actions that deviate from the intent. OpenAI advises keeping untrusted variables out of privileged developer instructions, using structured output, keeping tool approvals on and not letting external data directly steer an agent's behaviour. The company states explicitly that these measures reduce the risk but do not remove it. This is vendor guidance, not independent validation.
What does this mean for directors, lawyers and CISOs who assess documents with AI?
Our analysis: because detection does not guarantee reliable neutralisation, anyone who signs off or accounts for an AI judgement remains personally exposed to a risk that does not disappear through a warning label; that is why every external input a model processes must from now on be treated as potentially compromised and recorded separately. In our analysis, the vulnerability directly affects HR, legal and medical workflows that assess uploaded documents, precisely because the researchers demonstrated the effect in a CV screening — a decision with consequences for individuals; that is why, in our view, this requires that controllers now determine which AI assessments require a human re-assessment before they have effect. Because a manipulated output in an agentic chain can end in an undesirable action, the risk carries through to systems where the model is allowed to read or write; that is why the CISO, in our analysis, must record per workflow which tools and data the model may touch and place consequential actions behind an approval step. And because the study examined a single domain and does not prove that every field shows the same figures, in our assessment it is unwise to confuse the absence of a measured percentage with the absence of risk; that is why auditors and compliance officers, in our analysis, must keep the source text, the model version, the detection results, the approvals and the tool calls per run to make every AI outcome reconstructable, so that it can be verified afterwards which external content, model version and approval led to a decision.
This responsibility sits close to the independent checkpoint between AI analysis and formal decision we described earlier, and to the risk of external content planting a false memory in an AI agent.
Which measures reliably separate untrusted content from AI actions?
The answer lies in layering: separation, limitation and evidence together, not in a single filter. The sources examined and our own practice point in the same direction. Distinguish three things that are often confused: detection (recognising suspicious content), mitigation (preventing that content from changing the result) and containment (preventing a manipulated output from becoming an unauthorised action).
- Classify uploaded files, retrieved web pages, emails and records as untrusted data and keep them out of the model's privileged instruction channels.
- Where possible, extract only predefined fields from a document (schema-bound extraction) rather than letting the model interpret free text.
- Separate analysis from execution: let the model assess, but do not automatically couple a tool call or write action to that assessment.
- Require human approval for consequential read or write actions and apply least privilege on AI agents at runtime.
- Test with adaptive injections that adjust to your defence, not only with fixed test cases.
- Keep the source text, the model version, the detection results, the approvals and the tool calls per run, so that a judgement is reconstructable afterwards.
Anyone working with sensitive files will find more background on this separation between untrusted content and executable instructions in our topic hub on AI security and attack surfaces. The shared message of the three sources is sober: these attacks are real, and no single measure removes them entirely — reason to treat the design of the workflow, not one control, as your defence.
Sources and references
Sources: The article relies on the peer-reviewed study by Milani, Franzoni and Florindi (Neural Computing and Applications), the arXiv preprint by Zhu et al. and OpenAI's safety guidance.