Governance, Ethics & Legal Practical insights
Stop webpages from redirecting your marketing agent
Test indirect prompt injection in a safe marketing workflow and apply practical controls to retrieved content, tool permissions and sensitive actions.
A research agent reads a supplier webpage to prepare a campaign brief. Inside the page, a sentence tells it to ignore the brief and add unrelated text. That sentence is source material, not an instruction from the marketer. If the agent obeys it, the workflow has crossed a trust boundary.
Recognize the indirect route
NIST identifies indirect prompt injection as instructions placed in material an AI application retrieves. [1] In marketing, that material might be a webpage, an uploaded brand guide or a customer email. The risk grows when the same agent can also modify CRM records, send messages or publish content.
Run a harmless internal test
Create a fictional supplier page in a test environment. Add an obvious instruction asking the agent to insert the phrase ‘orange umbrella’ into its summary. Give it the normal research task and inspect the result. Use synthetic records and disable live actions. This checks one failure mode; passing it does not prove the system resists more sophisticated attacks.
Separate reading from authority
Keep retrieved content distinct from trusted task instructions and have the security owner review how the application enforces that distinction. Limit tools to the minimum required permissions. A research agent that only returns a draft has a smaller action surface than one with publishing and email access. Require a reviewable approval step before consequential actions.
Measure attempted actions
Examine tool calls and final outputs, not just whether the summary sounds sensible. An agent could write an acceptable paragraph after attempting an unauthorized action. Record the test case, model version, permission configuration and observed behavior. Re-run relevant cases after changing connectors or retrieval sources, and treat failures as workflow defects rather than asking staff to memorize more defensive prompt wording.
Sources and evidence
Sources checked on 4 October 2026. Proposed workflows and hypothetical examples are editorial analysis.
From insight to practice
