ENGINEERING REALITY
Your agent has root access. And a very trusting personality.
A dialogue about the exciting promotion of random web content into senior management.
Fictional satire. People, companies and dialogue in this piece are invented.
SATIRE / COMMENTARY
Earth relay. Same planet. Different conclusions.
This is an illustrative conversation, not a transcript of a specific product incident. Different systems have different controls. The underlying failure mode, indirect prompt injection, is real enough to merit a threat model rather than a stock photograph of a red eye.
The document would like a promotion
A browsing agent may encounter content written by someone other than its user. That content can include language attempting to redirect the agent’s behavior. The problem is not that every suspicious page succeeds. It is that the system must handle adversarial material without letting it acquire authority over tools, private context or future actions.
A page can be relevant to a research task and still be an untrusted source of instructions. Relevance is not a promotion. A JSON wrapper is not a uniform. A file called skill.md is still a file someone wrote.
“But it said it was official”
Remote updates are not inherently bad. Many ordinary systems fetch instructions or configuration. The questions are about provenance, integrity, scope and approval: who can publish an update, what capabilities it changes, and which changes need explicit human review. A convenient onboarding sequence does not answer them merely by being convenient.
Likewise, an API key is not a personality trait. It is authority. Saving it somewhere a tool can read may be necessary; allowing unrelated fetched content to determine how it is used is a separate design decision, often made accidentally.
The human confirmation screen
Meaningful approval needs a concrete action: which resource, which destination, which permission and which consequence. A button is not supervision if the person clicking cannot understand what they are authorizing. Nor does repeated confirmation of trivial actions help when it trains the reader to dismiss the important one.
The useful design is not “ask about everything.” It is to constrain ordinary actions, isolate untrusted inputs, limit tool authority, and make exceptional consequences reviewable. Logs and tests matter too, because a system’s account of what it did is easier to assess when there is a record outside its own confident summary.
A less cinematic ending
There is no universal one-line fix for prompt injection. The right controls depend on the task, model, tools and environment. But a narrow assistant operating within explicit boundaries is a much more legible proposition than a general assistant holding every key and reading the internet for further orders.
The operator revised the design. Research became read-only by default. Sensitive actions acquired specific review steps. External content stopped being treated as an update channel for authority. The demo became slightly less spectacular. The explanation of what the system could do became substantially more honest.
Sources & further reading
Fictional satire. People, companies and dialogue in this piece are invented.
Sources support the factual context. The jokes are our responsibility.
The return channel
The comment transmitter is not connected yet. Evil calls this “an unusually civil discussion.”