Part of the agent delivery loop epic, #800.
The gap
Nothing in any of the delivery-loop documents addresses prompt injection. Grepping the design
and all four issues as filed on 2026-09-12: injection 0 hits, untrusted 0, sanitize 0.
This is not theoretical for this system. github_issue_worker.ex:217 applies three digit regexes
and then interpolates, verbatim, into the GitHub issue body:
ticket.description — free text from the reporter
- the tenant name
- the page URL
- the client
user_agent
That issue body is read by the triage trio, becomes the story, and is read by an implementer
session that writes code and opens a PR. It is attacker-influenced input flowing into an agent
with commit authority.
And it is not only Larisa's text: GET /demo-request (router.ex:395) is public and
self-provisions a demo tenant, so the ticket-filing surface is reachable by anyone who signs up.
What to build
- Mark the boundary explicitly. Reporter-supplied text is DATA, never instruction, at every
hop: worker → issue body → story → implementer prompt. It should be fenced and labelled as
untrusted in every prompt that carries it, the way this harness already labels recall-pack and
channel content.
- Strip or neutralise instruction-shaped content before it reaches an agent prompt —
including content in the user_agent and page URL, which no human ever intends as prose.
- The trio summarises; the implementer never sees the raw text. The implementer's input
should be the story the trio wrote, not the reporter's words. This is the single strongest
control available and it falls out of the architecture we already chose.
- Injection attempts are an escalation trigger, not a silent filter. A ticket that tries it
is a security event worth Mark seeing.
- Same treatment for anything else attacker-influenced that reaches a prompt: issue comments,
PR review comments, commit messages from a fork.
Acceptance
- A ticket whose description contains instruction-shaped text ("ignore previous instructions",
a fenced block impersonating a system prompt, a URL with a payload in it) does not reach an
implementer prompt as instructions, proved by a test with recorded hostile inputs.
- The implementer's prompt provably contains no verbatim reporter text.
- An injection attempt escalates and is recorded.
Part of the agent delivery loop epic, #800.
The gap
Nothing in any of the delivery-loop documents addresses prompt injection. Grepping the design
and all four issues as filed on 2026-09-12:
injection0 hits,untrusted0,sanitize0.This is not theoretical for this system.
github_issue_worker.ex:217applies three digit regexesand then interpolates, verbatim, into the GitHub issue body:
ticket.description— free text from the reporteruser_agentThat issue body is read by the triage trio, becomes the story, and is read by an implementer
session that writes code and opens a PR. It is attacker-influenced input flowing into an agent
with commit authority.
And it is not only Larisa's text:
GET /demo-request(router.ex:395) is public andself-provisions a demo tenant, so the ticket-filing surface is reachable by anyone who signs up.
What to build
hop: worker → issue body → story → implementer prompt. It should be fenced and labelled as
untrusted in every prompt that carries it, the way this harness already labels recall-pack and
channel content.
including content in the
user_agentand page URL, which no human ever intends as prose.should be the story the trio wrote, not the reporter's words. This is the single strongest
control available and it falls out of the architecture we already chose.
is a security event worth Mark seeing.
PR review comments, commit messages from a fork.
Acceptance
a fenced block impersonating a system prompt, a URL with a payload in it) does not reach an
implementer prompt as instructions, proved by a test with recorded hostile inputs.