Skip to content

Treat reporter-supplied ticket text as untrusted input end to end (prompt injection) #804

Description

@mkreyman

Part of the agent delivery loop epic, #800.

The gap

Nothing in any of the delivery-loop documents addresses prompt injection. Grepping the design
and all four issues as filed on 2026-09-12: injection 0 hits, untrusted 0, sanitize 0.

This is not theoretical for this system. github_issue_worker.ex:217 applies three digit regexes
and then interpolates, verbatim, into the GitHub issue body:

  • ticket.description — free text from the reporter
  • the tenant name
  • the page URL
  • the client user_agent

That issue body is read by the triage trio, becomes the story, and is read by an implementer
session that writes code and opens a PR. It is attacker-influenced input flowing into an agent
with commit authority.

And it is not only Larisa's text: GET /demo-request (router.ex:395) is public and
self-provisions a demo tenant, so the ticket-filing surface is reachable by anyone who signs up.

What to build

  1. Mark the boundary explicitly. Reporter-supplied text is DATA, never instruction, at every
    hop: worker → issue body → story → implementer prompt. It should be fenced and labelled as
    untrusted in every prompt that carries it, the way this harness already labels recall-pack and
    channel content.
  2. Strip or neutralise instruction-shaped content before it reaches an agent prompt —
    including content in the user_agent and page URL, which no human ever intends as prose.
  3. The trio summarises; the implementer never sees the raw text. The implementer's input
    should be the story the trio wrote, not the reporter's words. This is the single strongest
    control available and it falls out of the architecture we already chose.
  4. Injection attempts are an escalation trigger, not a silent filter. A ticket that tries it
    is a security event worth Mark seeing.
  5. Same treatment for anything else attacker-influenced that reaches a prompt: issue comments,
    PR review comments, commit messages from a fork.

Acceptance

  • A ticket whose description contains instruction-shaped text ("ignore previous instructions",
    a fenced block impersonating a system prompt, a URL with a payload in it) does not reach an
    implementer prompt as instructions, proved by a test with recorded hostile inputs.
  • The implementer's prompt provably contains no verbatim reporter text.
  • An injection attempt escalates and is recorded.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions