Skip to content

Add AVE-2026-00078/79/80: multi-agent pipeline boundary records (arXiv:2608.00718) - #177

Merged
chaksaray merged 3 commits into
developfrom
feat/multi-agent-pipeline-boundary-records
Aug 14, 2026
Merged

chaksaray merged 3 commits into
developfrom
feat/multi-agent-pipeline-boundary-records

Conversation

@chaksaray

Copy link
Copy Markdown
Contributor

Summary

Three genuinely distinct behavioral classes extracted from Bappy et
al., "Adversarial Attacks in Multi-Agent LLM Pipelines: Unveiling
Structural Vulnerabilities in Agentic AI Architectures"
(arXiv:2608.00718, accepted IEEE
GLOBECOM 2026), empirically derived from 147 annotated TRAIL-benchmark
production traces (GAIA + SWE-Bench Lite) plus a controlled cross-model
evaluation (GPT-5-mini, Claude Sonnet 4.5, Kimi K2.5).

The paper actually describes four mechanisms. Its fourth (prompt
injection via retrieved content, the paper's "content boundary" class)
was checked against the live corpus by keyword sweep and field-level
provenance_vector comparison and confirmed already covered by
AVE-2026-00016 and related records — not drafted as new. The other
three cleared the distinctness bar and are drafted here, each with its
id confirmed via its own issue first.

Record Mechanism Boundary Distinct from Issue Severity
AVE-2026-00078 Consensus poisoning — orchestrator accepts one sub-agent's result with no quorum delegation AVE-2026-00020, AVE-2026-00018 #174 MEDIUM, 6.4
AVE-2026-00079 Plan hijacking via false completion signal — forced early termination delegation AVE-2026-00021, AVE-2026-00063 #175 MEDIUM, 6.2
AVE-2026-00080 Silent agent substitution (Sybil) — unverified identity at retry identity AVE-2026-00017, AVE-2026-00030 #176 MEDIUM, 6.8

All three scored MEDIUM — AARF rewards amplification breadth, not raw
impact, and each of these is an architectural, single/dual-vector
mechanism even though the underlying failure mode reads as severe.

Framework mappings

  • owasp_mcp: verified against the primary OWASP MCP Top 10 source
    docs (github.com/OWASP/www-project-mcp-top-10), not inferred from
    corpus usage. MCP06 (Intent Flow Subversion) for 00078/00079 — its
    own "Blind Planning" checklist criterion and Scenario B match
    closely; MCP07 (Insufficient Authentication & Authorization) for
    00080 — its own Impact list and Scenario 3 match closely.
  • mitre_atlas: swept the live 170-technique ATLAS.yaml directly
    for each mechanism. No fit found for any of the three — confirmed
    genuine gap, documented in each record's aivss.notes rather than
    forced.
  • nist_ai_rmf: verified against the actual NIST AI 100-1 primary
    text (Tables 1-4), not the corpus's own precedent alone.
  • owasp_asi: deliberately omitted — could not confirm a stable,
    independently-verifiable primary-source ASI01–10 category list at
    drafting time.

Attribution

researcher on all three credits the paper's actual authors (Faisal
Haque Bappy, Tahrim Hossain, Tarannum Shaila Zaman, Raiful Hasan,
Kamrul Hasan, Tariqul Islam), researcher_url points at the real
arXiv page — not defaulted to an AVE maintainer, per
docs/specs/researcher-process.md.

Checklist

Closes #174, closes #175, closes #176.

Three genuinely distinct behavioral classes extracted from Bappy et al.,
"Adversarial Attacks in Multi-Agent LLM Pipelines: Unveiling Structural
Vulnerabilities in Agentic AI Architectures" (arXiv:2608.00718, accepted
IEEE GLOBECOM 2026), empirically derived from 147 annotated TRAIL-
benchmark production traces (GAIA + SWE-Bench Lite) plus a controlled
cross-model evaluation (GPT-5-mini, Claude Sonnet 4.5, Kimi K2.5).

The paper's fourth mechanism (prompt injection via retrieved content,
its content-boundary class) was confirmed already covered by
AVE-2026-00016 and related records via keyword sweep and field-level
provenance_vector comparison -- not drafted as new.

- AVE-2026-00078: consensus poisoning -- an orchestrator accepts a
  single sub-agent's result as authoritative with no quorum or
  cross-verification across redundant sources. Distinct from
  AVE-2026-00020 and AVE-2026-00018 (issue #174). MEDIUM, AIVSS 6.4.
- AVE-2026-00079: plan hijacking via false completion signal -- a
  self-reported "task already completed" claim forces early
  termination of a declared plan with no plan-to-execution binding
  check. Distinct from AVE-2026-00021 and AVE-2026-00063 (issue #175).
  MEDIUM, AIVSS 6.2.
- AVE-2026-00080: silent agent substitution (Sybil) -- an unverified
  process responding at an agent's routing position during a retry is
  accepted as that agent with no credential check. Distinct from
  AVE-2026-00017 and AVE-2026-00030 (issue #176). MEDIUM, AIVSS 6.8.

owasp_mcp mappings verified against the primary OWASP MCP Top 10 source
docs (MCP06, MCP07). mitre_atlas swept against the live 170-technique
ATLAS.yaml -- no fit found for any of the three, confirmed genuine gap
rather than forced. nist_ai_rmf verified against the NIST AI 100-1
primary text. owasp_asi omitted -- no stable, independently-verifiable
primary-source category list could be confirmed.

researcher field credits the paper's actual authors, not an AVE
maintainer, per docs/specs/researcher-process.md.

Validated: scripts/validate_records.py, scripts/check_fixtures.py,
pytest tests/ (321 passed). Published via scripts/build-records.js.
README badge/stats/record-index and CHANGELOG updated.
Fixes a gap caught in PR #177 review: owasp_asi was silently absent
from AVE-2026-00078/00079/00080 despite each record's aivss.notes
already documenting the research that ruled it out. The key just
never got added. Set to [] on all three with the existing reasoning
kept in aivss.notes -- an absent key reads as "nobody checked", an
empty array reads as "checked, no fit found", and only the second is
honest.

- records/AVE-2026-00078/79/80.json: add "owasp_asi": [], reword the
  corresponding aivss.notes sentence from "omitted" to "left empty"
- docs/specs/researcher-process.md: replace the "optional, omit rather
  than force a fit" framing for owasp_asi/owasp_mcp/mitre_atlas/
  nist_ai_rmf (which incorrectly grouped owasp_mcp as optional even
  though the schema requires it) with a "governance and framework
  mappings" section stating owasp_mcp is schema-required and the other
  three must always have the key present (empty array when no fit),
  plus a Common Mistakes entry documenting this exact recurrence
- .claude/skills/add-ave-record/SKILL.md: same rule added to Step 3,
  with a pointer to the researcher-process.md detail

Making key-presence for owasp_asi/mitre_atlas/nist_ai_rmf an actual
schema constraint (not just a process-doc convention) is tracked
separately as a deliberate v1.2.0 minor version bump -- issue #178 --
not done here since schema/ave-record-1.1.0.schema.json is a frozen
snapshot that's never edited retroactively per
docs/specs/scaling-and-governance.md Section 2.

Revalidated: scripts/validate_records.py (80/80), check_fixtures.py,
pytest tests/ (321 passed). Republished via scripts/build-records.js.
@chaksaray

Copy link
Copy Markdown
Contributor Author

Fixed a gap caught in review: `owasp_asi` was silently missing (not
empty, just absent) from all three records, even though each
record's own `aivss.notes` already documented the research that
ruled it out — the key never got added to reflect that.

  • Added `"owasp_asi": []` to all three, reworded the corresponding
    `aivss.notes` sentence to match.
  • `docs/specs/researcher-process.md` and the `add-ave-record`
    skill now both say explicitly: `owasp_mcp` is schema-required,
    `owasp_asi`/`mitre_atlas`/`nist_ai_rmf` are not yet
    schema-required but the key must always be written — `[]` when
    nothing fits, never absent. Added a Common Mistakes entry citing
    this exact recurrence.
  • Opened Schema v1.2.0: require owasp_asi, mitre_atlas, nist_ai_rmf keys to always exist (empty array allowed) #178 to track making that a real schema constraint
    (required key, no minItems, so empty stays valid) as a
    deliberate v1.2.0 minor bump — not done in this PR since
    ave-record-1.1.0.schema.json is a frozen snapshot.

Re-ran validate_records.py (80/80), check_fixtures.py, and
pytest tests/ (321 passed) after the fix.

…us precedent

Prompted by discovering (issue #179) that owasp_asi's ASI01-ASI10
numbering, used consistently across ~65 records and hard-coded into
the schema's own regex, doesn't exist anywhere in OWASP's actual
Agentic Security Initiative document -- confirmed by fetching the
real PDF and grepping the full text, zero matches. The real taxonomy
uses T1-T17. The fabricated numbering traces to a third-party blog's
own retelling of the initiative, not OWASP's own document, and appears
to have propagated by each record copying the pattern already used by
prior records rather than any record tracing back to the source.

docs/specs/researcher-process.md and the add-ave-record skill both now
state explicitly: "primary source" means the actual document, fetched
and read -- not a search result, not a WebFetch summary, not a
third-party blog, and not corpus precedent no matter how many existing
records agree with each other. Internal agreement across many records
is not the same evidence as one primary-source document actually
opened and read once. Added as a Common Mistakes entry in
researcher-process.md alongside the existing ones (entry_class vs
enforcement_point confusion, aars/aarf mismatches, label-vs-field
duplicate checks).

No schema or record changes in this commit -- issue #179 scopes the
actual remediation (schema regex fix, per-record corpus audit) as
separate, larger work, per explicit direction to file a tracking issue
only for now.
@chaksaray

Copy link
Copy Markdown
Contributor Author

While deep-checking whether `owasp_asi`/`mitre_atlas` really had
no mapping for these three records (per request), fetched the actual
OWASP Agentic Security Initiative primary source PDF directly
(`genai.owasp.org/download/45674`, v1.1, Dec 2025) instead of
relying on search summaries. Found that its real taxonomy uses
`T1`-`T17`, not `ASI01`-`ASI10` — the numbering this corpus
(~65 records) and the schema's own `owasp_asi` regex both use
doesn't exist anywhere in the actual 47-page document (verified via
full-text grep, zero matches for `ASI0`).

That's a corpus-wide problem, not specific to this PR, so scoped it
separately rather than expanding this PR's diff: #179.

Added to this PR: `docs/specs/researcher-process.md` and the
`add-ave-record` skill now both say explicitly — primary source
means the real document, fetched and read, never a summary or
corpus precedent, citing this exact discovery as the cautionary
example. No schema or record value changes here; those are #179's
scope.

Re-validated after this addition: 80/80 schema-valid, 321 tests
passing.

@chaksaray
chaksaray merged commit d9409df into develop Aug 14, 2026
6 checks passed
@chaksaray
chaksaray deleted the feat/multi-agent-pipeline-boundary-records branch August 14, 2026 16:04
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant