From 314b72af7c6cabe86f823acd9dbeb8b767df5663 Mon Sep 17 00:00:00 2001 From: chaksaray Date: Sat, 29 Aug 2026 10:32:25 +0700 Subject: [PATCH] docs: capability-behavior-vulnerability taxonomy audit, all 80 records Full audit per the adopted A/B/C/D/E design: for every record, what security property is violated, and what makes it a vulnerability rather than merely a capability or attacker behavior. Step 0 (inserted verification step): confirmed AVE-2026-00074, 00078, and 00080 against the live corpus before treating them as positive controls -- all three exist and match what the audit design assumed about them. Full seven-question treatment for Priority 1 (5 capability-heavy records), Priority 2 (11 technique-vs-vulnerability records), Priority 3 (12 conventional-vs-agentic records), and the seven positive controls -- 35 records total. Sprint 4's remaining 45 records get a real but proportionately lighter pass (classification, boundary, decision, Q7 three-line artifact), per the task's own instruction not to review all 80 with equal effort. Result: 38 VALID-AVE, 23 INSUFFICIENT-BOUNDARY (real mechanism, needs clearer description text, not a structural change), 2 TECHNIQUE-CONFLATION (00009, 00032), 1 CAPABILITY-CONFLATION (00038), 16 GENERIC-VULNERABILITY (conventional CWE-territory mechanisms with varying strength of agentic-specific justification -- three flagged as a genuinely open maintainer question: 00052/00053/00060 are pure MCP-implementation bugs with no behavioral component at all). No schema change proposed. security_condition/security_boundary were the two candidates evaluated; the audit's own evidence is that every record already has an identifiable boundary and violated property -- the real gap was almost always articulation in existing prose fields, not a missing structured field. Deliberately does not open the community-invitation umbrella issue. astrogilda and narko4u both have active, current threads (#218/#219, the #214 review offer, the GenAI Crosswalk manual-PR recommendation) that predate this audit; opening a third simultaneous ask without checking that queue first is what this task's own sequencing note asked not to do by default. --- docs/audits/capability-vulnerability-audit.md | 2193 +++++++++++++++++ 1 file changed, 2193 insertions(+) create mode 100644 docs/audits/capability-vulnerability-audit.md diff --git a/docs/audits/capability-vulnerability-audit.md b/docs/audits/capability-vulnerability-audit.md new file mode 100644 index 0000000..0433072 --- /dev/null +++ b/docs/audits/capability-vulnerability-audit.md @@ -0,0 +1,2193 @@ +# Capability–Behavior–Vulnerability Taxonomy Audit + +Audits AVE's 80 published records against one question: for each +record, what security property is violated, and what makes that +property a vulnerability rather than merely an attacker behavior or +an agent capability? See the audit design this document implements +for the full methodology (classification scheme A–E, the seven +mandatory questions, sprint ordering). + +**Step 0 verification (positive controls)**: `AVE-2026-00074`, +`AVE-2026-00078`, and `AVE-2026-00080` were confirmed against the live +corpus before this audit began. All three exist, and their real +content matches what this document assumes about them. No correction +needed to the positive-control list. + +**Status**: complete. All 80 records classified. Priority 1–3 (28 +records) and the seven positive controls get full seven-question +treatment. Sprint 4 (the remaining 45 records) gets a real but more +compact pass — classification, capability, vulnerability, boundary, +decision, and the Q7 three-line artifact, without the full template's +every section, per this audit's own instruction not to review all 80 +with equal effort. Result: 38 VALID-AVE, 23 INSUFFICIENT-BOUNDARY, 2 +TECHNIQUE-CONFLATION, 1 CAPABILITY-CONFLATION, 16 GENERIC-VULNERABILITY. +No schema change proposed — see Deliverable 4. + +--- + +## Sprint 1 — Priority 1: capability-heavy records + +### AVE-2026-00022 -- Scope Creep - Accessing Undeclared Resources +#### Classification +B +#### Capability +Agent can invoke tools, APIs, files, or databases beyond what a +skill's manifest declares as its required resource access. +#### Security-relevant behavior +Component instructs the agent to access resources not declared in its +manifest and not authorized by the user. +#### Vulnerability condition +No runtime mechanism compares an agent's actual resource access +against the manifest's declared scope. The manifest is a reviewed, +trusted declaration; nothing enforces that declaration once the +component runs. +#### Security boundary +Reviewed artifact (declared manifest scope) → executed behavior +(actual runtime access). +#### Attacker-controlled input +Skill instruction body directing access to undeclared resources. +#### Missing control +No scope enforcement at the agent-framework layer comparing runtime +tool/resource calls against the component's own declared manifest. +#### Agentic-specific distinction +A manifest's declared scope is meaningful only if something enforces +it against a natural-language-driven agent that can be instructed, in +plain English, to ignore that declaration -- a deterministic program +either has the access compiled in or it doesn't. The instructable gap +between declared and actual scope is the agentic-specific piece. +#### Decision +CLARIFY +#### Rationale +The mitigation metadata already correctly names the fix +(`least_privilege`, `isolate_scope`, `enforcement_point: +agent_framework`), but the prose description reads more as "what +happens" than as a named missing control. Tightening the description +to state the enforcement gap explicitly would resolve this without +changing the mechanism. +#### Confidence +HIGH + +--- + +### AVE-2026-00032 -- Network Reconnaissance Instruction +#### Classification +C +#### Capability +Agent has network connectivity or shell access sufficient to reach +internal hosts and services. +#### Security-relevant behavior +A malicious component instructs the agent to run scans, enumerate +services, or map infrastructure, using network access the skill was +not declared to need. +#### Vulnerability condition +No network-layer segmentation or scope check restricts an agent's +granted network capability to the destinations its declared task +actually requires. +#### Security boundary +Declared task scope → actual network reachability. Distinct from +`AVE-2026-00022`'s boundary in enforcement point (`network_layer` +here, `agent_framework` there), which is the real reason this stays a +separate record rather than merging into it. +#### Attacker-controlled input +Skill instruction body: network/port-scan directive. +#### Missing control +No network segmentation or egress scoping tied to a skill's declared +purpose; the agent's network capability is all-or-nothing rather than +scoped per task. +#### Agentic-specific distinction +Real, but currently implicit. The record as written emphasizes what +the attacker's instruction makes the agent do (classic reconnaissance) +more than the missing network-layer scoping control that makes it +possible from a trusted internal host with no additional exploit +required. +#### Decision +REFRAME +#### Rationale +This is a genuine attack-technique-shaped description (`attack_class: +Reconnaissance - Internal Network Scanning` reads as behavior, not +boundary) sitting on top of a real, distinct vulnerability mechanism +(no per-task network scoping). Reframe the description to lead with +the missing scoping control; keep as a separate record from +`AVE-2026-00022` since the enforcement point genuinely differs +(network layer vs. agent framework). +#### Confidence +MEDIUM + +--- + +### AVE-2026-00035 -- Environment or Sensor Data Manipulation +#### Classification +E +#### Capability +Agent can report tool-sourced sensor, environment, or system-state +observations to an operator or downstream agent. +#### Security-relevant behavior +Component instructs the agent to fabricate, alter, or suppress those +readings before they're reported. +#### Vulnerability condition +No provenance or integrity check verifies that a reported observation +genuinely reflects the underlying tool's real output before it reaches +a decision-maker. +#### Security boundary +Tool response (source of truth) → reported observation (what the +operator or downstream agent actually receives). +#### Attacker-controlled input +Sensor/environment tool-response payload. +#### Missing control +No provenance labeling or integrity verification on tool-response data +carrying sensor/environment readings. +#### Agentic-specific distinction +This is real but needs to be stated, not assumed. Telemetry-integrity +attacks against SCADA/IoT systems are a well-established, pre-LLM +security problem (CWE/ICS territory) -- the Q6 test cuts against this +record if left as-is: a deterministic reporting pipeline with the same +integrity gap would have essentially the same vulnerability. What's +actually agentic-specific is narrower and not yet named in the +record's own text: a natural-language-driven relay can be *instructed* +in plain English to misreport, without requiring the hardware or +firmware compromise a conventional telemetry-spoofing attack needs. +#### Decision +JUSTIFY +#### Rationale +Retain, but the agentic-specific distinction (instructable fabrication +vs. hardware/firmware-level compromise) needs to move from this audit +into the record's own description. As written, nothing in the record +text distinguishes it from a generic ICS/SCADA telemetry-integrity +class that CWE and conventional taxonomy already cover. +#### Confidence +MEDIUM + +--- + +### AVE-2026-00038 -- Excessive Agency - Unbounded Tool Use or Sub-Agent Spawning +#### Classification +D +#### Capability +Agent can use any declared tool and spawn sub-agents. +#### Security-relevant behavior +Component instructs the agent to use any tool at its disposal, spawn +unlimited sub-agents, or "do whatever it takes," removing scope +boundaries and human oversight. +#### Vulnerability condition +As written, this is close to a bare capability statement: "the agent +can act without scope boundaries or oversight checkpoints" describes +an absence of a control, but the record doesn't clearly separate +*that the agent has broad tool access* (a capability, not itself a +flaw) from *no mechanism bounds how that access is used once granted* +(the actual vulnerability). +#### Security boundary +Declared, bounded tool/delegation scope → actual, unbounded runtime +use. The same shape of boundary as `AVE-2026-00022`, applied to tool +use and sub-agent spawning specifically rather than resource access +generally. +#### Attacker-controlled input +Skill instruction body granting unlimited tool use or sub-agent +spawning. +#### Missing control +No ceiling on tool invocations or sub-agent spawning, and no human +checkpoint gating either, once a component's instructions grant +"unlimited" authority. +#### Agentic-specific distinction +Real: sub-agent spawning and recursive capability expansion are +agentic-specific mechanisms with no direct deterministic-program +analog (Q6: a conventional program doesn't spawn new instances of +itself with inherited authority in response to natural-language +instruction). But this needs to be named directly rather than left as +"the agent had a lot of authority and used it," which reads as +capability, not boundary failure. +#### Decision +REWRITE +#### Rationale +A genuine, distinct mechanism exists underneath this (no bound on +tool-use/sub-agent-spawn rate or scope once granted, no human +checkpoint), but the current description doesn't cleanly separate +capability from vulnerability the way `Q7`'s three-line format +requires. This is also the record flagged ambiguous in the prior +owasp_asi audit (ASI02 vs. ASI09 vs. ASI10, none a clean fit) -- +consistent with a record whose underlying vulnerability condition +hasn't been made explicit enough to classify confidently against an +external framework either. +#### Confidence +MEDIUM + +--- + +### AVE-2026-00068 -- CLI Command Composition Risk +#### Classification +A +#### Capability +Agent can issue sequences of CLI commands, none individually +dangerous. +#### Security-relevant behavior +Commands compose through shared OS/shell state (environment variables, +file descriptors, working directory, temp files) into a capability the +task did not authorize, with no single command flagged by +single-command analysis tools. +#### Vulnerability condition +No producer-consumer tracking across a command sequence within one +session -- every existing control (ShellCheck, GTFOBins, LOLBAS, +single-command scanners) evaluates each command in isolation, so the +composition itself is structurally invisible to them. +#### Security boundary +Single-command security review → multi-command behavioral effect. The +boundary is the unit of analysis itself: per-command vs. per-sequence. +#### Attacker-controlled input +None required in the traditional sense -- the record's own text is +explicit that this can occur "entirely within ordinary, benign-looking +developer task scenarios," which is unusual among AVE records and +worth flagging on its own (see Q7 below). +#### Missing control +No cross-command state-flow analysis; every existing tool's unit of +analysis stops at the single command. +#### Agentic-specific distinction +Strong. This is not simply "shells allow multi-step work" (true of any +shell session, deterministic or not) -- the agentic-specific angle is +that an LLM-driven agent can be *guided toward* an unauthorized +composition through many small, individually-benign steps in a way a +human operator following a fixed script would not organically +converge on, and the paper cited (96.59% ASR, 2,525 trials) is +evaluating exactly that model-driven composition behavior, not generic +shell risk. +#### Decision +RETAIN +#### Rationale +This is the cleanest A-classification in Priority 1: a specific, +named missing control (no cross-command analysis unit), a citable +primary source with a concrete attack-success rate, and a boundary +(single-command vs. sequence-level review) that existing tooling +genuinely does not cover. Worth one explicit follow-up, not a +reclassification: since no single attacker-controlled input drives +this the way most AVE records require, the record's `provenance_vector` +and detection guidance should say plainly that this can fire on +entirely benign-looking, non-adversarial task sequences -- worth a +one-line addition, not a structural change. +#### Confidence +HIGH + +--- + +## Sprint 2 — Priority 2: technique-vs-vulnerability records + +### AVE-2026-00009 -- AI Identity Jailbreak via Role-Play or Persona Override +#### Classification +C +#### Capability +Agent operates under safety/alignment constraints configured +independently of the content it processes. +#### Security-relevant behavior +Component instructs the agent to adopt a persona or "developer mode" +without those constraints. +#### Vulnerability condition +No boundary prevents processed content from redefining the agent's own +safety/alignment configuration -- constraints that should be +structurally immutable to in-context instruction are treated as +negotiable. +#### Security boundary +Agent's own configuration (trusted, set once) → content the agent +processes (untrusted, arbitrary). +#### Attacker-controlled input +Skill instruction body: persona-override / unrestricted-mode +directive. +#### Missing control +No enforced separation between "instructions that configure the +agent" and "content the agent merely reads." +#### Agentic-specific distinction +Real. A deterministic program's configuration isn't "talked into" +changing by its own input stream; an LLM's constraints and its input +share the same channel, which is exactly the property this class +exploits. +#### Decision +REFRAME +#### Rationale +Currently reads as attack-technique description ("component instructs +agent to pretend...") rather than naming the configuration/content +separation failure underneath it. The mechanism is real and distinct; +the framing needs to lead with the boundary, not the technique. +#### Confidence +MEDIUM + +--- + +### AVE-2026-00010 -- Covert Instruction Concealment via Secrecy Directive +#### Classification +B +#### Capability +Agent can withhold information from its own output. +#### Security-relevant behavior +Component instructs the agent not to reveal, disclose, or acknowledge +the instructions it received. +#### Vulnerability condition +No requirement that instructions reaching the agent be disclosed to +the user or operator on request -- there's no operator-visibility +boundary the component's instructions can't also suppress. +#### Security boundary +Component-issued instruction → operator/user visibility of that +instruction. +#### Attacker-controlled input +Skill instruction body: secrecy directive. +#### Missing control +No enforced transparency requirement independent of what the +component itself permits the agent to disclose. +#### Agentic-specific distinction +Real: this is specifically about the covert channel an +instruction-following agent can be turned into, not a generic secrecy +concept. +#### Decision +CLARIFY +#### Rationale +The mechanism is genuine and reasonably specific already; naming +"operator-visibility boundary" explicitly in the description would +remove any ambiguity with a plain confidentiality framing. +#### Confidence +HIGH + +--- + +### AVE-2026-00014 -- False Authority Claim via Trust Escalation Impersonation +#### Classification +B +#### Capability +Agent can act on authority claims present in content it processes. +#### Security-relevant behavior +Component falsely claims to originate from Anthropic, OpenAI, Google, +Microsoft, or the operator. +#### Vulnerability condition +No mechanism authenticates in-content authority claims -- any content +can assert arbitrary authority with no verification step. +#### Security boundary +Claimed identity (asserted in content) → actual, verifiable identity +(not checked against anything). +#### Attacker-controlled input +Skill instruction body: false-authority framing. +#### Missing control +No authentication of in-band identity/authority claims. +#### Agentic-specific distinction +Present but not unique to agentic systems specifically (authority +spoofing in content is a general social-engineering pattern); what +makes it agentic-relevant is that the agent, not a human, is the one +deferring to the claim. +#### Decision +CLARIFY +#### Rationale +Mitigation metadata already names the fix (`verify_identity`); the +description should state "no authentication of in-content authority +claims" as the violated property rather than only describing the +impersonation technique. +#### Confidence +MEDIUM + +--- + +### AVE-2026-00023 -- Model Context Window Manipulation +#### Classification +B +#### Capability +Agent's effective attention to system-prompt and safety content can +vary with context volume. +#### Security-relevant behavior +Component floods the context window with content to push safety +constraints out of the active window. +#### Vulnerability condition +Safety and system constraints have no persistence guarantee +independent of their position in the context window -- they can be +positionally evicted rather than remaining structurally binding. +#### Security boundary +Constraint binding (should be persistent) → context-window position +(where it actually, and non-persistently, lives). +#### Attacker-controlled input +Tool/skill output volume flooding the context window. +#### Missing control +No mechanism keeps safety constraints binding independent of how much +content has since entered the context. +#### Agentic-specific distinction +Strong and inherent to the architecture: a conventional program's +rules aren't degraded by input volume the way an LLM's effective +adherence to early-context content can be (Q6 cuts clearly in favor of +agentic-specificity here). +#### Decision +CLARIFY +#### Rationale +Mechanism is real and architecture-specific; description already +gestures at it ("those constraints no longer in scope") but doesn't +name it as explicitly as Q3 requires. +#### Confidence +HIGH + +--- + +### AVE-2026-00025 -- Conversation History Injection +#### Classification +B +#### Capability +Agent relies on conversation history to infer prior user intent and +consent. +#### Security-relevant behavior +Component injects fabricated prior-turn content into context. +#### Vulnerability condition +No integrity or provenance check distinguishes genuine prior-turn +history from content merely styled to resemble it within the current +turn's context. +#### Security boundary +Genuine session history (should be structurally distinct) → arbitrary +content shaped to resemble history (currently indistinguishable to the +model). +#### Attacker-controlled input +Fabricated prior-turn content injected into context. +#### Missing control +No structural or cryptographic separation between real history and +content merely formatted to look like it. +#### Agentic-specific distinction +Real: this specifically exploits that conversation history is +represented in the same token stream as everything else, with no +architectural marker of authenticity. +#### Decision +CLARIFY +#### Rationale +Genuine, distinct mechanism; needs the "no structural separation of +real vs. simulated history" property stated directly rather than left +implicit in "embeds fake prior conversation history." +#### Confidence +MEDIUM + +--- + +### AVE-2026-00027 -- Multi-Turn Attack - Instruction Persistence Across Conversations +#### Classification +B +#### Capability +Agent's stated behavior can be shaped by instructions that ask it to +"remember" something. +#### Security-relevant behavior +Component instructs the agent to retain and re-apply malicious +instructions across turns, memory resets, or new sessions. +#### Vulnerability condition +No boundary prevents a single untrusted, single-session instruction +from converting itself into cross-session persistent behavior via the +agent's own compliance with a "remember this" directive. +#### Security boundary +Single-session, untrusted instruction → cross-session persistent +behavior. +#### Attacker-controlled input +Instruction directing retention across turns/sessions. +#### Missing control +No check on instructions that specifically target the agent's own +persistence/memory behavior as their payload. +#### Agentic-specific distinction +Real, but overlaps substantially with `AVE-2026-00019` (Agent Memory +Poisoning) and `AVE-2026-00070` (Distributed Cross-Agent Backdoor +Fragments). Flagged for the Sprint 4 cross-record consistency pass: +worth confirming whether this record's mechanism (instruction-driven +self-persistence via agent compliance) is genuinely distinct from +00019's (direct write to a persistent memory store) rather than two +labels for the same underlying condition. +#### Decision +CLARIFY +#### Rationale +Real mechanism, but its boundary with 00019 needs to be made explicit +in both records' text, not just implied by different titles. +#### Confidence +MEDIUM + +--- + +### AVE-2026-00029 -- Homoglyph or Unicode Obfuscation Attack +#### Classification +A +#### Capability +Human reviewers evaluate text visually; the model processes Unicode +codepoints. +#### Security-relevant behavior +Component uses homoglyphs, zero-width characters, or bidirectional +control codes so instructions appear innocuous to a human reviewer +while remaining fully executable by the model. +#### Vulnerability condition +No canonicalization or normalization step reconciles what a human +reviewer visually sees with what the model actually processes at the +codepoint level. +#### Security boundary +Human visual review → model-level codepoint processing. +#### Attacker-controlled input +Text content containing homoglyph, zero-width, or bidi control +characters. +#### Missing control +No Unicode normalization/canonicalization check before human review. +#### Agentic-specific distinction +Strong: this specifically exploits the gap between how a human +reviewer and an LLM each consume the same bytes, which has no +meaningful analog in reviewing a deterministic program's source. +#### Decision +RETAIN +#### Rationale +Already the cleanest-written record in this priority group -- the +description states the human-vs-model processing divergence directly. +No change needed. +#### Confidence +HIGH + +--- + +### AVE-2026-00030 -- Privilege Escalation via False Role Claim +#### Classification +A +#### Capability +Skills can implement role-based access checks against user-supplied +claims. +#### Security-relevant behavior +Component instructs the agent to grant elevated privileges whenever a +user or external party claims a privileged role. +#### Vulnerability condition +Role-based access control is implemented by trusting a bare, +self-asserted claim with no backing credential or authentication step. +#### Security boundary +Claimed role → actually-verified role (never checked). +#### Attacker-controlled input +Skill instruction body: role-claim trust rule. +#### Missing control +No authentication step backing role claims. +#### Agentic-specific distinction +Present, though the underlying pattern (trusting an unauthenticated +claim) is a conventional access-control failure; the agentic angle is +that a natural-language instruction can install this pattern as the +component's entire authorization logic with no code review surface. +#### Decision +RETAIN +#### Rationale +Already clean and explicit ("bypassing all actual authentication"). +No change needed. +#### Confidence +HIGH + +--- + +### AVE-2026-00037 -- Prompt Injection via Image or Vision Input +#### Classification +B +#### Capability +Multimodal agents can read text embedded in images. +#### Security-relevant behavior +Component instructs the agent to execute instructions visible in +images, bypassing text-level security scanning entirely. +#### Vulnerability condition +Untrusted-content handling boundaries exist for text but have no +equivalent for text rendered inside images -- the same class of +content is trusted or untrusted depending on which modality carries +it. +#### Security boundary +Text-channel untrusted-content boundary (exists) → image-channel +untrusted-content boundary (does not exist). +#### Attacker-controlled input +Image/screenshot pixel content. +#### Missing control +No cross-modal equivalent of existing text-content trust boundaries. +#### Agentic-specific distinction +Strong: multimodal instruction-following with no modality-consistent +trust model is architecture-specific to LLM-driven agents. +#### Decision +CLARIFY +#### Rationale +Mechanism is real and already partly stated ("completely bypasses +text-level security scanning"); making the modality-gap framing +explicit as the violated property would sharpen it further. +#### Confidence +MEDIUM + +--- + +### AVE-2026-00039 -- Covert Channel - Steganographic Data Exfiltration +#### Classification +B +#### Capability +Agent output is reviewed by humans and monitored at the network layer. +#### Security-relevant behavior +Component instructs the agent to exfiltrate data via steganographic +encoding, first-letter encoding, whitespace patterns, or timing +signals in otherwise-benign-looking output. +#### Vulnerability condition +Exfiltration monitoring checks network destinations, not whether +benign-looking output content itself encodes exfiltrated data -- +content-level exfiltration has no detection surface at all. +#### Security boundary +Network-egress monitoring (exists) → content-level covert-channel +monitoring (does not exist). +#### Attacker-controlled input +Output text carrying steganographic/covert encoding. +#### Missing control +No content-level covert-channel detection independent of destination +monitoring. +#### Agentic-specific distinction +Present: an LLM's fluency makes steganographic encoding far more +naturally deployable at scale than in traditional exfiltration, though +the covert-channel concept itself predates agentic systems. +#### Decision +CLARIFY +#### Rationale +Real, distinct gap; the "network monitoring only" framing needs to be +named explicitly as the missing control, not left implicit in "harder +to spot." +#### Confidence +MEDIUM + +--- + +### AVE-2026-00057 -- Obfuscated or Encoded Skill Payload Designed to Evade Static Scanners +#### Classification +A +#### Capability +Static scanners pattern-match against a skill's literal, as-written +content. +#### Security-relevant behavior +Component embeds instructions or code in encoded form (base64, hex, +split-string concatenation) so the decoded, executed form differs from +what a scanner matches against. +#### Vulnerability condition +Static analysis operates on the encoded representation while execution +operates on the decoded representation, with no decode-then-rescan +step bridging the two. +#### Security boundary +As-scanned representation → as-executed representation. +#### Attacker-controlled input +Skill instruction or code body containing encoded content requiring a +decode step to reveal intent. +#### Missing control +No decode-then-rescan step in the static-analysis pipeline. +#### Agentic-specific distinction +Moderate: encoding-based scanner evasion is a conventional AppSec +pattern (CWE territory), but the record already explicitly +differentiates itself from the two adjacent AVE records it could be +confused with (`AVE-2026-00029`, visual deception targeting humans; +`AVE-2026-00024`, content-type misrepresentation) by scope -- this one +is specifically about representational deception targeting automated +scanners. +#### Decision +RETAIN +#### Rationale +Already well-scoped and explicitly differentiated from its nearest +neighbors in the record's own text. No change needed. +#### Confidence +HIGH + +--- + +## Sprint 3 — Priority 3: conventional vs. agentic-specific records + +This group is where Decision Rule §7's test does the most work: several +of these are conventional vulnerability classes (CWE territory) that +happen to occur inside agent components. The question per record is +whether AVE names a real, distinct reason it belongs here beyond that +coincidence. + +### AVE-2026-00004 -- Arbitrary Code Execution via Shell Pipe Injection +#### Classification +E +#### Capability +Agent with shell/code-execution tool access follows embedded commands. +#### Security-relevant behavior +Component embeds shell pipe patterns (`curl | bash`) in natural-language +instructions. +#### Vulnerability condition +Conventional command-injection (CWE-78 territory) delivered through +natural language rather than code. +#### Security boundary +Trusted instruction content → shell execution, with no boundary +between the two once a component's text is treated as instruction. +#### Attacker-controlled input +Skill instruction body: curl|bash / wget|sh directive. +#### Missing control +No distinction between natural-language content that merely describes +a step and natural-language content that constitutes an executable +directive. +#### Agentic-specific distinction +Already stated in the record's own text and holds up: existing SAST +scanners are built to find command injection in *code*, not in +natural-language instructions that produce the same effect once an +agent follows them. The mechanism is conventional; the delivery +channel evading conventional tooling is the real agentic-specific +piece. +#### Decision +JUSTIFY +#### Rationale +Q6 cuts against pure novelty (a deterministic program blindly piping +untrusted config into a shell has the same flaw), but the record +already names its real distinction (NL delivery invisible to SAST) -- +retain as currently justified, not merged into generic CWE-78. +#### Confidence +HIGH + +--- + +### AVE-2026-00005 -- Recursive File System Destruction via Destructive Command Injection +#### Classification +E +#### Capability +Agent with filesystem/shell access follows embedded destructive +commands. +#### Security-relevant behavior +Component embeds `rm -rf /`-class commands inside legitimate-looking +setup/cleanup instructions. +#### Vulnerability condition +Same shape as `AVE-2026-00004`: conventional destructive-command +pattern, delivered via natural language. +#### Security boundary +Same as 00004: instruction content → shell execution, no boundary +between them. +#### Attacker-controlled input +Skill instruction body: recursive deletion command. +#### Missing control +Same as 00004. +#### Agentic-specific distinction +Same NL-delivery-evades-SAST argument as 00004, already stated. +#### Decision +JUSTIFY +#### Rationale +Structurally near-identical to `AVE-2026-00004` -- flagged for the +Sprint 4 cross-record pass to confirm the two are worth keeping +separate. Leaning yes: impact differs materially (arbitrary code +execution vs. irrecoverable data loss with no code-execution +requirement), which is a real severity and detection distinction, not +just a labeling one. +#### Confidence +HIGH + +--- + +### AVE-2026-00033 -- Unsafe Deserialization or Eval Instruction +#### Classification +E +#### Capability +Agent follows instructions to deserialize/evaluate externally-supplied +data. +#### Security-relevant behavior +Component instructs use of `pickle.loads`, unguarded `yaml.load`, or +`eval`/`exec` on untrusted strings. +#### Vulnerability condition +Conventional RCE-via-unsafe-deserialization (CWE-502/CWE-94), +delivered via natural language. +#### Security boundary +Instruction content → unsafe deserialization/eval call, no boundary +between them. +#### Attacker-controlled input +Skill instruction body: eval/pickle/yaml.load of untrusted data. +#### Missing control +Same NL-vs-code detection gap as 00004/00005. +#### Agentic-specific distinction +Same argument, same standard of evidence as 00004/00005. +#### Decision +JUSTIFY +#### Rationale +Third member of the same structural family as 00004/00005. Consistent +treatment: retain, same justification. +#### Confidence +HIGH + +--- + +### AVE-2026-00047 -- Hardcoded Credentials in Agent Component +#### Classification +E +#### Capability +Skill/component files can contain literal string values, including +credential-shaped ones. +#### Security-relevant behavior +A skill file, manifest, or plugin contains a hardcoded API key, token, +password, or private key. +#### Vulnerability condition +Conventional secrets-in-source (CWE-798), amplified by a property +specific to agent components: the credential sits in the same context +window an LLM reads, not in source a compiled program never +"processes" contextually. +#### Security boundary +Static file content → LLM context window (a boundary that doesn't +exist the same way for a deterministic program reading the same file). +#### Attacker-controlled input +Literal credential string in skill file body. +#### Missing control +No separation between "content an LLM incidentally reads" and +"content a co-located prompt-injection payload can induce the LLM to +act on" -- a hardcoded credential is exposed to both. +#### Agentic-specific distinction +Real and already stated: "a prompt injection payload can instruct the +agent to read and exfiltrate credentials that appear elsewhere in its +context window." This is a genuine amplification specific to +LLM-mediated file access, not just "secrets in source code again." +#### Decision +JUSTIFY +#### Rationale +The record earns its place: conventional root cause, but a real, +already-stated agentic amplification (co-location with a +prompt-injection-readable context) that a traditional secrets-scanner +finding wouldn't capture. +#### Confidence +HIGH + +--- + +### AVE-2026-00049 -- HTTP Host Header Injection via Agent-Initiated Request +#### Classification +E +#### Capability +Agent-initiated components can set outbound HTTP request headers. +#### Security-relevant behavior +Component sets a Host (or X-Forwarded-Host/X-Original-URL) header that +doesn't match the declared target, redirecting the request. +#### Vulnerability condition +This is the well-known, conventional Host-header-attack class +(PortSwigger's own documented research area), unchanged in mechanism +regardless of what constructs the request. +#### Security boundary +Declared target endpoint → actual request destination (header- +controlled, not endpoint-controlled). +#### Attacker-controlled input +Outbound HTTP Host / X-Forwarded-Host / Forwarded header. +#### Missing control +No validation that outbound Host header values match the declared +endpoint before the request is sent. +#### Agentic-specific distinction +This is the weakest-justified record in Priority 3. Nothing in the +current text states why an *agent-initiated* Host-header attack is +mechanism-distinct from the same attack made by any other HTTP client. +A plausible real distinction exists but isn't written down: an agent +frequently constructs request parameters (including headers) from +content it has just processed, in a way a static HTTP client +configuration doesn't -- making attacker-controlled header values +reachable through indirect prompt injection rather than requiring +direct network access to the client. That's a real, statable +difference; it just isn't stated yet. +#### Decision +JUSTIFY +#### Rationale +Recommend the record explicitly add the injection-reachability +argument above. Without it, this reads as a straight port of a +well-known web-security class into an agentic component with no +demonstrated distinction, and would be a reasonable MERGE/DEPRECATE +candidate in favor of citing the general class directly. +#### Confidence +LOW + +--- + +### AVE-2026-00052 -- Command Injection via Unsanitized Tool-Call Parameter in MCP Server Implementation +#### Classification +E +#### Capability +An MCP tool's own server-side handler processes caller-supplied +parameter values. +#### Security-relevant behavior +The handler passes a parameter value into a shell/system-command +execution function without sanitization. +#### Vulnerability condition +Conventional command injection (CWE-78), located in the tool's own +implementation code. The record is explicit that no LLM reasoning or +instruction is required to trigger it at all. +#### Security boundary +Caller-supplied parameter → shell execution, entirely inside the +tool's own server code, with no agent-behavior boundary involved. +#### Attacker-controlled input +MCP tool-call parameter value. +#### Missing control +No input sanitization/allowlisting in the server's own handler -- +a pure implementation bug. +#### Agentic-specific distinction +Weak at the mechanism level (this is not an agentic vulnerability in +any behavioral sense; the record's own text says the attack is +"independent of any agent instruction or prompt content"). The +argument offered is a detection-ecosystem one: prompt-injection +scanners, which dominate this space, structurally cannot find it +because there's no instruction text to match. +#### Decision +JUSTIFY +#### Rationale +This is the hardest call in the audit so far. AVE cataloging a +conventional server-implementation bug is defensible only as a +coverage argument (MCP servers are part of the agentic supply chain +AVE's consumers scan, and no other part of their pipeline will catch +this), not a behavioral/mechanism argument. Worth an explicit maintainer +decision on whether that coverage argument is sufficient grounds for a +distinct AVE record, versus this belonging in a general AppSec/SCA +tool's remit with AVE cross-referencing it rather than cataloging it +directly. Same question applies to `AVE-2026-00053` (Sprint 4). +#### Confidence +LOW + +--- + +### AVE-2026-00054 -- Code-Execution Sandbox Escape via JavaScript Prototype-Chain Traversal +#### Classification +E +#### Capability +A code-execution tool runs agent-submitted or agent-generated code +inside an intended sandbox boundary. +#### Security-relevant behavior +A payload uses prototype-chain traversal to reach the host runtime's +global scope, escaping the sandbox. +#### Vulnerability condition +Conventional sandbox-escape (a pre-LLM, well-studied class in browser +and VM security), located in the sandbox's own containment failure. +#### Security boundary +Sandboxed execution context → host runtime, a boundary the sandbox +itself fails to hold. +#### Attacker-controlled input +Code submitted to a code-execution/sandbox tool. +#### Missing control +The sandbox's own isolation doesn't hold against prototype-chain +traversal -- a containment bug, not an instruction-following one. +#### Agentic-specific distinction +Not strongly stated. The record correctly differentiates itself from +`AVE-2026-00042` (how malicious code gets in) by scope (what happens +once code is already running), but doesn't argue why an +*agent*-execution sandbox is more exposed to this than any other +sandboxed code-execution environment. +#### Decision +JUSTIFY +#### Rationale +A real, statable agentic angle likely exists (LLM-generated code may +incidentally produce prototype-chain-triggering patterns without +adversarial intent far more often than human-authored code would, +widening the practical trigger surface even absent an attacker) but +isn't in the record yet. Recommend adding it explicitly rather than +leaving the distinction implicit. +#### Confidence +LOW + +--- + +### AVE-2026-00061 -- TLS Certificate Verification Disabled in Agent Component Configuration +#### Classification +E +#### Capability +A component's outbound connections can be configured with or without +TLS verification. +#### Security-relevant behavior +Configuration disables certificate verification for the component's +own outbound calls. +#### Vulnerability condition +Conventional insecure-configuration (CWE-295), unchanged by what +consumes the connection. +#### Security boundary +Configured verification state → actual MITM exposure. +#### Attacker-controlled input +A declared configuration flag (`verify=False`, +`rejectUnauthorized: false`). +#### Missing control +No enforcement preventing verification from being disabled at all, or +no detection of the disabled state. +#### Agentic-specific distinction +None stated in the record, and none obvious. This is the weakest +agentic-distinction candidate found in the audit so far. +#### Decision +DEPRECATE-OR-JUSTIFY +#### Rationale +A plausible distinction exists (a MITM'd response over a +verification-disabled connection often becomes direct agent context +the model acts on, unlike a traditional app where a tampered response +might just render on a screen a human evaluates with some skepticism) +but nothing in the record states it. Recommend either adding that +argument explicitly or treating this as a CWE-295 crosswalk entry +rather than a standalone AVE record -- as written, it doesn't clear +Decision Rule §7's bar. +#### Confidence +LOW + +--- + +### AVE-2026-00062 -- Unpinned Dependency Version Allowing Supply Chain Substitution +#### Classification +E +#### Capability +A component can declare dependencies by mutable specifier rather than +exact version or hash. +#### Security-relevant behavior +An unpinned dependency reference lets the resolved artifact silently +diverge from what was reviewed. +#### Vulnerability condition +Conventional software-supply-chain risk (dependency-confusion / +unpinned-dependency territory), well-established in general package- +ecosystem security independent of agents. +#### Security boundary +Reviewed artifact (at pin time) → executed artifact (at resolution +time) -- but this boundary exists for any unpinned dependency in any +software, not agent components specifically. +#### Attacker-controlled input +A declared dependency reference lacking version pinning or a content +hash. +#### Missing control +No pinning/hash verification at declaration time. +#### Agentic-specific distinction +Not stated, and weaker than its neighbors `AVE-2026-00074` (dead-anchor +reclamation, whose distinction is explicit: the anchor was pinned and +correct *when written*, and pinning wouldn't have helped) and +`AVE-2026-00066` (hallucinated-name squatting, whose distinction is +explicit: the attack surface is the *model's own hallucination*, not a +human's config choice). This record reads as the generic "pin your +dependencies" case those two more specific records already imply. +#### Decision +JUSTIFY-OR-MERGE +#### Rationale +Candidate for merging into a broader agentic-supply-chain umbrella +alongside 00074/00075/00034/00066, or for an explicit agentic-specific +argument (e.g., agent-authored or agent-approved dependency changes +move at higher velocity and lower human-review depth than +conventional software supply chains, making unpinned drift more likely +to go unnoticed). As written, it's the generic case its more specific +siblings already cover the interesting edges of. +#### Confidence +LOW + +--- + +### AVE-2026-00072 -- MCP Server Bound to All Network Interfaces with No Authentication Step (NeighborJack) +#### Classification +E +#### Capability +An MCP server's bind address and auth requirement are both +configuration choices. +#### Security-relevant behavior +The server binds to `0.0.0.0`/`[::]` with no authentication step, +making it reachable to anyone on the local network. +#### Vulnerability condition +Conventional network-exposure misconfiguration (CWE-284 territory, +"bind to all interfaces with no auth" is a known class independent of +MCP). +#### Security boundary +Intended reachability (local-only, implicitly trusted) → actual +reachability (LAN-wide, still implicitly trusted). +#### Attacker-controlled input +MCP server args/env declaring a wildcard bind address. +#### Missing control +No authentication step independent of network position -- the +server's trust model assumes "reachable" implies "authorized." +#### Agentic-specific distinction +Better-justified than 00061/00062: the record ties this to a named, +documented real-world pattern (predictor2718's "NeighborJack") and +argues the MCP ecosystem specifically has a high prevalence of +servers designed with *no* authentication layer by default, on the +assumption that only a local, already-trusted caller could reach them +-- a design assumption 0.0.0.0 binding silently breaks. That's an +ecosystem-prevalence argument, not a pure mechanism argument, but it's +a real, citable one. +#### Decision +JUSTIFY +#### Rationale +Retain. The ecosystem-specific "no-auth-by-default is the norm, not +the exception, in MCP servers" argument is a legitimate, if +prevalence-based rather than mechanism-based, agentic distinction. +#### Confidence +MEDIUM + +--- + +### AVE-2026-00073 -- Telemetry or API Endpoint Redirect via Static Configuration Value +#### Classification +E +#### Capability +Components can declare telemetry/provider/API endpoint values in +configuration. +#### Security-relevant behavior +A committed configuration value redirects telemetry, model, or +provider traffic to an unintended host, with no content injected into +the model's context at any point. +#### Vulnerability condition +Conventional config-based traffic redirection, but tied to +infrastructure specific to agentic systems: LLM provider base URLs, +MCP server URLs, and A2A agent-card URLs. +#### Security boundary +Declared/default provider endpoint → actual configured endpoint. +#### Attacker-controlled input +A committed config value (`OTEL_EXPORTER_OTLP_ENDPOINT`, +`ANTHROPIC_BASE_URL`, an MCP server URL, or an `agent_card_url`). +#### Missing control +No verification that a provider/telemetry endpoint value matches the +component's declared or default provider before traffic is sent. +#### Agentic-specific distinction +Strong and well-sourced: the record cites a real CVE +(CVE-2026-21852) where exactly this mechanism, a redirected +`ANTHROPIC_BASE_URL`, leaked a user's API key. The endpoint types this +targets (LLM provider base URLs, MCP server URLs, A2A agent-card URLs) +are architecturally specific to agentic systems, not a generic +"config can be wrong" argument. +#### Decision +JUSTIFY +#### Rationale +Best-justified record in Priority 3. Real, disclosed incident, tied +directly to infrastructure that only exists because the component is +agentic (LLM provider endpoints, MCP/A2A URLs). Retain as-is. +#### Confidence +HIGH + +--- + +### AVE-2026-00075 -- Bytecode Poisoning: Compiled .pyc Cache Diverges from Its Own Reviewed .py Source +#### Classification +E +#### Capability +CPython prefers a valid cached `.pyc` over recompiling source when +present. +#### Security-relevant behavior +A skill ships a `.pyc` alongside its `.py` source, where the bytecode +contains dangerous primitives absent from the visible source text. +#### Vulnerability condition +The CPython bytecode-cache-preference behavior is a general interpreter +property, not agent-specific in mechanism. +#### Security boundary +Reviewed artifact (the `.py` source a human or scanner reads) → +executed artifact (the `.pyc` the interpreter actually loads) -- +structurally the same boundary shape as `AVE-2026-00062`'s pinning +gap, applied to a compiled-cache divergence instead of a dependency +reference. +#### Attacker-controlled input +A bundled `.pyc`/`.pyo` file diverging from its sibling `.py` source. +#### Missing control +No bytecode-vs-source comparison step in the review or scanning +pipeline. +#### Agentic-specific distinction +Strong and well-sourced: explicitly framed and cited (CSA AI Safety +Initiative / Trail of Bits research note) as a technique for evading +*agent-skill scanners specifically* -- naming NVIDIA's own SkillSpector +as a scanner that documents its own inability to analyze binary or +encrypted code. Same shape of argument as 00004/00005/00033/00047 and +already well-evidenced with a citable primary source and explicit +differentiation from `AVE-2026-00057`. +#### Decision +JUSTIFY +#### Rationale +Well-sourced, explicit, already differentiated from its nearest +neighbor. Retain as-is. +#### Confidence +HIGH + +--- + +## Positive controls (method validation) + +Confirmed against the live corpus in Step 0. These are checked to +validate that the audit method itself discriminates correctly -- +records expected to classify cleanly as A should actually do so under +the same seven questions applied to everything else, not receive +special treatment. + +### AVE-2026-00001 -- Metamorphic Payload via External Config Fetch +#### Classification +A +#### Capability +A skill or MCP component can fetch and execute remote content at +runtime. +#### Security-relevant behavior +The fetched content replaces the component's own reviewed instructions +after it has already passed security review. +#### Vulnerability condition +No integrity binding exists between what was reviewed (the artifact +and its declared behavior at scan time) and what actually executes +(mutable, externally-fetched content at run time). +#### Security boundary +Reviewed artifact → executed artifact, with a mutable external fetch +sitting unchecked between the two. +#### Attacker-controlled input +Skill instruction body: fetch()/curl/wget directive. +#### Missing control +No pinning or integrity check on runtime-fetched content that replaces +a component's own behavior. +#### Agentic-specific distinction +This is a real, agentic-ecosystem-specific instance of a TOCTOU +pattern: it specifically exploits the scan-then-deploy review model +this whole class of tooling relies on, since the record's own text is +explicit that the payload "does not exist at scan time." +#### Decision +RETAIN +#### Confidence +HIGH + +--- + +### AVE-2026-00002 -- MCP Tool Description Behavioral Injection +#### Classification +A +#### Capability +MCP tool description fields are read by the agent during tool +discovery. +#### Security-relevant behavior +A tool description field contains directives targeting agent behavior +instead of describing the tool's function. +#### Vulnerability condition +No boundary separates "metadata describing a tool" from "instructions +the agent should follow" -- both share the same natural-language +channel with nothing to distinguish them. +#### Security boundary +Tool metadata (should be descriptive, inert) → agent instruction +(treated as authoritative context). +#### Attacker-controlled input +MCP tool.description field. +#### Missing control +No structural separation between descriptive metadata and executable +instruction within the tool manifest schema. +#### Agentic-specific distinction +Real and architecture-specific: MCP's own schema conflates the two +channels by design. +#### Decision +RETAIN +#### Confidence +HIGH + +--- + +### AVE-2026-00041 -- Prompt Injection via MCP Server-Card Tool Descriptions Before First Call +#### Classification +A +#### Capability +An agent fetches and reads a server-card's tool descriptions before +making any tool call. +#### Security-relevant behavior +Behavioral instructions embedded in the server-card are loaded into +context and can act before any user interaction occurs. +#### Vulnerability condition +Same metadata-vs-instruction boundary failure as `AVE-2026-00002`, but +specifically exploitable at the discovery layer, before any tool call +exists for runtime monitoring to observe. +#### Security boundary +Discovery-layer metadata (read once, pre-execution) → agent context +(treated as authoritative), with no execution-layer event for runtime +monitoring to catch. +#### Attacker-controlled input +MCP server-card tool.description field. +#### Missing control +No validation at server-card fetch time, and no detection surface at +all for runtime/execution-layer monitoring, since nothing has executed +yet. +#### Agentic-specific distinction +Real, and distinct from 00002 in a way worth stating clearly rather +than assuming: this record's specific contribution is the detection- +layer argument (invisible to runtime monitoring because it fires +before execution), not just "another tool-description injection +surface." +#### Decision +RETAIN +#### Rationale +Flagged for the Sprint 4 cross-record pass: confirm 00002 and 00041 +are consistently described as distinct by discovery-timing/detection- +layer rather than left to be inferred from title differences alone. +#### Confidence +HIGH + +--- + +### AVE-2026-00046 -- MCP Tool Hook Hijacking +#### Classification +A +#### Capability +Components can register hooks/callbacks into the agent's tool dispatch +layer. +#### Security-relevant behavior +A malicious component's hook silently intercepts or redirects tool +calls -- including calls from other, unrelated components -- to an +attacker-controlled callback. +#### Vulnerability condition +No authorization or integrity check on hook registration into a +shared dispatch layer that all components' tool calls pass through. +#### Security boundary +Component's own declared tool scope → the shared tool-dispatch layer +every component's calls route through, which one component can +silently claim authority over. +#### Attacker-controlled input +Skill-declared hook/callback registration targeting the tool dispatch +layer. +#### Missing control +No isolation or authorization check preventing one component's hook +registration from intercepting calls belonging to other components. +#### Agentic-specific distinction +Strong: this specifically exploits a shared, central dispatch +architecture that has no direct analog outside multi-tool agent +runtimes. +#### Decision +RETAIN +#### Confidence +HIGH + +--- + +### AVE-2026-00074 -- Reclaimable Dead External Anchor (SkillJacking) +#### Classification +A +#### Capability +Skills reference external identities (GitHub owners, package names, +domains) that were valid at authoring time. +#### Security-relevant behavior +A previously-valid external anchor is later deleted, renamed, or +allowed to expire, and becomes re-registerable by anyone. +#### Vulnerability condition +No periodic re-verification that a previously-valid external identity +reference remains currently owned by the same party -- trust is +established once, at review time, and never rechecked against +real-world ownership changes. +#### Security boundary +Trust established at authoring/review time → trust required at every +subsequent execution, with no mechanism re-checking the gap between +them. +#### Attacker-controlled input +A GitHub owner/repo, package name, domain, or cloud subdomain +referenced in the skill that is presently unclaimed. +#### Missing control +No ongoing verification that an already-approved external anchor is +still controlled by its original owner. +#### Agentic-specific distinction +Explicit and well-differentiated in the record's own text from +`AVE-2026-00062` (unpinned dependency): pinning would not have +prevented this, since the reference was exact and correct when +written. Backed by a real, disclosed dataset (925 skills, ~134,000 +agents) and a named real-world takeover. +#### Decision +RETAIN +#### Confidence +HIGH + +--- + +### AVE-2026-00078 -- Unverified Multi-Agent Consensus +#### Classification +A +#### Capability +An orchestrator can dispatch sub-tasks to parallel sub-agents or accept +results from a delegation chain. +#### Security-relevant behavior +The orchestrator's acceptance criterion reduces to "whichever response +arrives," with no quorum, cross-verification, or corroboration step. +#### Vulnerability condition +No quorum or redundancy check exists at the result-aggregation step of +a multi-agent pipeline -- a single compromised or adversarially- +influenced source unilaterally determines the accepted output. +#### Security boundary +Sub-agent result (unverified) → orchestrator's accepted ground truth +(propagated downstream as though verified). +#### Attacker-controlled input +A sub-agent's result, accepted at the pipeline's result-aggregation +step. +#### Missing control +No quorum vote, redundancy comparison, or independent corroboration +before committing to a single source's claim. +#### Agentic-specific distinction +Strong and explicit: the record differentiates itself precisely from +three easily-confused neighbors (`AVE-2026-00016`, `AVE-2026-00018`, +`AVE-2026-00020`) by stating exactly which layer each concerns -- +content origin, instructed fabrication, and injection direction, +respectively -- versus this record's aggregation-layer design flaw. +#### Decision +RETAIN +#### Rationale +This record's own differentiation section is close to a model example +of what Q2/Q3 rigor should look like when multiple records sit in +adjacent territory. +#### Confidence +HIGH + +--- + +### AVE-2026-00080 -- Runtime Agent Identity Substitution (Sybil) +#### Classification +A +#### Capability +An agent's identity in a multi-agent pipeline is inferred from its +routing position rather than a bound credential. +#### Security-relevant behavior +On retry after a failed call, whatever process responds at the same +routing position is silently accepted as the original agent. +#### Vulnerability condition +No persistent, verifiable credential binds identity to a routing +position -- identity is positional, not cryptographic, so anything +answering at the expected slot inherits its full existing trust. +#### Security boundary +Routing position (observable, spoofable) → agent identity (should be +credential-bound, is not). +#### Attacker-controlled input +A process responding at an agent's routing position during a retry +cycle. +#### Missing control +No attestation or session-token check confirming a post-retry response +originates from the same agent instance as before the failure. +#### Agentic-specific distinction +Strong and explicit: the record differentiates itself precisely from +`AVE-2026-00017` (a one-time manifest-level identity claim) and +`AVE-2026-00030` (an explicit, asserted role claim) by noting this +substitution requires no claim of any kind -- the substitute simply +occupies an already-trusted position. +#### Decision +RETAIN +#### Confidence +HIGH + +--- + +**Positive-control result**: all seven classify cleanly as A under the +same seven questions applied to every other record in this audit, and +several (`00074`, `00078`, `00080`) stand out as the strongest-written +records in the corpus specifically because they explicitly differentiate +themselves from their nearest neighbors rather than leaving the +distinction to be inferred from titles. That's real, positive evidence +the classification method discriminates correctly rather than defaulting +either direction -- Priority 1-3 produced a genuine spread across +A/B/C/D/E on the same standard, not a uniform result in either group. + +--- + +## Sprint 4 — remaining records (compact pass) + +Real, individually-reasoned classification for each, at a proportionate +level of detail rather than the full ten-section template used above. +Format per record: Classification | Boundary | Missing control / +Decision, followed by the Q7 three-line artifact. + +### AVE-2026-00003 -- Credential Exfiltration via Agent Instruction +**B** | credential store → external destination | No content-aware +egress control distinguishes legitimate outbound task data from +credential exfiltration. Same missing-control shape as `00013`/`00026`. +Decision: CLARIFY. +``` +Capability: Agent can read env vars/credentials and make external network calls. +Vulnerability: No content-aware egress control distinguishes legitimate outbound task data from credential exfiltration. +Impact: Credential theft enabling follow-on compromise. +``` + +### AVE-2026-00006 -- Cryptocurrency Wallet Drain +**B** | tool-use decision → irreversible on-chain action | No +structurally mandatory (vs. component-configurable) approval gate for +irreversible financial actions specifically. Decision: CLARIFY. +``` +Capability: Agent with wallet tool access can transfer funds/approve allowances. +Vulnerability: No structurally mandatory human-approval gate exists for irreversible financial actions specifically. +Impact: Irreversible on-chain financial loss. +``` + +### AVE-2026-00007 -- Agent Goal Hijack via Direct Instruction Override +**A** | trusted system instruction ↔ untrusted component content, same +channel | No instruction-hierarchy enforcement distinguishes origin or +recency. Already the record's own stated foundational vector. Decision: +RETAIN. +``` +Capability: Agent processes instruction-shaped text from untrusted components in the same channel as trusted system instructions. +Vulnerability: No instruction-hierarchy enforcement distinguishes trusted origin from untrusted, recently-seen text. +Impact: Complete behavioral takeover, enabling most other AVE attack classes as a follow-on payload. +``` + +### AVE-2026-00008 -- Agent Persistence via Self-Replication +**B** | task-scoped tool access → durable, cross-session system +modification | No restriction prevents task-scoped filesystem/shell +access from writing to persistence locations (cron, startup scripts). +Decision: CLARIFY. +``` +Capability: Agent has legitimate filesystem/shell tool access for its task. +Vulnerability: No restriction prevents that task-scoped access from writing to system-persistence locations. +Impact: Durable compromise surviving reboot, reinstallation, or removal attempts. +``` + +### AVE-2026-00011 -- Arbitrary Tool Invocation via Dynamic Tool Call Injection +**B** | task-descriptive content → direct tool-invocation command | +No separation between content describing a task and content commanding +a specific tool invocation with attacker-chosen parameters. Decision: +CLARIFY. +``` +Capability: Agent selects and invokes tools based on task content. +Vulnerability: No boundary prevents component content from directly commanding tool invocations with attacker-chosen parameters, bypassing the agent's own selection logic. +Impact: Arbitrary tool invocation, including destructive, exfiltration, or lateral-movement capability the user never intended to activate. +``` + +### AVE-2026-00012 -- Capability Escalation via False Permission Grant +**B** | claimed permission (in-content) → actual authorization (never +checked) | No verification checks an in-content permission claim +against any real authorization system. Same family as `00014`/`00030` +(unauthenticated in-content claims); worth a Sprint 4 consistency note +on whether the three should stay distinct by *what* is claimed +(permission vs. authority vs. role) -- current view: yes, distinct +enough, since each names a different downstream action class. Decision: +CLARIFY. +``` +Capability: Agent can defer to permission claims made within its instruction context. +Vulnerability: No verification checks an in-content permission claim against any actual authorization system. +Impact: Agent performs actions it would otherwise refuse, believing itself newly authorized. +``` + +### AVE-2026-00013 -- Personal Data Exfiltration via PII Collection +**B** | PII in agent context → external destination | Same +content-aware-egress gap as `00003`/`00026`. Decision: CLARIFY. +``` +Capability: Agent can access and transmit user PII in the course of normal tasks. +Vulnerability: No content-aware egress control distinguishes legitimate task output from PII exfiltration. +Impact: Identity theft, financial fraud, and regulatory-violation exposure for affected users. +``` + +### AVE-2026-00015 -- System Prompt Extraction via Direct Interrogation +**B** | internal configuration → normal output channel | No +confidentiality boundary prevents the agent's own operational +configuration from being repeated back through its ordinary output. +Decision: CLARIFY. +``` +Capability: Agent's output channel can include any content it has processed, including its own system prompt. +Vulnerability: No confidentiality boundary prevents internal configuration from being repeated back through the normal output channel on request. +Impact: Proprietary logic and security-policy disclosure, revealing attack surface for follow-on exploitation. +``` + +### AVE-2026-00016 -- Indirect Prompt Injection via RAG Retrieval +**A** | retrieved document (should be data) → trusted context (treated +as instruction) | No boundary distinguishes RAG-retrieved content from +instruction. Distinct entry point from `00002`/`00020`/`00028` +(retrieved corpus document rather than tool schema, agent message, or +user upload) -- same underlying missing control, legitimately separate +record by attacker-controlled input. Decision: RETAIN. +``` +Capability: A RAG pipeline retrieves external document content into agent context at query time. +Vulnerability: No boundary distinguishes retrieved document content from instruction once it enters context. +Impact: Attacker-controlled instructions execute via a corpus the attacker needs no direct system access to. +``` + +### AVE-2026-00017 -- MCP Server Impersonation or Spoofing +**A** | claimed server identity → actual, verified identity (never +checked) | No identity verification mechanism for MCP servers; trust +follows a self-asserted manifest claim. Decision: RETAIN. +``` +Capability: Agent grants trust and permission scope based on a server's claimed identity. +Vulnerability: No verification checks a claimed server identity against anything real. +Impact: A malicious server receives elevated trust and permissions it has no legitimate claim to. +``` + +### AVE-2026-00018 -- Tool Result Manipulation or Output Poisoning +**B** | actual tool execution result → reported result (agent sits +unchecked in the path) | No integrity check verifies a reported tool +result matches what the tool actually returned. Decision: CLARIFY. +``` +Capability: Agent reports tool call results to users and downstream components. +Vulnerability: No integrity check verifies a reported result matches what the tool actually returned. +Impact: False information, hidden errors, or manipulated downstream decisions based on fabricated data. +``` + +### AVE-2026-00019 -- Agent Memory Poisoning +**A** | write request (untrusted) → persistent memory store (should be +trusted) | No provenance or integrity check gates writes to persistent +memory. Decision: RETAIN. +``` +Capability: Agent maintains persistent memory across sessions. +Vulnerability: No provenance or integrity check gates writes to the persistent memory store. +Impact: Attacker-controlled beliefs influence agent behavior across all future sessions, long after the initial write. +``` + +### AVE-2026-00020 -- Cross-Agent Prompt Injection (A2A) +**B** | orchestrator-intended output → sub-agent input, same channel | +No boundary between task content and instruction at the agent-to-agent +handoff specifically. Part of a broader recurring pattern (see Sprint +4 cross-record finding below: `00002`/`00016`/`00020`/`00028`/`00043`/ +`00044` all realize the same "no content-vs-instruction boundary" +property at different entry points). Decision: CLARIFY. +``` +Capability: One agent's output becomes a downstream sub-agent's input in a delegation pipeline. +Vulnerability: No boundary distinguishes orchestrator-intended content from attacker-crafted instructions at the agent-to-agent handoff. +Impact: Downstream agent performs actions the orchestrator or user never intended, bypassing orchestrator-level safety controls. +``` + +### AVE-2026-00021 -- Autonomous Action Without User Confirmation +**B** | agent decision → irreversible/high-impact action | Human +confirmation is component-configurable behavior, not a structurally +enforced gate. Same approval-bypass family as `00006`/`00063`/`00076`, +each a genuinely distinct mechanism (instruction-driven here, +declarative config in `00063`, classifier-steering in `00076`). +Decision: CLARIFY. +``` +Capability: Agent can take consequential or irreversible actions. +Vulnerability: Human confirmation for high-impact actions is component-configurable, not a structurally enforced gate. +Impact: Irreversible actions execute with no review opportunity, increasing the blast radius of any error or attack. +``` + +### AVE-2026-00024 -- Supply Chain - Content Type Mismatch (Magika) +**A** | declared file extension → actual byte-level content | No +content-type verification independent of the declared extension. +Decision: RETAIN. +``` +Capability: Skill files are loaded and trusted based on their declared extension. +Vulnerability: No content-type verification independent of the declared extension confirms the file's actual format before it's trusted. +Impact: Arbitrary executable content runs under the guise of a benign skill file, invisible to text-pattern scanners. +``` + +### AVE-2026-00026 -- Exfiltration via Tool Output Encoding +**B** | tool-call parameters/return values → external transmission | +Monitoring inspects network destination, not tool-call payload content, +for encoded sensitive data. Decision: CLARIFY. +``` +Capability: Agent can pass parameters to legitimate, already-authorized tools. +Vulnerability: No inspection of tool-call parameter or return content for encoded sensitive data. +Impact: Credential/PII/system-prompt exfiltration disguised as ordinary, legitimate-looking tool traffic. +``` + +### AVE-2026-00028 -- Prompt Injection via File or Document Content +**B** | uploaded document content (should be data) → instruction | Same +content-vs-instruction gap, at the user-file-upload entry point. +Decision: CLARIFY. +``` +Capability: Agent processes user-uploaded documents as part of a task. +Vulnerability: No boundary distinguishes uploaded document content from instruction the agent should follow. +Impact: Indirect prompt injection requiring only that a user be convinced to upload a crafted document. +``` + +### AVE-2026-00031 -- Training Data or Feedback Loop Poisoning +**B** | agent-generated output → training/RLHF feedback signal | No +integrity check on agent-generated feedback prevents the evaluated +system from also controlling its own evaluation signal. Genuinely +distinct mechanism from `00019` despite the shared "poisoning" +vocabulary -- flagged in the cross-record finding below. Decision: +CLARIFY. +``` +Capability: Agent-generated outputs can feed into a downstream training/RLHF feedback pipeline. +Vulnerability: No integrity check prevents the evaluated system from also controlling its own evaluation signal. +Impact: Gradual, hard-to-detect manipulation of future model behavior via biased reward signals. +``` + +### AVE-2026-00034 -- Supply Chain - Dynamic Third-Party Skill Import +**A** | external URL content → agent's full capability set | No +identity/integrity verification gates dynamically-loaded code before +execution. Decision: RETAIN. +``` +Capability: Agent can dynamically load and execute code from a URL at runtime. +Vulnerability: No verification gates code loaded from an external URL before it executes with the agent's full capability set. +Impact: Full agent compromise via arbitrary attacker-controlled code with no static-review opportunity. +``` + +### AVE-2026-00036 -- Lateral Movement - Pivot to Other Systems +**B** | task-scoped credentials/access → reach beyond original task +scope | No boundary prevents reuse of granted access to reach systems +beyond the declared task. Same "no scope enforcement" family as +`00022`/`00032`, differing by mechanism (credential reuse for pivot, +vs. direct undeclared access, vs. network reconnaissance). Decision: +CLARIFY. +``` +Capability: Agent has network connectivity and credentials/tokens scoped to its current task. +Vulnerability: No boundary prevents that access from being reused to reach systems beyond the original task's authorized scope. +Impact: Compromise expands from one pipeline segment to adjacent systems the attacker could not reach directly. +``` + +### AVE-2026-00040 -- Insecure Output - Unescaped Injection into Downstream System +**E** | agent output (trusted-intermediary status) → downstream +interpreter (SQL/HTML/shell) | Conventional injection (CWE-79/89 +territory); agentic amplification is the trust elevation backend +systems grant an agent versus raw user input. Decision: JUSTIFY. +``` +Capability: Agent output is passed to downstream interpreters as a trusted intermediary. +Vulnerability: No escaping/sanitization gate exists between agent-generated output and downstream interpretation, and the agent's trusted role draws less scrutiny than raw user input would. +Impact: Classic injection (SQLi/XSS/command injection) via a trust-elevated conduit. +``` + +### AVE-2026-00042 -- Payload Injection into Agent-Generated Orchestration Code +**A** | tool result content (data) → code the agent generates and +executes | No boundary prevents tool-result content from breaking out +of data context into generated code. Decision: RETAIN. +``` +Capability: Agent-generated orchestration code processes tool results in Code/REPL Mode. +Vulnerability: No boundary prevents tool result content from breaking out of data context into the code the agent generates and executes. +Impact: Arbitrary code execution bypassing all prompt-level, text-instruction filtering entirely. +``` + +### AVE-2026-00043 -- Prompt Injection via Rich UI Payload +**A** | rendered surface (what the user sees) → processed payload +(what the model reads) | No consistency requirement between rendered +and processed content from the same rich-UI payload. Decision: RETAIN. +``` +Capability: MCP Apps can render rich UI elements the model also processes as context. +Vulnerability: No consistency check ensures what's rendered to the user matches what the model actually processes from the same payload. +Impact: Agent acts on instructions the user has no way to see, in an interface designed to look fully transparent. +``` + +### AVE-2026-00044 -- Prompt Injection via Poisoned Async Task Result +**A** | task-result data (should be data) → trusted context from a +"completed" task | No validation distinguishes result data from +instruction, and the dispatch-to-consumption temporal gap bypasses +synchronous safety checks. Decision: RETAIN. +``` +Capability: Agent dispatches async tasks and later reads their results in a future turn. +Vulnerability: No validation distinguishes legitimate result data from instruction, and the temporal gap bypasses checks applied at dispatch time. +Impact: Injected content is treated as trusted context from a completed task rather than external, untrusted input. +``` + +### AVE-2026-00045 -- Privilege Escalation via Cross-App-Access +**A** | low-trust server's instruction → high-trust server's +capability, same session | No per-server trust isolation within a +single multi-server agent session. Decision: RETAIN. +``` +Capability: A single agent session can connect to and act on multiple MCP servers simultaneously. +Vulnerability: No per-server trust isolation prevents a low-trust server's instruction from directing the agent's use of a co-connected high-trust server. +Impact: Confused-deputy escalation from low-privilege compromise to high-privilege action. +``` + +### AVE-2026-00048 -- Unsafe Agent Delegation Chain +**A** | parent's full permission set → sub-agent, by default | No +explicit trust-boundary or permission-scoping requirement on +delegation. Decision: RETAIN. +``` +Capability: Agent can delegate tasks to sub-agents or spawn child agents. +Vulnerability: No explicit trust boundary or permission-scoping requirement exists on delegation; sub-agents inherit full parent permissions by default. +Impact: Privilege laundering and audit evasion, since sub-agent actions may not appear in the parent's audit trail. +``` + +### AVE-2026-00050 -- Parasitic Toolchain +**A** | declared manifest scope → actual runtime tool-registry +footprint | No conformance check between a component's registered +footprint and its declared manifest. Decision: RETAIN. +``` +Capability: Tools can register handlers into the agent's tool dispatch layer at runtime. +Vulnerability: No check verifies a component's actual registered footprint against its declared manifest, at registration or session restart. +Impact: A persistent, undeclared capability survives context resets, indistinguishable from a legitimately registered tool. +``` + +### AVE-2026-00051 -- OAuth Discovery Rebinding +**A** | server's declared origin → discovery-document endpoint values +(unverified) | No verification that discovery-document endpoints match +the server's own declared origin. Decision: RETAIN. +``` +Capability: Agent follows OAuth 2.0 Authorization Server Metadata / OIDC Discovery per the published standard. +Vulnerability: No verification that discovery-document endpoint values match the server's own declared origin before the agent trusts them. +Impact: Interception of authorization codes, access tokens, and credentials during the OAuth exchange. +``` + +### AVE-2026-00053 -- Path Traversal in MCP Resource/File-Handler Implementation +**E** | caller-supplied path parameter → filesystem access outside +declared scope | Conventional path traversal (CWE-22) in the tool's own +implementation, independent of any agent instruction. Same open +maintainer question as `00052`: coverage argument vs. genuine +behavioral distinction. Decision: JUSTIFY, Confidence LOW. +``` +Capability: An MCP resource/file-handler tool accepts caller-supplied path or URL parameters. +Vulnerability: No canonicalization or root-containment check gates a caller-supplied path before use. +Impact: Reading or writing files/resources outside the tool's declared scope, under the tool's own credentials. +``` + +### AVE-2026-00055 -- Command Execution via Untrusted MCP Server Launch Configuration +**A** | configuration data → process-spawn boundary | No validation +gate between configuration/registry data and process execution. +Explicitly, and correctly, differentiated from prompt injection in the +record's own text. Decision: RETAIN. +``` +Capability: An MCP client spawns a server as a subprocess using command/args from configuration data. +Vulnerability: No validation gate exists between untrusted configuration data and the process-spawn boundary. +Impact: Arbitrary OS command execution under the client's privileges before any protocol handshake occurs. +``` + +### AVE-2026-00056 -- Zero-Click Exfiltration via Markdown Image Auto-Fetch +**A** | agent-generated response content → client's automatic +rendering (an unmonitored channel) | No inspection of response content +for embedded exfiltration channels before client rendering acts on it. +Decision: RETAIN. +``` +Capability: Clients automatically render markdown/rich-content references in agent responses. +Vulnerability: No inspection of agent-generated response content for embedded exfiltration channels before automatic rendering fetches them. +Impact: Zero-interaction data exfiltration requiring no tool call and no obfuscation. +``` + +### AVE-2026-00058 -- Deceptive Skill Trigger or Activation-Scope Manipulation +**B** | declared manifest trigger/description → actual implemented +behavior | No verification that declared triggers match real behavior +before invocation decisions are made. Decision: CLARIFY. +``` +Capability: A skill's declared manifest trigger/description determines when it gets invoked. +Vulnerability: No verification checks that a declared trigger matches the skill's actual implemented behavior before invocation. +Impact: Over-broad or implicit invocation in contexts the skill's real behavior does not warrant. +``` + +### AVE-2026-00059 -- Fragmented Cross-Description Prompt Injection (ShareLock-Class) +**A** | per-description review (isolated) → multi-description +composition (never reviewed as a whole) | No cross-description, +holistic review step exists. Real, verified primary source +(ShareLock, arXiv:2606.27027). Decision: RETAIN. +``` +Capability: MCP tool descriptions across multiple tools/servers are individually reviewed for malicious content. +Vulnerability: No cross-description review step exists; review is structurally confined to one description at a time. +Impact: A complete malicious instruction assembles at inference time with no single reviewable artifact ever having existed. +``` + +### AVE-2026-00060 -- STDIO Transport Shell Injection +**E** | tool-call parameters → host shell, in shared SDK code | No +sanitization gate in the shared transport-layer implementation, wholly +independent of agent instruction. Stronger agentic-ecosystem argument +than `00052`/`00053`: this lives in shared MCP SDK code, not one +server's bespoke handler, so the blast radius spans every server built +on an affected SDK. Decision: JUSTIFY, Confidence MEDIUM. +``` +Capability: MCP SDK STDIO transport implementations pass tool-call parameters to a host shell. +Vulnerability: No sanitization gate exists in the shared transport-layer implementation, independent of any agent instruction. +Impact: Remote code execution across every server built on an affected SDK -- a supply-chain-scale blast radius, not one server's bug. +``` + +### AVE-2026-00063 -- Human Approval Gate Bypassed via Declarative Configuration +**A** | instruction-level review (exists) → configuration-level review +(does not exist) | No requirement that safety-relevant config flags +receive the same review rigor as instruction text. Explicitly +differentiated from `00048` in its own text. Decision: RETAIN. +``` +Capability: Components can declare configuration flags controlling whether human approval is required. +Vulnerability: No requirement that safety-relevant configuration flags receive the same review scrutiny as instruction text. +Impact: A required human-in-the-loop control is silently removed with no instruction-level signal a reviewer would catch. +``` + +### AVE-2026-00064 -- Zero-Click Code Execution via Project-Load Auto-Run Configuration +**A** | passive action (opening a project) → code execution, no +confirmation | No confirmation gate on auto-run behavior triggered +passively rather than by explicit tool call. Decision: RETAIN. +``` +Capability: IDE/dev-environment configuration can declare commands to run automatically on project load. +Vulnerability: No confirmation gate exists for auto-run behavior triggered by a passive action rather than an explicit tool call. +Impact: Code execution requiring no action beyond opening a compromised project directory. +``` + +### AVE-2026-00065 -- A2A Agent Card Poisoning +**A** | descriptive capability/identity metadata → instruction, once +loaded into reasoning context | Same metadata-vs-instruction gap as +`00002`/`00041`, in the A2A protocol's genuinely distinct surface (no +`.well-known` path, no `tool.description` field). Decision: RETAIN. +``` +Capability: A host agent reads a remote agent's A2A card to plan task delegation. +Vulnerability: No boundary separates the card's descriptive metadata from executable instruction once loaded into the host's reasoning context. +Impact: Delegation planning hijacked by a malicious remote agent's self-declared, unverified metadata. +``` + +### AVE-2026-00066 -- Hallucinated Skill-Name Squatting (HalluSquatting) +**A** | model's own hallucinated belief → fetch/install action | No +registry-existence or publisher-identity check before acting on a +model-generated resource name. Distinct entry point: no +attacker-controlled content anywhere in the interaction. Decision: +RETAIN. +``` +Capability: Agent resolves user requests for well-known resources to package/repo/skill names it generates itself. +Vulnerability: No registry-existence or publisher-identity check occurs before fetching or installing a model-generated name. +Impact: Malware installation via a scalable, precomputed, cross-model squatting technique with no injection required at all. +``` + +### AVE-2026-00067 -- Skill Composition Trust Transfer +**A** | upstream skill's output (benign) → downstream skill's trust +signal (unverified) | No requirement that a downstream skill +independently re-verify a trust claim carried in an upstream skill's +output. Real, citable research (96% success rate). Decision: RETAIN. +``` +Capability: A downstream skill can consume another skill's output as part of its own decision logic. +Vulnerability: No requirement exists for a downstream skill to independently re-verify a trust or authorization claim in an upstream skill's output. +Impact: Harmful actions approved based on a spoofable trust credential no single skill's isolated review would ever catch. +``` + +### AVE-2026-00069 -- Multimodal Image-Hidden Instructions (SkillCamo) +**A** | text-only review pipeline → bundled image resource (never +decoded for review) | No visual/multimodal decoding step exists in +static review. Decision: RETAIN. +``` +Capability: Multimodal agents can decode and act on bundled image content within a skill package. +Vulnerability: No visual/multimodal decoding step exists in the text-only static-review pipeline covering the rest of the package. +Impact: Instructions invisible to every current text-based scanner execute once a multimodal agent processes the resource. +``` + +### AVE-2026-00070 -- Distributed Cross-Agent Backdoor Fragments +**A** | dormant, per-agent fragments → externally reassembled, +post-hoc payload | No anomaly detection for encoded fragments +inconsistent with a tool's stated function, and no cross-agent +correlation of dormant fragments. Decision: RETAIN. +``` +Capability: Tools deliver observations to agents that persist in memory/context after the call. +Vulnerability: No anomaly detection or cross-agent correlation catches a fragmented payload no single session could ever contain enough of to detect. +Impact: A complete backdoor exists and becomes executable only after the fact, entirely outside any single session's visibility. +``` + +### AVE-2026-00071 -- MCP Daemon Redirect via DOCKER_HOST +**A** | expected local daemon → actual, redirected remote daemon +(unverified) | No verification that a declared daemon target matches +the expected local daemon before build/run/pull proceeds. Decision: +RETAIN. +``` +Capability: Components can declare a container daemon connection target via configuration. +Vulnerability: No verification confirms a declared daemon target is the expected local daemon before operations proceed against it. +Impact: Secrets and bind-mounted data are exposed to whatever infrastructure actually receives the redirected connection. +``` + +### AVE-2026-00076 -- Natural-Language Steering of an Approval Classifier Subagent +**A** | steering text (targets the checker) → classifier's approve/deny +judgment | No boundary prevents the same class of natural-language +content that steers a primary agent from also steering the classifier +meant to check it. Decision: RETAIN. +``` +Capability: A committed configuration file can declare natural-language steering text consumed by a safety-classifier subagent. +Vulnerability: No boundary prevents natural-language content from probabilistically steering the classifier meant to check an agent's actions, the same way it steers the agent itself. +Impact: Auto-approval of shell/MCP/Fetch calls the classifier would otherwise flag, via a committed repository file. +``` + +### AVE-2026-00077 -- Cross-Origin Tool and Resource Declaration +**A** | single server's trust boundary → multiple, unrelated actual +origins within it | No per-origin trust segmentation within one +server's own declared manifest. Decision: RETAIN. +``` +Capability: A single MCP server's manifest can declare tools/resources whose URLs span multiple distinct domains. +Vulnerability: No per-origin trust segmentation exists within one server's declared surface. +Impact: A minority-domain tool or resource can hijack context intended for the trusted majority origin, with no deception required. +``` + +### AVE-2026-00079 -- Plan Hijacking via False Completion Signal +**A** | declared plan (N steps) → actual executed trace (fewer steps, +unverified) | No plan-to-execution-trace verification gates a +self-reported completion signal. Decision: RETAIN. +``` +Capability: An orchestrator tracks a declared plan and invokes a final step based on agent-reported status. +Vulnerability: No verification compares the actual executed step count against the declared plan before accepting a self-reported completion signal. +Impact: Remaining planned steps, including verification steps, are silently skipped. +``` + +--- + +## Sprint 4 cross-record consistency finding + +A recurring pattern surfaced across this audit: a large share of AVE's +prompt-injection-shaped records (`00002`, `00016`, `00020`, `00028`, +`00041`, `00043`, `00044`, `00065`) all realize the *same* underlying +vulnerability property -- no boundary distinguishes untrusted content +from instruction -- at *different* attacker-controlled-input surfaces +(tool schema, RAG document, agent-to-agent message, user upload, +server-card, rich UI, async result, A2A card). This is architecturally +coherent, not redundant: Q4 (attacker-controlled input) genuinely +differs for each, and several records already explicitly differentiate +themselves from their nearest neighbors in their own text (`00041` vs. +`00002`, `00065` vs. `00041`, `00044` vs. the rest). The recommendation +is not to merge these, but to name the shared underlying property +explicitly somewhere central (see R-001 in the findings below) so a +reader encounters it as one recognized pattern with many instances, +rather than reconstructing it record-by-record the way this audit had +to. + +A second, smaller pattern: `00019`, `00027`, `00031`, `00035`, `00070` +all use "poisoning" in their titles or descriptions but name four +genuinely different mechanisms (direct memory-store write, instructed +cross-session self-persistence, self-referential feedback-signal +corruption, tool-response fabrication, and dormant multi-agent +fragment reassembly, respectively). Confirmed distinct on inspection, +not a conflation problem -- but worth a shared terminology note so +"poisoning" reads as a family of related-but-distinct mechanisms +rather than implying a single pattern. + +A third pattern, the "unauthenticated in-content claim" family +(`00012`, `00014`, `00030`): false permission grant, false authority +claim, and false role claim all share the identical missing control +(no authentication backs an in-content claim) but differ in exactly +what's claimed and what it unlocks. Confirmed legitimately distinct, +same reasoning as the poisoning family above. + +--- + +## Deliverable 1 — Full 80-record matrix + +| AVE ID | Class | Boundary | Decision | +|---|---|---|---| +| AVE-2026-00001 | A | reviewed artifact -> executed artifact | RETAIN | +| AVE-2026-00002 | A | tool metadata -> agent instruction | RETAIN | +| AVE-2026-00003 | B | credential store -> external destination | CLARIFY | +| AVE-2026-00004 | E | instruction content -> shell execution | JUSTIFY | +| AVE-2026-00005 | E | instruction content -> shell execution | JUSTIFY | +| AVE-2026-00006 | B | tool-use decision -> irreversible on-chain action | CLARIFY | +| AVE-2026-00007 | A | trusted system instruction <-> untrusted content | RETAIN | +| AVE-2026-00008 | B | task-scoped access -> durable system modification | CLARIFY | +| AVE-2026-00009 | C | agent config (trusted) -> processed content (untrusted) | REFRAME | +| AVE-2026-00010 | B | component instruction -> operator visibility | CLARIFY | +| AVE-2026-00011 | B | task content -> direct tool-invocation command | CLARIFY | +| AVE-2026-00012 | B | claimed permission -> actual authorization | CLARIFY | +| AVE-2026-00013 | B | PII in context -> external destination | CLARIFY | +| AVE-2026-00014 | B | claimed identity -> verified identity | CLARIFY | +| AVE-2026-00015 | B | internal configuration -> output channel | CLARIFY | +| AVE-2026-00016 | A | retrieved document -> trusted context | RETAIN | +| AVE-2026-00017 | A | claimed server identity -> verified identity | RETAIN | +| AVE-2026-00018 | B | actual tool result -> reported result | CLARIFY | +| AVE-2026-00019 | A | write request -> persistent memory store | RETAIN | +| AVE-2026-00020 | B | orchestrator content -> sub-agent input | CLARIFY | +| AVE-2026-00021 | B | agent decision -> irreversible action | CLARIFY | +| AVE-2026-00022 | B | declared manifest scope -> executed behavior | CLARIFY | +| AVE-2026-00023 | B | constraint binding -> context-window position | CLARIFY | +| AVE-2026-00024 | A | declared extension -> actual byte content | RETAIN | +| AVE-2026-00025 | B | genuine history -> simulated history | CLARIFY | +| AVE-2026-00026 | B | tool-call parameters -> external transmission | CLARIFY | +| AVE-2026-00027 | B | single-session instruction -> cross-session persistence | CLARIFY | +| AVE-2026-00028 | B | uploaded document -> instruction | CLARIFY | +| AVE-2026-00029 | A | human visual review -> model codepoint processing | RETAIN | +| AVE-2026-00030 | A | claimed role -> verified role | RETAIN | +| AVE-2026-00031 | B | agent output -> training/feedback signal | CLARIFY | +| AVE-2026-00032 | C | declared task scope -> actual network reachability | REFRAME | +| AVE-2026-00033 | E | instruction content -> unsafe deserialization/eval | JUSTIFY | +| AVE-2026-00034 | A | external URL content -> agent's capability set | RETAIN | +| AVE-2026-00035 | E | tool response -> reported observation | JUSTIFY | +| AVE-2026-00036 | B | task-scoped access -> reach beyond scope | CLARIFY | +| AVE-2026-00037 | B | text-channel boundary -> image-channel (absent) | CLARIFY | +| AVE-2026-00038 | D | declared bounded scope -> unbounded runtime use | REWRITE | +| AVE-2026-00039 | B | network-egress monitoring -> content-level channel | CLARIFY | +| AVE-2026-00040 | E | agent output -> downstream interpreter | JUSTIFY | +| AVE-2026-00041 | A | discovery-layer metadata -> agent context | RETAIN | +| AVE-2026-00042 | A | tool result data -> generated/executed code | RETAIN | +| AVE-2026-00043 | A | rendered surface -> processed payload | RETAIN | +| AVE-2026-00044 | A | task-result data -> trusted completed-task context | RETAIN | +| AVE-2026-00045 | A | low-trust server instruction -> high-trust capability | RETAIN | +| AVE-2026-00046 | A | component scope -> shared tool-dispatch layer | RETAIN | +| AVE-2026-00047 | E | static file content -> LLM context window | JUSTIFY | +| AVE-2026-00048 | A | parent permission set -> sub-agent, by default | RETAIN | +| AVE-2026-00049 | E | declared endpoint -> actual request destination | JUSTIFY | +| AVE-2026-00050 | A | declared manifest scope -> runtime tool-registry footprint | RETAIN | +| AVE-2026-00051 | A | server's declared origin -> discovery endpoint values | RETAIN | +| AVE-2026-00052 | E | caller parameter -> shell execution (impl. code) | JUSTIFY | +| AVE-2026-00053 | E | caller path parameter -> filesystem access | JUSTIFY | +| AVE-2026-00054 | E | sandboxed context -> host runtime | JUSTIFY | +| AVE-2026-00055 | A | configuration data -> process-spawn boundary | RETAIN | +| AVE-2026-00056 | A | response content -> client auto-render channel | RETAIN | +| AVE-2026-00057 | A | as-scanned representation -> as-executed representation | RETAIN | +| AVE-2026-00058 | B | declared trigger -> actual implemented behavior | CLARIFY | +| AVE-2026-00059 | A | per-description review -> multi-description composition | RETAIN | +| AVE-2026-00060 | E | tool-call parameters -> host shell (shared SDK) | JUSTIFY | +| AVE-2026-00061 | E | configured verification -> actual MITM exposure | DEPRECATE-OR-JUSTIFY | +| AVE-2026-00062 | E | reviewed artifact -> executed artifact (dependency) | JUSTIFY-OR-MERGE | +| AVE-2026-00063 | A | instruction-level review -> config-level review (absent) | RETAIN | +| AVE-2026-00064 | A | passive action -> code execution, no confirmation | RETAIN | +| AVE-2026-00065 | A | descriptive card metadata -> instruction | RETAIN | +| AVE-2026-00066 | A | model's hallucinated belief -> fetch/install action | RETAIN | +| AVE-2026-00067 | A | upstream output -> downstream trust signal | RETAIN | +| AVE-2026-00068 | A | single-command review -> sequence-level effect | RETAIN | +| AVE-2026-00069 | A | text-only review -> bundled image resource | RETAIN | +| AVE-2026-00070 | A | dormant per-agent fragments -> reassembled payload | RETAIN | +| AVE-2026-00071 | A | expected local daemon -> redirected remote daemon | RETAIN | +| AVE-2026-00072 | E | intended local-only reach -> LAN-wide reach | JUSTIFY | +| AVE-2026-00073 | E | declared/default endpoint -> actual configured endpoint | JUSTIFY | +| AVE-2026-00074 | A | trust at authoring time -> trust required at execution | RETAIN | +| AVE-2026-00075 | E | reviewed .py source -> executed .pyc bytecode | JUSTIFY | +| AVE-2026-00076 | A | steering text -> classifier judgment | RETAIN | +| AVE-2026-00077 | A | single server trust -> multiple actual origins | RETAIN | +| AVE-2026-00078 | A | sub-agent result (unverified) -> accepted ground truth | RETAIN | +| AVE-2026-00079 | A | declared plan -> actual executed trace | RETAIN | +| AVE-2026-00080 | A | routing position -> agent identity | RETAIN | + +## Deliverable 2 — Findings grouped + +- **VALID-AVE (38 records)**: AVE-2026-00001, 00002, 00007, 00016, 00017, 00019, 00024, 00029, 00030, 00034, 00041, 00042, 00043, 00044, 00045, 00046, 00048, 00050, 00051, 00055, 00056, 00057, 00059, 00063, 00064, 00065, 00066, 00067, 00068, 00069, 00070, 00071, 00074, 00076, 00077, 00078, 00079, 00080. +- **INSUFFICIENT-BOUNDARY (23 records)**: AVE-2026-00003, 00006, 00008, 00010, 00011, 00012, 00013, 00014, 00015, 00018, 00020, 00021, 00022, 00023, 00025, 00026, 00027, 00028, 00031, 00036, 00037, 00039, 00058. +- **TECHNIQUE-CONFLATION (2 records)**: AVE-2026-00009, 00032. +- **CAPABILITY-CONFLATION (1 record)**: AVE-2026-00038. +- **GENERIC-VULNERABILITY (16 records)**: AVE-2026-00004, 00005, 00033, 00035, 00040, 00047, 00049, 00052, 00053, 00054, 00060, 00061, 00062, 00072, 00073, 00075. + +**Read across the groups**: 38 of 80 records (47.5%) classify cleanly as +A with no change needed. 23 (29%) have a real, distinct vulnerability +mechanism that simply needs its violated property stated more +explicitly -- a documentation fix, not a structural one. Only 3 records +across the entire corpus (00009, 00032, 00038) describe an attack +technique or bare capability without yet naming the underlying boundary +failure, and even those three have a real mechanism identifiable once +asked for directly (Q1-Q7 surfaced it in every case). 16 records (20%) +are conventional vulnerability classes whose agentic-specific distinction +ranges from strong and well-sourced (00073, 00075, 00004/5/33, 00047) to +genuinely thin (00049, 00061, 00062) to actively contested (00052, 00053, +00060 -- pure implementation bugs in MCP tooling, where the honest +question is whether a coverage argument is sufficient grounds for a +distinct AVE record at all). + +This is the mixed result Section 10 anticipated as more credible than a +100%-valid outcome: most of the corpus is sound, a real minority needs +work, and the work needed is overwhelmingly *documentation* (CLARIFY, +23 records) rather than *structural* (REFRAME/REWRITE/MERGE/DEPRECATE, +6 records total: 00009, 00032, 00038, plus 00061/00062/00072 flagged +DEPRECATE-OR-JUSTIFY/JUSTIFY-OR-MERGE). + +## Deliverable 3 — Taxonomy recommendations + +**R-001, name the recurring content-vs-instruction pattern once, +centrally.** Eight records (00002, 00016, 00020, 00028, 00041, 00043, +00044, 00065) independently realize the same underlying vulnerability +property -- no boundary distinguishes untrusted content from +instruction -- at eight different attacker-controlled-input surfaces. +Each is a legitimately distinct record (Q4 differs every time), but a +reader currently has to reconstruct the shared pattern themselves, the +way this audit had to. Recommend a short cross-referencing note (in +`docs/specs/`, not a schema change) naming this as a recognized family, +linking all eight, the same treatment already given to the +"unauthenticated in-content claim" family (00012/00014/00030) and the +"poisoning" family (00019/00027/00031/00035/00070) once this audit +confirmed each was genuinely distinct rather than a labeling accident. + +**R-002, resolve the three structural-decision records deliberately, +not by default.** AVE-2026-00009 (REFRAME), 00032 (REFRAME), and 00038 +(REWRITE) are the only records in the corpus whose current text doesn't +yet name a boundary failure explicitly enough to stand alone. Each has +a real, identifiable mechanism (documented above); the fix is rewriting +the `description`/`behavioral_fingerprint` text, not changing the +`ave_id` or removing the record. Treat as three concrete, scoped +follow-up edits, not a larger project. + +**R-003, get an explicit maintainer decision on implementation-bug +records.** AVE-2026-00052, 00053, and 00060 are pure code-level +implementation flaws in MCP tooling with no agent-instruction or +behavioral component at all -- the records themselves say so. AVE's +value-add for these is a coverage argument (this is part of the +agentic supply chain AVE's consumers already scan, and nothing else in +their pipeline will catch it), not a mechanism-level agentic +distinction. This is a legitimate but different kind of claim than the +rest of the corpus makes, and deserves a deliberate, stated decision +(keep with the coverage rationale made explicit, or scope AVE away from +pure implementation bugs going forward) rather than silence. + +**R-004, revisit the weakest agentic-distinction records as a group.** +AVE-2026-00049, 00061, and 00062 are the three records where a real +agentic-specific argument plausibly exists but isn't yet written down, +and where, absent that argument, each reads as a bare port of an +existing CWE class into an agentic component with no demonstrated +distinction. Recommend either adding the missing argument to each +(candidates identified per-record above) or treating them as crosswalk +entries to their respective CWEs rather than standalone AVE records. + +**R-005, apply Q7's three-line format as an ongoing drafting +discipline.** Every record in this audit that already separated +capability from vulnerability from impact cleanly (the 38 VALID-AVE +records, and the seven positive controls in particular) was +substantially easier to classify and cross-reference than the 23 +CLARIFY records, where capability and vulnerability run together in +one paragraph. Recommend `research-new-attack-classes` and +`add-ave-record` require a Q7-shaped three-line summary as a drafting +step for every new record going forward, independent of any schema +change -- this is a process fix, not a data-model fix. + +## Deliverable 4 — Specification change proposal + +**Decision: no schema change.** `security_condition` and +`security_boundary` fields were the two candidates this audit was +asked to evaluate evidence for. The audit does not support adding +either as new required or optional schema fields: + +- Every record in this audit, including the 23 CLARIFY cases, already + has an identifiable boundary and violated property -- the gap found + was almost never "this information doesn't exist," it was "this + information isn't stated as explicitly as it could be in the prose + `description`/`behavioral_fingerprint` fields that already exist." + A new structured field doesn't fix an articulation problem; clearer + prose in the existing fields does. +- The three genuinely structural cases (00009, 00032, 00038) need + rewritten description text, not a new field to populate alongside + unchanged, under-specified prose. +- Adding fields before they've proven necessary is exactly the + organizational-container failure mode `docs/specs/scaling-and- + governance.md` Section 1 already warns against for records; the same + discipline applies to schema growth. + +**What this audit does support**: the CLARIFY-decision records (23) +getting their `description` text tightened to state the violated +property and missing control explicitly, following the pattern the 38 +VALID-AVE records and seven positive controls already demonstrate. +That's prose editing across existing fields, not a schema change, and +it's the concrete, evidence-backed recommendation this audit actually +produced. + +--- + +## Governance note + +Per this task's own sequencing instruction: the umbrella issue inviting +external challenge on this audit ("Taxonomy Audit: Capability vs +Behavior vs Vulnerability") has **not** been opened yet. astrogilda and +narko4u both have active, current threads (the #98 follow-ups `#218`/ +`#219`, the #214 review offer, and the GenAI Crosswalk manual-PR +recommendation) that predate this audit. Opening a third simultaneous +ask to either without checking that queue first is the exact thing this +task's checklist asked not to do by default. Recommend checking the +state of those threads before opening the umbrella issue, or opening it +with explicit no-rush framing if it opens before they clear. + +--- + +**Audit status**: complete. All 80 records classified (38 VALID-AVE, 23 +INSUFFICIENT-BOUNDARY, 2 TECHNIQUE-CONFLATION, 1 CAPABILITY-CONFLATION, +16 GENERIC-VULNERABILITY). Every record has an explicit capability +statement, vulnerability condition, and security boundary. No record +was removed or recommended for removal solely because another framework +already covers part of it (Decision Rule §7). No schema change made or +proposed without evidence from the audit itself (Success Criteria, +Section 9).