fix: mitre_atlas citation corrections per issue #127's audit - #162
Merged
Merged
Conversation
External-fetch instruction replacement at runtime matches AML.T0081 (Modify AI Agent Configuration) precisely -- config/instruction replacement persisting beyond one session. T0054 (LLM Jailbreak) doesn't describe guardrail circumvention here. Per issue #127's audit.
MCP tool-description injection is textbook LLM Prompt Injection, delivered via a separate metadata channel (Indirect, .001), not cited at all previously. T0054 doesn't fit -- no guardrail circumvention described. Per issue #127's audit.
Direct instruction to exfiltrate credentials isn't adversarial-input crafting (T0043's actual definition). T0048 (External Harms) already correctly covers the completed-impact endpoint. Per issue #127's audit.
Shell-pipe RCE matches AML.T0050 (Command and Scripting Interpreter) precisely. T0054 doesn't apply -- this is tool/code-exec abuse, not guardrail circumvention. T0011 (User Execution) already correctly kept. Per issue #127's audit.
Persistence mechanism (cron/systemd/shell-profile self-copy), not itself a completed external-harm outcome. Researched AML.TA0006's own technique list (T0020/T0018/T0061/T0070/T0080/T0081/T0093/T0099/T0110) directly -- T0061 'LLM Prompt Self-Replication' looked promising by name but is prompt-text propagating between LLMs via output, a different mechanism from OS-level persistence via an agent. No ATLAS technique covers this; confirmed gap, not left in place by default. Per issue #127's audit.
Dynamic tool-call injection hijacking tool selection matches AML.T0053 (AI Agent Tool Invocation) precisely -- 'adversaries may abuse this to gain access to tools they otherwise wouldn't be able to use.' Neither previously-cited ID fit. Per issue #127's audit.
PII exfiltration instruction doesn't describe guardrail circumvention. T0048 (User Harm) already correctly covers the endpoint. Per issue #127's audit.
Defeats a specific guardrail (system-prompt secrecy) but still a stretch vs. ATLAS's alignment-circumvention framing. JUDGMENT call, defaulted to drop; no cleaner-fitting technique found. Per issue #127's audit.
Fingerprint describes indirect injection via RAG-retrieved content, but the citation was .000 (Direct). Corrected to .001 (Indirect), matching ATLAS's own definition. Per issue #127's audit.
MCP server impersonation matches AML.T0073 (Impersonation) precisely -- 'impersonate a trusted person or organization in order to persuade and trick a target.' T0043 (adversarial-input crafting) never fit. Per issue #127's audit.
Tool-result manipulation/output poisoning matches AML.T0067 (LLM Trusted Output Components Manipulation) precisely -- manipulating response components to appear trustworthy. T0048 is a completed-impact category, not this mechanism. Per issue #127's audit.
Agent memory poisoning is an exact match for AML.T0080.000 (Memory) -- 'manipulate the memory of an LLM in order to persist changes... to future chat sessions.' T0054 never described this mechanism. Per issue #127's audit.
Cross-agent A2A prompt injection is LLM Prompt Injection delivered via inter-agent communication, a separate channel (Indirect, .001). Neither previously-cited ID fit; T0051 wasn't cited at all. Per issue #127's audit.
Bypasses an operational human-confirmation checkpoint, not the model's own safety alignment. Researched ATLAS broadly (human oversight, guardrail, security-control keywords) -- no technique covers this specific operational-oversight-bypass concept. Confirmed gap. Per issue #127's audit.
Accessing undeclared/out-of-scope resources matches AML.T0053 (AI Agent Tool Invocation) -- 'agents may be configured to have access to tools that are not directly accessible by users; adversaries may abuse this.' T0043 never fit. Per issue #127's audit.
Output encoding for exfiltration isn't adversarial-input crafting. T0048 (data/IP theft) already correctly covers the endpoint. Per issue #127's audit.
Multi-turn instruction persistence matches AML.T0080.001 (Thread) precisely -- 'introduce malicious instructions into a chat thread... to cause behavior changes which persist for the remainder of the thread.' Per issue #127's audit.
Indirect prompt injection via untrusted file/document content -- the record's own text already says 'indirect prompt injection.' Matches T0051.001 exactly; neither previously-cited ID fit and T0051 wasn't cited at all. Per issue #127's audit.
Unicode homoglyph/zero-width obfuscation matches AML.T0068 (LLM Prompt Obfuscation) precisely -- hiding instructions from human review while remaining processable by the model, including via encoding schemes. T0054 never described this. Per issue #127's audit.
Unverified self-reported role claim granting privilege -- checked AML.TA0012 (Privilege Escalation)'s full technique list (Valid Accounts, AI Agent Tool Invocation, LLM Jailbreak, Escape to Host) directly; none cover an unverified role-claim backdoor. Confirmed gap. Per issue #127's audit.
…ACE) Training/feedback-loop poisoning is an exact match for AML.T0020 (Poison Training Data), confirmed against ATLAS.yaml. T0054 never fit -- this corrupts a training pipeline, not runtime guardrails. T0011 (already correctly cited) stays. Per issue #127's audit.
…D-REPLACE) Network reconnaissance instruction is an exact match for AML.T0006 (Active Scanning), confirmed against ATLAS.yaml. Neither previously-cited ID fit -- this is network/port/service recon, not adversarial-input crafting or completed impact. Per issue #127's audit.
Unsafe deserialization/eval instruction is a code-exec/tool-abuse mechanism, not guardrail circumvention. T0011 (User Execution, already correctly cited) already covers this. Per issue #127's audit.
Dynamic third-party skill import is a supply-chain mechanism -- AML.T0010 (AI Supply Chain Compromise) fits its own description directly. T0054 never fit. T0011 (already correctly cited) stays. Per issue #127's audit.
Fabricating sensor/environment data to deceive operators. Checked AML.T0041 (Physical Environment Access, external physical tampering with data collection) and T0067 (LLM output-trustworthiness manipulation, already used for 00018) -- neither cleanly covers an agent instructed to lie about its own telemetry/observations specifically. Confirmed gap rather than force a stretch. Per issue #127's audit.
…D-DROP + follow-up) Lateral movement pivot matches AML.T0091 (Use Alternate Authentication Material) precisely -- 'AI services commonly use alternate authentication material... making them vulnerable to this technique,' i.e. using the agent's existing credentials/tokens to move laterally. Confirmed AML.TA0015 (Lateral Movement) is a distinct tactic with only two techniques (Phishing, already used elsewhere for the confused-deputy records; and this one); neither previously-cited ID was under it. Per issue #127's audit.
Prompt injection via image/vision input bypassing text-level filters -- the fingerprint explicitly describes this. Matches T0051.001 (Indirect, ATLAS's own definition names multimedia explicitly). T0054 never fit. Per issue #127's audit.
Unbounded tool use + unrestricted sub-agent spawning. T0053 (AI Agent Tool Invocation) covers only the tool-use half cleanly; no ATLAS technique for unrestricted sub-agent/multi-agent spawning was found. Confirmed gap rather than force a partial fit. Per issue #127's audit.
MCP server-card tool-description injection is LLM Prompt Injection via a metadata channel (Indirect, .001). Neither previously-cited ID fit; T0051 wasn't cited at all. Per issue #127's audit.
Poisoned tool-result content injected into orchestration code is LLM Prompt Injection via a tool-result channel (Indirect, .001). Same pattern as 00041/00044. Per issue #127's audit.
Rich-UI payload injection via non-rendered elements (hidden divs, ARIA labels, SVG metadata) is Indirect Prompt Injection -- the user doesn't see it, the model processes it. T0043 never fit. Per issue #127's audit.
Poisoned async task result from an external queue/webhook is LLM Prompt Injection via that channel (Indirect, .001). Same pattern as 00041/00042. Per issue #127's audit.
T0052 (Phishing) was already the correctly-fitting citation here -- the templated T0043+T0048 pair added nothing to a confused-deputy cross-app-access mechanism. Per issue #127's audit.
Same pattern as 00045 -- T0052 (Phishing) already fits MCP tool-hook hijacking; the templated pair is dropped. Per issue #127's audit.
Same pattern as 00045/00046 -- T0052 (Phishing) already fits unsafe agent delegation; the templated pair is dropped. Per issue #127's audit.
T0010 (AI Supply Chain Compromise, already correctly cited) already fits silent tool registration; T0043 (adversarial-input crafting) added nothing. Per issue #127's audit.
Summarizes the 43-record correction across the preceding commits. dist/ave-records-latest.json regenerated to match; frozen versioned snapshot untouched.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #127. Implements the full scorecard, record by record (44 commits: 43 individual record fixes + one wrap-up for CHANGELOG/dist).
Summary
ATLAS.yamlresearch against each record's actualbehavioral_fingerprint(not inferred or reused): AVE-2026-00001 (T0081), 00002/00020/00028/00037/00041/00042/00043/00044 (T0051.001 -- determined per-record which prompt-injection sub-technique fits, all landed on Indirect since none are a user directly typing into chat), 00004 (T0050), 00011/00022 (T0053), 00017 (T0073), 00018 (T0067), 00019 (T0080.000), 00027 (T0080.001), 00029 (T0068), 00036 (T0091).No score, severity, or mechanism-description changes anywhere -- citation accuracy only.
last_updatedbumped on all 43 touched records.Test plan
python3 scripts/validate_records.py-- 76/76 valid, exit 0 (same 6 pre-existing, unrelated researcher-warning notices as before this PR)python3 scripts/check_fixtures.py-- all records have positive + negative fixturespython3 scripts/validate_crosswalks.py-- 6/6 validpytest tests/ -x -q-- 305 passednode scripts/build-records.js-- dist regenerated, frozen v1.1.0 snapshot untouched