Skip to content

fix: mitre_atlas citation corrections per issue #127's audit - #162

Merged
chaksaray merged 44 commits into
developfrom
fix/mitre-atlas-audit-127
Aug 9, 2026
Merged

chaksaray merged 44 commits into
developfrom
fix/mitre-atlas-audit-127

Conversation

@chaksaray

Copy link
Copy Markdown
Contributor

Closes #127. Implements the full scorecard, record by record (44 commits: 43 individual record fixes + one wrap-up for CHANGELOG/dist).

Summary

  • 9 records got a source-verified replacement technique, found via fresh ATLAS.yaml research against each record's actual behavioral_fingerprint (not inferred or reused): AVE-2026-00001 (T0081), 00002/00020/00028/00037/00041/00042/00043/00044 (T0051.001 -- determined per-record which prompt-injection sub-technique fits, all landed on Indirect since none are a user directly typing into chat), 00004 (T0050), 00011/00022 (T0053), 00017 (T0073), 00018 (T0067), 00019 (T0080.000), 00027 (T0080.001), 00029 (T0068), 00036 (T0091).
  • 2 already-verified replacements from the issue itself: 00031 -> T0020 (Poison Training Data), 00032 -> T0006 (Active Scanning).
  • 5 records got a confirmed straight drop, no forced replacement: 00008, 00021, 00030, 00035, 00038 -- researched each against the relevant ATLAS tactic's full technique list directly (not just keyword search), genuinely no technique covers the mechanism. Reasoning for each is in its own commit.
  • 8 "defensible either way" judgment calls defaulted to drop: 00007, 00010, 00012, 00014, 00015, 00025, 00040, 00056 -- per this project's verify-don't-infer framework-mapping standard, holding to ATLAS's strict technique definitions rather than keeping a stretch.
  • 1 subtype fix: 00016, T0051.000 (Direct) -> T0051.001 (Indirect), matching its own fingerprint.
  • The rest: straightforward drops of the templated T0043+T0048 pair where a correct citation (T0052, T0010, T0011) was already present alongside it (00003, 00013, 00026, 00033, 00045, 00046, 00048, 00050), and 00034 additionally gains T0010 per discussion.

No score, severity, or mechanism-description changes anywhere -- citation accuracy only. last_updated bumped on all 43 touched records.

Test plan

  • python3 scripts/validate_records.py -- 76/76 valid, exit 0 (same 6 pre-existing, unrelated researcher-warning notices as before this PR)
  • python3 scripts/check_fixtures.py -- all records have positive + negative fixtures
  • python3 scripts/validate_crosswalks.py -- 6/6 valid
  • pytest tests/ -x -q -- 305 passed
  • node scripts/build-records.js -- dist regenerated, frozen v1.1.0 snapshot untouched

External-fetch instruction replacement at runtime matches AML.T0081 (Modify AI Agent Configuration) precisely -- config/instruction replacement persisting beyond one session. T0054 (LLM Jailbreak) doesn't describe guardrail circumvention here.

Per issue #127's audit.
MCP tool-description injection is textbook LLM Prompt Injection, delivered via a separate metadata channel (Indirect, .001), not cited at all previously. T0054 doesn't fit -- no guardrail circumvention described.

Per issue #127's audit.
Direct instruction to exfiltrate credentials isn't adversarial-input crafting (T0043's actual definition). T0048 (External Harms) already correctly covers the completed-impact endpoint.

Per issue #127's audit.
Shell-pipe RCE matches AML.T0050 (Command and Scripting Interpreter) precisely. T0054 doesn't apply -- this is tool/code-exec abuse, not guardrail circumvention. T0011 (User Execution) already correctly kept.

Per issue #127's audit.
T0054 was a defensible-but-not-clean stretch per issue #127's own scorecard; defaulted to drop rather than keep per this project's verify-don't-infer standard. T0051 (already correctly cited) stays.

Per issue #127's audit.
Persistence mechanism (cron/systemd/shell-profile self-copy), not itself a completed external-harm outcome. Researched AML.TA0006's own technique list (T0020/T0018/T0061/T0070/T0080/T0081/T0093/T0099/T0110) directly -- T0061 'LLM Prompt Self-Replication' looked promising by name but is prompt-text propagating between LLMs via output, a different mechanism from OS-level persistence via an agent. No ATLAS technique covers this; confirmed gap, not left in place by default.

Per issue #127's audit.
Secrecy directive (don't reveal instructions to the user) targets human review, not the model's own guardrails. JUDGMENT call per issue #127, defaulted to drop; no ATLAS technique found for this specific mechanism.

Per issue #127's audit.
Dynamic tool-call injection hijacking tool selection matches AML.T0053 (AI Agent Tool Invocation) precisely -- 'adversaries may abuse this to gain access to tools they otherwise wouldn't be able to use.' Neither previously-cited ID fit.

Per issue #127's audit.
Same pattern as 00007 -- defensible stretch, defaulted to drop per issue #127. T0051 (already correctly cited) stays.

Per issue #127's audit.
PII exfiltration instruction doesn't describe guardrail circumvention. T0048 (User Harm) already correctly covers the endpoint.

Per issue #127's audit.
Same pattern as 00007/00012 -- defensible stretch, defaulted to drop per issue #127. T0051 (already correctly cited) stays.

Per issue #127's audit.
Defeats a specific guardrail (system-prompt secrecy) but still a stretch vs. ATLAS's alignment-circumvention framing. JUDGMENT call, defaulted to drop; no cleaner-fitting technique found.

Per issue #127's audit.
Fingerprint describes indirect injection via RAG-retrieved content, but the citation was .000 (Direct). Corrected to .001 (Indirect), matching ATLAS's own definition.

Per issue #127's audit.
MCP server impersonation matches AML.T0073 (Impersonation) precisely -- 'impersonate a trusted person or organization in order to persuade and trick a target.' T0043 (adversarial-input crafting) never fit.

Per issue #127's audit.
Tool-result manipulation/output poisoning matches AML.T0067 (LLM Trusted Output Components Manipulation) precisely -- manipulating response components to appear trustworthy. T0048 is a completed-impact category, not this mechanism.

Per issue #127's audit.
Agent memory poisoning is an exact match for AML.T0080.000 (Memory) -- 'manipulate the memory of an LLM in order to persist changes... to future chat sessions.' T0054 never described this mechanism.

Per issue #127's audit.
Cross-agent A2A prompt injection is LLM Prompt Injection delivered via inter-agent communication, a separate channel (Indirect, .001). Neither previously-cited ID fit; T0051 wasn't cited at all.

Per issue #127's audit.
Bypasses an operational human-confirmation checkpoint, not the model's own safety alignment. Researched ATLAS broadly (human oversight, guardrail, security-control keywords) -- no technique covers this specific operational-oversight-bypass concept. Confirmed gap.

Per issue #127's audit.
Accessing undeclared/out-of-scope resources matches AML.T0053 (AI Agent Tool Invocation) -- 'agents may be configured to have access to tools that are not directly accessible by users; adversaries may abuse this.' T0043 never fit.

Per issue #127's audit.
Fabricated conversation history to manipulate perceived consent -- defensible stretch, defaulted to drop per issue #127's JUDGMENT call.

Per issue #127's audit.
Output encoding for exfiltration isn't adversarial-input crafting. T0048 (data/IP theft) already correctly covers the endpoint.

Per issue #127's audit.
Multi-turn instruction persistence matches AML.T0080.001 (Thread) precisely -- 'introduce malicious instructions into a chat thread... to cause behavior changes which persist for the remainder of the thread.'

Per issue #127's audit.
Indirect prompt injection via untrusted file/document content -- the record's own text already says 'indirect prompt injection.' Matches T0051.001 exactly; neither previously-cited ID fit and T0051 wasn't cited at all.

Per issue #127's audit.
Unicode homoglyph/zero-width obfuscation matches AML.T0068 (LLM Prompt Obfuscation) precisely -- hiding instructions from human review while remaining processable by the model, including via encoding schemes. T0054 never described this.

Per issue #127's audit.
Unverified self-reported role claim granting privilege -- checked AML.TA0012 (Privilege Escalation)'s full technique list (Valid Accounts, AI Agent Tool Invocation, LLM Jailbreak, Escape to Host) directly; none cover an unverified role-claim backdoor. Confirmed gap.

Per issue #127's audit.
…ACE)

Training/feedback-loop poisoning is an exact match for AML.T0020 (Poison Training Data), confirmed against ATLAS.yaml. T0054 never fit -- this corrupts a training pipeline, not runtime guardrails. T0011 (already correctly cited) stays.

Per issue #127's audit.
…D-REPLACE)

Network reconnaissance instruction is an exact match for AML.T0006 (Active Scanning), confirmed against ATLAS.yaml. Neither previously-cited ID fit -- this is network/port/service recon, not adversarial-input crafting or completed impact.

Per issue #127's audit.
Unsafe deserialization/eval instruction is a code-exec/tool-abuse mechanism, not guardrail circumvention. T0011 (User Execution, already correctly cited) already covers this.

Per issue #127's audit.
Dynamic third-party skill import is a supply-chain mechanism -- AML.T0010 (AI Supply Chain Compromise) fits its own description directly. T0054 never fit. T0011 (already correctly cited) stays.

Per issue #127's audit.
Fabricating sensor/environment data to deceive operators. Checked AML.T0041 (Physical Environment Access, external physical tampering with data collection) and T0067 (LLM output-trustworthiness manipulation, already used for 00018) -- neither cleanly covers an agent instructed to lie about its own telemetry/observations specifically. Confirmed gap rather than force a stretch.

Per issue #127's audit.
…D-DROP + follow-up)

Lateral movement pivot matches AML.T0091 (Use Alternate Authentication Material) precisely -- 'AI services commonly use alternate authentication material... making them vulnerable to this technique,' i.e. using the agent's existing credentials/tokens to move laterally. Confirmed AML.TA0015 (Lateral Movement) is a distinct tactic with only two techniques (Phishing, already used elsewhere for the confused-deputy records; and this one); neither previously-cited ID was under it.

Per issue #127's audit.
Prompt injection via image/vision input bypassing text-level filters -- the fingerprint explicitly describes this. Matches T0051.001 (Indirect, ATLAS's own definition names multimedia explicitly). T0054 never fit.

Per issue #127's audit.
Unbounded tool use + unrestricted sub-agent spawning. T0053 (AI Agent Tool Invocation) covers only the tool-use half cleanly; no ATLAS technique for unrestricted sub-agent/multi-agent spawning was found. Confirmed gap rather than force a partial fit.

Per issue #127's audit.
Insecure output handling is plausible as an eventual-impact tag but is more precisely a mechanism-stage injection vector than a completed T0048 impact. JUDGMENT call, defaulted to drop per issue #127.

Per issue #127's audit.
MCP server-card tool-description injection is LLM Prompt Injection via a metadata channel (Indirect, .001). Neither previously-cited ID fit; T0051 wasn't cited at all.

Per issue #127's audit.
Poisoned tool-result content injected into orchestration code is LLM Prompt Injection via a tool-result channel (Indirect, .001). Same pattern as 00041/00044.

Per issue #127's audit.
Rich-UI payload injection via non-rendered elements (hidden divs, ARIA labels, SVG metadata) is Indirect Prompt Injection -- the user doesn't see it, the model processes it. T0043 never fit.

Per issue #127's audit.
Poisoned async task result from an external queue/webhook is LLM Prompt Injection via that channel (Indirect, .001). Same pattern as 00041/00042.

Per issue #127's audit.
T0052 (Phishing) was already the correctly-fitting citation here -- the templated T0043+T0048 pair added nothing to a confused-deputy cross-app-access mechanism.

Per issue #127's audit.
Same pattern as 00045 -- T0052 (Phishing) already fits MCP tool-hook hijacking; the templated pair is dropped.

Per issue #127's audit.
Same pattern as 00045/00046 -- T0052 (Phishing) already fits unsafe agent delegation; the templated pair is dropped.

Per issue #127's audit.
T0010 (AI Supply Chain Compromise, already correctly cited) already fits silent tool registration; T0043 (adversarial-input crafting) added nothing.

Per issue #127's audit.
Zero-click markdown-image auto-fetch is a client-side auto-render side channel, not the LLM being induced to ignore instructions -- the weakest T0051 citation in the corpus per issue #127. JUDGMENT call, defaulted to drop.

Per issue #127's audit.
Summarizes the 43-record correction across the preceding commits.
dist/ave-records-latest.json regenerated to match; frozen versioned
snapshot untouched.
@chaksaray
chaksaray merged commit 243b19d into develop Aug 9, 2026
6 checks passed
@chaksaray
chaksaray deleted the fix/mitre-atlas-audit-127 branch August 9, 2026 15:57
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant