fix: owasp_mcp audit corrections (39 records) and framework_sources backfill - #259
Merged
Merged
Conversation
…ework_sources Implements #257, after a deep re-check that found and fixed a real arithmetic error in that issue's own summary counts (23+22+21+8=74, not 80) and 11 additional corrections a fuller read of each record's full description (not just the truncated behavioral_fingerprint excerpt used for the first pass) resolved more precisely. Real final disposition, all 80 records: 39 corrected (applied here), 26 confirmed correct as-is (no change), 15 left unchanged with no confidently-better alternative (8 genuinely ambiguous between two defensible options, 7 where none of the 10 current MCP Top 10 categories covers the record's mechanism at all -- forcing a tag onto either group risked a worse error than leaving the existing one). 39 + 26 + 15 = 80. Two systemic patterns corrected, matching the shape #196's owasp_asi audit found: MCP01 (Token Mismanagement and Secret Exposure) was applied to 7+ records with no credential/token mechanism at all (shell injection, goal hijack, jailbreak, hidden instruction, dynamic tool-call injection, STDIO shell injection, zero-click auto-run); MCP10 (Context Injection and Over-Sharing) was applied to 10+ single-session records with no cross-session, cross-tenant, or persistent-sharing component, several of which are near-exact fits for MCP06 (Intent Flow Subversion) instead -- confirmed against AVE's own corpus, which already cross-references AVE-2026-00002/00016/00020/00028/00041/00043/00044/00065 as sharing one underlying mechanism ("the same missing content-versus-instruction boundary"), just realized at different delivery vectors; that grouping maps cleanly onto MCP03 (the tool's own declared surface, read at discovery) versus MCP06 (retrieved or received content, read in-flow) once checked against each category's real, current text. Corrections found and applied during the deep re-check that #257's first pass had marked ambiguous or missed: 00003 (drop MCP05, no command construction described), 00030/00061/00067 (a false-claim/trust-bypass mechanism is MCP07's own text almost verbatim in each case, not MCP02 or MCP05), 00038/00063 (MCP02 alone, dropping a weak secondary), 00041/00059 (MCP03 alone -- both explicitly happen at tool-discovery time per their own record text, not in retrieved context), 00044/00048/00052 (drop a weak secondary once the dominant real mechanism -- MCP06, MCP07, MCP05 respectively -- is identified), 00050 (MCP02's own "not declared in the manifest" language is a near-verbatim match, replacing a weak MCP04), 00076 (natural-language content steering a decision-making classifier is MCP03's territory, not MCP02's). One record (AVE-2026-00073) was upgraded from ambiguous to confirmed correct after the fuller text surfaced a concrete, cited case (CVE-2026-21852) where the described mechanism directly caused a real API-key leak, grounding the MCP01 tag rather than leaving it a stretch. framework_sources.owasp_mcp backfilled for all 80 records, pinned to OWASP/www-project-mcp-top-10 commit 165fe0f78ef104459237b4a8e0f6e78db9b02391 (2026-07-29, confirmed live via the API before use) and today's date (2026-09-05) -- this audit is exactly the dedicated, per-record verification issue #255 deferred this field pending. Found, not fixed here (separate, pre-existing, out of scope for this PR): crosswalks/ave-to-owasp-mcp.md claims in its own header to be 'generated directly from each record's owasp_mcp field' but has no generator script anywhere in the repo and was already stale before this change -- it covers only 56 of the corpus's 80 records. Flagging in the PR body rather than folding a second, unrelated repair into this one. Validated: python scripts/validate_records.py (80/80 valid), python scripts/check_fixtures.py (all pass), pytest tests/ -x -q (463 passed, unchanged from before this PR since no test asserts specific owasp_mcp values), python scripts/check_framework_sources.py (owasp_mcp no longer appears in any finding; only nist_ai_rmf and the 10 mitre_atlas records already deferred by #255 remain), node scripts/build-records.js (dist/ regenerated, diff is only the intended additions). Every touched record diffed by hand for both the tag corrections and the framework_sources insertions -- minimal, byte-exact changes, no unrelated reformatting.
chaksaray
force-pushed
the
fix/owasp-mcp-audit-corrections
branch
from
September 4, 2026 23:30
e0750a1 to
b3d4abd
Compare
This was referenced Sep 4, 2026
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Implements #257, after a deep re-check per the follow-up request. That
re-check found and fixed a real problem in #257's own summary counts
(23 + 22 + 21 + 8 = 74, not 80 — an arithmetic error in my own audit,
the same kind of self-reported-count mistake this project's other audits
have caught in external sources) and 11 additional corrections a fuller
read of each record's complete
descriptionfield (not just thetruncated
behavioral_fingerprintexcerpt the first pass used) resolvedmore precisely.
Real final disposition, all 80 records
genuinely ambiguous between two defensible options, 7 where none of
the 10 current MCP Top 10 categories covers the record's mechanism at
all. Forcing a tag onto either group risked a worse error than
leaving the existing one; per this project's own standing rule, a
weak forced fit is worse than an honest gap.
39 + 26 + 15 = 80.
Two systemic patterns, same shape as #196's owasp_asi audit
records with no credential/token mechanism at all — shell injection,
goal hijack, jailbreak, hidden instruction, dynamic tool-call
injection, STDIO shell injection, zero-click auto-run.
single-session records with no cross-session, cross-tenant, or
persistent-sharing component. Several are near-exact fits for MCP06
(Intent Flow Subversion) instead — confirmed against AVE's own corpus,
which already cross-references AVE-2026-00002/00016/00020/00028/
00041/00043/00044/00065 as sharing one underlying mechanism ("the same
missing content-versus-instruction boundary"), just realized at
different delivery vectors. That grouping maps cleanly onto MCP03 (the
tool's own declared surface, read at discovery) versus MCP06 (retrieved
or received content, read in-flow) once checked against each
category's real, current text.
Corrections the deep re-check surfaced beyond #257's first pass
00003 (drop MCP05, no command construction described), 00030/00061/00067
(a false-claim/trust-bypass mechanism is MCP07's own text almost
verbatim in each case, not MCP02 or MCP05), 00038/00063 (MCP02 alone,
dropping a weak secondary), 00041/00059 (MCP03 alone — both explicitly
happen at tool-discovery time per their own record text, not in
retrieved context), 00044/00048/00052 (drop a weak secondary once the
dominant real mechanism is identified), 00050 (MCP02's own "not declared
in the manifest" language is a near-verbatim match, replacing a weak
MCP04), 00076 (natural-language content steering a decision-making
classifier is MCP03's territory, not MCP02's).
One record, AVE-2026-00073, was upgraded from ambiguous to confirmed
correct after the fuller text surfaced a concrete, cited case
(CVE-2026-21852) directly grounding the MCP01 tag rather than leaving it
a stretch.
framework_sources.owasp_mcp backfill, all 80 records
Pinned to
OWASP/www-project-mcp-top-10commit165fe0f78ef104459237b4a8e0f6e78db9b02391(2026-07-29, confirmed live via the API before use) and today's date —
this audit is exactly the dedicated, per-record verification issue #255
deferred this field pending.
Found, not fixed here
crosswalks/ave-to-owasp-mcp.mdclaims in its own header to be"generated directly from each record's owasp_mcp field" but has no
generator script anywhere in the repo, and was already stale before this
change — it covers only 56 of the corpus's 80 records. Flagging here
rather than folding a second, unrelated repair into this PR; worth its
own follow-up issue.
Validated
python scripts/validate_records.py: 80/80 validpython scripts/check_fixtures.py: all passpytest tests/ -x -q: 463 passed (unchanged — no test asserts specificowasp_mcpvalues)python scripts/check_framework_sources.py:owasp_mcpno longer appears in any finding; onlynist_ai_rmfand the 10mitre_atlasrecords already deferred by framework_sources backfill, and a real gap found: pin_status has no 'unknown' state #255 remainnode scripts/build-records.js:dist/regenerated, diff is only the intended additionsframework_sourcesinsertions — minimal, byte-exact changes, no unrelated reformatting