Skip to content

fix: benchmark-2026-06.md citation audit (closes #241) - #244

Merged
chaksaray merged 1 commit into
developfrom
fix/benchmark-citation-audit
Sep 2, 2026
Merged

chaksaray merged 1 commit into
developfrom
fix/benchmark-citation-audit

Conversation

@chaksaray

Copy link
Copy Markdown
Contributor

Closes #241. Full citation audit prompted by two confirmed-wrong citations found during the resource-exhaustion roadmap check (MCPSecBench class 13, MCP-SafetyBench class 6 — neither exists). Two wrong out of an unplanned check was treated as a sample, not the finding, per the same reasoning that made the owasp_asi audit (#196) worth running after a smaller sample turned up problems.

Real numbers: 102 claims checked, 30 confirmed correct, 72 confirmed wrong, 0 unverifiable.

All six cited sources are real papers/studies that exist — MCPSecBench (arXiv:2508.13220), the Formal Security Framework for MCP / MCPSHIELD (arXiv:2604.05969), Hou et al. (arXiv:2503.23278), MCP-SafetyBench (arXiv:2512.15163), MCPTox (arXiv:2508.14925), and the OpenClaw/ClawHub concordance study (arXiv:2606.01494). The fabrication is in the specific class names and coverage numbers this document attributed to them:

Dataset Real classes matched
MCPSecBench 4 of 17
FSF-MCP / MCPSHIELD 5 of 23
Hou et al. 2025 5 of 16 (also: its cited paper title was fabricated)
MCP-SafetyBench 5 of 20 (also: dated a year wrong, 2025 vs. real 2026)
MCPTox 0 of 11 — the entire section misdescribed the paper's actual subject (real paper: tool-poisoning detection against 45 live MCP servers; claimed here: content-toxicity/safety-alignment categories, which the paper has nothing to do with)
OpenClaw study Headline numbers (10.4% pairwise overlap, 0.69% all-three agreement) confirmed correct; one detail (which three systems were compared) was wrong and fixed

Every per-dataset coverage table, the consolidated gap analysis, and the overall coverage summary rested on these fabricated per-class tables and are marked retracted rather than re-asserted — recomputing them honestly against the real taxonomies (included in full in this PR) is a fresh mechanism-level research task, out of scope for a citation audit. The one operationally load-bearing conclusion (the resource-exhaustion new-record recommendation) is removed, consistent with #241.

Every retracted claim is kept visible with a [RETRACTED, ORIGINAL] or [AUDITED ...] marker rather than deleted, and a dated audit note sits at the top of the file — same publish-negative-results standard as every other correction this project has made.

Validated: pytest tests/ (435 passed), scripts/validate_records.py (80/80 valid, unrelated to this change but run for hygiene).

Issue #241 found two wrong citations (MCPSecBench class 13, MCP-SafetyBench
class 6, neither exists). Two wrong was treated as a sample, not the
finding, so every checkable claim in the file was audited against its
real primary source rather than patching those two lines.

Result: 102 claims checked (87 per-class, 15 dataset-level), 30 confirmed
correct, 72 confirmed wrong, 0 unverifiable. All six cited sources are real,
existing papers/studies -- the fabrication is in the specific class names
and coverage claims attributed to them, not in whether the sources exist.

Per dataset: MCPSecBench 4/17 real, FSF-MCP/MCPSHIELD 5/23 real, Hou et al.
2025 5/16 real (also: its cited title was fabricated -- corrected), MCP-
SafetyBench 5/20 real (also: dated a year wrong -- corrected), MCPTox 0/11 --
the entire dataset section misdescribed the paper's actual subject (it is a
tool-poisoning benchmark, not a content-toxicity one). The OpenClaw/ClawHub
concordance numbers (10.4% pairwise, 0.69% all-three) are confirmed correct;
one detail (which three systems were compared) was wrong and is fixed.

Every per-dataset coverage table, the consolidated gap analysis, and the
overall coverage summary rested on the fabricated per-class tables and are
marked retracted rather than re-asserted -- recomputing them properly is a
fresh mechanism-level research task, not something to improvise inside a
citation audit. The one derived recommendation that mattered operationally
(the resource-exhaustion new-record candidate) is removed, consistent with
#241's disposition. Real, verified taxonomies for all five numbered datasets
are included in place of the fabricated tables, for whoever picks up that
follow-on research task.

Kept every retracted claim visible with a [RETRACTED, ORIGINAL] or
[AUDITED ...] marker rather than deleting it, and added a dated audit note
at the top of the file recording what was found -- the same
publish-negative-results standard already applied to every other error this
project has fixed.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

benchmark-2026-06.md's resource-exhaustion gap claim does not hold up against its own cited primary sources

1 participant