Skip to content

Withdraw M4's null as evidence about the guard (#122) - #123

Merged
MongLong0214 merged 1 commit into
devfrom
bug-issue-122-verdict
Jul 28, 2026
Merged

Withdraw M4's null as evidence about the guard (#122)#123
MongLong0214 merged 1 commit into
devfrom
bug-issue-122-verdict

Conversation

@MongLong0214

Copy link
Copy Markdown
Owner

No run in either arm received injected context. Both arms are defined by receiving records; neither did.

Established by restoring the pinned harness 081d858c and probing it — the saved transcript contains no occurrence of commitlore, Ruled-out, Limit: or active records anywhere in its text. A recording gap would leave the context in the transcript and empty only the field.

transcripts-final:  30/60  have injected_context
transcripts-m4:      0/112

What changes

The verdict's headline. Its null stops being an answer about the guard and becomes an observation about data that measured no treatment.

What does not

Provenance is clean — one harness commit, one dist digest, 112 rows, no mid-run rebuild. The failure is upstream of provenance: the harness recorded faithfully what it did, and what it did was run both arms without records.

The corrected analysis stays. McNemar p=0.1094, ICC 0.581, DEFF 8.56, effective n≈13, the four saturated tasks — still true about the data, irrelevant as evidence about the guard. The verdict says both.

M4 is not retracted and not called invalid. Its data is valid; what it measured was not the treatment. Recorded in Ruled-out:.

The harness fix is a separate branch (bug-issue-122) so the correction to the record and the correction to the code are reviewed apart.

Tests 1385 / 39 files. check-readme-numbers.mjs exit 0.

No run in either arm received injected context. Both arms are defined by
receiving records; neither did. The comparison was nothing against nothing, and
its null is not a weak result about the guard — it is not a result about the
guard.

Established by restoring the pinned harness 081d858 and probing it: the saved
transcript carries no occurrence of commitlore, Ruled-out, Limit: or active
records anywhere in its text. A recording gap would leave the context in the
transcript and empty only the field. It was never delivered.

The corrected statistical analysis stays. McNemar p=0.1094, ICC 0.581,
DEFF 8.56, effective n 13, four saturated tasks — all still true about the data
and all irrelevant as evidence about the guard. Both things need saying and the
verdict now says both.

Provenance is unchanged and still clean: one harness commit, one dist digest,
112 rows, no mid-run rebuild. The failure sits upstream of provenance. The
harness recorded faithfully what it did, and what it did was run both arms
without records.

Ruled-out: retracting the dataset or calling M4 invalid | the data is valid and its provenance is clean; what it measured was not the treatment, and those are different words
Limit: the guard question is now unanswered rather than answered null
Blast: system
Undo: costly
Certainty: firm
Record-Id: r-m4withdraw
@github-actions

Copy link
Copy Markdown

CommitLore — record lint

Trailers: clean — 1 commit in origin/dev..e5f9b73d55705d329daf4f750c5be0cd6273bbd6
Active constraints: 25 limits · 57 ruled-out · 30 warnings — from 47 records over 7 changed paths

Active constraints for the paths this PR touches

Limits (25)

  • r-m4withdraw e5f9b73 — the guard question is now unanswered rather than answered null
  • r-protoreg31 1bd02c1 — the calibration arm cannot be used to select confirmatory tasks, only to size them
  • r-readmeux1 b664205 — interactive record building does not exist, so the honest answer is still "an agent writes it or you do"
  • r-expreadme1 9e69abe — bench/VERDICT-M4.md still cites the Fisher figure; the two disagree until the verdict records why the number was withdrawn from the README
  • r-expomerge1 d6ad014 — M4's existing rows have no exposure field and must read as unknown, not as not-exposed — backfilling by inference would erase the finding
  • r-exposure52 ba69411 — legacy JSONL artifacts predate model and guard-exposure fields | their absence remains unknown and is never inferred or backfilled
  • r-rdme96a 9c9371c — scripts/check-readme-numbers.mjs's withdrawal-notice and stray-statistic checks constrain what can appear outside the (absent, here) generated benchmark block — re-checked after every edit, not just at the end
  • r-relinstall c6e1d04 — never tested against the real GitHub release infrastructure (no release exists yet — that is the owner's action) — verified against a locally built SEA binary, a hand-made SHA256SUMS, and a local HTTP server standing in for GitHub's release-asset redirects, which is everything this repository lets a change verify before a tag exists.
  • r-seabin39 9e9cd0e — doctor's PreToolUse hook runtime check still shells to scripts/commitlore-run.sh via bash for its own probe; a binary install with the Claude Code plugin hook already wired reports a plain ENOENT-style fail there rather than trying the binary directly -- not one of B-09 · Single static binary — remove the Node runtime dependency #39's six required commands, not fixed here
  • r-fix70a1 d707fc7 — one encoding layer and explicit lexical forms in the four published languages; semantic paraphrases, nested encodings, and split payloads remain outside coverage
  • r-shallow66 60a8659 — a depth-1 clone can only inspect its reachable commit history
  • r-7a3e91 cf859e4 — better-sqlite3 stays external because it is native — the bundle degrades to --no-index without it, which only works because r-6f2a08 made that load lazy first
  • r-9c07e2 9c4d25a — the plugin still needs Node for the CLI — the protocol does not, but guard, the index and the MCP server do (T-706 · Bundle the CLI as a single file — run from a clone alone #38)
  • r-3e8a41 2f0a8a0 — scoping is not implemented, so off-path records reach every arm
  • r-9c2f74 d653153 — the ablation arms cannot discriminate on these fixtures -- no-grade and no-lifecycle are byte-identical to the treatment in 9 of 10 tasks, because the seeds carry one reconstructed record and one task with a lifecycle trailer between them
  • r-9c2f74 d653153 — the harness assembles its own projection rather than calling the shipped injector, so what is measured is the harness's rendering of the records, not src/core/inject.ts (issue B-08 · Replace the benchmark harness injector with the actual src/core/inject.ts #36)
  • r-6e1a72 5e09846npx commitlore is the first thing a reader will try, and it fails until the package is published
  • r-3f7a29 49817dc — reconstruction reads text written before the protocol existed, so the evidence is thinner than a harvest and the discard rate is expected to be high
  • r-5c8b31 60ddc39 — the agent CLI exposes no in-flight turn limit, so a per-task turn budget can only ever be observed with this driver
  • r-0b7c44 d2b2ce3 — a command is only real once --help names it, because that is where users look before they read source
  • r-6c2b95 5da793c — the CLI holds no API key, so any real driver runs through the user's own agent session and cannot be exercised in this environment
  • r-7f0e39 76f3f2d — literal substitution only catches the exact strings you list, so the same term written with a different separator survives
  • r-9d31b7 4ac6e30 — the example lives in four translated files, so any fix that is not mechanically enforced will drift again on the next edit
  • r-c0f4e2 3d249cd — npm gitlore is held by an active same-domain CLI, so the owner's first-choice name was not available
  • r-a8f3c1 ef48843 — Rename must land before any code exists -- after 27 tickets it would touch spec, fixtures, index, hooks and every doc

Ruled out (57)

  • r-m4withdraw e5f9b73 — retracting the dataset or calling M4 invalid | the data is valid and its provenance is clean; what it measured was not the treatment, and those are different words
  • r-protoreg31 1bd02c1 — keeping GEE with a small-sample correction | permutation is exact rather than corrected, and the correction literature stops short of an ICC this high
  • r-protoreg31 1bd02c1 — a per-task qualification threshold read off outcomes | any such threshold reintroduces the selection that saturated M4
  • r-protoreg31 1bd02c1 — leaving the duplicate id and exempting it in the test | the check is right and the history was wrong
  • r-readmeux1 b664205 — describing harvest as automatic record creation | it drafts from a transcript and a human still commits, and claiming otherwise is the overreach this project keeps closing issues about
  • r-exposure52 ba69411 — compute a treatment effect with unknown guard exposure | an old row that never recorded whether the guard reached the run cannot distinguish no treatment from an ignored treatment
  • r-m4pair1 490f8f6 — reporting the raw "8.5x too narrow" figure a collaborator first proposed | that number applies the design effect to standard error instead of variance; verifying it independently (sqrt(8.556) = 2.92, not 8.556) caught the conflation before it reached a durable document, and the collaborator confirmed the correction after checking it themselves
  • r-m4docs1 32d5bf1 — keeping the withdrawal notice and only landing the verdict document | bench/report.ts already draws this line -- a provenanced dataset that still shows a withdrawal is a hard failure in check-readme-numbers.mjs (checked here), not a style choice left open
  • r-m4vrd01 21d8bf3 — re-cutting the result by task after seeing it | PREREGISTRATION.md §4 forbids subset analysis regardless of how the result came out, and applies with more force to the run built to give the hypothesis a fair matrix
  • r-rel0200a a074754 — bumping ci.yml's "v0.1.0 was published with zero attached assets" comments | those describe a historical fact about the actual v0.1.0 release, not a version this project declares; the check they document (releases/latest/download/SHA256SUMS returning 200) is written to start exercising the real path automatically the day any release ships assets, v0.2.0 included, with no workflow edit
  • r-rel0200a a074754 — touching docs/adr/ADR-0001-scope-v010.md, docs/tickets/release.md, bench/VERDICT-M1.md, HANDOFF.md, bench/ROUTE-GAP.md | planning and historical-record prose that names v0.1.0 as a past decision or measurement subject, not a live version carrier
  • r-rel0200a a074754 — changing test/mcp.test.ts's CommitLore-Version: 0.1.0 fixture trailer | that's protocol-version content inside a synthetic seed commit (what an old commit's trailer looked like), unrelated to and never asserted against package.json's version
  • r-rdme96a 9c9371c — dropping the git-clone / source-build paths from the top entirely | install.sh is the fast path, not a universal one (no Windows binary yet, per ADR-0015) — the detailed section has to stay reachable, just not first
  • r-relinstall c6e1d04 — guessing the current version to build the asset URL directly | would need either the GitHub API (rate-limited, needs no-auth headers handled correctly) or trusting a redirect's final Location header parsing. Downloading the fixed-URL SHA256SUMS first and reading the real asset name back out of it needs neither and is what the checksum step has to fetch anyway.
  • r-relinstall c6e1d04local for scoping — not POSIX per se, but supported by dash, bash, and every shell this script is realistically piped into (verified directly, see Verified) | not used in the end; the script has few enough variables that scoping was not needed, only noted here because it was considered.
  • r-seabin39 9e9cd0e — mainFormat: "module" (an ESM SEA main) | verified to fail both blob generation and runtime on this Node line, not merely documented as unsupported (see above)
  • r-seabin39 9e9cd0e — pkg / nexe | third-party bundlers embedding a separate, forked Node runtime this project does not control the patch cadence of; pkg is archived upstream. Trades the Node runtime dependency this ticket removes for a different, less-maintained one
  • r-seabin39 9e9cd0e — Deno compile / Bun compile | a different runtime. node:sqlite, the TypeScript, and NodeNext resolution are all Node-specific; retargeting them is a second runtime port, not a build step, and issue B-09 · Single static binary — remove the Node runtime dependency #39's own first option ("Node SEA -- no source rewrite") needs none
  • r-seabin39 9e9cd0e — reimplement in Go/Rust | issue B-09 · Single static binary — remove the Node runtime dependency #39's own second option, and a real one via spec/fixtures + spec/contract-cases, but an order of magnitude more work than this ticket and not needed to solve either problem (latency, no-Node-on-PATH) this ticket opens with
  • r-seabin39 9e9cd0e — committing dist/commitlore next to dist/commitlore.mjs | breaks ADR-0011's committed-dist/-matches-src/ invariant at ~115 MiB per platform/arch, and a pushed blob that size is not removable from git history again
  • r-seabin39 9e9cd0e — Windows (commitlore.exe) in this PR | Node's docs describe a signtool path this repository has no CI runner to verify; shipping an unverified platform claim is what this project's numbers-or-silence discipline exists to refuse. classifyBinTarget and the resolution order are written so it is a small additive follow-up, not a redesign
  • r-fix70a1 d707fc7 — exhaustive per-language phrase enumeration | unbounded phrase lists cannot provide semantic coverage, so this fix documents a bounded lexical policy and independent corpus
  • r-3b57e2 30f2d5f — converting the three translated READMEs for consistency | they are the product, not the record, and two checks exist specifically to keep them
  • r-5a83e9 1c683df — requiring a positive benchmark for release | a tool that answers honestly ships whether or not the effect turns out to be large, and pretending otherwise is what the withdrawal of the published numbers was about
  • r-6f92c4 3cebb89 — restating the withdrawn numbers as prose ("we measured a reduction") | it is the same claim with the evidence removed
  • r-6f92c4 3cebb89 — leading with the protocol's features and putting the measurement record near the bottom | that is the arrangement of someone hoping it is not read
  • r-2b58d4 4842356 — exempting datasets written before the fields existed | it is one line and it deletes the guarantee
  • r-6b83f2 4c1a503 — deleting both sentences | the clone-runs-without-installing claim is true for validate, context, guard and the MCP server, and dropping it would understate what a clone gives you as badly as the old text overstated it
  • r-3d92a8 f85101a — keeping the searches first and fixing the shim | the shim belongs to npm, not to us, and the version-skew problem survives the fix
  • r-3d92a8 f85101a — a config-only hook check | it was written, it reported ok, and the hook failed on the next commit
  • r-7a3e91 cf859e4 — inlining spec/SPEC.md and the schema into the bundle | SPEC.md would need a codegen step that itself needs a drift guard, and the package-root walk removes the reason to want it
  • r-7a3e91 cf859e4 — replacing the tsc output with the bundle | test/cli.test.ts, test/hooks.test.ts and test/mcp.test.ts import dist internals by path
  • r-c53d19 110be8c — leaving the claim and letting T-706 · Bundle the CLI as a single file — run from a clone alone #38 make it true later | the README is what someone reads while deciding to adopt this, and a claim that is false today does not become honest because it is scheduled
  • r-9c07e2 9c4d25a — invoking npx on every Edit | it puts a registry round trip on the hot path of every tool call, which is how a hook earns being uninstalled
  • r-9c07e2 9c4d25a — committing dist/ so the plugin is self-contained from a git clone | it puts build output in review diffs forever to save one background install
  • r-0d4b81 8005227 — a longer quickstart that demonstrates context, limits, ruled-out, warnings and stale | an agent calls those itself once the MCP server is registered, so listing them teaches the human a workflow that is not theirs
  • r-7f31c9 750ab17 — reporting both datasets from one source list | readSources groups by condition and cannot separate repositories, so any second dataset with a commitlore-on arm silently corrupts the headline test
  • r-3e8a41 2f0a8a0 — off-path records that advocate the ruled-out option | with no scoping they reach all three arms, raising re-proposal everywhere and compressing the grading and lifecycle contrasts the set is built to isolate
  • r-9c2f74 d653153 — resume the pilot into the same file | a new process would load the edited code and create the mixing that had not happened
  • r-9c2f74 d653153 — run the ablation arms as they stand | three nulls from comparing identical inputs read as "these guarantees do not matter"
  • r-9c2f74 d653153 — keep the tasks that showed an effect and rewrite only the rest | the property is the criterion, not the direction of the result
  • r-6e1a72 5e09846 — leave the banner until release | it understates for weeks and readers leave rather than build from source
  • r-6e1a72 5e09846 — update English only and translate later | the lag is itself a wrong answer for whoever reads the other three
  • r-3f7a29 49817dc — repair a draft that fails verification | backfill's source material is weak enough that a repair loop would mostly be inventing
  • r-3f7a29 49817dc — write reconstructed records into commit messages | history rewriting is irreversible and reaches every existing clone
  • r-3f7a29 49817dc — post a fresh comment per push | it turns the signal into noise and the check gets muted
  • r-5c8b31 60ddc39 — keep turns and explain it in prose | the JSONL outlives the prose, and whoever reads the rows later will not have it
  • r-5c8b31 60ddc39 — drop the turn budget since it cannot be enforced | the overrun is still the signal that a run went off the rails
  • r-6c2b95 5da793c — skip the dry-run driver | then nothing exercises the harness until a key exists, and the first real run debugs the harness instead of measuring anything
  • r-6c2b95 5da793c — emit dry-run rows without a marker | indistinguishable from measurements the moment they leave the terminal
  • r-9d31b7 4ac6e30 — fix the values and move on | the same drift already happened once through a rename, and prose review did not catch it either time
  • r-9d31b7 4ac6e30 — parse the README at runtime in the CLI | the check belongs in the conformance suite, not in shipped code
  • r-c0f4e2 3d249cd — GitLore published as git-lore | the binary and search results still collide with the existing gitlore tool
  • r-c0f4e2 3d249cd — keep Annals | the sound problem does not decay, and with code near zero this is the cheapest moment the project will ever have
  • r-c0f4e2 3d249cd — rename code and spec first, documents later | the drift window makes every artifact written in it wrong
  • r-a8f3c1 ef48843 — keep name, change vocabulary only | vocabulary is the protocol, so half the change leaves the substance untouched
  • r-a8f3c1 ef48843 — drop Certainty as a dead field | a real route exists -- stale sweep prioritizes guess-level records for review

Warnings (30)

  • r-exposure52 ba69411 (claim) — guard exposure is instrumented in the benchmark hook adapter, not by changing guard scoring | M4 rows remain unexposed and metrics refuses their effect estimate
  • r-rel0200a a074754 (claim) — scripts/commitlore-bootstrap.sh is orphaned -- no hooks.json entry invokes it, and its npm-install strategy contradicts ADR-0011. It still carries a live version default, now bumped for consistency, but nothing exercises it. Worth a follow-up issue: either wire it up correctly or delete it.
  • r-seabin39 9e9cd0e (claim) — node:sea is "Active development" per Node's own docs; its schema or CommonJS-only constraint could change between Node versions. core/paths.ts's readInstalledFile/isSea split and build-binary.mjs's asset map are the one place that assumption is absorbed, same posture as ADR-0012 already committed to for node:sqlite
  • r-fix70a1 d707fc7 (claim) — add malicious and benign fixtures together when extending scanner patterns; false positives can make the defence unusable
  • r-shallow66 60a8659 (claim) — shallow history remains advisory; query and guard exit-code semantics are unchanged
  • r-3b57e2 30f2d5f (claim)bench/PREREGISTRATION.md is append-only and was translated in place. Its section numbering and order are unchanged, but a translation is still an edit to a file whose whole discipline is that it is not edited. Recorded here rather than left to be noticed
  • r-5a83e9 1c683df (claim) — the claim that gitseed's rejected alternatives will produce a higher control base rate is a prediction, not a measurement. The qualification round exists to test it, and it may say no
  • r-6f92c4 3cebb89 (claim) — with no numbers, the first-impression case now rests entirely on the test links. If M3-b also comes back null, this framing is what the project has
  • r-2b58d4 4842356 (claim) — this leaves the README with no measured numbers at all until M3-b runs. That is the honest state and it is also a worse first impression. The alternative was publishing numbers produced by a binary nobody recorded
  • r-1a63f5 2bb4993 (claim) — "CI is green" was said five times today against a red CI, including in the commit that introduced the rule saying to check CI before saying it. The rule is in docs/RELEASE-GATE.md §5 and it was not followed by its own author. This commit is not claiming CI is green; that claim comes after the run reports
  • r-6b83f2 4c1a503 (claim) — the second claim will become true when ADR-0012 lands and false again if the notes refspec story changes. A README sentence about distribution needs a test, and there is none — scripts/check-readme-numbers.mjs checks numbers
  • r-3d92a8 f85101a (claim)hook-runtime executes the hook on every doctor run. The probe message is valid so nothing is written, but it is no longer a read-only command
  • r-3d92a8 f85101a (claim) — the check pins PATH to /usr/bin:/bin, which assumes git is there. On a system where it is not, this reports a hook failure that is really a probe failure
  • r-7a3e91 cf859e4 (claim) — hardcoding ../ counts back to the package root is what broke this — new code reads assets through installedPath(), never through import.meta.url
  • r-c53d19 110be8c (claim) — r-6f2a08's message says the clone gap closes with B-09 · Single static binary — remove the Node runtime dependency #39; it closes with T-706 · Bundle the CLI as a single file — run from a clone alone #38. The commit message is history and stays as written
  • r-9c07e2 9c4d25a (claim) — hooks here must exit 0 on every path — a non-zero exit from a PreToolUse hook blocks the edit, and no record is worth that
  • r-0d4b81 8005227 (claim)claude mcp add commitlore -- commitlore mcp is Claude Code's syntax — other MCP clients register a stdio server their own way
  • r-7f31c9 750ab17 (claim) — adding a file to README_SOURCES pools it into every aggregate in the block, including the significance test — check the arm names first
  • r-8b41e6 0bbad66 (claim) — the payload share and the fixture share are not interchangeable — the payload adds framing and drops what grading and lifecycle withhold, so a reader who swaps one for the other will be wrong by a small, plausible margin
  • r-3e8a41 2f0a8a0 (claim) — the ablation set is a different synthetic repository from bench/tasks, so its commitlore-on arm must never be pooled with the primary matrix's
  • r-9c2f74 d653153 (claim) — after the measurement, check that git status is clean and the recorded sha is still HEAD -- an edit mid-run breaks reproducibility silently, and that check is the only thing that catches it
  • r-6e1a72 5e09846 (claim) — the moment v0.1.0 is published this banner is wrong again -- it names npx not working, which release makes false
  • r-3f7a29 49817dc (claim) — every backfilled record is Provenance: reconstructed, which the trust model always renders as a claim -- do not add a path that lets a draft override that field
  • r-5c8b31 60ddc39 (claim) — a driver that gains a real turn limit should stop emitting over-turns and start emitting an enforced label -- do not reuse over-turns for something the harness actually stopped
  • r-6c2b95 5da793c (claim) — never cite a row with simulated:true -- README numbers come from bench/results logs and those rows are not results
  • r-7f0e39 76f3f2d (claim) — when retiring a command name, grep the bare word too, not just the prefixed form -- the prefix is what the substitution keyed on
  • r-9d31b7 4ac6e30 (claim) — the four READMEs must keep the example block byte-identical -- translating the code block will fail spec/verify.sh
  • r-c0f4e2 3d249cd (claim) — ADR-0008 and ADR-0009 keep the literal string Annals on purpose -- mechanical substitution there destroys the decision trail
  • r-c0f4e2 3d249cd (claim) — the residual grep for lore_query reports a false positive because commitlore_query contains it as a substring, so check the prefix
  • r-a8f3c1 ef48843 (claim) — docs/adr/ADR-0008 is the canonical vocabulary -- do not reintroduce old terms from memory

git log --follow accepts exactly one pathspec, so renames are not followed for 7 paths; query one path at a time to follow its rename chain

withheld the content of 2 record(s) graded blocked: a Verified trailer matching an injection pattern is reported, never quoted (SPEC §7)

Trailer violations fail this check. Active constraints are informational — they are what the repository already decided, not a verdict on this PR.

@MongLong0214
MongLong0214 merged commit 4fc4836 into dev Jul 28, 2026
8 checks passed
MongLong0214 added a commit that referenced this pull request Jul 28, 2026
Co-authored-by was ingested as a record and classified [claim], so on a
repository that uses it routinely — which is any repository with
AI-assisted commits — a path's whole projection filled with attribution
lines and crowded out the decision context the command exists to deliver.
The reported repository had 175 trailers over 173 commits and returned
three co-author lines for Package.swift and nothing else.

CONVENTIONAL_TRAILER_KEYS names the set in one place: Co-authored-by,
Signed-off-by, Reviewed-by, Acked-by, Tested-by, Reported-by, Suggested-by,
Cc and Change-Id. Matching is on the lowercased key, because all three of
Co-authored-by, Co-Authored-By and Co-authored-By occur in the same real
history and a case-sensitive list would have fixed a third of the problem.

Fixes and Closes are deliberately not in the set. They name the issue a
change addresses, which is closer to decision context than to attribution —
an agent reading "Fixes #123" learns something a co-author's name never
tells it — so they still land in "other" rather than being discarded with
pure attribution.

The exclusion is counted rather than silent, so a user who wonders where an
attribution line went can find out instead of concluding the trailer was
never parsed.

Verified end to end rather than by unit: a scratch repository whose only
trailer is Co-authored-by now reports no active records, and one carrying
Co-Authored-By, Signed-off-by and a real Limit reports the Limit alone.

Ruled-out: excluding Fixes and Closes with the rest | they carry decision context an agent can use, and discarding them would trade a projection full of attribution for one missing the issue a change answers
Limit: the denylist answers a different question from isRecordKey's allowlist, so a conventional trailer this protocol later claims would need removing from one and adding to the other
Blast: system
Undo: easy
Certainty: firm
Record-Id: r-convtrail150
MongLong0214 added a commit that referenced this pull request Jul 28, 2026
Co-authored-by was ingested as a record and classified [claim], so on a
repository that uses it routinely — which is any repository with
AI-assisted commits — a path's whole projection filled with attribution
lines and crowded out the decision context the command exists to deliver.
The reported repository had 175 trailers over 173 commits and returned
three co-author lines for Package.swift and nothing else.

CONVENTIONAL_TRAILER_KEYS names the set in one place: Co-authored-by,
Signed-off-by, Reviewed-by, Acked-by, Tested-by, Reported-by, Suggested-by,
Cc and Change-Id. Matching is on the lowercased key, because all three of
Co-authored-by, Co-Authored-By and Co-authored-By occur in the same real
history and a case-sensitive list would have fixed a third of the problem.

Fixes and Closes are deliberately not in the set. They name the issue a
change addresses, which is closer to decision context than to attribution —
an agent reading "Fixes #123" learns something a co-author's name never
tells it — so they still land in "other" rather than being discarded with
pure attribution.

The exclusion is counted rather than silent, so a user who wonders where an
attribution line went can find out instead of concluding the trailer was
never parsed.

Verified end to end rather than by unit: a scratch repository whose only
trailer is Co-authored-by now reports no active records, and one carrying
Co-Authored-By, Signed-off-by and a real Limit reports the Limit alone.

Ruled-out: excluding Fixes and Closes with the rest | they carry decision context an agent can use, and discarding them would trade a projection full of attribution for one missing the issue a change answers
Limit: the denylist answers a different question from isRecordKey's allowlist, so a conventional trailer this protocol later claims would need removing from one and adding to the other
Blast: system
Undo: easy
Certainty: firm
Record-Id: r-convtrail150
MongLong0214 added a commit that referenced this pull request Jul 28, 2026
Co-authored-by was ingested as a record and classified [claim], so on a
repository that uses it routinely — which is any repository with
AI-assisted commits — a path's whole projection filled with attribution
lines and crowded out the decision context the command exists to deliver.
The reported repository had 175 trailers over 173 commits and returned
three co-author lines for Package.swift and nothing else.

CONVENTIONAL_TRAILER_KEYS names the set in one place: Co-authored-by,
Signed-off-by, Reviewed-by, Acked-by, Tested-by, Reported-by, Suggested-by,
Cc and Change-Id. Matching is on the lowercased key, because all three of
Co-authored-by, Co-Authored-By and Co-authored-By occur in the same real
history and a case-sensitive list would have fixed a third of the problem.

Fixes and Closes are deliberately not in the set. They name the issue a
change addresses, which is closer to decision context than to attribution —
an agent reading "Fixes #123" learns something a co-author's name never
tells it — so they still land in "other" rather than being discarded with
pure attribution.

The exclusion is counted rather than silent, so a user who wonders where an
attribution line went can find out instead of concluding the trailer was
never parsed.

Verified end to end rather than by unit: a scratch repository whose only
trailer is Co-authored-by now reports no active records, and one carrying
Co-Authored-By, Signed-off-by and a real Limit reports the Limit alone.

Ruled-out: excluding Fixes and Closes with the rest | they carry decision context an agent can use, and discarding them would trade a projection full of attribution for one missing the issue a change answers
Limit: the denylist answers a different question from isRecordKey's allowlist, so a conventional trailer this protocol later claims would need removing from one and adding to the other
Blast: system
Undo: easy
Certainty: firm
Record-Id: r-convtrail150
MongLong0214 added a commit that referenced this pull request Aug 18, 2026
A blind refutation round on this branch broke three of the four claims I put
to it, and verifying one of them found a defect it had not been looking for.

`gh pr create` renders "GitHub closes #123 as merged" into the canonical pull
request's body. GitHub binds a closing keyword to the number straight after it,
and a pull request closed by keyword is recorded closed with `mergedAt` null --
the opposite of the sentence containing it, and the opposite of what T-1502
accepts. Measured on #752 six hours ago: an integration body said "GitHub
closes #752, #755, #756 ... as merged", the keyword bound to #752 alone, and
that one was recorded closed while the five with no keyword were recorded
merged. There is no API to convert it afterwards. This workflow would have
reproduced it on every run, and no test read the body.

Two ticket statements were also wrong against the file. "The job never checks
out or executes a pull request's head" was borrowed from the rule #723 fixed
for `preserve`, which only reads a pull request; this one rebuilds it, and
rebuilding somebody's change means running it. Unsatisfiable as written, so it
would have been dropped rather than met -- what the job split actually holds is
that the runner executing that code has no credential. And the negative control
the ticket named, skipping `artifact:manifest`, cannot be performed from a pull
request: the step is hard-coded in a workflow loaded from the default branch and
the source-only filter refuses workflow edits. A negative control nobody can run
is the defect it was written to prevent, so it is replaced with one that can be:
edit `dist/` on the pushed canonical branch and watch `ci.yml` go red.

Limit: the canonical pull request asks for a merge commit and cannot enforce one -- squash and rebase are both enabled and the button remembers the last method used, which is how #760 closed five of six as merged
Blast: module
Undo: easy
Certainty: firm
Record-Id: r-t1502body
Provenance: authored
Verified: restored the keyword and watched the new test fail naming `closes #123`, then restored the fix and saw 21 tests pass across both workflow test files
CommitLore-Version: 2.0.0
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant