Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions README.ja.md
Original file line number Diff line number Diff line change
Expand Up @@ -162,7 +162,7 @@ text search ではなく Git trailer parser を使います。本文の `Key:`

## Evidence: より狭い製品上の主張

112 回の実験を行いましたが、測定した agent behavior の主張は裏付けられませんでした。そのため CommitLore は上記の、より狭い製品上の主張をします。全ての限界は [M4 verdict](bench/VERDICT-M4.md) で読めます
112 回の実験は記録されましたが、M4 はどちらの arm にも record を届けませんでした。したがって agent behavior の主張を検証も支持も反証もしていません。上記のより狭い製品上の主張は独立して検証可能な動作に基づきます。クリーンなデータセットと撤回については [M4 verdict](bench/VERDICT-M4.md) を読んでください

<details>
<summary>完全な benchmark record(112 回の実験)</summary>
Expand Down Expand Up @@ -224,7 +224,7 @@ node ~/.commitlore/dist/commitlore.mjs context src/auth
- Windows は未対応です: [#95](https://github.com/MongLong0214/commitlore/issues/95)。
- Alpine および他の musl Linux host は未対応です: [#99](https://github.com/MongLong0214/commitlore/issues/99)。
- cryptographic author verification、repository-wide record coverage、symbol anchor、interactive record builder は未実装です: [#28](https://github.com/MongLong0214/commitlore/issues/28)、[#32](https://github.com/MongLong0214/commitlore/issues/32)、[#33](https://github.com/MongLong0214/commitlore/issues/33)、[#34](https://github.com/MongLong0214/commitlore/issues/34)。
- benchmarkagent behavior に対する guard の効果を実証していません: [#37](https://github.com/MongLong0214/commitlore/issues/37)。
- M4guard の効果を検証していません。どちらの arm にも injected record が届きませんでした: [#122](https://github.com/MongLong0214/commitlore/issues/122)。

## コントリビュート

Expand Down
4 changes: 2 additions & 2 deletions README.ko.md
Original file line number Diff line number Diff line change
Expand Up @@ -162,7 +162,7 @@ git log --follow --format='%h %(trailers:key=Limit,valueonly)' -- src/auth/

## 근거: 더 좁은 제품 주장

112회 실험을 했지만, 측정한 에이전트 행동 주장은 뒷받침되지 않았다. 그래서 CommitLore는 위의 더 좁은 제품 주장을 한다. 전체 한계는 [M4 verdict](bench/VERDICT-M4.md)에서 읽을 수 있다.
112회 실험은 기록됐지만 M4는 어느 arm에도 record를 전달하지 않았다. 따라서 에이전트 행동 주장을 시험하거나 뒷받침하거나 반박하지 못한다. 위의 더 좁은 제품 주장은 독립적으로 검증 가능한 동작에 근거한다. 깨끗한 데이터셋과 철회 내용은 [M4 verdict](bench/VERDICT-M4.md)에서 읽을 수 있다.

<details>
<summary>전체 benchmark 기록 (112회 실험)</summary>
Expand Down Expand Up @@ -224,7 +224,7 @@ node ~/.commitlore/dist/commitlore.mjs context src/auth
- Windows는 지원하지 않는다: [#95](https://github.com/MongLong0214/commitlore/issues/95).
- Alpine 및 다른 musl Linux host는 지원하지 않는다: [#99](https://github.com/MongLong0214/commitlore/issues/99).
- 암호학적 작성자 검증, 저장소 전체 record coverage, symbol anchor, interactive record builder는 아직 구현되지 않았다: [#28](https://github.com/MongLong0214/commitlore/issues/28), [#32](https://github.com/MongLong0214/commitlore/issues/32), [#33](https://github.com/MongLong0214/commitlore/issues/33), [#34](https://github.com/MongLong0214/commitlore/issues/34).
- benchmark는 에이전트 행동에 대한 guard 효과를 입증하지 못한다: [#37](https://github.com/MongLong0214/commitlore/issues/37).
- M4는 guard 효과를 시험하지 못했다. 어느 arm도 injected record를 받지 못했다: [#122](https://github.com/MongLong0214/commitlore/issues/122).

## 기여하기

Expand Down
4 changes: 2 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -162,7 +162,7 @@ These are product claims about Git-bound, human-verifiable decision history. The

## Evidence: a narrower product claim

112 experiments were run, but the measured agent-behavior claim was not supported. CommitLore therefore makes the narrower product claim above. Read the [M4 verdict](bench/VERDICT-M4.md) for the full limits.
112 experiments were recorded, but M4 delivered records in neither arm. It did not test, support, or refute the agent-behavior claim. The narrower product claim above rests on independently testable behavior; read the [M4 verdict](bench/VERDICT-M4.md) for the clean dataset and withdrawal.

<details>
<summary>Full benchmark record (112 experiments)</summary>
Expand Down Expand Up @@ -224,7 +224,7 @@ node ~/.commitlore/dist/commitlore.mjs context src/auth
- Windows is unsupported: [#95](https://github.com/MongLong0214/commitlore/issues/95).
- Alpine and other musl Linux hosts are unsupported: [#99](https://github.com/MongLong0214/commitlore/issues/99).
- Cryptographic author verification, repository-wide record coverage, symbol anchors, and an interactive record builder are not implemented yet: [#28](https://github.com/MongLong0214/commitlore/issues/28), [#32](https://github.com/MongLong0214/commitlore/issues/32), [#33](https://github.com/MongLong0214/commitlore/issues/33), [#34](https://github.com/MongLong0214/commitlore/issues/34).
- The benchmark does not demonstrate a guard effect on agent behavior: [#37](https://github.com/MongLong0214/commitlore/issues/37).
- M4 did not test a guard effect: neither arm received injected records ([#122](https://github.com/MongLong0214/commitlore/issues/122)).

## Contributing

Expand Down
4 changes: 2 additions & 2 deletions README.zh-CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -162,7 +162,7 @@ git log --follow --format='%h %(trailers:key=Limit,valueonly)' -- src/auth/

## Evidence:更窄的产品主张

完成了 112 次实验,但所测 agent behavior 的主张未获支持。因此 CommitLore 提出的是上面更窄的产品主张。完整限制见 [M4 verdict](bench/VERDICT-M4.md)。
已记录 112 次实验,但 M4 没有向任何 arm 交付 record。因此它没有检验、支持或反驳 agent behavior 的主张。上面更窄的产品主张基于可独立验证的行为。关于干净的数据集和撤回,请见 [M4 verdict](bench/VERDICT-M4.md)。

<details>
<summary>完整 benchmark 记录(112 次实验)</summary>
Expand Down Expand Up @@ -224,7 +224,7 @@ node ~/.commitlore/dist/commitlore.mjs context src/auth
- 不支持 Windows:[#95](https://github.com/MongLong0214/commitlore/issues/95)。
- 不支持 Alpine 与其他 musl Linux host:[#99](https://github.com/MongLong0214/commitlore/issues/99)。
- 尚未实现 cryptographic author verification、repository-wide record coverage、symbol anchor 和 interactive record builder:[#28](https://github.com/MongLong0214/commitlore/issues/28)、[#32](https://github.com/MongLong0214/commitlore/issues/32)、[#33](https://github.com/MongLong0214/commitlore/issues/33)、[#34](https://github.com/MongLong0214/commitlore/issues/34)。
- benchmark 未能证明 guard 对 agent behavior 有效果:[#37](https://github.com/MongLong0214/commitlore/issues/37)
- M4 没有检验 guard 效果:没有任何 arm 收到 injected record([#122](https://github.com/MongLong0214/commitlore/issues/122))

## 贡献

Expand Down
38 changes: 22 additions & 16 deletions bench/README.md
Original file line number Diff line number Diff line change
@@ -1,19 +1,19 @@
# CommitLoreBench

**M4 is the citable dataset.** `bench/results/t702-m4-final.jsonl` records the
harness commit and the `dist/` digest for every row, `bench/report.ts` summarizes
it, and the README's numbers block is generated from it. M3 was voided for
lacking that provenance (§15); M4 was designed, registered and run to supply it,
and its result is null. The historical executed report is
`bench/VERDICT-M4.md`; the canonical paired-and-clustered correction is
`docs/VERDICT-M4.md`. Every earlier dataset
(M1, M1-b, M2) still lacks the fields and is not pooled into the generated block
for that reason, not because it was withdrawn as a record; each has its own
verdict document.

Measures the one thing that decides whether CommitLore is worth building: does an
agent that can see recorded decisions stop re-proposing the approaches a team
already rejected?
**M4 is a citable, clean-provenance dataset, not a guard test.**
`bench/results/t702-m4-final.jsonl` records the harness commit and `dist/`
digest for every row, and `bench/report.ts` summarizes it. M3 was voided for
lacking that provenance (§15); M4 has it. But M4's 112 transcripts contain no
injected context in either arm, so its valid data do not answer the guard
question. The withdrawal and the corrected statistics as observations about
the data are in `bench/VERDICT-M4.md`. Every earlier dataset (M1, M1-b, M2)
still lacks the fields and is not pooled into the generated block for that
reason, not because it was withdrawn as a record; each has its own verdict
document.

The benchmark is designed to ask whether an agent that receives recorded
decisions stops re-proposing approaches a team already rejected. M4 did not
deliver those decisions, so it does not answer that question.

- Design: `docs/adr/ADR-0007-commitlorebench.md`
- Requirements: `docs/prd/PRD-F7-commitlorebench.md`
Expand Down Expand Up @@ -81,6 +81,11 @@ Conditions are an open string enum, so M4 adds arms without touching the runner.

Planned arms are rejected at CLI parse time with a pointer to their ticket.

**M4 correction:** although its labels were `commitlore-on` and
`commitlore-guard`, all 112 stored M4 transcripts have `injected_context: null`.
The condition table describes the intended harness behavior; M4 did not receive
the treatment and is not evidence about it.

`--cond both` is the two arms of the primary comparison and `--cond all` is
every supported arm. Until T-703 those were the same list; they are not any
more, and against a live driver the difference is two arms or five. The runner
Expand Down Expand Up @@ -938,8 +943,9 @@ Fisher exact also treats runs as independent, while the design is paired by
(task, seed). The output says so. The test is the one ADR-0007 and T-702
registered, but it does not provide a valid hypothesis test for paired data.
The original number remains part of the historical report; the registered
replacement and M4 correction are in `docs/MEASUREMENT-PROTOCOL.md` and
`docs/VERDICT-M4.md`.
replacement is in `docs/MEASUREMENT-PROTOCOL.md`. M4's paired/clustered
statistics are preserved in `bench/VERDICT-M4.md` as descriptions of rows that
did not receive the treatment, not as a correction of a guard estimate.

### What makes the measured effect a floor — one thing, not two

Expand Down
Loading
Loading