Skip to content

test(chat): cover the tool-replay reconciliation algorithm directly - #3443

Merged
kwakayama merged 2 commits into
refactor/chat-tool-replay-extractionfrom
test/replay-reconciliation-coverage
Aug 7, 2026
Merged

kwakayama merged 2 commits into
refactor/chat-tool-replay-extractionfrom
test/replay-reconciliation-coverage

Conversation

@kojiwakayama

@kojiwakayama kojiwakayama commented Aug 6, 2026 •

Copy link
Copy Markdown
Contributor

Stacked on #3439 (refactor/chat-tool-replay-extraction) — the module under test does not exist on main. Merge #3439 first.

Writes the first real test suite for findProviderVisibleToolReplayMatches. Tests only — no source changes.

Why

#3439 extracted a ~100-line algorithm that decides which tool-call/result occurrences in replay history are authoritative for provider conversion — matching by part object identity via seven WeakSet/WeakMap collections. It shipped with 2 smoke tests (isTransientToolState classification, and empty history). The algorithm itself was covered only indirectly, through provider conversion in conversation.test.ts / message-prep.test.ts.

That was the point of extracting it. This is the follow-through.

What was added

10 cases (2 existing smoke tests kept; 12 total): pair matching, supersession via both paths, transient preservation, name-mismatch rejection, unmatched pending call, batch start, toolCallsById supersession after pending-queue eviction, the user-message batch boundary, and object identity.

The case that matters most

Identity discrimination. It builds a matched call/result pair, then checks structurally identical but distinct clones against matchedToolCallParts, matchedToolResultParts, and matchedToolResultNames — asserting the real object reads true/has-a-name and the clone reads false/undefined.

Its exclusive value: it is the only case that fails if the returned WeakSet/WeakMap collections are swapped for structural-equality lookups. No consumer's compile would catch that — callers only use .has() and .get(). Without this case, a change that "optimized" by comparing parts structurally, or by cloning them internally, would pass every other test while silently breaking replay matching.

Test quality was verified by mutation, not by reading

The suite was reviewed by building 9 mutants of the implementation (as copies — tracked source was never modified) and mapping which cases catch which:

Mutation Caught by
Clone parts internally before adding to the sets cases 1, 2, 3, 5, 8
Return structural-equality collections instead of WeakSet/WeakMap case 2 only
Drop supersession propagation to the earlier matched result case 3 only
isCompatibleToolResultName → always true case 6 only
Preserve every transient call unconditionally cases 5, 6, 9
Batch-start fires on any pending call case 8 only
Self-contained path also adds to matchedToolCallParts cases 4, 9

That exercise found a real defect in the first draft: a case titled as fencing the user-message boundary survived both boundary mutants. Its assertions were correct, but they pinned something else — its fixture shared a toolCallId, so removePendingCallsWithId evicted the stale entry regardless and the boundary was never exercised. It has been retitled to what it actually pins, and a genuine boundary case added alongside it (verified to flip under the mutant).

A documented limitation

The new boundary case fences the "user text stops counting as provider-visible content" regression. It does not fence deleting the earlier-message flush inside the visible-content branch: for a call-less user message, pendingCountBeforeSameMessageVisibleContent still captures the pre-message length, so the end-of-loop splice fallback evicts exactly the same entries — the two paths are provably equivalent in that shape. Discriminating it would need a boundary message with calls interleaved with visible content. This is stated in the case's comment rather than left implied.

Behaviour pinned, not fixed

A call already evicted from pendingCalls as stale still lands in supersededToolCallParts when a later same-toolCallId occurrence arrives — toolCallsById accumulates every occurrence and is never pruned, while pendingCalls is pruned by three separate paths.

Reviewed as harmless, and arguably the safer behaviour: the flag only fires when another occurrence with the same toolCallId appears later, and emitting two tool-call blocks under one id is exactly what this module exists to prevent. Consumers checked: conversation.ts:816 (pushToolCall, reached only after the transient-skip guard), :889/:928/:949 (provider tool messages), and message-prep.ts:883-885 (stripPendingToolParts). Pinned as characterization; no source change.

Evidence

  • deno task test:unit: branch baseline 3804 passed / 27918 steps → 3804 / 27928, zero regressions.
  • Target file green; deno lint, deno task verify:quick, deno task lint:chat-ratchets all exit 0.
  • Commit touches only src/chat/tool-replay-reconciliation.test.ts.

Summary by CodeRabbit

  • Tests
    • Added comprehensive coverage for reconciling assistant, user, tool, and dynamic tool messages.
    • Verified matching behavior, duplicate calls, transient data preservation, unresolved calls, batch boundaries, stale entries, and user-message boundaries.

Add 9 direct cases for the replay reconciliation matcher: identity vs
structural equality, supersession via both the pending-call-match path
and the self-contained-call path, transient preservation timing, batch
starts, name-mismatch rejection, and unresolved pending calls. Only the
two prior smoke cases previously exercised this module directly; the
rest of its coverage came indirectly through conversation.test.ts.
Case 9 was titled as pinning that a user-message boundary drops stale
pending tool calls, but its staleCall/freshCall shared a toolCallId, so
id-based eviction (removePendingCallsWithId) and toolCallsById-based
supersession masked the boundary logic entirely — freshCall was also
self-contained, so no result ever needed to match against pendingCalls.
Mutants deleting the boundary flush or the user-visible-content check left
the whole suite green.

Kept the old fixture under an honest title (it does correctly pin
toolCallsById supersession of an already-evicted call) and added a new
case with a fixture that actually needs pendingCalls to retain a boundary-
crossing entry: staleCall and its late result share a toolCallId used
nowhere else, so same-id eviction can't do the work for it.

Also added a one-line positive control to the isCompatibleToolResultName
mismatch case so it can't pass vacuously against a rotted fixture.
@kojiwakayama
kojiwakayama requested a review from kwakayama as a code owner August 6, 2026 22:06
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@coderabbitai

coderabbitai Bot commented Aug 6, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 50d9b10e-96ed-48c5-9079-39039f4f9478

📥 Commits

Reviewing files that changed from the base of the PR and between f14619d and d32d1f9.

📒 Files selected for processing (1)
  • src/chat/tool-replay-reconciliation.test.ts

📝 Walkthrough

Walkthrough

The pull request adds test fixtures and comprehensive coverage for tool replay reconciliation, including matching, identity checks, supersession, transient calls, stale-call eviction, batch boundaries, and user-message boundaries.

Changes

Tool replay reconciliation

Layer / File(s) Summary
Fixtures and basic matching
src/chat/tool-replay-reconciliation.test.ts
Adds builders for assistant, tool, user, raw tool-call/result, and dynamic-tool messages. Tests empty history, basic matching, recorded tool names, and object identity.
Supersession and transient calls
src/chat/tool-replay-reconciliation.test.ts
Tests same-ID supersession, dynamic-tool calls, transient-call preservation, tool-name mismatches, and unresolved calls.
Batch eviction and message boundaries
src/chat/tool-replay-reconciliation.test.ts
Tests batch-start detection, stale pending-call eviction, call-ID index supersession, and user-message termination of matching windows.

Estimated code review effort: 3 (Moderate) | ~20 minutes

Possibly related PRs

Suggested reviewers: kwakayama

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely describes the pull request's main change: adding direct test coverage for the tool-replay reconciliation algorithm.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch test/replay-reconciliation-coverage

Comment @coderabbitai help to get the list of available commands.

@kwakayama
kwakayama merged commit 1eba258 into refactor/chat-tool-replay-extraction Aug 7, 2026
1 check passed
@kwakayama
kwakayama deleted the test/replay-reconciliation-coverage branch August 7, 2026 02:42
@kwakayama

Copy link
Copy Markdown
Contributor

Closing as superseded by #3439.

The two commits from this branch were squashed onto refactor/chat-tool-replay-extraction as 1eba258, so #3439 now contains everything here plus the extraction this test suite exercises. At the time of the squash the two trees were byte-identical (git diff between the heads was empty); #3439 has since moved ahead with the review nitpicks addressed in 67ffa89.

Both PRs target main, so merging both would conflict. #3439 is the one to land: it carries the full check suite, while this branch never triggered a CI/CD run.

No work is lost. findProviderVisibleToolReplayMatches coverage ships in #3439.

@kwakayama

Copy link
Copy Markdown
Contributor

Correction to my previous comment: this was already merged, not superseded, and my note misread the state.

It merged into refactor/chat-tool-replay-extraction (#3439) as 1eba258, not into main. So the coverage ships when #3439 lands, which is the same outcome, but nothing here needed closing and no action is required.

The rest stands: #3439 carries this work plus the review nitpicks addressed in 67ffa89.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants