Skip to content

feat(agents): request summarized Claude thinking and label empty rows - #8051

Merged
wpfleger96 merged 5 commits into
mainfrom
hayt/empty-thinking-label
Oct 2, 2026
Merged

wpfleger96 merged 5 commits into
mainfrom
hayt/empty-thinking-label

Conversation

@wpfleger96

@wpfleger96 wpfleger96 commented Oct 2, 2026 •

Copy link
Copy Markdown
Member

Claude Code agents on Opus showed Thinking rows that opened to nothing. Opus returns no readable thinking unless Claude Code is asked for summaries, and Buzz never asked. Sonnet returns readable thinking either way.

Request summaries from Claude Code. session_new_full now adds _meta.claudeCode.options.extraArgs: {"thinking-display": "summarized"} for the Claude adapter. claude-agent-acp spreads extraArgs into the CLI's arguments; 0.36.1, 0.60.0 and 0.66.0 all read the same shape. The value is merged into the existing _meta, so systemPrompt and sessionTitle are kept. session_new_full is the only place Buzz sends session/new, so agent sessions and model discovery both get it.

Only for a CLI that accepts the flag. --thinking-display first shipped in Claude Code 2.1.94, and older CLIs exit on unknown flags, which would fail every session start. When the adapter is spawned, Buzz runs $CLAUDE_CODE_EXECUTABLE --version once with a 5-second limit, reading at most 256 bytes of output and killing its whole process group on timeout, and sends the flag only on 2.1.94 or newer. The flag is left off, and sessions start exactly as before, when:

  • the CLI is older than 2.1.94;
  • CLAUDE_CODE_EXECUTABLE is unset (the adapter then runs its own bundled CLI, which Buzz can't inspect) or empty;
  • the binary hangs, fails, or prints anything other than a clean major.minor.patch version;
  • BUZZ_ACP_LAUNCH_PREFIX is set, since the wrapper can change which CLI the adapter actually runs.

The result is logged once at debug level.

Label rows that still have no text. A Thinking row whose text is empty or only whitespace keeps its header and shows a muted "No readable reasoning tokens", with nothing to expand. This covers models or CLIs that still return no readable thinking. If text arrives later in the same row, it renders as an expandable section as before. ThoughtActivity is the only component that draws Thinking rows.

Claude models can return thinking sealed (signature only, no readable
text), and the feed rendered that as an expandable Thinking row with a
blank body. Keep the row, since the model did think, but show a muted
"No readable reasoning tokens" notice with nothing to expand.

Co-authored-by: Will Pfleger <pfleger.will@gmail.com>
Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
@wpfleger96
wpfleger96 requested a review from a team as a code owner October 2, 2026 14:57
@wpfleger96
wpfleger96 deployed to codex-review October 2, 2026 14:57 — with GitHub Actions Active
@github-actions github-actions Bot added the codex-security-review-current The posted Codex security review matches its recorded range. label Oct 2, 2026
@github-actions

github-actions Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

🔐 Codex Security Review

Note: This is an automated, security-focused review generated by Codex.
Use it as a supplement to human review; false positives are possible.

Scope

  • Exact PR diff: 133fb9869228dab4c00cb2b9e259bf834a55dbf7...c27444a35fb52be53588054bfaa7067ac5ba8308
  • Model: gpt-5.6-sol

💡 Click "edited" above to see earlier reviews for this PR.


Review Summary

Overall Risk: HIGH

The Claude version probe bypasses the adapter subprocess environment sanitization and can expose an ambient Nostr signing key.

Findings

[HIGH] Claude version probe inherits a private key removed from the adapter environment

  • Category: Cryptography
  • Location: crates/buzz-acp/src/acp.rs:2457 (source)
  • Description: The adapter command explicitly removes NOSTR_PRIVATE_KEY before spawning, and existing integration coverage treats that as a required secret boundary. The new capability check creates an independent Command for CLAUDE_CODE_EXECUTABLE. That command inherits the harness process environment rather than the sanitized adapter environment, so an ambient NOSTR_PRIVATE_KEY is available during execution of --version. This occurs before the protected adapter is spawned and can expose a distinct operator or owner key to the configured Claude executable.
  • Impact: A compromised or substituted Claude executable can steal the ambient Nostr private key and use it to sign events, impersonate that identity, or access resources authorized to it.
  • Recommendation: Run the probe with the same effective sanitized environment as the adapter command, including all env removals, or use env_clear and copy only the minimal variables required by the version executable. Add a regression test that places a canary in NOSTR_PRIVATE_KEY and verifies the version probe cannot observe it.

Notes

  • Review used read-only source and history inspection as required; repository tests and binaries were not executed.

Generated by Codex Security Review |
Requested by: @wpfleger96 |
Workflow run

Opus omits readable thinking unless Claude Code is asked for summaries, so
Thinking rows arrived empty. session/new now passes --thinking-display
summarized through claude-agent-acp extraArgs, but only when the CLI in
CLAUDE_CODE_EXECUTABLE reports 2.1.94 or newer: older CLIs exit on the
unknown flag, and an unreadable version leaves sessions unchanged.

Co-authored-by: Will Pfleger <pfleger.will@gmail.com>
Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
@wpfleger96 wpfleger96 changed the title fix(desktop): label textless thinking rows instead of an empty section feat(agents): request summarized Claude thinking and label empty rows Oct 2, 2026
@github-actions github-actions Bot removed the codex-security-review-current The posted Codex security review matches its recorded range. label Oct 2, 2026
@wpfleger96
wpfleger96 deployed to codex-review October 2, 2026 16:02 — with GitHub Actions Active
Leave summaries off when a launch prefix is set, since the wrapper can
change which CLI the adapter runs, and require a clean version string.
Tests now drive a Claude-named adapter through session/new with fake CLIs.

Co-authored-by: Will Pfleger <pfleger.will@gmail.com>
Signed-off-by: Will Pfleger <pfleger.will@gmail.com>

@wesbillman wesbillman left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Carl, an automated reviewer, commenting via Wes’s GitHub account.

No blocking findings. Reviewed head 5e6b88a78346a704bec4927df3c06bde0e25b08d against base 448407a972ca9da0c1e13d49ee2c2170821be8a2, including independent UI and Claude-adapter review. One optional subprocess-hardening note is inline. This is a comment, not approval.

Validation: hosted Rust checks passed; Desktop passed 6,779 ordinary JS tests and 133 jsdom tests, including all three new renderer cases. CI tested merge cfa3b4143575768425644682184f188344bf16d6; its entire Desktop tree and all five changed files match the reviewed head. Confirmed the extraArgs contract against upstream adapter versions 0.36.1, 0.60.0 and 0.66.0.

Remaining gates: smoke shard 2 failed the unchanged empty-edit-delete.spec.ts:95 message-edit test on every attempt; I found no connection to this PR’s changed paths, but CI is not green. Security review is still running. Real authenticated Claude/Opus output and human acceptance remain unverified. Resolve or explicitly disposition the CI failure and verify the live thinking flow before treating this as ready to merge.

Comment thread crates/buzz-acp/src/acp.rs Outdated
A hung launcher could leave its child running after the timeout, and the
version output had no size limit.

Co-authored-by: Will Pfleger <pfleger.will@gmail.com>
Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
@wpfleger96
wpfleger96 deployed to codex-review October 2, 2026 17:49 — with GitHub Actions Active
Co-authored-by: Will Pfleger <pfleger.will@gmail.com>
Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
@github-actions github-actions Bot added the codex-security-review-current The posted Codex security review matches its recorded range. label Oct 2, 2026
@wpfleger96
wpfleger96 enabled auto-merge (squash) October 2, 2026 17:58
@github-actions github-actions Bot removed the codex-security-review-current The posted Codex security review matches its recorded range. label Oct 2, 2026
@wpfleger96
wpfleger96 deployed to codex-review October 2, 2026 18:02 — with GitHub Actions Active
@github-actions github-actions Bot added the codex-security-review-current The posted Codex security review matches its recorded range. label Oct 2, 2026
@wpfleger96
wpfleger96 merged commit cc528db into main Oct 2, 2026
81 checks passed
@wpfleger96
wpfleger96 deleted the hayt/empty-thinking-label branch October 2, 2026 18:46
ArnaudLafosse92100 added a commit to ArnaudLafosse92100/buzz that referenced this pull request Oct 4, 2026
Brings 16 upstream block/buzz commits (a14107a) into the fork
integration branch: ACP mention/edit steering (block#6131, block#6132), quiet-host
recovery wakes (block#7459), relay NIP-FI shadow mode (block#8034, block#8062), writer
lock foundations (block#7706), Goose MCP handshake (block#8037), Claude model names
(block#8053), summarized thinking (block#8051) and mobile iOS changes.

Merged cleanly without textual conflicts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Arnoldinh0 <arnaudlafosse92100@gmail.com>

This branch was successfully deployed

1 active deployment
codex-review — c27444a3 Deployed Oct 2, 2026 by wpfleger96 via Run Codex Security Review #6594
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

codex-security-review-current The posted Codex security review matches its recorded range.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants