Skip to content

refactor(agent): drop opaque orchestration prose from load_skill results - #3473

Merged
kwakayama merged 2 commits into
mainfrom
fix/load-skill-next-step
Aug 8, 2026
Merged

kwakayama merged 2 commits into
mainfrom
fix/load-skill-next-step

Conversation

@kwakayama

@kwakayama kwakayama commented Aug 8, 2026 •

Copy link
Copy Markdown
Contributor

Fixes veryfront/veryfront-issue-inbox#5.

load_skill returned a nextStep string carrying internal orchestration policy alongside the skill's actual instructions:

IMPORTANT: load_skill only loads instructions. It does not perform the task or finish the turn. Continue the same turn after calling it. Keep the root assistant visibly owning the work. For multi-step or isolated work, call invoke_agent; otherwise keep working directly with the available tools. Delegate only when isolation, parallelism, or a different tool/model budget materially helps. Pass through any returned model, thinking, or maxSteps overrides to invoke_agent when delegating.

The finding that makes this safe

Every clause of that string is already in the <available_skills> system prompt, word for word. Rendered from current main:

You have access to these skills. Use load_skill to load full instructions when needed.
load_skill only loads instructions plus metadata. Continue the same turn after calling it.
Keep the root assistant visibly owning the work. When delegating, use the available legacy
`invoke_agent` tool. Delegate only when isolation, parallelism, or a different tool/model
budget materially helps. Pass through any returned model, thinking, or maxSteps overrides
to invoke_agent when delegating. ...
nextStep clause In the system prompt?
"Continue the same turn after calling it." verbatim
"Keep the root assistant visibly owning the work." verbatim
delegation advice, threshold, override forwarding verbatim
"It does not perform the task or finish the turn." covered by "only loads instructions plus metadata"

So the result was restating trusted system policy as untrusted tool output — exactly the duplication the issue objects to. Removing it deletes a copy, not a signal.

What the result looks like now

{
  "skillId": "deploy",
  "instructions": "---\nname: deploy\n...\n---\nDo the deploy.",
  "model": "sonnet",
  "references": ["references/guide.md"]
}

Every field is structured, and every field has a documented consumer:

Field Consumer
skillId caller, to correlate the result
instructions model — the only instruction content in the result
references model — loadable via load_skill's file parameter
model / thinking / maxSteps caller — overrides to forward when delegating

Removed

Prose fields:

  • nextStep — the orchestration policy above
  • overrideNote — "Pass through any returned model, thinking, or maxSteps overrides…", verbatim in the system prompt
  • referenceNote — "After this skill is loaded, use load_skill with the file parameter…", which duplicates the load_skill tool description

Dead since #3464, still on the type and still being copied by copyLoadedSkillResponse:

  • allowedTools, note, delegationTools, unavailableCurrentRunTools, delegationNote
  • assertRuntimeResponseMetadata, which validated delegationTools that nothing emitted

Configuration surface that existed only to override the deleted prose:

  • RuntimeLoadedSkillResponseMessages, RuntimeLoadSkillToolMessages, RUNTIME_LOAD_SKILL_CONTINUATION_NOTE, and the nextStep / messages tool options
  • Three public re-exports from src/agent/index.ts

availableToolNames is also dropped from buildStrictRuntimeLoadedSkillResponse — it was read and never used once overrideNote went. Its bound is still enforced, independently, by assertRuntimeBoundaryCollections.

Acceptance criteria

  • The public load_skill schema documents every returned field and its consumer — each field on RuntimeLoadedSkillResponse now carries a doc comment naming who reads it
  • Internal root-agent/delegation policy is not emitted as opaque result text
  • Studio does not expose implementation-policy text as a normal tool result
  • Regression tests prove the post-load turn does not rely on an ad hoc prose nextStep
  • The model still continues after loading a skill — see below

The one criterion I cannot prove in tests

The continuation instruction now appears once (system prompt) instead of twice (system prompt + every load_skill result). Unit tests can prove the signal still exists; they cannot prove a model still obeys it.

This is worth watching, because veryfront/veryfront-issue-inbox#392's residual is precisely "a child stopped after load_skill". I measured that baseline while triaging it: 2 runs in 30 days across agent_run_event matched the stop-after-load_skill signature (the raw figure was 561, but 558 of those were one retired control-agent burst that ended 2026-07-15).

So there is a concrete before-number to compare against. If the rate climbs above ~2/month after this ships, this change is the first suspect and reverting the nextStep removal is the obvious probe.

Testing

  • src/skill/ + src/agent/ — 1186 passed, 2050 steps, 0 failed
  • deno task typecheck — 0 errors
  • deno lint src/ — clean
  • deno task docs:api-reference:check — current
  • deno fmt --check — clean

Tests asserting the removed fields were deleted rather than weakened (6 in load-skill-tool.test.ts, 1 in skill-metadata.test.ts), all of which existed only to pin prose or its configurability. Where a test's real subject survived — accessor rejection, input bounds, iterator snapshotting — the removed-field half was dropped and the rest kept.

On the full-tree run: src/ cli/ tests/ shows 26 failures against a 8-failure baseline, but the difference is entirely e2e/dev-server suites (CSS, MDX Pages, Static Files, Relative Import Resolution, skill-capabilities). I verified tests/e2e/features/skill-capabilities.test.ts fails identically with my changes stashed — 403 vs expected 200, a local admission failure — and that file contains zero references to nextStep. These suites are non-deterministic under parallel dev servers in this checkout; CI runs them in dedicated jobs.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • Updates

    • Simplified loaded-skill responses by removing legacy continuation, delegation, tool-availability, and message fields.
    • Loaded skills now provide continuation guidance directly within their instructions.
    • Preserved skill loading, authorization, caching, references, cancellation, metadata, and inventory behavior.
    • Retired legacy runtime skill response exports and related configuration options.
  • Documentation

    • Refreshed API reference links and removed entries for retired runtime skill response types.

@kwakayama
kwakayama requested a review from kojiwakayama as a code owner August 8, 2026 12:57
@coderabbitai

coderabbitai Bot commented Aug 8, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 5385a148-fba8-4a46-90f2-36179ff72dea

📥 Commits

Reviewing files that changed from the base of the PR and between 8c30128 and 3f4d221.

📒 Files selected for processing (3)
  • docs/api-reference/veryfront/agent.md
  • src/agent/runtime/load-skill-tool.test.ts
  • src/agent/runtime/load-skill-tool.ts
🚧 Files skipped from review as they are similar to previous changes (3)
  • docs/api-reference/veryfront/agent.md
  • src/agent/runtime/load-skill-tool.test.ts
  • src/agent/runtime/load-skill-tool.ts

📝 Walkthrough

Walkthrough

The change simplifies runtime loaded-skill responses by removing continuation, message, tool, and note fields. It updates load-skill options, public exports, tests, and API reference links.

Changes

Runtime skill response simplification

Layer / File(s) Summary
Loaded-skill response contract
src/agent/runtime/skill-metadata.ts
RuntimeLoadedSkillResponse and response construction retain instructions, metadata overrides, and normalized references. Legacy response and tool fields are removed.
Load-skill tool and public exports
src/agent/runtime/load-skill-tool.ts, src/agent/index.ts, docs/api-reference/veryfront/agent.md
Load-skill options and exports no longer expose removed metadata. The shared tool description contains delegation-policy guidance. API links and catalog entries are updated.
Validation tests and fixtures
src/agent/runtime/load-skill-tool.test.ts, src/agent/runtime/skill-metadata.test.ts
Tests and fixtures no longer expect removed fields. Tests retain coverage for metadata validation, references, accessors, loading, and tool inventory behavior.

Estimated code review effort: 3 (Moderate) | ~20 minutes

Possibly related PRs

Suggested reviewers: kojiwakayama

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely describes the main change: removing opaque orchestration prose from load_skill results.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/load-skill-next-step

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
src/agent/runtime/load-skill-tool.test.ts (1)

141-142: 📐 Maintainability & Code Quality | 🟠 Major | 🏗️ Heavy lift

Use the required BDD test API.

This changed test remains registered with Deno.test. Convert the touched test cases to describe() and it() from #veryfront/testing/bdd.ts. Use assertions from #veryfront/testing/assert.ts.

As per coding guidelines, "**/*.{test,spec}.ts: Use describe() and it() from #veryfront/testing/bdd.ts, use assertions from #veryfront/testing/assert.ts."

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/agent/runtime/load-skill-tool.test.ts` around lines 141 - 142, Convert
the affected tests in load-skill-tool.test.ts from Deno.test to describe() and
it() imported from `#veryfront/testing/bdd.ts`, and replace their assertions with
imports from `#veryfront/testing/assert.ts`. Preserve the existing test cases and
expectations while registering them through the required BDD API.

Source: Coding guidelines

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Outside diff comments:
In `@src/agent/runtime/load-skill-tool.test.ts`:
- Around line 141-142: Convert the affected tests in load-skill-tool.test.ts
from Deno.test to describe() and it() imported from `#veryfront/testing/bdd.ts`,
and replace their assertions with imports from `#veryfront/testing/assert.ts`.
Preserve the existing test cases and expectations while registering them through
the required BDD API.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 6260f4d4-6723-4231-96fd-f78393d2365e

📥 Commits

Reviewing files that changed from the base of the PR and between 48a2bca and 2c5d84e.

📒 Files selected for processing (6)
  • docs/api-reference/veryfront/agent.md
  • src/agent/index.ts
  • src/agent/runtime/load-skill-tool.test.ts
  • src/agent/runtime/load-skill-tool.ts
  • src/agent/runtime/skill-metadata.test.ts
  • src/agent/runtime/skill-metadata.ts
💤 Files with no reviewable changes (3)
  • src/agent/index.ts
  • src/agent/runtime/skill-metadata.test.ts
  • src/agent/runtime/load-skill-tool.ts

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 2c5d84edc9

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

delegationNote?: string;
/** Reference files this skill advertises, loadable via load_skill's `file` parameter. */
references?: string[];
/** Model override the skill declares, for the caller to forward when delegating. */

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Preserve override forwarding for authored skill blocks

When an agent's instructions already contain a complete <available_skills> element, buildAgentCallContext deliberately skips the generated block (src/agent/runtime/call-context.ts:168-178), and RUNTIME_LOAD_SKILL_DESCRIPTION does not contain LOAD_SKILL_OVERRIDE_FORWARDING. In a legacy invoke_agent run where a skill declares model, thinking, or maxSteps, the old nextStep and overrideNote supplied that missing instruction, but this response now emits only the raw values, so the model is never told to forward them and can delegate with the wrong configuration. Keep the forwarding guidance in trusted prompt or tool-description text outside the replaceable skills block.

Useful? React with 👍 / 👎.

load_skill returned a nextStep string carrying internal orchestration
policy - "Continue the same turn after calling it", "Keep the root
assistant visibly owning the work", delegation guidance - alongside the
skill's actual instructions. It also returned overrideNote and
referenceNote, two more prose fields, plus five fields left dead by
veryfront/veryfront-issue-inbox#406.

Every clause of nextStep is already in the <available_skills> system
prompt, word for word. The result was restating trusted system policy as
untrusted tool output, which is what this issue objects to.

The result is now only structured data with a documented consumer:

  { skillId, instructions, references?, model?, thinking?, maxSteps? }

Removed: nextStep, overrideNote, referenceNote, allowedTools, note,
delegationTools, unavailableCurrentRunTools, delegationNote, the
RuntimeLoadedSkillResponseMessages configuration surface, and the
nextStep/messages tool options that existed to override the prose.

The continuation signal now lives in exactly one place - the system
prompt - rather than being duplicated into every load_skill result.

Fixes veryfront/veryfront-issue-inbox#5.

Claude-Session: https://claude.ai/code/session_01Xo93b6StAu691YV9g8Fm53
@kwakayama
kwakayama force-pushed the fix/load-skill-next-step branch from 2c5d84e to 8c30128 Compare August 8, 2026 13:06
buildAgentCallContext skips the generated <available_skills> block when
an agent's own instructions already contain one (call-context.ts:168).
For those agents the system prompt carries whatever the author wrote, so
it cannot be relied on to hold runtime orchestration policy.

The tool description already carried continue-same-turn, root ownership,
and the delegation threshold, but not override forwarding. Dropping
nextStep and overrideNote therefore left an authored-block agent with no
instruction to pass a skill's model, thinking, or maxSteps through to a
legacy invoke_agent delegation.

Move that clause into the tool description, which is always sent. It
belongs in the trusted tool contract rather than in the result payload,
which is what veryfront/veryfront-issue-inbox#5 asks for.

Found by Codex review on #3473.

Claude-Session: https://claude.ai/code/session_01Xo93b6StAu691YV9g8Fm53
@kwakayama

Copy link
Copy Markdown
Contributor Author

Addressed both reviews. Pushed 3f4d221c5.

Codex — override forwarding for authored skill blocks: valid, fixed

This was a real gap and it undermined this PR's central claim, so thank you for catching it.

My argument was "every clause of nextStep is already in the <available_skills> system prompt." I verified that against the generated block. But call-context.ts:168 only appends the generated block when the agent's own instructions do not already contain one:

if (input.skills?.length && !hasBlock(input.instructions, AVAILABLE_SKILLS_BLOCK_NAME)) {

For an agent with an authored block, the system prompt carries whatever the author wrote. My premise does not hold there.

Checked what the always-present tool description carries:

Clause In RUNTIME_LOAD_SKILL_DESCRIPTION?
LOAD_SKILL_CONTINUE_SAME_TURN yes
LOAD_SKILL_ROOT_OWNERSHIP yes
LOAD_SKILL_DELEGATION_THRESHOLD yes
LOAD_SKILL_OVERRIDE_FORWARDING no

So continuation, ownership, and the delegation threshold survived an authored block; override forwarding was the single clause that did not. Exactly as you described.

Fix: moved LOAD_SKILL_OVERRIDE_FORWARDING into the tool description. That keeps it in the trusted tool contract rather than the result payload, which is what #5 asks for, and it is sent regardless of whether the skills block is generated or authored.

Added a regression test pinning that the description carries all four clauses, so this cannot silently reopen. Verified non-vacuous — removing the clause fails it (0 passed | 1 failed).

CodeRabbit — convert touched tests to describe()/it(): declining, with reason

The guideline is real, but applying it to only the lines this PR touches makes the file worse:

  • load-skill-tool.test.ts is 74 Deno.test blocks, 0 describe. Converting the handful I touched produces a file with two competing registration styles.
  • It is not lint-enforced — deno task lint and lint:test-typecheck both pass, and scripts/lint's own tests use Deno.test.
  • Converting all 74 is a mechanical diff several times the size of this change, and would bury a behavioural PR under test-harness churn.

Worth doing as its own PR; not worth doing halfway here.

Verification after the fix

  • src/skill/ + src/agent/ — 1187 passed, 0 failed
  • deno task lint, lint:test-typecheck, docs:api-reference:check — all OK
  • deno task typecheck — 0 errors

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant