Repository navigation
[M9] send() and execute() with a listener stream model calls like stream() (breaking) - #310
Merged
Merged
Conversation
…am() RunEvents splits its one flag in two: `streamModelCalls` (stream each model call when the provider can) and `iterated` (the caller iterates a stream()). A run with listeners but no iteration streams its model calls; output guardrails read `iterated`, so send() keeps checking only the final reply. New ExecuteOptions.streamModelCalls (default true) restores whole steps. MockLLMProvider.stream() now streams the step generate() returns (tool calls, finish reason, usage, exact text): its old stream() dropped tool calls and added a trailing space, which is what changed Agent Forge's results in round 1. Forge's abortable provider checks stop() between streamed chunks. Closes #232 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
1 task
This was referenced Oct 2, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #232
Breaking.
send()andAgentExecutor.execute()with a listener (createAgent({ onEvent }),onAgentEvent, or the deprecatedonEventoption) now stream model calls. A step's text arrives as severaltext.deltaevents, as it does instream().What changed
RunEvents(src/execution/agentRun.ts) now has two flags instead ofstreamed:streamModelCalls: stream each model call when the provider can.iterated: the caller iterates astream().stream()sets both to true.observeRun(listeners, no iteration) setsstreamModelCalls(defaulttrue) anditerated: false.RunEventSink.streamedwas renamed toiterated. It is internal: no entry point exportsRunEventSink.AgentExecutor.guardOutputreadsiterated, so output guardrails onsend()still check only the final reply.New
ExecuteOptions.streamModelCalls?: boolean. It defaults totruewith listeners, andfalsegenerates whole steps.stream()ignores it, and it is not added tocreateAgent().The
ctx.emitdoc inhooks.tsnow reads "set when the run has listeners or is streamed". The notes increateAgent.ts,AgentExecutor.ts,agentRun.tsandgenerateStep.tsare updated too.MockLLMProvider.stream()now streams the same step thatgenerate()returns. This is why round 1 changed Agent Forge's results. The oldstream()never emitted tool calls, used different usage numbers, and its chunks wereword + ' ', so the text ended in an extra space ('This is a mock response. '). Before this fix, Forge's mock runs stopped calling tools once they streamed. I checked this: with the new run loop and the oldmock.ts, 9 Forge server tests fail (runRegistry ×6, chat, localTools, timeTravel). With the fixed mock, all of them pass. Two pinned expectations changed because of it:llm.test.ts('Hello, world! '→'Hello, world!') andcloudflare.test.ts("This is a mock response. "→"This is a mock response.").Agent Forge:
abortableProvider.stream()now checksstop()between streamed chunks and after the last one. That is the same point at whichgenerate()checked after its call returned.ChatPanel.tsxcomment is updated.text.delta. The logs usetext.doneand the chat uses the reconciled messages, so no rendering change was needed.Tests
src/execution/sendStreams.test.ts(new) covers:send()withonEventemits more than onetext.deltaper step, and their concatenation equals thetext.donetext.result.usage,toolCalls,stepsand thestep.doneusage equal those of the same run without a listener (no double counting).execute({ onAgentEvent })streams, andexecute({ streamModelCalls: false })gives one delta per step.stream(), and a provider withsupportsStreaming: false, each give one delta per step.send()with a listener and a tool-calling step does not check the intermediate text, whilestream()does and trips.createAgent({ retry, fallbackModels })with a listener retries a failed streamed call and then falls back. The listener getsprovider.retryandprovider.fallback.legacyEvents.test.ts: thesupportsStreaming: () => falseoverride is gone, and the assertion is unchanged.compactionEvents.test.ts: tests the correctedctx.emitrule (send()with a listener gets the hook's event).cancellation.test.ts: the scripted provider stubsstream: vi.fn()(it returns undefined). ItssupportsStreamingis nowfalse("generate-only"). The assertions are unchanged.llm.test.ts: a new case checks that the mock's stream equals itsgenerate(): text, tool calls, finish reason and usage.runRegistry.test.tshas a new test that uses a realwsclient onattachWebSocketServer. Each of the two mock steps has more than onetext.deltaand exactly onetext.done, and the deltas add up to thetext.donetext.abortableProvider.test.ts(new) covers a stream passed through, and a stream stopped while being read, which throwsRunAbortedError.reasoning.test.ts,agentStream.test.ts, guardrails, sub-agents,agent_await([M4] Background sub-agents that need approval pause the lead at agent_await and resume #305), drift (M10c: compare a paused sub-agent's fingerprint fully on resume #300), remote usage ([M10b] Add remote sub-agent token usage to the lead's totals #297), the exporter and trace files ([M5a] lousho traces: file trace exporter, createAgent exporter, terminal viewer #286, [M5b] Agent Forge: persisted trace history in the Trace tab #301) andtodo.updated([N12] todo.updated stream event and useTodos() for React, Vue and Svelte #299) all pass unchanged.Agent Forge, verified for real
I built the studio (
npm run build:studio) and startednode bin/lousho.js studio --port 4793from an empty scratch directory. A script then drove it over HTTP andWS /agents/:id/stream. For the baseline, I ran the same script against a detachedorigin/main(a8711e1) studio on port 4794.PUT /agents/m9check(mock,current-date),POST /run"please use current-date""This is a mock response."run.start, step.start, text.delta, text.done, tool.start, tool.done, step.done, step.start, text.delta, text.done, step.done, run.donetext.deltais now 5text.deltaevents ("This ","is ", …) that join to thetext.donetextlog)POST /agents/m9check/message"hello there")text.doneGET /agents/m9check/tracesGET /runs/m9check/history[1, finished, stop, [], 39 tokens]demo-approvalagent): run → pausedawaiting_approvalwithpendingApproval→POST /approve202 → stopped"This is a mock response."… tool.start, approval.requested, step.done, run.done, run.start, tool.start, tool.done, …), except for the deltasGET /served the client (200). Stop is covered by the existing "stop() then run() resumes from the last checkpoint" test and the newabortableProvidertest.Verification (on the merged head 37be264, after merging origin/main twice: a8711e1, then 59f2a75 / #303)
Earlier full runs on d8d7415 also hit failures under load, and each was a different test. All of them pass when run alone, and none of them touches the run loop or the mock:
http.test.ts: timeout.sandbox-wiring.test.ts: timeout.NodeWorkspace.test.tsshell: timeout.SubprocessSandbox.test.tsDocker integration: "no such container", because another agent is using the daemon.backgroundSubagents.test.ts"reports a sub-agent paused for approval…": the test waits a fixed 10 ms. It sends without a listener, so it takes the unchangedgenerate()path.Peer matrix, run locally (the CI
peersjob steps), for the provider and execution tests (src/providers src/execution src/createAgent src/context src/testing src/subagents):ai@6.0.300(ai@6 @ai-sdk/openai@3 @ai-sdk/anthropic@3): tsc 0, test:types 0, vitest 1168 passed.ai@7.0.127(ai@7 @ai-sdk/openai@4 @ai-sdk/anthropic@4): tsc 0, test:types 0, build 0, vitest 1170 passed.After the matrix,
npm cirestored the defaults (ai4.3.19), and the lockfile is unchanged. I did not runpack-smoke: this change does not touch exports, bin, package.json or the build.Live test
Skipped: the OpenRouter account is out of credit (
/credits:total_usage10.20 ≥total_credits10). Live test spend: before 0, after 0 (no calls).src/execution/sendStreams.live.test.tsis committed. It skips withoutOPENROUTER_API_KEYand runs only withnpm run test:live. It recordssrc/execution/__fixtures__/cassettes/send-streams.json. The replay test (sendStreams.replay.test.ts) is not included, because it needs that cassette. A checklist line was added to #260.Docs
These pages were edited, with no heading changes:
docs/streaming.md: "Listening without iterating", and "Token streaming and providers".docs/compaction.md,docs/hooks.md: thectx.emitrule.docs/guardrails.md: the output row says "astream()run (one you iterate)".docs/executor-api.md: a newstreamModelCallsrow.There are no new pages. CHANGELOG has a BREAKING entry with the migration note, plus entries for the mock provider and Agent Forge.
llms.txt/llms-full.txtare regenerated.🤖 Generated with Claude Code