Skip to content

Fix: retry a step that ends without a finish reason, and never send Gemini a trailing model turn #121

Description

@byapparov

Context

On Vertex, gemini-3.8-flash sometimes ends a model step with no finishReason: the stream carries thought chunks only, then {"error":{"code":500,"status":"INTERNAL"}} inside an HTTP 200 stream. The AI SDK reports the step as finish "unknown" and raises no error. @aictrl/cli@0.4.3 treats unknown like tool-calls and continues the loop, rebuilding the history with the reasoning-only assistant message as the last turn. Vertex rejects that request with 400 Requests ending with a model turn are not supported, which is not retryable, so the session errors and the process exits 1. Root cause and captured request: aictrl-dev/aictrl#5645 (root-cause comment).

New evidence from fault injection with real Vertex (2026-09-29): the fault happened on its own in 3 of 14 local runs, at model requests 20–29, i.e. mid-run. At that point no output exists, so no harness-side output contract (a harness-side output contract) can rescue the run. Code Review is exposed too when a mid-run fault kills the CLI. This fix is the primary fix for #5645 and is independent of ADR-022.

Problem / Goal

A provider step that ends with no finish reason and no tool call must be treated as a failed, retryable step, not as "continue". No request to a Gemini provider may end with a model turn.

Code locations (from the #5645 analysis, CLI 0.4.3):

  • session/processor.ts:251 stores finish = "unknown"; :429-431 returns "continue".
  • session/prompt.ts:336-341 (top-of-loop exit) and :711 (modelFinished) exclude ["tool-calls","unknown"]; :725-734 continues.
  • session/message-v2.ts:549-611 and :663 rebuild history keeping the reasoning-only assistant message last.
  • No guard in provider/transform.ts:252 or session/llm.ts:241.

Reproduction

  1. Put a logging proxy in front of Vertex (Vertex, location global) and run aictrl run --model google-vertex/gemini-3.8-flash --thinking --format json on a PR-Explanation-style prompt.
  2. Have the proxy answer one main-loop request with HTTP 200 SSE: three thought-only chunks, then the bare {"error":{"code":500,"status":"INTERNAL"}} body and no finishReason (same bytes as the captured runs/11.proxy.jsonl).
  3. Observed: NDJSON tail step_finish(unknown), message_complete(finish=unknown), then the next request has lastRole: "model", Vertex returns 400, session_error, exit 1.
  4. Expected: the step is retried with a history that ends in a user or tool-result turn; the run continues.

Environment

  • @aictrl/cli@0.4.3 (pinned in aictrl docker/executor/Dockerfile).
  • Two production runs of a PR-Explanation-style task failed this way (details in the internal tracker).

Proposed Approach

  1. unknown with no tool call = failed step. In processor.ts, set a retryable error (for example "provider ended stream without finishReason") and route it through the existing SessionRetry path, so the retry resends the same history ending in a user turn. Drop reasoning-only assistant messages from the rebuilt history (message-v2.ts). Cap the retries (for example 2 per step), then fail the invocation through the normal failure lifecycle (non-zero exit, consistent with Fix successful exit on provider error finish reasons #108).
  2. Gemini guard (defence in depth). In provider/transform.ts message(), for @ai-sdk/google and @ai-sdk/google-vertex, drop trailing assistant messages that contain no tool-call. This alone would spin on empty steps, so it ships with (1), not instead of it.
  3. Leave auxiliary requests alone. The CLI sends a concurrent auxiliary request with empty contents (title generation; Vertex answers 404). The retry counter, the guard and any request indexing must not count or alter it.

Acceptance Criteria

  • A deterministic provider-stream fixture (thought-only chunks, embedded 500, no finishReason) triggers a retry whose request ends with a user or tool-result turn; the run then completes when the retried step succeeds.
  • After the retry cap, the invocation fails with a structured, retryable-classified error; message_complete, session_complete, invocation_complete and the exit code agree (non-zero).
  • For Google/Vertex providers, no outgoing request has a trailing assistant message without a tool call (unit test on transform.ts).
  • The auxiliary empty-contents request is not counted by the retry logic and is sent unmodified (test).
  • NDJSON records the retry (reason, attempt) so aictrl can measure it (reuse the diagnostics shape from Preserve provider finish reasons and bounded diagnostics #109 if merged first).
  • Released as a new @aictrl/cli version; aictrl #5645 bumps the pin and re-runs PR Explanation and Code Review on gemini-3.8-flash.

Out of Scope

Roadmap Alignment

  • Pillar: EXEC (aictrl executor harness)
  • Quarter: Q4 2026 (Growth Engine)
  • Theme fit: Gemini 3.8 Flash is the managed default model, so this failure hits new self-serve users first.
  • Decision gate impact: indirect: "self-serve signups converting" depends on first runs succeeding.

References

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

bugSomething isn't working

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions