Skip to content

[Go plan] Empty completions (finish_reason=other) and mid-stream cutoffs on long-reasoning requests — reproduced via direct API calls across multiple models #44148

Description

@Duanzhoutao

Description

Severity: high — this breaks the Go plan's core use case during affected windows. Please treat it with priority.

On the Go plan (https://opencode.ai/zen/go/v1), requests that trigger extended reasoning (30–110s of thinking) frequently come back empty: the model streams a large reasoning payload, then the stream ends with zero text output and zero tool calls. Logged 18 times across 3 models over 4 days (deepseek-v4-flash ×10, deepseek-v4-flash-vision-exp ×7, glm-5.3 ×1), and reproduced with plain HTTPS calls directly against the API (no coding tool in the loop), which rules out client-side causes.

Two failure shapes, same user-visible result:

  1. Complete stream, empty completion — the stream arrives intact including a finish chunk, but finish_reason: "other", 0 content chars, 0 tool calls, after up to ~22k chars of reasoning.
  2. Mid-stream cutoff — the reasoning stream is terminated by a clean FIN (not RST) after 30–109s: no finish chunk, no [DONE], zero text.

Reproduction via direct API calls (Aug 22, 10:42–10:52 UTC; same request body, stream: true, max_tokens: 32000, a single user message asking for deep step-by-step reasoning):

Egress path Requests Outcome
Direct HTTPS from Node.js (no proxy, no gateway, no coding tool) 6 5 cut off mid-reasoning (clean FIN, no finish chunk, no [DONE], 0 text)
Via local CONNECT proxy 6 6 cut off, identical signature
Short-prompt control (same connection setup) 3 3/3 complete: text + finish_reason=stop + [DONE]

11 of 12 consecutive long-reasoning probes failed within a ~10-minute window, while short prompts over the same connections worked 3/3 — so the termination happens server-side (the Zen routing layer or the model backend, most likely an idle/stall timeout on in-flight generations), not in the client, local network, or proxy.

Impact: the empty result arrives inside a 2xx "successful" stream, so clients classify it as a provider-level result and do not retry (observed in our client: 0 retries, attempt=1/maxAttempts=11; the whole turn fails). User-visible behavior: the model "thinks" for 35–60 seconds, then returns nothing. Long-reasoning requests are the norm for agent/coding workloads — the Go plan's core use case — so this hits exactly the traffic the plan is sold for.

Not a duplicate of #43565: that issue is one specific model (gpt-5.6-luna) failing consistently (empty completion on stream=false, 403 on stream=true) while other models work — a model-specific routing/adapter problem. This issue is multi-model, intermittent, time-dependent (works during some hours, ~90% failure during others), and about server-side termination of in-flight long-reasoning streams. Both are real, but they are different failure modes; cross-referencing #43565 here.

Related upstream reports: #33714 ("Upstream idle timeout exceeded"), #37852, #40010, #38541 (the latter two are client-side mitigations, not root-cause fixes).

Ask:

  1. Investigate the server-side handling of long-running generations on the Go plan — specifically whether an idle/stall timeout between the Zen gateway and the model backends terminates in-flight reasoning streams (the clean-FIN cutoff shape).
  2. If generation is aborted server-side, return a distinguishable, documented signal (SSE error event or a semantic finish_reason) instead of a 2xx stream with an empty or other completion, so clients can retry meaningfully.

Happy to share the probe script or additional logs if useful.

Plugins

None — verified with raw HTTPS calls and a third-party OpenAI-compatible client (ZCode); no OpenCode CLI/server plugins involved.

OpenCode version

n/a — reproduced via direct HTTP against https://opencode.ai/zen/go/v1/chat/completions, no CLI/server in the loop.

Steps to reproduce

  1. POST https://opencode.ai/zen/go/v1/chat/completions with Authorization: Bearer $GO_API_KEY, "stream": true, "max_tokens": 32000, and a single user message asking for deep step-by-step reasoning (a long-reasoning prompt; short prompts complete fine and were used as controls).
  2. Repeat ~6 times with a 2s pause.
  3. Expected: a reasoning stream that ends with finish_reason=stop and text content.
  4. Actual during affected windows: 5–6/6 requests stream reasoning for 30–109s then end early — either a clean FIN with no finish chunk and no [DONE], or a complete stream ending finish_reason: "other" with 0 text and 0 tool calls. HTTP status stays 200 in both cases.

Screenshot and/or share link

n/a

Operating System

Windows 11

Terminal

n/a

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions