Skip to content

Apply logs timeout and count globally across concurrent targets - #60058

Merged
pelikhan merged 6 commits into
mainfrom
copilot/review-logs-command-timeout
Sep 10, 2026
Merged

pelikhan merged 6 commits into
mainfrom
copilot/review-logs-command-timeout

Conversation

Copilot AI commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor

Multi-target logs downloads applied --timeout and --count independently per target. These limits should bound the combined concurrent operation.

  • Timeout

    • Use one deadline for the full multi-target download.
    • Prevent queued targets from receiving a fresh timeout.
    • Keep final API operations within the same deadline.
  • Count

    • Share one atomic run budget across targets.
    • Stop collection when the global count is reached.
    • Globally cap and sort merged results.
  • Partial results

    • Preserve runs collected before a target fails or reaches a limit.
    • Emit resumable continuations for queued date-range targets that make no progress.
  • Documentation

    • Clarify global multi-target semantics in CLI help and reference documentation.
gh aw logs workflow-a workflow-b --count 100 --timeout 10

This returns up to 100 runs combined and limits the overall download to 10 minutes.

Copilot AI and others added 2 commits September 10, 2026 21:31
Co-authored-by: pelikhan <4175913+pelikhan@users.noreply.github.com>
Co-authored-by: pelikhan <4175913+pelikhan@users.noreply.github.com>
Copilot AI changed the title Apply logs timeout and count globally across targets Apply logs timeout and count globally across concurrent targets Sep 10, 2026
Copilot AI requested a review from pelikhan September 10, 2026 21:42
@pelikhan
pelikhan marked this pull request as ready for review September 10, 2026 21:42
Copilot AI balanced review requested due to automatic review settings September 10, 2026 21:42
@github-actions

github-actions Bot commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor

✅ Ponytail Reviewer completed successfully!

Warning

Firewall blocked 1 domain

The following domain was blocked by the firewall during workflow execution:

  • ab.chatgpt.com

To allow these domains, add them to the network.allowed list in your workflow frontmatter:

network:
  allowed:
    - defaults
    - "ab.chatgpt.com"

See Network Configuration for more information.

Generated by Ponytail Reviewer for #60058

@github-actions

github-actions Bot commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor

⚠️ PR Code Quality Reviewer failed during code quality review.

Warning

Threat Detection Engine Failure — The analysis engine could not complete. This is a tooling failure, not a security finding.

What happened

The threat detection engine failed to produce results.

Review the workflow run logs for details.

🔎 Code quality review by PR Code Quality Reviewer

@github-actions

github-actions Bot commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor

✅ Test Quality Sentinel completed test quality analysis.

Test Quality Sentinel skipped because pre-fetch PR data was unavailable: unable to fetch test file diff

🧪 Test quality analysis by Test Quality Sentinel

@github-actions

github-actions Bot commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor

⚠️ Security scanning failed for Design Decision Gate 🏗️. Review the logs for details.

🏗️ ADR gate enforced by Design Decision Gate 🏗️

@github-actions

github-actions Bot commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor

🧠 Matt Pocock Skills Reviewer failed during the skills-based review.

Warning

Threat Detection Engine Failure — The analysis engine could not complete. This is a tooling failure, not a security finding.

What happened

The threat detection engine failed to produce results.

Review the workflow run logs for details.

🧠 Reviewed using Matt Pocock's skills by Matt Pocock Skills Reviewer

@pelikhan

Copy link
Copy Markdown
Collaborator

@copilot make the sure cancelation context is propagated correctly through all the layers and workers

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ponytail review: a few opportunities to trim wrapper and duplicate-limit plumbing without changing behavior.

net: -12 lines possible.

Warning

Firewall blocked 1 domain

The following domain was blocked by the firewall during workflow execution:

  • ab.chatgpt.com

To allow these domains, add them to the network.allowed list in your workflow frontmatter:

network:
  allowed:
    - defaults
    - "ab.chatgpt.com"

See Network Configuration for more information.

Generated by ✂️ Ponytail Reviewer for #60058 · codex · gpt53codex · 3.13 AIC · ⌖ 4.12 AIC · ⊞ 12.8K
Comment /ponytail to run again

return shouldStopAfterWorkflowRunBatch(batch, opts.Verbose), nil
}

func markSharedLogsCountReached(state *logsCollectionState, fetchAllInRange bool, limit *logsCountLimit) bool {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

L340: shrink: markSharedLogsCountReached only wraps limit.isReached() and assignment. Inline the two lines at call sites; delete helper.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Keeping markSharedLogsCountReached as a helper: it's called from two sites (collectProcessedWorkflowRuns and fetchAndProcessLogsBatch/finishLogsBatch), so inlining would duplicate the isReached()/assignment logic rather than simplify it.

if len(processedRuns) >= opts.count {
continue
}
if !opts.countLimit.tryAdd() {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

L582: shrink: double-guard (len(processedRuns) < opts.count and countLimit.tryAdd) duplicates budget checks. Keep shared countLimit.tryAdd gate only; remove local length check.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not removing the local length check: opts.countLimit is nil in single-target mode (only set via newLogsCountLimit in the multi-target path), and countLimit.tryAdd() returns true unconditionally for a nil receiver. The len(processedRuns) < opts.count guard is the only enforcement of --count for single-target downloads, so dropping it would break that case.

rateLimitFirstRequest bool
maxConcurrentDownloads int
storageLimit *logsStorageLimit
countLimit *logsCountLimit

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

L64: yagni: carrying both Count and countLimit in options creates parallel limit mechanisms. Keep one source of truth (countLimit) and derive Count at boundary only.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Keeping both fields: Count is the public per-command option set for every download (including single-target, where countLimit stays nil), while countLimit is an internal cross-goroutine atomic budget only allocated for multi-target runs. They aren't parallel copies of the same limit — countLimit doesn't exist unless multiple targets share one budget.

@github-actions

Copy link
Copy Markdown
Contributor
🏗️ ADR required — draft added for PR #60058

I enforced the design-decision gate for this PR because it adds more than 100 lines in business-logic directories (pkg/).

Evidence used

  • Prefetch summary: default_business_additions = 192, has_implementation_label = false, requires_adr_by_default_volume = true
  • PR title: Apply logs timeout and count globally across concurrent targets
  • Diff evidence:
    • adds a shared wall-clock timeout context for multi-target logs downloads
    • adds a shared atomic count limiter across concurrent targets
    • preserves partial results and resumable continuations when shared limits or failures stop progress
    • updates CLI/docs text so --count and --timeout are global across all targets

Outcome

I added a draft ADR at:

  • docs/adr/60058-apply-global-limits-to-multi-target-logs.md

Next action

Please review and refine the draft ADR so the decision rationale and trade-offs match maintainer intent before merge.

🏗️ ADR gate enforced by Design Decision Gate 🏗️ · pi · gpt54 · 14.6 AIC · ⊞ 10.1K · ◷
Comment /review to run again

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Global selection and continuation handling have correctness gaps under concurrent limits.

Once you've addressed the issues Copilot identified, you can request another Copilot review.

Pull request overview

Applies count and timeout limits globally across concurrent multi-target log downloads.

Changes:

  • Adds shared deadline and atomic run-count budgeting.
  • Preserves partial results and adds continuation handling.
  • Updates tests, CLI help, and documentation.
File summaries
File Description
pkg/cli/logs_orchestrator.go Preserves partial runs and expands continuation behavior.
pkg/cli/logs_orchestrator_unit_test.go Tests zero-progress count continuations.
pkg/cli/logs_orchestrator_types.go Adds shared count-limit state.
pkg/cli/logs_orchestrator_download.go Enforces the shared count budget during collection.
pkg/cli/logs_multi.go Coordinates global timeout, count, sorting, and partial results.
pkg/cli/logs_multi_test.go Tests global limits and partial-result merging.
pkg/cli/logs_command.go Clarifies global flag semantics.
docs/src/content/docs/setup/cli.md Documents multi-target limits.
Review details
  • Files reviewed: 9/9 changed files
  • Comments generated: 3
  • Review effort level: Balanced

💡 Add a code-review agent skill for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +582 to +583
if !opts.countLimit.tryAdd() {
continue

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is a real trade-off, not a bug we're silently accepting: enforcing strict CreatedAt ordering across concurrent targets would require buffering/merging candidate metadata before any artifact download starts, which conflicts with the concurrent early-stop design this PR relies on to bound total work under the shared timeout. That trade-off (global admission-order fairness vs. concurrent early-stopping) is captured in docs/adr/60058-apply-global-limits-to-multi-target-logs.md for maintainer review. Not changing the admission strategy in this pass; happy to revisit if the ADR discussion lands on stricter ordering.

Comment thread pkg/cli/logs_multi.go
Comment on lines +82 to +83
allAPIRateLimits := startGitHubAPIRateLimitReports(activeCtx, logsTargetRateLimitHosts(targets))
results := collectLogsTargets(activeCtx, opts, targets)

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in b56d1da: a target still waiting on the semaphore when the shared context is canceled no longer returns a bare error. queuedLogsTargetResult builds a resumable continuation from the target's own options (preserving its cursor and timeout flag) instead. Covered by TestCollectLogsTargetsEmitsContinuationForQueuedTarget.

Comment thread pkg/cli/logs_orchestrator_download.go Outdated
Comment on lines +329 to +330
if markSharedLogsCountReached(state, runtime.fetchAllInRange, opts.countLimit) {
return true, nil

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in e572154: the allRunsConsumed cursor advance now runs before the shared count-limit check in fetchAndProcessLogsBatch (extracted into finishLogsBatch), so a continuation built after the shared budget is exhausted mid-batch resumes from the correct pagination boundary instead of re-scanning the batch. Added TestFetchAndProcessLogsBatchAdvancesCursorWhenSharedCountLimitReached to cover it.

@pelikhan

Copy link
Copy Markdown
Collaborator

@copilot run pr-finisher skill

Co-authored-by: pelikhan <4175913+pelikhan@users.noreply.github.com>

Copilot AI commented Sep 10, 2026

Copy link
Copy Markdown
Contributor Author

@copilot make the sure cancelation context is propagated correctly through all the layers and workers

Fixed in 30ddcb5 (or the latest push): a target that is still waiting on the worker semaphore when the shared deadline/cancellation fires no longer returns a bare ctx.Err(). It now builds a resumable continuation from its own options (preserving its BeforeRunID cursor) so it isn't silently dropped from the report. Also relaxed buildContinuationIfNeeded so a target that made zero progress can still emit that continuation. Verified downstream API calls (gh CLI invocations, rate-limit checks, artifact download pool) already thread the same context/deadline through; added TestCollectLogsTargetsEmitsContinuationForQueuedTarget to cover the queued-target path.

Co-authored-by: pelikhan <4175913+pelikhan@users.noreply.github.com>
@github-actions

Copy link
Copy Markdown
Contributor

🎉 This pull request is included in a new release.

Release: v0.89.3

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants