Skip to content

fix: manage staging health pings with Cloudflare (#1009) - #673

Merged
thomasluizon merged 2 commits into
mainfrom
fix/ticket-1009-staging-pinger
Oct 1, 2026
Merged

thomasluizon merged 2 commits into
mainfrom
fix/ticket-1009-staging-pinger

Conversation

@thomasluizon

@thomasluizon thomasluizon commented Oct 1, 2026 •

Copy link
Copy Markdown
Owner

GitHub's scheduled keepalive can leave the staging API asleep during the review window. This adds a Terraform-managed Cloudflare Worker with five-minute UTC cron triggers covering 08:00 to 24:00 America/Sao_Paulo and persisted run logs.

Refs thomasluizon/orbit-tickets#1009.

Changes

  • infra/staging-pinger.tf declares only the Worker script and its Cron Trigger resource. Both use the existing account ID and have no dependency on DNS, web services, or either API service.
  • infra/workers/staging-pinger.mjs calls only staging /health, checks the scheduled and execution windows, bypasses cache, rejects redirects, and aborts after 90 seconds. Failures reject the scheduled invocation; structured logs retain timestamps, outcome, HTTP status when available, and duration.
  • infra/staging-pinger.test.mjs checks all 192 daily cron slots, UTC boundaries, delayed events, request destination and options, successful logging, HTTP and network failures, timeout, and timer cleanup. .github/workflows/terraform.yml runs it with the existing infrastructure tests.
  • infra/README.md documents targeted deployment, the exact Cloudflare screens for triggers and run history, and seven independent probes over one hour. The GitHub keepalive workflow remains until the orchestrator completes deployment and live proof, as the work order's final note requires.

Assumptions

  • Keep the pinger in its own Terraform file and Worker module rather than expanding cloudflare.tf, so its targeted deployment stays independent of zone and Render resources.
  • Use a scheduled handler without an HTTP route rather than exposing a manual ping endpoint; a Cron Trigger does not need a route.
  • Keep the existing keepalive's 90-second request allowance rather than applying the two-second acceptance threshold to initial wake-up requests; independent acceptance probes still require strictly under two seconds.
  • Check actual execution time as well as scheduled time rather than allowing a delayed event to wake staging outside the specified window.

External interface evidence

infra/.terraform.lock.hcl pins cloudflare/cloudflare to 5.26.0. Downloaded those locked providers with terraform init -backend=false -input=false -lockfile=readonly in an isolated copy of the Terraform root. For schema inspection only, removed the S3 backend block from that temporary copy, then ran terraform providers schema -json. No real state or credentials were read.

The installed schema confirms cloudflare_workers_script.account_id and .script_name as required strings; .content and .main_module as optional strings; .compatibility_date as an optional/computed string; .observability.enabled as required boolean, .head_sampling_rate as optional number, and .logs.enabled/.invocation_logs as required booleans with optional .head_sampling_rate and optional/computed .persist. It confirms cloudflare_workers_cron_trigger.account_id and .script_name as required strings and .schedules as a required list with required string .cron and computed .created_on/.modified_on strings. Only the required cron field is supplied.

The versioned Worker schema and Cron Trigger schema provide a way to recheck the arguments. The cron example uses body, but the installed schema requires schedules; the implementation follows the schema.

The scheduled handler contract defines scheduledTime as UTC epoch milliseconds. The Response contract defines the consumed status, ok, and text() interfaces. The Request contract confirms no-store, redirect rejection, and an AbortSignal. Response and request tests use native Response and Request objects rather than invented response fields. Curl's probe output format was checked through a real request to the Cloudflare Response documentation, which returned http_status=200 duration_seconds=0.237335; this was not a staging acceptance probe.

Test evidence

  • Before implementation, node --test infra/check-web-plan.test.mjs infra/ses-isolation.test.mjs passed all 12 unchanged tests with the unreliable GitHub workflow present. They do not exercise scheduler delivery. A strengthened pre-fix test reproducing dropped GitHub runs was not obtained: this is external scheduler behavior, and deployment and live observation belong to the orchestrator. No unit test result is claimed as proof of Cloudflare delivery cadence.
  • node --test infra/staging-pinger.test.mjs: 14 passed.
  • node --test infra/check-web-plan.test.mjs infra/ses-isolation.test.mjs infra/staging-pinger.test.mjs: 26 passed.
  • terraform fmt -check -recursive infra: passed. terraform validate in the isolated provider-initialized root containing the new files: passed.
  • dotnet build Orbit.slnx: zero errors, 13 warnings in unchanged C# projects.
  • The first dotnet test run had 7,032 passes and one failure in unchanged ReminderSchedulerServiceTests.CheckAndSendReminders_TwoSameDayScheduledReminders_PersistsBothWithoutUniqueViolation. It ran during the first UTC minute, before its 00:01 reminder was due, and observed one reminder instead of two. That test and the scheduler have no diff against main. dotnet test tests/Orbit.Infrastructure.Tests --no-build --filter FullyQualifiedName~CheckAndSendReminders_TwoSameDayScheduledReminders_PersistsBothWithoutUniqueViolation passed unchanged after 00:01 UTC.
  • Full dotnet test rerun: 7,033 passed, zero failed, zero skipped, with no source or test changes between runs.
  • Commit hooks passed; no DTO, minimum-version, user-date, notification, or application module changes.
    The unrelated reminder-test finding needs its own ticket. tools/create-ticket.mjs, the only filing interface authorized by this work order, is absent from this checkout; no alternative issue-creation interface was used.
  • SonarCloud at c9b57dd2 failed only on new-code coverage: the pinger's node tests ran in terraform.yml, but sonarcloud.yml passed only coverage/web-plan.lcov. The workflow now also runs infra/staging-pinger.test.mjs with --experimental-test-coverage scoped to infra/workers/staging-pinger.mjs and passes coverage/staging-pinger.lcov, the same way it already measures infra/check-web-plan.mjs.
  • Local run of that exact command: exit 0; lcov for infra/workers/staging-pinger.mjs reports 50 of 50 lines and 14 of 14 branches.

Manual steps

  1. After review, the orchestrator loads CLOUDFLARE_API_TOKEN from Keychain entry orbit-cloudflare-api-token without printing it. The Cloudflare account token must permit Workers Scripts Write for account 29945c90bc934c629c8e5a11cbfd146b. Initialize the existing infra/ backend and use the existing infra/local.tfvars. Plan with terraform -chdir=infra plan -var-file=local.tfvars -target=cloudflare_workers_script.staging_pinger -target=cloudflare_workers_cron_trigger.staging_pinger. Proceed only if the plan changes those two pinger resources. Apply with terraform -chdir=infra apply -var-file=local.tfvars -target=cloudflare_workers_script.staging_pinger -target=cloudflare_workers_cron_trigger.staging_pinger. This worker did not run a real-state plan or apply.
  2. In Cloudflare Workers & Pages > orbit-staging-pinger > Settings > Triggers > Cron Triggers, confirm */5 11-23 * * * and */5 0-2 * * *. Allow trigger propagation, then use the Worker's Observability screen to prove scheduled and actual invocations every five minutes for the observation hour.
  3. Follow infra/README.md to record seven staging /health probes ten minutes apart over at least one hour entirely inside the window. Record each timestamp, HTTP status, curl exit code, and duration. Every probe must succeed in strictly under two seconds. Record intervening Worker runs as well, because independent probes and the retained GitHub workflow can themselves keep staging awake. Live proof remains pending.
  4. Only after deployment and both proofs pass, the orchestrator deletes .github/workflows/staging-keepalive.yml and commits that removal. Its retention in this PR is deliberate and required by the work order's final deployment note.

pullfrog[bot]
pullfrog Bot previously approved these changes Oct 1, 2026

@pullfrog pullfrog Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ No new issues found.

Reviewed changes Reviewed all five changed files, the existing staging keepalive and Cloudflare configuration, and the pinned provider’s module-upload behavior.

  • Managed scheduler: Terraform adds an isolated Cloudflare Worker and two five-minute UTC cron schedules covering the Sao Paulo review window, with persisted logs.
  • Health requests: The scheduled handler restricts requests to staging, checks scheduled and execution windows, bypasses cache, rejects redirects, enforces a 90-second timeout, and records success or failure.
  • Verification and cutover: Fourteen new unit tests join the infrastructure CI suite; deployment instructions retain GitHub keepalive until live cadence and latency evidence is recorded.

All 26 infrastructure unit tests passed in this review run. Cloudflare runtime documentation and provider 5.26.0 source support the interfaces used. Live deployment, cron delivery, and staging latency were not tested here; those remain explicit operator verification steps in the PR.

Pullfrog  | View workflow run | Using openai/gpt-6.1-sol | 𝕏

The pinger's node tests ran in the Terraform workflow only, so SonarCloud read
0% coverage on the new Worker. Run them with coverage beside the web plan
guard and pass both lcov reports.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@pullfrog pullfrog Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ No new issues found.

Reviewed changes Reviewed the SonarCloud coverage update since the previous Pullfrog review at c9b57dd2, with the full PR diff for context.

  • Added pinger coverage: Ran the existing staging-pinger tests with coverage scoped to infra/workers/staging-pinger.mjs and a dedicated LCOV output.
  • Expanded coverage import: Passed both infrastructure LCOV reports to SonarScanner while preserving the existing web-plan coverage command.

All 26 infrastructure unit tests passed. The new workflow coverage command also succeeded and produced the correct source path with 50/50 lines and 14/14 branches covered. Remote SonarCloud analysis was still running when checked. Worker runtime behavior and deployment requirements are unchanged; live cadence and latency proof remain operator steps.

Pullfrog  | View workflow run | Using openai/gpt-6.1-sol | 𝕏

@sonarqubecloud

sonarqubecloud Bot commented Oct 1, 2026

Copy link
Copy Markdown

@thomasluizon
thomasluizon merged commit d879947 into main Oct 1, 2026
27 checks passed
@thomasluizon
thomasluizon deleted the fix/ticket-1009-staging-pinger branch October 1, 2026 00:56
thomasluizon added a commit that referenced this pull request Oct 1, 2026
* fix: manage staging health pings with Cloudflare (#1009) (#673)

* fix: add Cloudflare staging health pinger for ticket 1009

* ci: report staging pinger coverage to SonarCloud

The pinger's node tests ran in the Terraform workflow only, so SonarCloud read
0% coverage on the new Worker. Run them with coverage beside the web plan
guard and pass both lcov reports.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit d879947)

* fix: isolate staging Google Play package (#674)

(cherry picked from commit 2b50ca8)

* fix: correct Stripe checkout success routes (#675)

* fix: correct Stripe checkout success routes

* fix: bind Stripe session idempotency keys to request inputs

(cherry picked from commit 6444752)

* fix: send the staging pinger with manual redirects

Cloudflare's runtime rejects redirect: "error" when it builds the Request,
so every scheduled ping failed before reaching staging /health. Manual
mode is accepted, and the existing response.ok check still fails the
invocation on a 3xx response, now covered by its own test.

Refs thomasluizon/orbit-tickets#1009

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 819165f)

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant