fix: manage staging health pings with Cloudflare (#1009) - #673
Conversation
There was a problem hiding this comment.
✅ No new issues found.
Reviewed changes Reviewed all five changed files, the existing staging keepalive and Cloudflare configuration, and the pinned provider’s module-upload behavior.
- Managed scheduler: Terraform adds an isolated Cloudflare Worker and two five-minute UTC cron schedules covering the Sao Paulo review window, with persisted logs.
- Health requests: The scheduled handler restricts requests to staging, checks scheduled and execution windows, bypasses cache, rejects redirects, enforces a 90-second timeout, and records success or failure.
- Verification and cutover: Fourteen new unit tests join the infrastructure CI suite; deployment instructions retain GitHub keepalive until live cadence and latency evidence is recorded.
All 26 infrastructure unit tests passed in this review run. Cloudflare runtime documentation and provider 5.26.0 source support the interfaces used. Live deployment, cron delivery, and staging latency were not tested here; those remain explicit operator verification steps in the PR.
openai/gpt-6.1-sol | 𝕏
The pinger's node tests ran in the Terraform workflow only, so SonarCloud read 0% coverage on the new Worker. Run them with coverage beside the web plan guard and pass both lcov reports. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
There was a problem hiding this comment.
✅ No new issues found.
Reviewed changes Reviewed the SonarCloud coverage update since the previous Pullfrog review at c9b57dd2, with the full PR diff for context.
- Added pinger coverage: Ran the existing staging-pinger tests with coverage scoped to
infra/workers/staging-pinger.mjsand a dedicated LCOV output. - Expanded coverage import: Passed both infrastructure LCOV reports to SonarScanner while preserving the existing web-plan coverage command.
All 26 infrastructure unit tests passed. The new workflow coverage command also succeeded and produced the correct source path with 50/50 lines and 14/14 branches covered. Remote SonarCloud analysis was still running when checked. Worker runtime behavior and deployment requirements are unchanged; live cadence and latency proof remain operator steps.
openai/gpt-6.1-sol | 𝕏
|
* fix: manage staging health pings with Cloudflare (#1009) (#673) * fix: add Cloudflare staging health pinger for ticket 1009 * ci: report staging pinger coverage to SonarCloud The pinger's node tests ran in the Terraform workflow only, so SonarCloud read 0% coverage on the new Worker. Run them with coverage beside the web plan guard and pass both lcov reports. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> (cherry picked from commit d879947) * fix: isolate staging Google Play package (#674) (cherry picked from commit 2b50ca8) * fix: correct Stripe checkout success routes (#675) * fix: correct Stripe checkout success routes * fix: bind Stripe session idempotency keys to request inputs (cherry picked from commit 6444752) * fix: send the staging pinger with manual redirects Cloudflare's runtime rejects redirect: "error" when it builds the Request, so every scheduled ping failed before reaching staging /health. Manual mode is accepted, and the existing response.ok check still fails the invocation on a 3xx response, now covered by its own test. Refs thomasluizon/orbit-tickets#1009 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> (cherry picked from commit 819165f) --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>




GitHub's scheduled keepalive can leave the staging API asleep during the review window. This adds a Terraform-managed Cloudflare Worker with five-minute UTC cron triggers covering 08:00 to 24:00 America/Sao_Paulo and persisted run logs.
Refs thomasluizon/orbit-tickets#1009.
Changes
infra/staging-pinger.tfdeclares only the Worker script and its Cron Trigger resource. Both use the existing account ID and have no dependency on DNS, web services, or either API service.infra/workers/staging-pinger.mjscalls only staging/health, checks the scheduled and execution windows, bypasses cache, rejects redirects, and aborts after 90 seconds. Failures reject the scheduled invocation; structured logs retain timestamps, outcome, HTTP status when available, and duration.infra/staging-pinger.test.mjschecks all 192 daily cron slots, UTC boundaries, delayed events, request destination and options, successful logging, HTTP and network failures, timeout, and timer cleanup..github/workflows/terraform.ymlruns it with the existing infrastructure tests.infra/README.mddocuments targeted deployment, the exact Cloudflare screens for triggers and run history, and seven independent probes over one hour. The GitHub keepalive workflow remains until the orchestrator completes deployment and live proof, as the work order's final note requires.Assumptions
cloudflare.tf, so its targeted deployment stays independent of zone and Render resources.External interface evidence
infra/.terraform.lock.hclpinscloudflare/cloudflareto 5.26.0. Downloaded those locked providers withterraform init -backend=false -input=false -lockfile=readonlyin an isolated copy of the Terraform root. For schema inspection only, removed the S3 backend block from that temporary copy, then ranterraform providers schema -json. No real state or credentials were read.The installed schema confirms
cloudflare_workers_script.account_idand.script_nameas required strings;.contentand.main_moduleas optional strings;.compatibility_dateas an optional/computed string;.observability.enabledas required boolean,.head_sampling_rateas optional number, and.logs.enabled/.invocation_logsas required booleans with optional.head_sampling_rateand optional/computed.persist. It confirmscloudflare_workers_cron_trigger.account_idand.script_nameas required strings and.schedulesas a required list with required string.cronand computed.created_on/.modified_onstrings. Only the required cron field is supplied.The versioned Worker schema and Cron Trigger schema provide a way to recheck the arguments. The cron example uses
body, but the installed schema requiresschedules; the implementation follows the schema.The scheduled handler contract defines
scheduledTimeas UTC epoch milliseconds. The Response contract defines the consumedstatus,ok, andtext()interfaces. The Request contract confirmsno-store, redirect rejection, and an AbortSignal. Response and request tests use nativeResponseandRequestobjects rather than invented response fields. Curl's probe output format was checked through a real request to the Cloudflare Response documentation, which returnedhttp_status=200 duration_seconds=0.237335; this was not a staging acceptance probe.Test evidence
node --test infra/check-web-plan.test.mjs infra/ses-isolation.test.mjspassed all 12 unchanged tests with the unreliable GitHub workflow present. They do not exercise scheduler delivery. A strengthened pre-fix test reproducing dropped GitHub runs was not obtained: this is external scheduler behavior, and deployment and live observation belong to the orchestrator. No unit test result is claimed as proof of Cloudflare delivery cadence.node --test infra/staging-pinger.test.mjs: 14 passed.node --test infra/check-web-plan.test.mjs infra/ses-isolation.test.mjs infra/staging-pinger.test.mjs: 26 passed.terraform fmt -check -recursive infra: passed.terraform validatein the isolated provider-initialized root containing the new files: passed.dotnet build Orbit.slnx: zero errors, 13 warnings in unchanged C# projects.dotnet testrun had 7,032 passes and one failure in unchangedReminderSchedulerServiceTests.CheckAndSendReminders_TwoSameDayScheduledReminders_PersistsBothWithoutUniqueViolation. It ran during the first UTC minute, before its 00:01 reminder was due, and observed one reminder instead of two. That test and the scheduler have no diff againstmain.dotnet test tests/Orbit.Infrastructure.Tests --no-build --filter FullyQualifiedName~CheckAndSendReminders_TwoSameDayScheduledReminders_PersistsBothWithoutUniqueViolationpassed unchanged after 00:01 UTC.dotnet testrerun: 7,033 passed, zero failed, zero skipped, with no source or test changes between runs.The unrelated reminder-test finding needs its own ticket.
tools/create-ticket.mjs, the only filing interface authorized by this work order, is absent from this checkout; no alternative issue-creation interface was used.c9b57dd2failed only on new-code coverage: the pinger's node tests ran interraform.yml, butsonarcloud.ymlpassed onlycoverage/web-plan.lcov. The workflow now also runsinfra/staging-pinger.test.mjswith--experimental-test-coveragescoped toinfra/workers/staging-pinger.mjsand passescoverage/staging-pinger.lcov, the same way it already measuresinfra/check-web-plan.mjs.infra/workers/staging-pinger.mjsreports 50 of 50 lines and 14 of 14 branches.Manual steps
CLOUDFLARE_API_TOKENfrom Keychain entryorbit-cloudflare-api-tokenwithout printing it. The Cloudflare account token must permit Workers Scripts Write for account29945c90bc934c629c8e5a11cbfd146b. Initialize the existinginfra/backend and use the existinginfra/local.tfvars. Plan withterraform -chdir=infra plan -var-file=local.tfvars -target=cloudflare_workers_script.staging_pinger -target=cloudflare_workers_cron_trigger.staging_pinger. Proceed only if the plan changes those two pinger resources. Apply withterraform -chdir=infra apply -var-file=local.tfvars -target=cloudflare_workers_script.staging_pinger -target=cloudflare_workers_cron_trigger.staging_pinger. This worker did not run a real-state plan or apply.*/5 11-23 * * *and*/5 0-2 * * *. Allow trigger propagation, then use the Worker's Observability screen to prove scheduled and actual invocations every five minutes for the observation hour.infra/README.mdto record seven staging/healthprobes ten minutes apart over at least one hour entirely inside the window. Record each timestamp, HTTP status, curl exit code, and duration. Every probe must succeed in strictly under two seconds. Record intervening Worker runs as well, because independent probes and the retained GitHub workflow can themselves keep staging awake. Live proof remains pending..github/workflows/staging-keepalive.ymland commits that removal. Its retention in this PR is deliberate and required by the work order's final deployment note.