Skip to content

feat(selfhost): add automatic dead-letter job retry for self-host queues #2534

Description

@JSONbored

Parent: #1936

Problem

Self-host's SQLite and Postgres queue backends have no mechanism to move a job out of status='dead' once it lands there. There is a dead-letter redrive path (src/queue/dlq.ts) but it is wired exclusively to the Cloudflare Workers Queues consumer entrypoint and is not reachable from the self-host queue implementations at all. In practice this means: a job that dies from a transient or since-fixed bug (a bad deploy, a backend-compatibility gap, a flaky third-party dependency) requires a manual database UPDATE or DELETE to ever run again — indefinitely, even after the underlying cause is resolved and redeployed.

Requirements

  • Add a bounded, safe automatic retry path for dead-lettered self-host jobs: e.g. a scheduled maintenance pass that re-queues dead jobs whose job_key/type hasn't already been retried within some cooldown window, capped to a small number of automatic attempts before requiring manual intervention again.
  • Must not create a retry storm: a job that fails the SAME way immediately after requeue should not be retried indefinitely — a hard ceiling on auto-retries per job is required, after which it stays dead and relies on the alerting from the sibling alerting issue.
  • Should reuse existing queue primitives (attempts counter, job_key) rather than inventing a parallel bookkeeping table.

Deliverables

  • A retry/redrive mechanism in src/selfhost/sqlite-queue.ts and src/selfhost/pg-queue.ts (or a shared helper both call into), gated by a bounded auto-retry ceiling and cooldown.
  • Metrics distinguishing an auto-recovered job from a fresh enqueue.
  • Tests covering: a dead job recovers automatically after its cooldown, a job that dies repeatedly stops auto-retrying at the ceiling, and the ceiling/cooldown are configurable.

Acceptance criteria

  • A job dead-lettered by a bug that gets fixed and redeployed recovers automatically within the next maintenance cycle, without manual database intervention.
  • A job that is genuinely broken (fails identically every time) does not retry forever and does not starve other queue throughput.
  • No regression to existing dead-letter counting/alerting.

Expected outcome

Self-host operators don't need direct database access to recover from a transient bug that's already been fixed upstream.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

maintainer-onlyOwner-only work — yields no Gittensor points.roadmapOn the Wave-2 agent-layer roadmap board (project 9)

Projects

Relationships

None yet

Development

No branches or pull requests

Issue actions