Skip to content

Count duplicate tasks in the cost estimate [ENG-286] - #23

Open
emirlan-zen wants to merge 6 commits into
mainfrom
eng-286-duplicate-tasks-in-estimate
Open

emirlan-zen wants to merge 6 commits into
mainfrom
eng-286-duplicate-tasks-in-estimate

Conversation

@emirlan-zen

Copy link
Copy Markdown
Contributor

The client-side half of ENG-286. Server side is ZenRows/conveyor#312.

Batch does not deduplicate URLs within a job: identical URLs are separate scrapes, separate charges and separate results. Idempotency-Key looks like it should cover it but hashes the whole request body and never inspects the task list. The estimate is the last point where a caller can see this before spending anything, so it now reports it.

What changed

CostEstimate.duplicate_tasks — how many tasks repeat a request an earlier task already makes, counted in the loop the estimate already runs. Duplicates are still priced like any other task; nothing is collapsed.

>>> client.estimate_cost(body).format()
4 tasks → 20 credits (4 tasks)
       4 x js_render (5) = 20
       2 duplicate tasks — scraped and charged separately

The key matches the server's: URL, canonical HTTP method, body, and merged zenrows_params (job-level overlaid with per-task). external_id and metadata are excluded — they don't change what is fetched or what it costs. URLs are compared exactly, with no normalisation. Booleans are coerced the way the server coerces them, so {"js_render": True} and {"js_render": "true"} are one request.

One deliberate divergence: the body is compared by JSON value rather than by the bytes the server will see, so two bodies differing only in key order count as one request here where the server may read them as distinct. That's the only case the two counts can disagree, and it errs toward telling the caller about redundant work. Documented on _request_key.

Spec + models

Only the ENG-286 delta is applied to docs/openapi.yaml, not a full resync from conveyor — the local copy is behind on unrelated rerun and export wording, and sweeping that in would ship other people's changes under this PR.

duplicate_tasks lands on SubmitJobResponse / AddTasksResponse as optional with default: 0 rather than required, so this parses responses from a server that predates the field — including an idempotency snapshot replayed from inside its 24h TTL. That means the two repos can ship in either order.

Verified

make check (ty + ruff) clean, 109 tests pass, 13 of them new.

Against live dev (async.api.zrdev.co, running conveyor#312): the regenerated SubmitJobResponse parses the real response, and the client-side estimate and the server's count independently return 2 for the same body — including a task whose only difference is external_id, and a task restating the job-level js_render as a boolean against the job's string.

🤖 Generated with Claude Code

asyschikov and others added 4 commits July 13, 2026 12:14
Add ZenRowsBatchClient — a synchronous, typed client for the async-job
Batch API — and migrate the package to a src/ layout.

- Typestate resource handles: JobRef/JobHandle, RunRef/RunHandle,
  ExportRef/ExportHandle (id-only refs vs loaded handles with .data),
  with job.run.* (current run) and job.schedule.* facets that
  disambiguate the two "pause" endpoints.
- submit_regular / submit_open / submit_scheduled, typed schedule
  builders (At/Rate/Calendar), CSV file-input upload, and cursor-free
  scanners (iter_jobs / iter_runs / iter_results).
- Waiters (run / ingest / export) with jittered backoff; transport-level
  retries of transient failures (429/502/503/504 + network errors) on
  idempotent requests, honoring Retry-After.
- Downloads pulled straight from each result's presigned result_url (no
  /content endpoint): bulk download_to_dir / download_to_memory,
  server-side zip download_all_results, and per-task
  download_task_to_file / download_task_to_memory.
- Webhook management (get/put/delete job webhook, test_webhook), HMAC key
  lifecycle, and offline cost estimation (client.estimate_cost).
- RFC 7807 error mapping to BatchAPIError; pydantic v2 models generated
  from the OpenAPI spec; generated markdown API reference and runnable
  examples.
Batch does not deduplicate. Identical URLs in one job are separate
scrapes, separate charges and separate results, and `Idempotency-Key`
does not help -- it deduplicates whole submits, never URLs inside one.
The estimate is the last place a caller can see that before spending
anything, so count it there.

`CostEstimate.duplicate_tasks` reports how many tasks repeat a request
an earlier task already makes, keyed the same way the server keys it:
URL, canonical method, body, and merged params, with `external_id` and
`metadata` excluded and no URL normalisation. Booleans are coerced the
way the server coerces them, so `js_render: True` and `"true"` are one
request. Duplicates are still priced like any other task -- nothing is
collapsed -- and `format()` names them only when there are some.

The body is compared by JSON value rather than by the bytes the server
will see, so two bodies differing only in key order count as one
request here where the server may read them as distinct. That is the
only case where the two counts can disagree, and it errs toward
telling the caller about redundant work.

Spec: only the ENG-286 delta is applied to the local openapi copy, not
a full resync -- the copy is behind conveyor on unrelated rerun and
export wording, and sweeping that in would ship other changes under
this one. `duplicate_tasks` lands on the response models as optional
with default 0, so this parses responses from a server that predates
the field.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@linear

linear Bot commented Sep 24, 2026

Copy link
Copy Markdown

ENG-286

emirlan-zen and others added 2 commits September 24, 2026 13:40
Conveyor no longer counts duplicates on the submit request path — a
100k upload takes the carrier path, where the 202 is written before
anything has read the task list, so the count is done by the ingest
worker instead. The number therefore lives on the run, not on the
submit response: present immediately for a small submit, null on a
202 until `ingest_status` is `done`.

Only the schema move is applied here. The client-side estimate is
unaffected: it holds the task list in memory, so it still reports
`duplicate_tasks` directly, and it remains the only place a caller can
see the number before spending anything.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The reference named SubmitJobResponse.duplicate_tasks, which stopped
existing when the count moved onto the run. Caught in review of
ZenRows/conveyor#312.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants