Count duplicate tasks in the cost estimate [ENG-286] - #23
Open
emirlan-zen wants to merge 6 commits into
Open
emirlan-zen wants to merge 6 commits into
emirlan-zen wants to merge 6 commits into
Conversation
Add ZenRowsBatchClient — a synchronous, typed client for the async-job Batch API — and migrate the package to a src/ layout. - Typestate resource handles: JobRef/JobHandle, RunRef/RunHandle, ExportRef/ExportHandle (id-only refs vs loaded handles with .data), with job.run.* (current run) and job.schedule.* facets that disambiguate the two "pause" endpoints. - submit_regular / submit_open / submit_scheduled, typed schedule builders (At/Rate/Calendar), CSV file-input upload, and cursor-free scanners (iter_jobs / iter_runs / iter_results). - Waiters (run / ingest / export) with jittered backoff; transport-level retries of transient failures (429/502/503/504 + network errors) on idempotent requests, honoring Retry-After. - Downloads pulled straight from each result's presigned result_url (no /content endpoint): bulk download_to_dir / download_to_memory, server-side zip download_all_results, and per-task download_task_to_file / download_task_to_memory. - Webhook management (get/put/delete job webhook, test_webhook), HMAC key lifecycle, and offline cost estimation (client.estimate_cost). - RFC 7807 error mapping to BatchAPIError; pydantic v2 models generated from the OpenAPI spec; generated markdown API reference and runnable examples.
Batch does not deduplicate. Identical URLs in one job are separate scrapes, separate charges and separate results, and `Idempotency-Key` does not help -- it deduplicates whole submits, never URLs inside one. The estimate is the last place a caller can see that before spending anything, so count it there. `CostEstimate.duplicate_tasks` reports how many tasks repeat a request an earlier task already makes, keyed the same way the server keys it: URL, canonical method, body, and merged params, with `external_id` and `metadata` excluded and no URL normalisation. Booleans are coerced the way the server coerces them, so `js_render: True` and `"true"` are one request. Duplicates are still priced like any other task -- nothing is collapsed -- and `format()` names them only when there are some. The body is compared by JSON value rather than by the bytes the server will see, so two bodies differing only in key order count as one request here where the server may read them as distinct. That is the only case where the two counts can disagree, and it errs toward telling the caller about redundant work. Spec: only the ENG-286 delta is applied to the local openapi copy, not a full resync -- the copy is behind conveyor on unrelated rerun and export wording, and sweeping that in would ship other changes under this one. `duplicate_tasks` lands on the response models as optional with default 0, so this parses responses from a server that predates the field. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Conveyor no longer counts duplicates on the submit request path — a 100k upload takes the carrier path, where the 202 is written before anything has read the task list, so the count is done by the ingest worker instead. The number therefore lives on the run, not on the submit response: present immediately for a small submit, null on a 202 until `ingest_status` is `done`. Only the schema move is applied here. The client-side estimate is unaffected: it holds the task list in memory, so it still reports `duplicate_tasks` directly, and it remains the only place a caller can see the number before spending anything. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The reference named SubmitJobResponse.duplicate_tasks, which stopped existing when the count moved onto the run. Caught in review of ZenRows/conveyor#312. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The client-side half of ENG-286. Server side is ZenRows/conveyor#312.
Batch does not deduplicate URLs within a job: identical URLs are separate scrapes, separate charges and separate results.
Idempotency-Keylooks like it should cover it but hashes the whole request body and never inspects the task list. The estimate is the last point where a caller can see this before spending anything, so it now reports it.What changed
CostEstimate.duplicate_tasks— how many tasks repeat a request an earlier task already makes, counted in the loop the estimate already runs. Duplicates are still priced like any other task; nothing is collapsed.The key matches the server's: URL, canonical HTTP method, body, and merged
zenrows_params(job-level overlaid with per-task).external_idandmetadataare excluded — they don't change what is fetched or what it costs. URLs are compared exactly, with no normalisation. Booleans are coerced the way the server coerces them, so{"js_render": True}and{"js_render": "true"}are one request.One deliberate divergence: the body is compared by JSON value rather than by the bytes the server will see, so two bodies differing only in key order count as one request here where the server may read them as distinct. That's the only case the two counts can disagree, and it errs toward telling the caller about redundant work. Documented on
_request_key.Spec + models
Only the ENG-286 delta is applied to
docs/openapi.yaml, not a full resync from conveyor — the local copy is behind on unrelated rerun and export wording, and sweeping that in would ship other people's changes under this PR.duplicate_taskslands onSubmitJobResponse/AddTasksResponseas optional withdefault: 0rather than required, so this parses responses from a server that predates the field — including an idempotency snapshot replayed from inside its 24h TTL. That means the two repos can ship in either order.Verified
make check(ty + ruff) clean, 109 tests pass, 13 of them new.Against live dev (
async.api.zrdev.co, running conveyor#312): the regeneratedSubmitJobResponseparses the real response, and the client-side estimate and the server's count independently return 2 for the same body — including a task whose only difference isexternal_id, and a task restating the job-leveljs_renderas a boolean against the job's string.🤖 Generated with Claude Code