Repository navigation
Conversation
download_resource() short-circuits on `dest_path.exists()` but streamed straight into that path, so a connection dropped mid-transfer left a truncated file behind. The next attempt returned it as a complete download. The sync retry loop could not recover from this: _download_raw_with_retry names its temp file `<uuid8>-<basename>`, while CKAN/Saude ignores the requested filename and writes `filename_for(resource, ...)` inside `output.parent` (as _download_once's own docstring notes). The retry's cleanup deleted a path that never existed, leaving the partial file in place for the next attempt to pick up. The engine then hashed it, converted it and uploaded it to S3 as the official artifact. Two changes: download_resource() streams into a sibling `.<name>.<uuid8>.part` file and os.replace()s it on success, removing it on any BaseException. The rename is also atomic, so two Saude packages publishing a resource with the same name -- which resolve to the same dest_path in a shared tmp dir -- can no longer interleave writes into one file. _download_raw_with_retry() gives each attempt its own scratch directory and removes the whole directory on retry, which is the only reliable cleanup when the origin chose the filename. No directory is created for the final attempt, so nothing is left behind. _cleanup_stale_tmp() also sweeps these directories, recognised by _is_attempt_dir(), so a killed run does not leak them into the next one.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
download_resource()short-circuits ondest_path.exists()but streamed directly into that path:The sync retry loop could not recover from this.
_download_raw_with_retrynames its temp file<uuid8>-<basename>, but CKAN/Saude ignores the requested filename and writesfilename_for(resource, ...)insideoutput.parent— as_download_once's own docstring states:So the retry's
self._cleanup_local(output)deleted a path that never existed, anddownload_resourcethen immediately returned the truncated file. The engine hashed it, converted it, and uploaded it to S3 as the official artifact.A second, related hazard: all Saude workers share
management/tmpasdest_dir, andfilename_forderives the name purely fromresource.name. Two packages publishing a resource with the same name resolve to the samedest_path, and two concurrent workers interleaveopen("wb")writes into one file.Fix
download_resource()streams into a sibling.<name>.<uuid8>.partfile andos.replace()s it on success, unlinking it on anyBaseException. The rename is atomic, so concurrent writers can no longer interleave._download_raw_with_retry()gives each attempt its own scratch directory and removes the whole directory on retry — the only reliable cleanup when the origin chose the filename itself. No directory is created for the final attempt, so nothing is left behind._cleanup_stale_tmp()also sweeps these directories (recognised by the new_is_attempt_dir(): 8 hex chars, which nothing else in that tmp area uses), so a killed run does not leak them into the next one.Tests
test_download.py— new_FlakyTransportyields real bytes and then raisesReadError, so the failure happens after the destination has been written to:.partscratchtest_sync_more.py— newTestDownloadRawWithRetry:_is_attempt_diraccepts only hex-8 names_cleanup_stale_tmpsweeps leftover attempt directoriesVerified the two download tests fail against the previous code.
1741 passed, 6 skippedblackandisortclean