Proofread: transcribe and translate a video into editable subtitle files - #45
Merged
Merged
Conversation
Adds POST /v1/proofread to the core client and the CLI, shaped on dubbing, which it feeds: proofread returns one editable .srt per language plus the source-language transcript, and those corrected files go back to /v1/dubbing as subtitles[<language>] so the dub speaks the approved wording. Core (sonilo): - client.proofread / AsyncSonilo.proofread with submit() and generate(), the same verbs the dubbing resource exposes. - build_proofread_parts: exactly one of video / video_url (https required client-side, as for dubbing), languages as a JSON-array string field omitted when unset, optional source_language. Language codes are not validated client-side; the server owns that list. - ProofreadResult (subtitles map, source_language, cue_count, warnings, duration_seconds) with save/asave and save_all/asave_all writing <prefix>.<language>.srt, mirroring DubbingResult. ProofreadIssue models one non-blocking warning and keeps unknown per-code fields in extras. - parse_proofread_result wired into tasks.get / tasks.wait polling. CLI (sonilo-cli): - `sonilo proofread` with --video/--video-url, --languages, --source-language, --out-dir, --prefix and the usual 600s --timeout; writes every language's .srt and prints the detected source language, the cue count and any warnings. Docs: proofread sections in both READMEs covering the proofread-then-dub workflow and the billing rule (video seconds x target languages, a transcript-only request counting as one, 2 free calls), the free-trial tables, and four context7 rules. Versions: sonilo 0.19.0 -> 0.20.0, sonilo-cli 0.18.0 -> 0.19.0, and sonilo-cli's core pin widened to >=0.20.0,<0.21 in the same change so a single editable install of all three packages still resolves.
The root lock carries the editable core's own version, the same one-line refresh the 0.19.0 release made. sonilo-cli/uv.lock is deliberately left alone: it resolves the PUBLISHED core from PyPI, and 0.20.0 is not there yet — it needs a refresh after the release, exactly as the previous one did. CI installs all three packages with pip, not uv, so neither lock gates it.
Replaces --out-dir plus --prefix with the template flag every other file-producing command in this CLI already takes. `--output scripts/clip.srt` writes scripts/clip.en.srt, scripts/clip.fr.srt, ... through the same _language_path transform dubbing uses for one video per language, so the flag vocabulary stays one idiom rather than two, and two runs into the same directory can be told apart without an extra flag. Default: proofread.srt. _language_path gains a default_suffix keyword, defaulting to dubbing's ".mp4" so that call site is unchanged; proofread passes ".srt" so an extension-less template does not name subtitles after a video container, which the new test_proofread_output_without_an_extension_still_gets_srt covers. cmd_proofread creates the template's parent directory, which --out-dir used to get from save_all and this template routinely needs. ProofreadResult.save_all / asave_all keep the directory-plus-prefix signature: that one mirrors DubbingResult.save_all and belongs to the core SDK, not to this CLI. Also makes test_proofread_requires_a_video_source assert the stderr text rather than the exit code alone, matching its two neighbours: as written it would have passed had argparse exited 1 for an unrelated reason.
README.md said proofread takes "the same 17 codes client.dubbing takes" while the dubbing list it points at carries 24, and the same paragraph then names pa_in and sd_in, which no 17-code set contains. Drop the number instead of correcting it: it is duplicated across two sections and would go stale again on the next language addition, which is why the CLI README and context7.json already say "the same codes as dubbing" with no count. Both READMEs now state the two contract requirements they were missing: the video must have an audio track (the likeliest caller mistake here, and TRANSCRIPTION_EMPTY does not explain itself), and billing has a 10-second floor, which the repo already documents for video-analysis. The CLI README also gains the 300 MB cap the core README had — the CLI is the surface where a local file is actually uploaded. Drops "since translation is cue by cue" from README.md: the contract supports the fact that every language has the same cue count, not the reason. Adds the asave_all test that was the one untested half of the sync/async result pair.
Lightsage docs evalsWaiting for the staging docs URL before running evals. Lightsage will start the selected PR evals automatically when GitHub reports a successful docs deployment for this PR. This usually happens within 15 minutes. Commit: |
This was referenced Sep 18, 2026
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds the new
POST /v1/proofreadendpoint to the core SDK and the CLI. Proofread transcribes a video and translates the transcript into editable.srtfiles, one per requested language plus the detected source language. It is the step beforedubbing: correct the wording, then pass the files todubbingassubtitles[<language>].client.proofread.submit(...)/generate(...)(sync + async), multipart with exactly one ofvideo/video_url(https), optionallanguagesandsource_language; language codes are not validated client-side (server-owned, like dubbing).ProofreadResultwithsubtitles(language → URL),source_language,cue_count,warnings(language →ProofreadIssuelist), plussave/save_alland async twins.tasks.get/tasks.waitparsetype == "proofread"into it.sonilo proofread --video clip.mp4 --languages ja,zh_cn --output scripts/clip.srtwritesscripts/clip.<lang>.srtthrough the same--outputtemplate asdubbing; default timeout 600 sec.context7.jsonupdated.sonilo0.19.0 → 0.20.0,sonilo-cli0.18.0 → 0.19.0, CLI pin widened to>=0.20.0,<0.21.sonilo-cli/uv.lockis left as is until 0.20.0 is on PyPI, as in the previous release.Test plan
high_text_speedwarning).