Skip to content

Proofread: transcribe and translate a video into editable subtitle files - #45

Merged
spencer-zqian merged 5 commits into
mainfrom
feat/proofread
Sep 18, 2026
Merged

spencer-zqian merged 5 commits into
mainfrom
feat/proofread

Conversation

@spencer-zqian

Copy link
Copy Markdown
Contributor

Summary

Adds the new POST /v1/proofread endpoint to the core SDK and the CLI. Proofread transcribes a video and translates the transcript into editable .srt files, one per requested language plus the detected source language. It is the step before dubbing: correct the wording, then pass the files to dubbing as subtitles[<language>].

  • Core: client.proofread.submit(...) / generate(...) (sync + async), multipart with exactly one of video / video_url (https), optional languages and source_language; language codes are not validated client-side (server-owned, like dubbing).
  • Result: ProofreadResult with subtitles (language → URL), source_language, cue_count, warnings (language → ProofreadIssue list), plus save / save_all and async twins. tasks.get / tasks.wait parse type == "proofread" into it.
  • CLI: sonilo proofread --video clip.mp4 --languages ja,zh_cn --output scripts/clip.srt writes scripts/clip.<lang>.srt through the same --output template as dubbing; default timeout 600 sec.
  • Docs: README and CLI README gain a Proofread section (billing: video seconds × target languages, transcript-only counts as one, $0.001/sec, 10 sec floor, 2 free trial calls; 300 sec / 300 MB limits; must have an audio track). context7.json updated.
  • Versions: sonilo 0.19.0 → 0.20.0, sonilo-cli 0.18.0 → 0.19.0, CLI pin widened to >=0.20.0,<0.21. sonilo-cli/uv.lock is left as is until 0.20.0 is on PyPI, as in the previous release.

Test plan

  • Fresh venv, CI's single editable install of all three packages: 357 core + 53 video-kit + 182 CLI pass on Python 3.12 and 3.9.
  • Parser verified against a real finished-task envelope from prod (7 subtitles, 65 cues, one high_text_speed warning).
  • No vendor names in the diff.

Adds POST /v1/proofread to the core client and the CLI, shaped on dubbing,
which it feeds: proofread returns one editable .srt per language plus the
source-language transcript, and those corrected files go back to
/v1/dubbing as subtitles[<language>] so the dub speaks the approved wording.

Core (sonilo):
- client.proofread / AsyncSonilo.proofread with submit() and generate(),
  the same verbs the dubbing resource exposes.
- build_proofread_parts: exactly one of video / video_url (https required
  client-side, as for dubbing), languages as a JSON-array string field
  omitted when unset, optional source_language. Language codes are not
  validated client-side; the server owns that list.
- ProofreadResult (subtitles map, source_language, cue_count, warnings,
  duration_seconds) with save/asave and save_all/asave_all writing
  <prefix>.<language>.srt, mirroring DubbingResult. ProofreadIssue models
  one non-blocking warning and keeps unknown per-code fields in extras.
- parse_proofread_result wired into tasks.get / tasks.wait polling.

CLI (sonilo-cli):
- `sonilo proofread` with --video/--video-url, --languages,
  --source-language, --out-dir, --prefix and the usual 600s --timeout;
  writes every language's .srt and prints the detected source language,
  the cue count and any warnings.

Docs: proofread sections in both READMEs covering the proofread-then-dub
workflow and the billing rule (video seconds x target languages, a
transcript-only request counting as one, 2 free calls), the free-trial
tables, and four context7 rules.

Versions: sonilo 0.19.0 -> 0.20.0, sonilo-cli 0.18.0 -> 0.19.0, and
sonilo-cli's core pin widened to >=0.20.0,<0.21 in the same change so a
single editable install of all three packages still resolves.
The root lock carries the editable core's own version, the same one-line
refresh the 0.19.0 release made. sonilo-cli/uv.lock is deliberately left
alone: it resolves the PUBLISHED core from PyPI, and 0.20.0 is not there
yet — it needs a refresh after the release, exactly as the previous one
did. CI installs all three packages with pip, not uv, so neither lock
gates it.
Replaces --out-dir plus --prefix with the template flag every other
file-producing command in this CLI already takes. `--output scripts/clip.srt`
writes scripts/clip.en.srt, scripts/clip.fr.srt, ... through the same
_language_path transform dubbing uses for one video per language, so the flag
vocabulary stays one idiom rather than two, and two runs into the same
directory can be told apart without an extra flag. Default: proofread.srt.

_language_path gains a default_suffix keyword, defaulting to dubbing's
".mp4" so that call site is unchanged; proofread passes ".srt" so an
extension-less template does not name subtitles after a video container,
which the new test_proofread_output_without_an_extension_still_gets_srt
covers. cmd_proofread creates the template's parent directory, which
--out-dir used to get from save_all and this template routinely needs.

ProofreadResult.save_all / asave_all keep the directory-plus-prefix
signature: that one mirrors DubbingResult.save_all and belongs to the core
SDK, not to this CLI.

Also makes test_proofread_requires_a_video_source assert the stderr text
rather than the exit code alone, matching its two neighbours: as written it
would have passed had argparse exited 1 for an unrelated reason.
README.md said proofread takes "the same 17 codes client.dubbing takes"
while the dubbing list it points at carries 24, and the same paragraph then
names pa_in and sd_in, which no 17-code set contains. Drop the number
instead of correcting it: it is duplicated across two sections and would go
stale again on the next language addition, which is why the CLI README and
context7.json already say "the same codes as dubbing" with no count.

Both READMEs now state the two contract requirements they were missing: the
video must have an audio track (the likeliest caller mistake here, and
TRANSCRIPTION_EMPTY does not explain itself), and billing has a 10-second
floor, which the repo already documents for video-analysis. The CLI README
also gains the 300 MB cap the core README had — the CLI is the surface
where a local file is actually uploaded.

Drops "since translation is cue by cue" from README.md: the contract
supports the fact that every language has the same cue count, not the
reason.

Adds the asave_all test that was the one untested half of the sync/async
result pair.
@lightsage-app

lightsage-app Bot commented Sep 18, 2026

Copy link
Copy Markdown

Lightsage docs evals

Waiting for the staging docs URL before running evals.

Lightsage will start the selected PR evals automatically when GitHub reports a successful docs deployment for this PR. This usually happens within 15 minutes.

Commit: 83fe1a6
Status: waiting for staging docs URL

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant