From 1ddd43cb9f019ea56b9da1759229359cd574acc9 Mon Sep 17 00:00:00 2001 From: Spencer Qian Date: Fri, 18 Sep 2026 11:06:49 -0700 Subject: [PATCH 1/5] Proofread: transcribe and translate a video into editable subtitle files Adds POST /v1/proofread to the core client and the CLI, shaped on dubbing, which it feeds: proofread returns one editable .srt per language plus the source-language transcript, and those corrected files go back to /v1/dubbing as subtitles[] so the dub speaks the approved wording. Core (sonilo): - client.proofread / AsyncSonilo.proofread with submit() and generate(), the same verbs the dubbing resource exposes. - build_proofread_parts: exactly one of video / video_url (https required client-side, as for dubbing), languages as a JSON-array string field omitted when unset, optional source_language. Language codes are not validated client-side; the server owns that list. - ProofreadResult (subtitles map, source_language, cue_count, warnings, duration_seconds) with save/asave and save_all/asave_all writing ..srt, mirroring DubbingResult. ProofreadIssue models one non-blocking warning and keeps unknown per-code fields in extras. - parse_proofread_result wired into tasks.get / tasks.wait polling. CLI (sonilo-cli): - `sonilo proofread` with --video/--video-url, --languages, --source-language, --out-dir, --prefix and the usual 600s --timeout; writes every language's .srt and prints the detected source language, the cue count and any warnings. Docs: proofread sections in both READMEs covering the proofread-then-dub workflow and the billing rule (video seconds x target languages, a transcript-only request counting as one, 2 free calls), the free-trial tables, and four context7 rules. Versions: sonilo 0.19.0 -> 0.20.0, sonilo-cli 0.18.0 -> 0.19.0, and sonilo-cli's core pin widened to >=0.20.0,<0.21 in the same change so a single editable install of all three packages still resolves. --- README.md | 88 ++++++- context7.json | 8 +- pyproject.toml | 2 +- sonilo-cli/README.md | 41 ++- sonilo-cli/pyproject.toml | 4 +- sonilo-cli/src/sonilo_cli/__init__.py | 2 +- sonilo-cli/src/sonilo_cli/__main__.py | 76 ++++++ sonilo-cli/tests/test_cli.py | 138 ++++++++++ src/sonilo/__init__.py | 4 + src/sonilo/_async_client.py | 2 + src/sonilo/_client.py | 2 + src/sonilo/_requests.py | 57 ++++ src/sonilo/_version.py | 2 +- src/sonilo/resources/proofread.py | 129 ++++++++++ src/sonilo/resources/tasks.py | 79 ++++++ src/sonilo/types.py | 123 +++++++++ tests/test_proofread.py | 358 ++++++++++++++++++++++++++ 17 files changed, 1106 insertions(+), 9 deletions(-) create mode 100644 src/sonilo/resources/proofread.py create mode 100644 tests/test_proofread.py diff --git a/README.md b/README.md index b6284bc..57fc94e 100644 --- a/README.md +++ b/README.md @@ -499,6 +499,92 @@ task: every video is still delivered and `subtitles` simply lacks that language. Values inside both reports may arrive as strings rather than numbers, so read them defensively. +## Proofread + +`client.proofread` transcribes a video and translates the transcript into +editable subtitle files — one `.srt` per language plus the source-language +transcript. Nothing is dubbed and nothing is spoken: this is the step *before* +`client.dubbing`, so you can read and correct the wording before any voice is +rendered. + +Pass exactly one of `video` / `video_url` (`video_url` must be **https**), plus +optional `languages` — the target languages to translate into, the same 17 +codes `client.dubbing` takes (see [Dubbing](#dubbing) for the list, and for +what `pt_br`, `es_419`, `pa_in` and `sd_in` mean), so a proofread script can go +straight into a dub. Omit `languages`, or pass `[]`, for the source-language +transcript alone. `source_language` is an optional hint telling transcription +which language to expect, which helps on short, noisy or mixed-language audio; +without it the language is detected. Either way the finished task reports the +language the transcript is in. Language codes are not checked client-side — the +server owns that list, exactly as it does for dubbing. + +Source videos may be at most 300 seconds long and 300 MB. Billing is per second +of video multiplied by the number of target languages at $0.001/second, and a +transcript-only request counts as one language. Self-serve accounts get 2 free +calls — see [Free trial](#free-trial). + +```python +from sonilo import Sonilo + +with Sonilo() as client: + result = client.proofread.generate( + video_url="https://example.com/clip.mp4", + languages=["ja", "zh_cn"], + ) + print(result.source_language, result.cue_count) + for language, path in result.save_all("./scripts").items(): + print(language, path) +``` + +`ProofreadResult.subtitles` is a language → presigned `.srt`-URL map and always +includes the **detected** source language alongside the requested targets, so +even a request with no `languages` comes back with one file. Use +`result.save(language, path)` for one language or `save_all(dir)` for all of +them (`asave`/`asave_all` on `AsyncSonilo`), which write +`{prefix}.{language}.srt` with `prefix` defaulting to `proofread`. Use +`submit()` instead of `generate()` to get a `task_id` back immediately and poll +it yourself with +`client.tasks.wait(task_id, parser=parse_proofread_result)`. + +`cue_count` is the number of subtitle cues in the source script; every language +has the same count, since translation is cue by cue. `warnings` maps a language +to the non-blocking issues its script raised and is empty when there are none — +nothing in it fails the task or withholds a file. Each issue carries `cue` (the +1-based cue it is about), `code` and `severity`, plus whatever measurement the +code brought with it, kept verbatim in `extras`: + +```python +for language, issues in result.warnings.items(): + for issue in issues: + print(language, issue.cue, issue.code, issue.get("characters_per_second")) +``` + +### Proofread, then dub + +The two endpoints are two halves of one workflow: correct the `.srt` files +proofread returned, then hand them to `client.dubbing` as +`subtitles[]` so the dub speaks exactly the approved wording. + +```python +scripts = client.proofread.generate( + video_url="https://example.com/clip.mp4", languages=["es", "fr"] +).save_all("./scripts") + +# ... edit ./scripts/proofread.es.srt and ./scripts/proofread.fr.srt ... + +dub = client.dubbing.generate( + video_url="https://example.com/clip.mp4", + languages=["es", "fr"], + subtitles={"es": scripts["es"], "fr": scripts["fr"]}, + timeout=7200, +) +dub.save_all("./dubbed") +``` + +Drop the source-language entry from `scripts` before passing it on: dubbing's +`subtitles` set must match its `languages` exactly, and proofread always +returns the source language too. + ## Video analysis `client.video_analysis` analyzes a video and returns a **creative brief** for @@ -619,7 +705,7 @@ endpoints — no card required: | Free runs | Endpoints | | --- | --- | -| 2 each | text-to-music, text-to-sfx, audio-ducking, video-analysis | +| 2 each | text-to-music, text-to-sfx, audio-ducking, video-analysis, proofread | | 1 each | video-to-music, video-to-sfx, video-to-video-music, video-to-video-sfx, video-to-sound, video-to-video-sound | | 0 | dubbing | diff --git a/context7.json b/context7.json index 2c1dbd3..9f2f27f 100644 --- a/context7.json +++ b/context7.json @@ -27,12 +27,16 @@ "Result media (.url) is a short-lived presigned URL, not the API's own domain \u2014 download it with the result's .save() helper; do not send the Authorization header to it.", "Catch AuthenticationError (401), PaymentRequiredError (402), RateLimitError (429) and TaskFailedError (status \"failed\") separately rather than one generic except \u2014 callers usually handle these differently. All extend SoniloError.", "video / video_url accept exactly one of the two, never both and never neither \u2014 validate before constructing a request.", - "Self-serve accounts start with free runs per endpoint (2 each for text-to-music, text-to-sfx, audio-ducking, video-analysis; 1 each for other video endpoints; none for dubbing), then bill normally. A first call succeeding is not proof billing works.", + "Self-serve accounts start with free runs per endpoint (2 each for text-to-music, text-to-sfx, audio-ducking, video-analysis, proofread; 1 each for other video endpoints; none for dubbing). A first call succeeding is not proof billing works.", "Before a paid call, read client.account.services().get(\"trial\", {}) and degrade gracefully when a service's remaining is 0: that call raises TrialExhaustedError (402 trial_exhausted), which no retry fixes \u2014 ask for a payment method. trial may be absent.", "client.audio_ducking ducks an EXISTING music bed under an EXISTING voice track; nothing is generated. One of voice/voice_url, one of music/music_url. Voice may be audio or video (video returns a .mp4); music must be audio. Result: output_url, no stems.", "client.dubbing dubs one video into many languages in one async call; languages: en, zh_cn, ja, ko, pt, pt_br, es, es_419, de, fr, it, ru, th, ar, tr, vi, id, ta, ml, kn, gu, pa_in, sd_in, hi; default zh_cn,es,fr. Billed per language, no free trial.", "A DubbingResult has no audio/video/output_url. Its results live in result.outputs, a map of language code to dubbed .mp4 URL: use result.save(lang, path) or result.save_all(dir). dubbing's video_url must be https.", "client.video_analysis returns a creative BRIEF, not media: analyze() (not generate()), no save(). Music brief: result.segments (start/end/label/prompt) + result.variations[i].prompt; mode both (default) adds result.sfx_segments + result.sfx_prompt.", - "Pass a video_analysis variation's prompt straight to video_to_music / video_to_sfx / video_to_sound as their prompt. One of video/video_url plus optional prompt, variants_num (1-5, billed per brief), mode (both/music/sfx, same price); max 480s, 10s floor." + "Pass a video_analysis variation's prompt straight to video_to_music / video_to_sfx / video_to_sound as their prompt. One of video/video_url plus optional prompt, variants_num (1-5, billed per brief), mode (both/music/sfx, same price); max 480s, 10s floor.", + "client.proofread transcribes a video and translates the transcript into editable .srt files. One of video/video_url (https), optional languages (the same codes as dubbing) and source_language (a hint; omit it and the language is detected). Max 300s.", + "Omit languages (or send []) to get the source-language transcript alone. ProofreadResult.subtitles maps language to .srt URL and ALWAYS includes the detected source_language, so even a transcript-only request comes back with one file.", + "Proofread is the step before dubbing: save the scripts (result.save(lang, path) / save_all(dir)), correct the wording, then pass the files to client.dubbing as subtitles[] so the dub speaks exactly the approved lines.", + "proofread bills video seconds x the number of target languages ($0.001/sec; a transcript-only request counts as one), 2 free trial calls. result.warnings is language -> non-blocking issues (cue/code/severity + extras) and withholds nothing." ] } diff --git a/pyproject.toml b/pyproject.toml index 1cce918..2cc3412 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -4,7 +4,7 @@ build-backend = "hatchling.build" [project] name = "sonilo" -version = "0.19.0" +version = "0.20.0" description = "Official Python client for the Sonilo API" readme = "README.md" license = "MIT" diff --git a/sonilo-cli/README.md b/sonilo-cli/README.md index 2965996..73b8965 100644 --- a/sonilo-cli/README.md +++ b/sonilo-cli/README.md @@ -84,6 +84,8 @@ production sign-in coexist without overwriting each other. # as a new .mp4 with the ducked mix muxed in sonilo video-analysis --video clip.mp4 --variants 2 # prints a creative brief as JSON; generates nothing + sonilo proofread --video clip.mp4 --languages ja,zh_cn --out-dir scripts + # writes scripts/proofread..srt, including the detected source language sonilo dubbing --video-url https://example.com/clip.mp4 --languages es,fr --output dubbed.mp4 # writes dubbed.es.mp4 and dubbed.fr.mp4 sonilo tasks get @@ -268,6 +270,43 @@ command that produces no media file — nothing is generated: sonilo video-analysis --video clip.mp4 --output brief.json sonilo video-to-music --video clip.mp4 --prompt "$(jq -r '.variations[0].prompt' brief.json)" +### Proofread + +`proofread` transcribes a video and translates the transcript into editable `.srt` files — one per +language, plus the source-language transcript. Nothing is dubbed: this is the step **before** +`dubbing`, so the wording can be corrected before any voice is rendered. + + sonilo proofread --video clip.mp4 --languages ja,zh_cn --out-dir scripts + # writes scripts/proofread.en.srt, scripts/proofread.ja.srt, scripts/proofread.zh_cn.srt + +- `--languages` is comma-separated and takes the same codes as `dubbing` (see [Dubbing](#dubbing) + below for the list), so a proofread script can go straight into a dub. Omit it for the + source-language transcript alone. +- `--source-language` tells transcription which language to expect, which helps on short, noisy or + mixed-language audio. Omit it to have the language detected; either way the detected code is + printed and names the source-language file. +- `--out-dir` is where the files land (default `.`, created if missing) and `--prefix` names them: + `..srt`, with `--prefix` defaulting to `proofread`. Every language is always + written — the URLs on the result are presigned and expire, and the files are the point. +- The source language is **always** returned alongside the requested targets, so a one-language + request writes two files. +- After the files, the command prints the detected source language, the cue count, and one line per + non-blocking warning (`Warning fr: high_text_speed (warning) at cue 33 — ...`). A warning never + withholds a file. +- Source videos may be at most 300 seconds long. Billing is per second of video multiplied by the + number of target languages at $0.001/second, and a transcript-only request counts as one; there + are 2 free runs — see [Free trial](#free-trial) below. +- `--timeout` defaults to 600 seconds, the usual default: a proofread job typically finishes in well + under a minute. If the wait does time out, the task keeps running server-side — resume it with + `sonilo tasks wait `. +- Edit the files, then feed them straight into `dubbing`, which makes the dub speak your exact + wording (drop the source-language file: `--subtitle` must match `--languages`): + + sonilo proofread --video clip.mp4 --languages es,fr --out-dir scripts + # ... correct scripts/proofread.es.srt and scripts/proofread.fr.srt ... + sonilo dubbing --video clip.mp4 --languages es,fr \ + --subtitle es=scripts/proofread.es.srt --subtitle fr=scripts/proofread.fr.srt + ### Dubbing `dubbing` dubs a video into one or more target languages in a single async call: @@ -317,7 +356,7 @@ required: | Free runs | Endpoints | | --- | --- | -| 2 each | text-to-music, text-to-sfx, audio-ducking, video-analysis | +| 2 each | text-to-music, text-to-sfx, audio-ducking, video-analysis, proofread | | 1 each | video-to-music, video-to-sfx, video-to-video-music, video-to-video-sfx, video-to-sound, video-to-video-sound | | 0 | dubbing | diff --git a/sonilo-cli/pyproject.toml b/sonilo-cli/pyproject.toml index 09e3d5a..233a692 100644 --- a/sonilo-cli/pyproject.toml +++ b/sonilo-cli/pyproject.toml @@ -4,13 +4,13 @@ build-backend = "hatchling.build" [project] name = "sonilo-cli" -version = "0.18.0" +version = "0.19.0" description = "Command-line interface for the Sonilo API: generate music and sound effects from text or video" readme = "README.md" license = "MIT" requires-python = ">=3.9" authors = [{ name = "Sonilo AI" }] -dependencies = ["sonilo>=0.19.0,<0.20"] +dependencies = ["sonilo>=0.20.0,<0.21"] keywords = ["sonilo", "cli", "music", "sfx", "text-to-music", "video-to-music", "ai"] [project.urls] diff --git a/sonilo-cli/src/sonilo_cli/__init__.py b/sonilo-cli/src/sonilo_cli/__init__.py index 247c7f8..dbd8a6b 100644 --- a/sonilo-cli/src/sonilo_cli/__init__.py +++ b/sonilo-cli/src/sonilo_cli/__init__.py @@ -1,3 +1,3 @@ -__version__ = "0.18.0" +__version__ = "0.19.0" __all__ = ["__version__"] diff --git a/sonilo-cli/src/sonilo_cli/__main__.py b/sonilo-cli/src/sonilo_cli/__main__.py index c154b34..e4a4512 100644 --- a/sonilo-cli/src/sonilo_cli/__main__.py +++ b/sonilo-cli/src/sonilo_cli/__main__.py @@ -726,6 +726,48 @@ def cmd_dubbing(client: Sonilo, args: argparse.Namespace) -> None: print(_export_line(language, result.subtitle_export.get(language, {}))) +def _warning_line(language: str, issue: Any) -> str: + """One line per non-blocking issue. The measurement that came with the + code (`characters_per_second` on `high_text_speed`) is appended verbatim, + because the codes are server-owned and each brings its own.""" + where = f"cue {issue.cue}" if issue.cue is not None else "script" + line = f"Warning {language}: {issue.code} ({issue.severity}) at {where}" + if issue.extras: + details = ", ".join(f"{k}={v}" for k, v in sorted(issue.extras.items())) + line += f" — {details}" + return line + + +def cmd_proofread(client: Sonilo, args: argparse.Namespace) -> None: + """Proofread writes one .srt per language into --out-dir, the same way + dubbing writes one video per language: the URLs on the result are + presigned and expire, and the whole point of the endpoint is the files you + then edit. The source language always comes back too, whether or not any + target languages were asked for.""" + languages = None + if args.languages is not None: + languages = [code.strip() for code in args.languages.split(",") if code.strip()] + if not languages: + _fail("--languages needs at least one language code, e.g. --languages es,fr") + result = client.proofread.generate( + video=args.video, + video_url=args.video_url, + languages=languages, + source_language=args.source_language, + timeout=args.timeout, + ) + if not result.subtitles: + _fail("task succeeded but returned no subtitle files") + for language, path in result.save_all(args.out_dir, prefix=args.prefix).items(): + _wrote(path, path.stat().st_size) + print(f"Source language: {result.source_language or 'unknown'}") + if result.cue_count is not None: + print(f"Cues: {result.cue_count}") + for language in sorted(result.warnings): + for issue in result.warnings[language]: + print(_warning_line(language, issue)) + + def _identity(body: Any) -> Any: return body @@ -1154,6 +1196,40 @@ def build_parser() -> argparse.ArgumentParser: ) p_dub.set_defaults(func=cmd_dubbing) + p_pr = sub.add_parser( + "proofread", + help="Transcribe a video and translate the transcript into editable .srt files", + ) + _add_global(p_pr) + _add_video_source(p_pr) + p_pr.add_argument( + "--languages", default=None, + help="Comma-separated target languages to translate the transcript into. " + "Omit it for the source-language transcript alone. Same codes as " + "`sonilo dubbing`, so a proofread script can go straight into a dub.", + ) + p_pr.add_argument( + "--source-language", dest="source_language", default=None, + help="Tell transcription which language to expect, which helps on short, " + "noisy or mixed-language audio. One of the same codes. Omit it to " + "have the language detected.", + ) + p_pr.add_argument( + "--out-dir", dest="out_dir", default=".", + help="Directory to write the .srt files into, one per language, named " + "..srt. Created if it does not exist. Default: .", + ) + p_pr.add_argument( + "--prefix", default="proofread", + help="Filename stem for the written .srt files. Default: proofread", + ) + p_pr.add_argument( + "--timeout", type=float, default=DEFAULT_WAIT_TIMEOUT, + help="Give up waiting after this many seconds. Default: 600. A timed-out " + "task may still finish — resume it with `sonilo tasks wait `.", + ) + p_pr.set_defaults(func=cmd_proofread) + p_tasks = sub.add_parser("tasks", help="Inspect async tasks") _add_global(p_tasks) tsub = p_tasks.add_subparsers(dest="tasks_command", metavar="") diff --git a/sonilo-cli/tests/test_cli.py b/sonilo-cli/tests/test_cli.py index 5439db3..c7a4acf 100644 --- a/sonilo-cli/tests/test_cli.py +++ b/sonilo-cli/tests/test_cli.py @@ -963,6 +963,144 @@ def test_dubbing_rejects_a_subtitle_that_is_not_srt_or_vtt(tmp_path, capsys): assert not route.called +# --- proofread ------------------------------------------------------------- + +PROOFREAD_BODY = { + "task_id": "pr1", + "type": "proofread", + "status": "succeeded", + "duration_seconds": 206.32, + "source_language": "en", + "subtitles": {"en": "https://r2/en.srt", "fr": "https://r2/fr.srt"}, + "cue_count": 65, + "warnings": { + "fr": [ + { + "cue": 33, + "code": "high_text_speed", + "severity": "warning", + "characters_per_second": 26.92, + } + ] + }, +} + + +def _stub_proofread(body=None): + route = respx.post(f"{BASE}/v1/proofread").mock( + return_value=httpx.Response(202, json={"task_id": "pr1", "status": "processing"}) + ) + respx.get(f"{BASE}/v1/tasks/pr1").mock( + return_value=httpx.Response(200, json=body or PROOFREAD_BODY) + ) + for language in (body or PROOFREAD_BODY)["subtitles"]: + respx.get(f"https://r2/{language}.srt").mock( + return_value=httpx.Response(200, content=f"{language}-bytes".encode()) + ) + return route + + +@respx.mock +def test_proofread_writes_one_srt_per_language(tmp_path, capsys): + _stub_proofread() + run([ + "proofread", + "--video-url", "https://x/v.mp4", + "--languages", "fr", + "--out-dir", str(tmp_path / "scripts"), + ]) + # The detected source language comes back alongside the requested target, + # so a one-language request still writes two files. + assert (tmp_path / "scripts" / "proofread.en.srt").read_bytes() == b"en-bytes" + assert (tmp_path / "scripts" / "proofread.fr.srt").read_bytes() == b"fr-bytes" + out = capsys.readouterr().out + assert "Source language: en" in out + assert "Cues: 65" in out + + +@respx.mock +def test_proofread_prints_the_warnings(tmp_path, capsys): + _stub_proofread() + run([ + "proofread", "--video-url", "https://x/v.mp4", + "--languages", "fr", "--out-dir", str(tmp_path), + ]) + out = capsys.readouterr().out + assert "Warning fr: high_text_speed (warning) at cue 33" in out + # The code's own measurement is printed with it, not dropped. + assert "characters_per_second=26.92" in out + + +@respx.mock +def test_proofread_sends_languages_as_a_json_array(tmp_path): + route = _stub_proofread() + run([ + "proofread", "--video-url", "https://x/v.mp4", + "--languages", " fr , en ", "--out-dir", str(tmp_path), + ]) + body = unquote_plus(route.calls.last.request.content.decode()) + assert '["fr", "en"]' in body + + +@respx.mock +def test_proofread_without_languages_omits_the_field(tmp_path): + route = _stub_proofread({ + "task_id": "pr1", "status": "succeeded", + "source_language": "en", + "subtitles": {"en": "https://r2/en.srt"}, + }) + run(["proofread", "--video-url", "https://x/v.mp4", "--out-dir", str(tmp_path)]) + # Omitting --languages asks for the transcript alone; sending the field at + # all would be a different request. + assert b"languages" not in route.calls.last.request.content + assert (tmp_path / "proofread.en.srt").exists() + + +@respx.mock +def test_proofread_sends_source_language(tmp_path): + route = _stub_proofread() + run([ + "proofread", "--video-url", "https://x/v.mp4", + "--source-language", "en", "--out-dir", str(tmp_path), + ]) + assert "source_language=en" in unquote_plus(route.calls.last.request.content.decode()) + + +@respx.mock +def test_proofread_prefix_names_the_files(tmp_path): + _stub_proofread() + run([ + "proofread", "--video-url", "https://x/v.mp4", "--languages", "fr", + "--out-dir", str(tmp_path), "--prefix", "interview", + ]) + assert (tmp_path / "interview.fr.srt").exists() + + +def test_proofread_requires_a_video_source(): + with pytest.raises(SystemExit) as exc: + main(["--api-key", "sk-test", "proofread"]) + assert exc.value.code == 1 + + +@respx.mock +def test_proofread_non_https_url_exits_1(capsys): + with pytest.raises(SystemExit) as exc: + main(["--api-key", "sk-test", "proofread", "--video-url", "http://x/v.mp4"]) + assert exc.value.code == 1 + assert "https" in capsys.readouterr().err + + +@respx.mock +def test_proofread_empty_languages_value_exits_1(capsys): + with pytest.raises(SystemExit) as exc: + main([ + "--api-key", "sk-test", "proofread", + "--video-url", "https://x/v.mp4", "--languages", " , ", + ]) + assert exc.value.code == 1 + assert "--languages" in capsys.readouterr().err + + # --- --segments ----------------------------------------------------------- # # The two shapes are not interchangeable: music segments are diff --git a/src/sonilo/__init__.py b/src/sonilo/__init__.py index d6c055a..b8f848f 100644 --- a/src/sonilo/__init__.py +++ b/src/sonilo/__init__.py @@ -23,6 +23,8 @@ MusicResult, MusicStems, MusicTitle, + ProofreadIssue, + ProofreadResult, Segment, SfxMedia, SfxResult, @@ -51,6 +53,8 @@ "MusicStems", "MusicTitle", "PaymentRequiredError", + "ProofreadIssue", + "ProofreadResult", "RateLimitError", "Segment", "SfxMedia", diff --git a/src/sonilo/_async_client.py b/src/sonilo/_async_client.py index be7fcc6..eb09df7 100644 --- a/src/sonilo/_async_client.py +++ b/src/sonilo/_async_client.py @@ -10,6 +10,7 @@ from sonilo.resources.account import AsyncAccount from sonilo.resources.audio_ducking import AsyncAudioDucking from sonilo.resources.dubbing import AsyncDubbing +from sonilo.resources.proofread import AsyncProofread from sonilo.resources.video_analysis import AsyncVideoAnalysis from sonilo.resources.tasks import AsyncTasks from sonilo.resources.text_to_music import AsyncTextToMusic @@ -54,6 +55,7 @@ def __init__( self.video_to_video_sound = AsyncVideoToVideoSound(self) self.audio_ducking = AsyncAudioDucking(self) self.dubbing = AsyncDubbing(self) + self.proofread = AsyncProofread(self) self.video_analysis = AsyncVideoAnalysis(self) self.account = AsyncAccount(self) self.tasks = AsyncTasks(self) diff --git a/src/sonilo/_client.py b/src/sonilo/_client.py index 614100a..0742d94 100644 --- a/src/sonilo/_client.py +++ b/src/sonilo/_client.py @@ -11,6 +11,7 @@ from sonilo.resources.account import Account from sonilo.resources.audio_ducking import AudioDucking from sonilo.resources.dubbing import Dubbing +from sonilo.resources.proofread import Proofread from sonilo.resources.tasks import Tasks from sonilo.resources.text_to_music import TextToMusic from sonilo.resources.text_to_sfx import TextToSfx @@ -84,6 +85,7 @@ def __init__( self.video_to_video_sound = VideoToVideoSound(self) self.audio_ducking = AudioDucking(self) self.dubbing = Dubbing(self) + self.proofread = Proofread(self) self.video_analysis = VideoAnalysis(self) self.account = Account(self) self.tasks = Tasks(self) diff --git a/src/sonilo/_requests.py b/src/sonilo/_requests.py index 2ec6a04..653bedd 100644 --- a/src/sonilo/_requests.py +++ b/src/sonilo/_requests.py @@ -299,6 +299,63 @@ def build_dubbing_parts( return data, files or None, MultiClose(opened) if opened else None +def build_proofread_parts( + video: Any, + video_url: Optional[str], + languages: Optional[List[str]] = None, + source_language: Optional[str] = None, +) -> Tuple[Dict[str, str], Optional[Dict[str, tuple]], bool]: + """Build the multipart parts for POST /v1/proofread. + + Proofread is dubbing's sibling — it transcribes the video and translates + the transcript so the `.srt` files can be corrected before they are sent + back to /v1/dubbing as `subtitles[]` — so the two builders agree + field for field wherever they overlap. `languages` travels as one opaque + form field holding a JSON array string, exactly as it does for dubbing, + and is omitted entirely when unset (which asks for the source-language + transcript alone). Unlike dubbing's, an empty list IS meaningful here and + is sent as `[]`: it is the explicit spelling of "transcript only". + + `source_language` is a hint for the transcription, sent only when given; + without it the language is detected and reported back on the finished + task. Language codes are deliberately NOT checked here — the backend owns + that list, exactly as it does for dubbing, and a hardcoded copy would make + this SDK reject codes added later. + + The https check on `video_url` is local for the same reason as dubbing's: + this pipeline fetches the source URL itself and requires https + specifically, so a plain-http URL is a guaranteed server-side 422. + + Only one file can ever be opened here (there are no subtitle uploads on + the way in), so this returns the plain `opened` bool the single-input + builders use rather than dubbing's `MultiClose`. + """ + if (video is None) == (video_url is None): + raise SoniloError("Provide exactly one of video or video_url") + + # Assemble data dict completely before opening any files + data: Dict[str, str] = {} + if video_url is not None: + if not video_url.lower().startswith("https://"): + raise SoniloError( + "video_url must use https — the proofread pipeline requires an https URL" + ) + data["video_url"] = video_url + if languages is not None: + data["languages"] = json.dumps(languages) + if source_language is not None: + data["source_language"] = source_language + + # Now open files (only after data is fully assembled) + files: Optional[Dict[str, tuple]] = None + opened = False + if video is not None: + filename, fileobj, opened = normalize_video(video) + files = {"video": (filename, fileobj, "video/mp4")} + + return data, files, opened + + def build_video_analysis_parts( video: Any, video_url: Optional[str], diff --git a/src/sonilo/_version.py b/src/sonilo/_version.py index 11ac8e1..5f4bb0b 100644 --- a/src/sonilo/_version.py +++ b/src/sonilo/_version.py @@ -1 +1 @@ -__version__ = "0.19.0" +__version__ = "0.20.0" diff --git a/src/sonilo/resources/proofread.py b/src/sonilo/resources/proofread.py new file mode 100644 index 0000000..77c15a0 --- /dev/null +++ b/src/sonilo/resources/proofread.py @@ -0,0 +1,129 @@ +from __future__ import annotations + +from typing import TYPE_CHECKING, Any, List, Optional + +from sonilo._requests import build_proofread_parts +from sonilo.resources.tasks import ( + DEFAULT_POLL_INTERVAL, + DEFAULT_WAIT_TIMEOUT, + parse_proofread_result, + parse_sfx_task, +) +from sonilo.types import ProofreadResult, SfxTask + +if TYPE_CHECKING: + from sonilo._async_client import AsyncSonilo + from sonilo._client import Sonilo + +PATH = "/v1/proofread" + + +class Proofread: + """Transcribe one video and translate the transcript into editable + subtitle files. Async only; the result carries a language → `.srt`-URL map + under `subtitles`. + + This is the step before `client.dubbing`, not a replacement for it: + proofread returns one `.srt` per language plus the source-language + transcript, you review or correct the wording, and the corrected files go + to `client.dubbing` as `subtitles[]` so the dub speaks exactly + the approved lines. The language codes are the same on both endpoints, so + a proofread script can go straight into a dub. + + Pass exactly one of `video` / `video_url` (`video_url` must be **https**), + plus optional `languages` — the target languages to translate into, sent + as a JSON array. Omit it, or pass `[]`, for the source-language transcript + alone. `source_language` is an optional hint telling transcription which + language to expect, which helps on short, noisy or mixed-language audio; + without it the language is detected, and either way the finished task + reports what the transcript is in. + + Billing is per second of video multiplied by the number of target + languages; a transcript-only request counts as one. + """ + + def __init__(self, client: "Sonilo") -> None: + self._client = client + + def submit( + self, + *, + video: Any = None, + video_url: Optional[str] = None, + languages: Optional[List[str]] = None, + source_language: Optional[str] = None, + ) -> SfxTask: + data, files, opened = build_proofread_parts( + video, video_url, languages, source_language + ) + close_after = files["video"][1] if files is not None and opened else None + return parse_sfx_task( + self._client._post_json(PATH, data=data, files=files, close_after=close_after) + ) + + def generate( + self, + *, + video: Any = None, + video_url: Optional[str] = None, + languages: Optional[List[str]] = None, + source_language: Optional[str] = None, + poll_interval: float = DEFAULT_POLL_INTERVAL, + timeout: float = DEFAULT_WAIT_TIMEOUT, + ) -> ProofreadResult: + task = self.submit( + video=video, video_url=video_url, languages=languages, + source_language=source_language, + ) + return self._client.tasks.wait( + task.task_id, + poll_interval=poll_interval, + timeout=timeout, + parser=parse_proofread_result, + ) + + +class AsyncProofread: + """Async twin of Proofread; same parameters and same result shape.""" + + def __init__(self, client: "AsyncSonilo") -> None: + self._client = client + + async def submit( + self, + *, + video: Any = None, + video_url: Optional[str] = None, + languages: Optional[List[str]] = None, + source_language: Optional[str] = None, + ) -> SfxTask: + data, files, opened = build_proofread_parts( + video, video_url, languages, source_language + ) + close_after = files["video"][1] if files is not None and opened else None + return parse_sfx_task( + await self._client._post_json( + PATH, data=data, files=files, close_after=close_after + ) + ) + + async def generate( + self, + *, + video: Any = None, + video_url: Optional[str] = None, + languages: Optional[List[str]] = None, + source_language: Optional[str] = None, + poll_interval: float = DEFAULT_POLL_INTERVAL, + timeout: float = DEFAULT_WAIT_TIMEOUT, + ) -> ProofreadResult: + task = await self.submit( + video=video, video_url=video_url, languages=languages, + source_language=source_language, + ) + return await self._client.tasks.wait( + task.task_id, + poll_interval=poll_interval, + timeout=timeout, + parser=parse_proofread_result, + ) diff --git a/src/sonilo/resources/tasks.py b/src/sonilo/resources/tasks.py index 7d1367d..6a93f2a 100644 --- a/src/sonilo/resources/tasks.py +++ b/src/sonilo/resources/tasks.py @@ -15,6 +15,8 @@ MusicResult, MusicStems, MusicTitle, + ProofreadIssue, + ProofreadResult, SfxMedia, SfxResult, SfxTask, @@ -286,6 +288,83 @@ def parse_dubbing_result(body: Dict[str, Any]) -> "DubbingResult": raise SoniloError(f"Malformed task response: missing {e.args[0]!r}") from e +_ISSUE_NAMED_FIELDS = ("cue", "code", "severity") + + +def _proofread_issue_from(data: Any) -> Optional[ProofreadIssue]: + """Coerce one warning entry. Everything the check reports beyond the three + named fields is kept in `extras` rather than dropped: the issue codes are + server-owned and each brings its own measurement (`high_text_speed` brings + `characters_per_second`), so a code added later must still arrive whole. + `cue` is read leniently — the pipeline's store can hand a number back as a + string — and reads as None when it is neither.""" + if not isinstance(data, dict): + return None + try: + cue: Optional[int] = int(data["cue"]) + except (KeyError, TypeError, ValueError): + cue = None + return ProofreadIssue( + code=str(data.get("code") or "unknown"), + severity=str(data.get("severity") or "warning"), + cue=cue, + extras={k: v for k, v in data.items() if k not in _ISSUE_NAMED_FIELDS}, + ) + + +def _proofread_warnings_from(data: Any) -> Dict[str, List[ProofreadIssue]]: + """Coerce the language → issue-list map, dropping malformed entries, for + the same reason _url_map_from coerces dubbing's outputs: a differently + shaped entry from a backend change should surface here and not as an + AttributeError deep inside the caller's loop.""" + if not isinstance(data, dict): + return {} + warnings: Dict[str, List[ProofreadIssue]] = {} + for language, issues in data.items(): + if not isinstance(issues, list): + continue + warnings[str(language)] = [ + issue for issue in map(_proofread_issue_from, issues) if issue is not None + ] + return warnings + + +def _int_or_none(value: Any) -> Optional[int]: + """`cue_count` may arrive as a number or, from a store that keeps numbers + as strings, as a string. Read either; anything else is None.""" + try: + return int(value) + except (TypeError, ValueError): + return None + + +def parse_proofread_result(body: Dict[str, Any]) -> "ProofreadResult": + """Map a GET /v1/tasks/{id} body for a proofread task to ProofreadResult; + unknown fields are ignored. + + `subtitles` is coerced exactly as dubbing's `outputs` is — see + _url_map_from — and always carries the detected source language alongside + the requested targets. `warnings` is `{}` when the scripts raised nothing, + which is the common case; it never gates delivery of a file. + """ + try: + return ProofreadResult( + task_id=body["task_id"], + status=body["status"], + type=body.get("type"), + subtitles=_url_map_from(body.get("subtitles")), + source_language=body.get("source_language"), + cue_count=_int_or_none(body.get("cue_count")), + warnings=_proofread_warnings_from(body.get("warnings")), + duration_seconds=body.get("duration_seconds"), + cost=body.get("cost"), + error=body.get("error"), + refunded=body.get("refunded"), + ) + except KeyError as e: + raise SoniloError(f"Malformed task response: missing {e.args[0]!r}") from e + + def _analysis_segment_from(data: Any) -> Optional[AnalysisSegment]: if not isinstance(data, dict): return None diff --git a/src/sonilo/types.py b/src/sonilo/types.py index 5dc5e35..f1bf098 100644 --- a/src/sonilo/types.py +++ b/src/sonilo/types.py @@ -777,6 +777,129 @@ async def asave_all_subtitles( } +@dataclass +class ProofreadIssue: + """One non-blocking validation issue on a proofread script. + + The fields every issue carries are named; anything else the check reports + is kept verbatim in `extras` rather than dropped, because the issue codes + are server-owned and grow — `high_text_speed` comes with + `characters_per_second`, and a code added later will come with fields this + SDK has never heard of. `cue` is the 1-based index of the subtitle cue the + issue is about, and is None only when the report omitted it. + """ + + code: str + severity: str + cue: Optional[int] = None + extras: Dict[str, Any] = field(default_factory=dict) + + def get(self, key: str, default: Any = None) -> Any: + """Read one pass-through field, e.g. `issue.get("characters_per_second")`.""" + return self.extras.get(key, default) + + +@dataclass +class ProofreadResult: + """State of a proofread task (`tasks.get`) or its final result + (`wait`/`generate`). + + Shaped like DubbingResult, because the two are two halves of one workflow: + proofread transcribes a video and translates the transcript, you correct + the wording, and the corrected files go back to `client.dubbing` as + `subtitles[]` so the dub speaks exactly what was approved. + + `subtitles` is a language → presigned `.srt`-URL map and always includes + the DETECTED source language (reported in `source_language`) alongside one + entry per requested target language, so a request with no `languages` at + all still comes back with one file. `save(language, path)` fetches one and + `save_all(dir)` fetches every one of them, mirroring DubbingResult. + + `cue_count` is the number of subtitle cues in the source script; every + language has the same count, since translation is cue-by-cue. `warnings` + maps a language to the non-blocking issues its script raised and is empty + when there are none — nothing in it fails the task or withholds a file. + """ + + task_id: str + status: str + type: Optional[str] = None + subtitles: Dict[str, str] = field(default_factory=dict) + source_language: Optional[str] = None + cue_count: Optional[int] = None + warnings: Dict[str, List[ProofreadIssue]] = field(default_factory=dict) + duration_seconds: Optional[float] = None + cost: Optional[float] = None + error: Optional[Dict[str, Any]] = None + refunded: Optional[bool] = None + + def _url(self, language: str) -> str: + if language not in self.subtitles: + available = ", ".join(sorted(self.subtitles)) or "none" + raise SoniloError( + f"No subtitle for language {language!r} on this result " + f"(status={self.status}; available: {available})" + ) + return self.subtitles[language] + + def save( + self, + language: str, + path: Union[str, Path], + *, + timeout: float = DOWNLOAD_TIMEOUT, + ) -> Path: + """Download one language's `.srt` to `path` and return it. The URL is + presigned — no API key is sent.""" + return _download_to(self._url(language), path, timeout) + + async def asave( + self, + language: str, + path: Union[str, Path], + *, + timeout: float = DOWNLOAD_TIMEOUT, + ) -> Path: + """Async variant of save().""" + return await _adownload_to(self._url(language), path, timeout) + + def save_all( + self, + directory: Union[str, Path], + *, + prefix: str = "proofread", + timeout: float = DOWNLOAD_TIMEOUT, + ) -> Dict[str, Path]: + """Download every language into `directory` as + `{prefix}.{language}.srt`, returning the language → path map. The + directory is created if it does not exist.""" + target = Path(directory) + target.mkdir(parents=True, exist_ok=True) + return { + language: self.save( + language, target / f"{prefix}.{language}.srt", timeout=timeout + ) + for language in sorted(self.subtitles) + } + + async def asave_all( + self, + directory: Union[str, Path], + *, + prefix: str = "proofread", + timeout: float = DOWNLOAD_TIMEOUT, + ) -> Dict[str, Path]: + """Async variant of save_all().""" + target = Path(directory) + target.mkdir(parents=True, exist_ok=True) + return { + language: await self.asave( + language, target / f"{prefix}.{language}.srt", timeout=timeout + ) + for language in sorted(self.subtitles) + } + + @dataclass class AnalysisSegment: """One time-aligned section of the analyzed video, with the creative diff --git a/tests/test_proofread.py b/tests/test_proofread.py new file mode 100644 index 0000000..d86e072 --- /dev/null +++ b/tests/test_proofread.py @@ -0,0 +1,358 @@ +from pathlib import Path +from urllib.parse import unquote_plus + +import httpx +import pytest +import respx + +from sonilo import AsyncSonilo, ProofreadIssue, ProofreadResult, Sonilo +from sonilo._requests import build_proofread_parts +from sonilo.errors import SoniloError +from sonilo.types import SfxTask +from sonilo.resources.tasks import parse_proofread_result + +# The real prod body (task 4288764d..., 2026-09-18), URLs shortened. Keeping +# the actual shape means the parser is checked against what the API sends — +# including `warnings` carrying a code with its own measurement field. +SUCCESS_BODY = { + "task_id": "4288764d-0057-4a77-885e-03c0391c7c1d", + "type": "proofread", + "status": "succeeded", + "duration_seconds": 206.32, + "source_language": "en", + "subtitles": { + "en": "https://r2/en.srt", + "ko": "https://r2/ko.srt", + "fr": "https://r2/fr.srt", + "de": "https://r2/de.srt", + "ar": "https://r2/ar.srt", + "th": "https://r2/th.srt", + "ru": "https://r2/ru.srt", + }, + "cue_count": 65, + "warnings": { + "fr": [ + { + "cue": 33, + "code": "high_text_speed", + "severity": "warning", + "characters_per_second": 26.92, + } + ] + }, +} + +ACK = {"task_id": "pr1", "status": "processing"} + + +# --- result parsing -------------------------------------------------------- + + +def test_parse_proofread_result_reads_the_subtitles_map(): + result = parse_proofread_result(SUCCESS_BODY) + assert result.task_id == "4288764d-0057-4a77-885e-03c0391c7c1d" + assert result.status == "succeeded" + assert result.type == "proofread" + assert result.source_language == "en" + assert result.cue_count == 65 + assert result.duration_seconds == 206.32 + # The detected source language is always in the map, alongside the + # requested targets. + assert "en" in result.subtitles + assert result.subtitles["fr"] == "https://r2/fr.srt" + assert len(result.subtitles) == 7 + + +def test_parse_proofread_result_reads_the_warnings(): + result = parse_proofread_result(SUCCESS_BODY) + assert set(result.warnings) == {"fr"} + issue = result.warnings["fr"][0] + assert isinstance(issue, ProofreadIssue) + assert issue.cue == 33 + assert issue.code == "high_text_speed" + assert issue.severity == "warning" + # Everything past the three named fields is kept verbatim, because each + # server-owned code brings its own measurement. + assert issue.extras == {"characters_per_second": 26.92} + assert issue.get("characters_per_second") == 26.92 + assert issue.get("nothing_like_this") is None + + +def test_parse_proofread_result_defaults_warnings_to_empty(): + result = parse_proofread_result( + {"task_id": "pr1", "status": "succeeded", "subtitles": {"en": "https://r2/en.srt"}} + ) + assert result.warnings == {} + assert result.source_language is None + assert result.cue_count is None + + +def test_parse_proofread_result_defaults_subtitles_to_empty(): + result = parse_proofread_result({"task_id": "pr1", "status": "processing"}) + assert result.subtitles == {} + + +def test_parse_proofread_result_tolerates_stringy_numbers(): + """A store that keeps numbers as strings is what made dubbing's reports + arrive as strings; read either shape here rather than crash.""" + result = parse_proofread_result( + { + "task_id": "pr1", + "status": "succeeded", + "cue_count": "65", + "warnings": {"fr": [{"cue": "7", "code": "high_text_speed"}]}, + } + ) + assert result.cue_count == 65 + assert result.warnings["fr"][0].cue == 7 + # severity is defaulted rather than dropped, so callers can print it. + assert result.warnings["fr"][0].severity == "warning" + + +def test_parse_proofread_result_drops_malformed_warning_entries(): + result = parse_proofread_result( + { + "task_id": "pr1", + "status": "succeeded", + "subtitles": "nope", + "warnings": {"fr": "not-a-list", "de": ["not-an-issue", {"code": "x"}]}, + } + ) + assert result.subtitles == {} + assert "fr" not in result.warnings + assert [issue.code for issue in result.warnings["de"]] == ["x"] + + +def test_parse_proofread_result_rejects_a_body_without_a_task_id(): + with pytest.raises(SoniloError): + parse_proofread_result({"status": "succeeded"}) + + +# --- saving ---------------------------------------------------------------- + + +@respx.mock +def test_save_downloads_one_language(tmp_path): + respx.get("https://r2/fr.srt").mock( + return_value=httpx.Response(200, content=b"1\nbonjour\n") + ) + result = parse_proofread_result(SUCCESS_BODY) + path = result.save("fr", tmp_path / "clip.fr.srt") + assert path.read_bytes() == b"1\nbonjour\n" + + +def test_save_rejects_a_language_the_task_did_not_produce(tmp_path): + result = parse_proofread_result(SUCCESS_BODY) + with pytest.raises(SoniloError): + result.save("es", tmp_path / "clip.es.srt") + + +@respx.mock +def test_save_all_writes_one_srt_per_language(tmp_path): + for language in SUCCESS_BODY["subtitles"]: + respx.get(f"https://r2/{language}.srt").mock( + return_value=httpx.Response(200, content=f"{language}-bytes".encode()) + ) + result = parse_proofread_result(SUCCESS_BODY) + paths = result.save_all(tmp_path / "out") + assert set(paths) == set(SUCCESS_BODY["subtitles"]) + assert (tmp_path / "out" / "proofread.en.srt").read_bytes() == b"en-bytes" + assert (tmp_path / "out" / "proofread.ko.srt").read_bytes() == b"ko-bytes" + + +@respx.mock +async def test_asave_downloads_one_language(tmp_path): + respx.get("https://r2/de.srt").mock( + return_value=httpx.Response(200, content=b"1\nhallo\n") + ) + result = parse_proofread_result(SUCCESS_BODY) + path = await result.asave("de", tmp_path / "clip.de.srt") + assert path.read_bytes() == b"1\nhallo\n" + + +# --- request shape --------------------------------------------------------- + + +@respx.mock +def test_submit_posts_to_v1_proofread(): + route = respx.post("https://api.sonilo.com/v1/proofread").mock( + return_value=httpx.Response(202, json=ACK) + ) + with Sonilo(api_key="sk-test") as client: + task = client.proofread.submit( + video_url="https://x/v.mp4", languages=["ja", "zh_cn"] + ) + assert isinstance(task, SfxTask) + assert task.task_id == "pr1" and task.status == "processing" + # video_url-only submissions have no file part, so httpx sends this as + # application/x-www-form-urlencoded (not multipart) and percent-escapes + # the JSON array; unquote before checking for the raw JSON substring. + sent = unquote_plus(route.calls.last.request.content.decode()) + assert "video_url=https://x/v.mp4" in sent + assert '["ja", "zh_cn"]' in sent + + +@respx.mock +def test_submit_omits_languages_when_unset(): + """Omitting the field is what asks for the transcript alone; sending + `languages=None` as a string would be a 422.""" + route = respx.post("https://api.sonilo.com/v1/proofread").mock( + return_value=httpx.Response(202, json=ACK) + ) + with Sonilo(api_key="sk-test") as client: + client.proofread.submit(video_url="https://x/v.mp4") + assert b"languages" not in route.calls.last.request.content + + +@respx.mock +def test_submit_sends_an_explicit_empty_language_list(): + """Unlike dubbing, `[]` is meaningful here — it is the explicit spelling + of "transcript only" — so it goes on the wire rather than being dropped.""" + route = respx.post("https://api.sonilo.com/v1/proofread").mock( + return_value=httpx.Response(202, json=ACK) + ) + with Sonilo(api_key="sk-test") as client: + client.proofread.submit(video_url="https://x/v.mp4", languages=[]) + sent = unquote_plus(route.calls.last.request.content.decode()) + assert "languages=[]" in sent + + +@respx.mock +def test_submit_sends_source_language_only_when_given(): + route = respx.post("https://api.sonilo.com/v1/proofread").mock( + return_value=httpx.Response(202, json=ACK) + ) + with Sonilo(api_key="sk-test") as client: + client.proofread.submit(video_url="https://x/v.mp4", source_language="en") + client.proofread.submit(video_url="https://x/v.mp4") + bodies = [unquote_plus(c.request.content.decode()) for c in route.calls] + assert "source_language=en" in bodies[0] + # Unset → omitted entirely so the server detects the language itself. + assert "source_language" not in bodies[1] + + +@respx.mock +def test_submit_uploads_a_local_video_as_a_file_part(tmp_path): + route = respx.post("https://api.sonilo.com/v1/proofread").mock( + return_value=httpx.Response(202, json=ACK) + ) + clip = tmp_path / "interview.mp4" + clip.write_bytes(b"fake-mp4") + with Sonilo(api_key="sk-test") as client: + client.proofread.submit(video=str(clip), languages=["ja"]) + body = route.calls.last.request.content.decode(errors="replace") + assert 'name="video"' in body + assert 'filename="interview.mp4"' in body + # multipart, so the JSON array travels as a plain field value. + assert '["ja"]' in body + + +@respx.mock +def test_submit_rejects_a_non_https_url_before_sending(): + route = respx.post("https://api.sonilo.com/v1/proofread") + with Sonilo(api_key="sk-test") as client: + with pytest.raises(SoniloError): + client.proofread.submit(video_url="http://x/v.mp4") + assert not route.called + + +@respx.mock +def test_submit_requires_exactly_one_video_input(): + route = respx.post("https://api.sonilo.com/v1/proofread") + with Sonilo(api_key="sk-test") as client: + with pytest.raises(SoniloError): + client.proofread.submit() + with pytest.raises(SoniloError): + client.proofread.submit(video=b"bytes", video_url="https://x/v.mp4") + assert not route.called + + +@respx.mock +def test_submit_closes_a_local_video_even_when_the_api_rejects_it(tmp_path, monkeypatch): + respx.post("https://api.sonilo.com/v1/proofread").mock( + return_value=httpx.Response(422, json={"error": {"code": "unprocessable_entity"}}) + ) + clip = tmp_path / "clip.mp4" + clip.write_bytes(b"fake-mp4") + opened = [] + real_open = Path.open + + def spy(self, *args, **kwargs): + handle = real_open(self, *args, **kwargs) + opened.append(handle) + return handle + + monkeypatch.setattr(Path, "open", spy) + with Sonilo(api_key="sk-test") as client: + with pytest.raises(SoniloError): + client.proofread.submit(video=str(clip)) + assert opened and all(handle.closed for handle in opened) + + +def test_build_proofread_parts_passes_unknown_codes_through(): + """Language codes are server-owned, exactly as on dubbing: a code this + SDK has never heard of must reach the API and earn its own 422.""" + data, _, _ = build_proofread_parts(None, "https://x/v.mp4", ["klingon"], "klingon") + assert data["languages"] == '["klingon"]' + assert data["source_language"] == "klingon" + + +# --- polling --------------------------------------------------------------- + + +@respx.mock +def test_generate_polls_to_a_proofread_result(): + respx.post("https://api.sonilo.com/v1/proofread").mock( + return_value=httpx.Response(202, json=ACK) + ) + respx.get("https://api.sonilo.com/v1/tasks/pr1").mock( + return_value=httpx.Response(200, json=dict(SUCCESS_BODY, task_id="pr1")) + ) + with Sonilo(api_key="sk-test") as client: + result = client.proofread.generate( + video_url="https://x/v.mp4", languages=["fr", "de"], poll_interval=0 + ) + assert isinstance(result, ProofreadResult) + assert result.source_language == "en" + assert result.subtitles["fr"] == "https://r2/fr.srt" + assert result.warnings["fr"][0].code == "high_text_speed" + + +@respx.mock +def test_tasks_get_parses_a_proofread_task(): + respx.get("https://api.sonilo.com/v1/tasks/pr1").mock( + return_value=httpx.Response(200, json=dict(SUCCESS_BODY, task_id="pr1")) + ) + with Sonilo(api_key="sk-test") as client: + result = client.tasks.get("pr1", parser=parse_proofread_result) + assert result.type == "proofread" + assert result.cue_count == 65 + + +@respx.mock +async def test_async_generate_polls_to_a_proofread_result(): + respx.post("https://api.sonilo.com/v1/proofread").mock( + return_value=httpx.Response(202, json=ACK) + ) + respx.get("https://api.sonilo.com/v1/tasks/pr1").mock( + return_value=httpx.Response(200, json=dict(SUCCESS_BODY, task_id="pr1")) + ) + async with AsyncSonilo(api_key="sk-test") as client: + result = await client.proofread.generate( + video_url="https://x/v.mp4", source_language="en", poll_interval=0 + ) + assert result.subtitles["en"] == "https://r2/en.srt" + + +@respx.mock +async def test_async_submit_sends_the_same_fields(): + route = respx.post("https://api.sonilo.com/v1/proofread").mock( + return_value=httpx.Response(202, json=ACK) + ) + async with AsyncSonilo(api_key="sk-test") as client: + await client.proofread.submit( + video_url="https://x/v.mp4", languages=["ja"], source_language="en" + ) + sent = unquote_plus(route.calls.last.request.content.decode()) + assert '["ja"]' in sent + assert "source_language=en" in sent From 94d154aa5bc748135299d9f63ee774d6345b0ce8 Mon Sep 17 00:00:00 2001 From: Spencer Qian Date: Fri, 18 Sep 2026 11:07:51 -0700 Subject: [PATCH 2/5] uv.lock: record sonilo 0.20.0 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The root lock carries the editable core's own version, the same one-line refresh the 0.19.0 release made. sonilo-cli/uv.lock is deliberately left alone: it resolves the PUBLISHED core from PyPI, and 0.20.0 is not there yet — it needs a refresh after the release, exactly as the previous one did. CI installs all three packages with pip, not uv, so neither lock gates it. --- uv.lock | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/uv.lock b/uv.lock index c936885..c9c254a 100644 --- a/uv.lock +++ b/uv.lock @@ -267,7 +267,7 @@ wheels = [ [[package]] name = "sonilo" -version = "0.19.0" +version = "0.20.0" source = { editable = "." } dependencies = [ { name = "httpx" }, From 992539db80439cd9293a80c36774bd2fdaae654b Mon Sep 17 00:00:00 2001 From: Spencer Qian Date: Fri, 18 Sep 2026 11:20:07 -0700 Subject: [PATCH 3/5] CLI: proofread names its files with --output, like dubbing Replaces --out-dir plus --prefix with the template flag every other file-producing command in this CLI already takes. `--output scripts/clip.srt` writes scripts/clip.en.srt, scripts/clip.fr.srt, ... through the same _language_path transform dubbing uses for one video per language, so the flag vocabulary stays one idiom rather than two, and two runs into the same directory can be told apart without an extra flag. Default: proofread.srt. _language_path gains a default_suffix keyword, defaulting to dubbing's ".mp4" so that call site is unchanged; proofread passes ".srt" so an extension-less template does not name subtitles after a video container, which the new test_proofread_output_without_an_extension_still_gets_srt covers. cmd_proofread creates the template's parent directory, which --out-dir used to get from save_all and this template routinely needs. ProofreadResult.save_all / asave_all keep the directory-plus-prefix signature: that one mirrors DubbingResult.save_all and belongs to the core SDK, not to this CLI. Also makes test_proofread_requires_a_video_source assert the stderr text rather than the exit code alone, matching its two neighbours: as written it would have passed had argparse exited 1 for an unrelated reason. --- sonilo-cli/src/sonilo_cli/__main__.py | 39 +++++++++++++---------- sonilo-cli/tests/test_cli.py | 45 ++++++++++++++++++++------- 2 files changed, 57 insertions(+), 27 deletions(-) diff --git a/sonilo-cli/src/sonilo_cli/__main__.py b/sonilo-cli/src/sonilo_cli/__main__.py index e4a4512..5c666df 100644 --- a/sonilo-cli/src/sonilo_cli/__main__.py +++ b/sonilo-cli/src/sonilo_cli/__main__.py @@ -629,13 +629,17 @@ def cmd_video_analysis(client: Sonilo, args: argparse.Namespace) -> None: DUBBING_WAIT_TIMEOUT = 7200.0 -def _language_path(out: str, language: str) -> str: +def _language_path(out: str, language: str, default_suffix: str = ".mp4") -> str: """Turn one --output value into a per-language path: `clip.mp4` + `es` becomes `clip.es.mp4`. A dubbing task returns one video per language, so a single literal destination cannot express the result. This is the same - transform _stem_path applies for --stem, so both flags read the same way.""" + transform _stem_path applies for --stem, so both flags read the same way. + + `default_suffix` only decides what an extension-less template gets, and + defaults to dubbing's `.mp4`; proofread passes `.srt` so `--output clip` + does not name its subtitles after a video container.""" base = Path(out) - return str(base.with_name(f"{base.stem}.{language}{base.suffix or '.mp4'}")) + return str(base.with_name(f"{base.stem}.{language}{base.suffix or default_suffix}")) def _subtitles(values: Optional[List[str]]) -> Optional[Dict[str, str]]: @@ -739,11 +743,12 @@ def _warning_line(language: str, issue: Any) -> str: def cmd_proofread(client: Sonilo, args: argparse.Namespace) -> None: - """Proofread writes one .srt per language into --out-dir, the same way - dubbing writes one video per language: the URLs on the result are - presigned and expire, and the whole point of the endpoint is the files you - then edit. The source language always comes back too, whether or not any - target languages were asked for.""" + """Proofread writes one .srt per language, the same way dubbing writes one + video per language and through the same --output template: the URLs on the + result are presigned and expire, and the whole point of the endpoint is the + files you then edit. The source language always comes back too, whether or + not any target languages were asked for.""" + out = args.output if args.output is not None else "proofread.srt" languages = None if args.languages is not None: languages = [code.strip() for code in args.languages.split(",") if code.strip()] @@ -758,7 +763,11 @@ def cmd_proofread(client: Sonilo, args: argparse.Namespace) -> None: ) if not result.subtitles: _fail("task succeeded but returned no subtitle files") - for language, path in result.save_all(args.out_dir, prefix=args.prefix).items(): + # Unlike dubbing's, this template routinely names a directory of its own + # ("--output scripts/clip.srt"), so create it rather than fail on the write. + Path(out).parent.mkdir(parents=True, exist_ok=True) + for language in sorted(result.subtitles): + path = result.save(language, _language_path(out, language, ".srt")) _wrote(path, path.stat().st_size) print(f"Source language: {result.source_language or 'unknown'}") if result.cue_count is not None: @@ -1215,13 +1224,11 @@ def build_parser() -> argparse.ArgumentParser: "have the language detected.", ) p_pr.add_argument( - "--out-dir", dest="out_dir", default=".", - help="Directory to write the .srt files into, one per language, named " - "..srt. Created if it does not exist. Default: .", - ) - p_pr.add_argument( - "--prefix", default="proofread", - help="Filename stem for the written .srt files. Default: proofread", + "--output", default=None, + help="Filename template, not a single destination: one .srt is written " + "per language with the code inserted before the extension " + "(scripts/clip.srt -> scripts/clip.en.srt). Missing directories are " + "created. Default: proofread.srt", ) p_pr.add_argument( "--timeout", type=float, default=DEFAULT_WAIT_TIMEOUT, diff --git a/sonilo-cli/tests/test_cli.py b/sonilo-cli/tests/test_cli.py index c7a4acf..0990d5e 100644 --- a/sonilo-cli/tests/test_cli.py +++ b/sonilo-cli/tests/test_cli.py @@ -1002,17 +1002,19 @@ def _stub_proofread(body=None): @respx.mock def test_proofread_writes_one_srt_per_language(tmp_path, capsys): + """--output is a template, the same one dubbing uses: the language code is + inserted before the extension, and the directory is created.""" _stub_proofread() run([ "proofread", "--video-url", "https://x/v.mp4", "--languages", "fr", - "--out-dir", str(tmp_path / "scripts"), + "--output", str(tmp_path / "scripts" / "clip.srt"), ]) # The detected source language comes back alongside the requested target, # so a one-language request still writes two files. - assert (tmp_path / "scripts" / "proofread.en.srt").read_bytes() == b"en-bytes" - assert (tmp_path / "scripts" / "proofread.fr.srt").read_bytes() == b"fr-bytes" + assert (tmp_path / "scripts" / "clip.en.srt").read_bytes() == b"en-bytes" + assert (tmp_path / "scripts" / "clip.fr.srt").read_bytes() == b"fr-bytes" out = capsys.readouterr().out assert "Source language: en" in out assert "Cues: 65" in out @@ -1023,7 +1025,7 @@ def test_proofread_prints_the_warnings(tmp_path, capsys): _stub_proofread() run([ "proofread", "--video-url", "https://x/v.mp4", - "--languages", "fr", "--out-dir", str(tmp_path), + "--languages", "fr", "--output", str(tmp_path / "clip.srt"), ]) out = capsys.readouterr().out assert "Warning fr: high_text_speed (warning) at cue 33" in out @@ -1036,7 +1038,7 @@ def test_proofread_sends_languages_as_a_json_array(tmp_path): route = _stub_proofread() run([ "proofread", "--video-url", "https://x/v.mp4", - "--languages", " fr , en ", "--out-dir", str(tmp_path), + "--languages", " fr , en ", "--output", str(tmp_path / "clip.srt"), ]) body = unquote_plus(route.calls.last.request.content.decode()) assert '["fr", "en"]' in body @@ -1049,11 +1051,14 @@ def test_proofread_without_languages_omits_the_field(tmp_path): "source_language": "en", "subtitles": {"en": "https://r2/en.srt"}, }) - run(["proofread", "--video-url", "https://x/v.mp4", "--out-dir", str(tmp_path)]) + run([ + "proofread", "--video-url", "https://x/v.mp4", + "--output", str(tmp_path / "clip.srt"), + ]) # Omitting --languages asks for the transcript alone; sending the field at # all would be a different request. assert b"languages" not in route.calls.last.request.content - assert (tmp_path / "proofread.en.srt").exists() + assert (tmp_path / "clip.en.srt").exists() @respx.mock @@ -1061,25 +1066,43 @@ def test_proofread_sends_source_language(tmp_path): route = _stub_proofread() run([ "proofread", "--video-url", "https://x/v.mp4", - "--source-language", "en", "--out-dir", str(tmp_path), + "--source-language", "en", "--output", str(tmp_path / "clip.srt"), ]) assert "source_language=en" in unquote_plus(route.calls.last.request.content.decode()) @respx.mock -def test_proofread_prefix_names_the_files(tmp_path): +def test_proofread_output_without_an_extension_still_gets_srt(tmp_path): + """The shared template transform defaults an extension-less value to + dubbing's .mp4; proofread has to override that or name subtitles after a + video container.""" _stub_proofread() run([ "proofread", "--video-url", "https://x/v.mp4", "--languages", "fr", - "--out-dir", str(tmp_path), "--prefix", "interview", + "--output", str(tmp_path / "interview"), ]) assert (tmp_path / "interview.fr.srt").exists() + assert not (tmp_path / "interview.fr.mp4").exists() -def test_proofread_requires_a_video_source(): +@respx.mock +def test_proofread_default_output_template(tmp_path, monkeypatch): + """With no --output the files land beside the caller as proofread..srt, + the way dubbing defaults to output.mp4.""" + _stub_proofread() + monkeypatch.chdir(tmp_path) + run(["proofread", "--video-url", "https://x/v.mp4", "--languages", "fr"]) + assert (tmp_path / "proofread.en.srt").exists() + assert (tmp_path / "proofread.fr.srt").exists() + + +def test_proofread_requires_a_video_source(capsys): with pytest.raises(SystemExit) as exc: main(["--api-key", "sk-test", "proofread"]) assert exc.value.code == 1 + # Assert the message too: a bare exit code would also pass if argparse + # bailed out for some unrelated reason. + assert "--video" in capsys.readouterr().err @respx.mock From d0b0b4aa26df9475385c8e1ee32d807921766d3f Mon Sep 17 00:00:00 2001 From: Spencer Qian Date: Fri, 18 Sep 2026 11:20:17 -0700 Subject: [PATCH 4/5] Proofread docs: two contract requirements, and drop a stale code count MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit README.md said proofread takes "the same 17 codes client.dubbing takes" while the dubbing list it points at carries 24, and the same paragraph then names pa_in and sd_in, which no 17-code set contains. Drop the number instead of correcting it: it is duplicated across two sections and would go stale again on the next language addition, which is why the CLI README and context7.json already say "the same codes as dubbing" with no count. Both READMEs now state the two contract requirements they were missing: the video must have an audio track (the likeliest caller mistake here, and TRANSCRIPTION_EMPTY does not explain itself), and billing has a 10-second floor, which the repo already documents for video-analysis. The CLI README also gains the 300 MB cap the core README had — the CLI is the surface where a local file is actually uploaded. Drops "since translation is cue by cue" from README.md: the contract supports the fact that every language has the same cue count, not the reason. Adds the asave_all test that was the one untested half of the sync/async result pair. --- README.md | 18 +++++++++--------- sonilo-cli/README.md | 34 +++++++++++++++++++--------------- tests/test_proofread.py | 13 +++++++++++++ 3 files changed, 41 insertions(+), 24 deletions(-) diff --git a/README.md b/README.md index 57fc94e..7ab1a07 100644 --- a/README.md +++ b/README.md @@ -507,11 +507,11 @@ transcript. Nothing is dubbed and nothing is spoken: this is the step *before* `client.dubbing`, so you can read and correct the wording before any voice is rendered. -Pass exactly one of `video` / `video_url` (`video_url` must be **https**), plus -optional `languages` — the target languages to translate into, the same 17 -codes `client.dubbing` takes (see [Dubbing](#dubbing) for the list, and for -what `pt_br`, `es_419`, `pa_in` and `sd_in` mean), so a proofread script can go -straight into a dub. Omit `languages`, or pass `[]`, for the source-language +Pass exactly one of `video` / `video_url` (`video_url` must be **https**); the +video must have an audio track. Plus optional `languages` — the target +languages to translate into, the same codes `client.dubbing` takes (see +[Dubbing](#dubbing) for the list, and for what `pt_br`, `es_419`, `pa_in` and +`sd_in` mean), so a proofread script can go straight into a dub. Omit `languages`, or pass `[]`, for the source-language transcript alone. `source_language` is an optional hint telling transcription which language to expect, which helps on short, noisy or mixed-language audio; without it the language is detected. Either way the finished task reports the @@ -519,9 +519,9 @@ language the transcript is in. Language codes are not checked client-side — th server owns that list, exactly as it does for dubbing. Source videos may be at most 300 seconds long and 300 MB. Billing is per second -of video multiplied by the number of target languages at $0.001/second, and a -transcript-only request counts as one language. Self-serve accounts get 2 free -calls — see [Free trial](#free-trial). +of video multiplied by the number of target languages at $0.001/second, a +transcript-only request counts as one language, and billing has a 10-second +floor. Self-serve accounts get 2 free calls — see [Free trial](#free-trial). ```python from sonilo import Sonilo @@ -547,7 +547,7 @@ it yourself with `client.tasks.wait(task_id, parser=parse_proofread_result)`. `cue_count` is the number of subtitle cues in the source script; every language -has the same count, since translation is cue by cue. `warnings` maps a language +has the same count. `warnings` maps a language to the non-blocking issues its script raised and is empty when there are none — nothing in it fails the task or withholds a file. Each issue carries `cue` (the 1-based cue it is about), `code` and `severity`, plus whatever measurement the diff --git a/sonilo-cli/README.md b/sonilo-cli/README.md index 73b8965..5759123 100644 --- a/sonilo-cli/README.md +++ b/sonilo-cli/README.md @@ -84,8 +84,8 @@ production sign-in coexist without overwriting each other. # as a new .mp4 with the ducked mix muxed in sonilo video-analysis --video clip.mp4 --variants 2 # prints a creative brief as JSON; generates nothing - sonilo proofread --video clip.mp4 --languages ja,zh_cn --out-dir scripts - # writes scripts/proofread..srt, including the detected source language + sonilo proofread --video clip.mp4 --languages ja,zh_cn --output scripts/clip.srt + # writes scripts/clip..srt, including the detected source language sonilo dubbing --video-url https://example.com/clip.mp4 --languages es,fr --output dubbed.mp4 # writes dubbed.es.mp4 and dubbed.fr.mp4 sonilo tasks get @@ -273,11 +273,12 @@ command that produces no media file — nothing is generated: ### Proofread `proofread` transcribes a video and translates the transcript into editable `.srt` files — one per -language, plus the source-language transcript. Nothing is dubbed: this is the step **before** -`dubbing`, so the wording can be corrected before any voice is rendered. +language, plus the source-language transcript. The video must have an audio track. Nothing is +dubbed: this is the step **before** `dubbing`, so the wording can be corrected before any voice is +rendered. - sonilo proofread --video clip.mp4 --languages ja,zh_cn --out-dir scripts - # writes scripts/proofread.en.srt, scripts/proofread.ja.srt, scripts/proofread.zh_cn.srt + sonilo proofread --video clip.mp4 --languages ja,zh_cn --output scripts/clip.srt + # writes scripts/clip.en.srt, scripts/clip.ja.srt, scripts/clip.zh_cn.srt - `--languages` is comma-separated and takes the same codes as `dubbing` (see [Dubbing](#dubbing) below for the list), so a proofread script can go straight into a dub. Omit it for the @@ -285,27 +286,30 @@ language, plus the source-language transcript. Nothing is dubbed: this is the st - `--source-language` tells transcription which language to expect, which helps on short, noisy or mixed-language audio. Omit it to have the language detected; either way the detected code is printed and names the source-language file. -- `--out-dir` is where the files land (default `.`, created if missing) and `--prefix` names them: - `..srt`, with `--prefix` defaulting to `proofread`. Every language is always - written — the URLs on the result are presigned and expire, and the files are the point. +- `--output` is a filename template, not a single destination, exactly as it is for `dubbing`: one + `.srt` is written per language with the code inserted before the extension, so + `--output scripts/clip.srt` writes `scripts/clip.en.srt`, `scripts/clip.fr.srt`, etc. Missing + directories are created. Default: `proofread.srt`. Every language is always written — the URLs on + the result are presigned and expire, and the files are the point. - The source language is **always** returned alongside the requested targets, so a one-language request writes two files. - After the files, the command prints the detected source language, the cue count, and one line per non-blocking warning (`Warning fr: high_text_speed (warning) at cue 33 — ...`). A warning never withholds a file. -- Source videos may be at most 300 seconds long. Billing is per second of video multiplied by the - number of target languages at $0.001/second, and a transcript-only request counts as one; there - are 2 free runs — see [Free trial](#free-trial) below. +- Source videos may be at most 300 seconds long and 300 MB. Billing is per second of video + multiplied by the number of target languages at $0.001/second, a transcript-only request counts as + one, and billing has a 10-second floor; there are 2 free runs — see [Free trial](#free-trial) + below. - `--timeout` defaults to 600 seconds, the usual default: a proofread job typically finishes in well under a minute. If the wait does time out, the task keeps running server-side — resume it with `sonilo tasks wait `. - Edit the files, then feed them straight into `dubbing`, which makes the dub speak your exact wording (drop the source-language file: `--subtitle` must match `--languages`): - sonilo proofread --video clip.mp4 --languages es,fr --out-dir scripts - # ... correct scripts/proofread.es.srt and scripts/proofread.fr.srt ... + sonilo proofread --video clip.mp4 --languages es,fr --output scripts/clip.srt + # ... correct scripts/clip.es.srt and scripts/clip.fr.srt ... sonilo dubbing --video clip.mp4 --languages es,fr \ - --subtitle es=scripts/proofread.es.srt --subtitle fr=scripts/proofread.fr.srt + --subtitle es=scripts/clip.es.srt --subtitle fr=scripts/clip.fr.srt ### Dubbing diff --git a/tests/test_proofread.py b/tests/test_proofread.py index d86e072..b02bba1 100644 --- a/tests/test_proofread.py +++ b/tests/test_proofread.py @@ -160,6 +160,19 @@ def test_save_all_writes_one_srt_per_language(tmp_path): assert (tmp_path / "out" / "proofread.ko.srt").read_bytes() == b"ko-bytes" +@respx.mock +async def test_asave_all_writes_one_srt_per_language(tmp_path): + for language in SUCCESS_BODY["subtitles"]: + respx.get(f"https://r2/{language}.srt").mock( + return_value=httpx.Response(200, content=f"{language}-bytes".encode()) + ) + result = parse_proofread_result(SUCCESS_BODY) + paths = await result.asave_all(tmp_path / "out", prefix="clip") + assert set(paths) == set(SUCCESS_BODY["subtitles"]) + assert (tmp_path / "out" / "clip.en.srt").read_bytes() == b"en-bytes" + assert (tmp_path / "out" / "clip.th.srt").read_bytes() == b"th-bytes" + + @respx.mock async def test_asave_downloads_one_language(tmp_path): respx.get("https://r2/de.srt").mock( From 83fe1a6d5f7ab952c21b189187e4cdbb0b5874d6 Mon Sep 17 00:00:00 2001 From: Spencer Qian Date: Fri, 18 Sep 2026 11:21:47 -0700 Subject: [PATCH 5/5] README: rewrap the proofread languages paragraph --- README.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/README.md b/README.md index 7ab1a07..6e6fdf0 100644 --- a/README.md +++ b/README.md @@ -511,8 +511,8 @@ Pass exactly one of `video` / `video_url` (`video_url` must be **https**); the video must have an audio track. Plus optional `languages` — the target languages to translate into, the same codes `client.dubbing` takes (see [Dubbing](#dubbing) for the list, and for what `pt_br`, `es_419`, `pa_in` and -`sd_in` mean), so a proofread script can go straight into a dub. Omit `languages`, or pass `[]`, for the source-language -transcript alone. `source_language` is an optional hint telling transcription +`sd_in` mean), so a proofread script can go straight into a dub. Omit +`languages`, or pass `[]`, for the source-language transcript alone. `source_language` is an optional hint telling transcription which language to expect, which helps on short, noisy or mixed-language audio; without it the language is detected. Either way the finished task reports the language the transcript is in. Language codes are not checked client-side — the