diff --git a/README.md b/README.md index b6284bc..6e6fdf0 100644 --- a/README.md +++ b/README.md @@ -499,6 +499,92 @@ task: every video is still delivered and `subtitles` simply lacks that language. Values inside both reports may arrive as strings rather than numbers, so read them defensively. +## Proofread + +`client.proofread` transcribes a video and translates the transcript into +editable subtitle files — one `.srt` per language plus the source-language +transcript. Nothing is dubbed and nothing is spoken: this is the step *before* +`client.dubbing`, so you can read and correct the wording before any voice is +rendered. + +Pass exactly one of `video` / `video_url` (`video_url` must be **https**); the +video must have an audio track. Plus optional `languages` — the target +languages to translate into, the same codes `client.dubbing` takes (see +[Dubbing](#dubbing) for the list, and for what `pt_br`, `es_419`, `pa_in` and +`sd_in` mean), so a proofread script can go straight into a dub. Omit +`languages`, or pass `[]`, for the source-language transcript alone. `source_language` is an optional hint telling transcription +which language to expect, which helps on short, noisy or mixed-language audio; +without it the language is detected. Either way the finished task reports the +language the transcript is in. Language codes are not checked client-side — the +server owns that list, exactly as it does for dubbing. + +Source videos may be at most 300 seconds long and 300 MB. Billing is per second +of video multiplied by the number of target languages at $0.001/second, a +transcript-only request counts as one language, and billing has a 10-second +floor. Self-serve accounts get 2 free calls — see [Free trial](#free-trial). + +```python +from sonilo import Sonilo + +with Sonilo() as client: + result = client.proofread.generate( + video_url="https://example.com/clip.mp4", + languages=["ja", "zh_cn"], + ) + print(result.source_language, result.cue_count) + for language, path in result.save_all("./scripts").items(): + print(language, path) +``` + +`ProofreadResult.subtitles` is a language → presigned `.srt`-URL map and always +includes the **detected** source language alongside the requested targets, so +even a request with no `languages` comes back with one file. Use +`result.save(language, path)` for one language or `save_all(dir)` for all of +them (`asave`/`asave_all` on `AsyncSonilo`), which write +`{prefix}.{language}.srt` with `prefix` defaulting to `proofread`. Use +`submit()` instead of `generate()` to get a `task_id` back immediately and poll +it yourself with +`client.tasks.wait(task_id, parser=parse_proofread_result)`. + +`cue_count` is the number of subtitle cues in the source script; every language +has the same count. `warnings` maps a language +to the non-blocking issues its script raised and is empty when there are none — +nothing in it fails the task or withholds a file. Each issue carries `cue` (the +1-based cue it is about), `code` and `severity`, plus whatever measurement the +code brought with it, kept verbatim in `extras`: + +```python +for language, issues in result.warnings.items(): + for issue in issues: + print(language, issue.cue, issue.code, issue.get("characters_per_second")) +``` + +### Proofread, then dub + +The two endpoints are two halves of one workflow: correct the `.srt` files +proofread returned, then hand them to `client.dubbing` as +`subtitles[]` so the dub speaks exactly the approved wording. + +```python +scripts = client.proofread.generate( + video_url="https://example.com/clip.mp4", languages=["es", "fr"] +).save_all("./scripts") + +# ... edit ./scripts/proofread.es.srt and ./scripts/proofread.fr.srt ... + +dub = client.dubbing.generate( + video_url="https://example.com/clip.mp4", + languages=["es", "fr"], + subtitles={"es": scripts["es"], "fr": scripts["fr"]}, + timeout=7200, +) +dub.save_all("./dubbed") +``` + +Drop the source-language entry from `scripts` before passing it on: dubbing's +`subtitles` set must match its `languages` exactly, and proofread always +returns the source language too. + ## Video analysis `client.video_analysis` analyzes a video and returns a **creative brief** for @@ -619,7 +705,7 @@ endpoints — no card required: | Free runs | Endpoints | | --- | --- | -| 2 each | text-to-music, text-to-sfx, audio-ducking, video-analysis | +| 2 each | text-to-music, text-to-sfx, audio-ducking, video-analysis, proofread | | 1 each | video-to-music, video-to-sfx, video-to-video-music, video-to-video-sfx, video-to-sound, video-to-video-sound | | 0 | dubbing | diff --git a/context7.json b/context7.json index 2c1dbd3..9f2f27f 100644 --- a/context7.json +++ b/context7.json @@ -27,12 +27,16 @@ "Result media (.url) is a short-lived presigned URL, not the API's own domain \u2014 download it with the result's .save() helper; do not send the Authorization header to it.", "Catch AuthenticationError (401), PaymentRequiredError (402), RateLimitError (429) and TaskFailedError (status \"failed\") separately rather than one generic except \u2014 callers usually handle these differently. All extend SoniloError.", "video / video_url accept exactly one of the two, never both and never neither \u2014 validate before constructing a request.", - "Self-serve accounts start with free runs per endpoint (2 each for text-to-music, text-to-sfx, audio-ducking, video-analysis; 1 each for other video endpoints; none for dubbing), then bill normally. A first call succeeding is not proof billing works.", + "Self-serve accounts start with free runs per endpoint (2 each for text-to-music, text-to-sfx, audio-ducking, video-analysis, proofread; 1 each for other video endpoints; none for dubbing). A first call succeeding is not proof billing works.", "Before a paid call, read client.account.services().get(\"trial\", {}) and degrade gracefully when a service's remaining is 0: that call raises TrialExhaustedError (402 trial_exhausted), which no retry fixes \u2014 ask for a payment method. trial may be absent.", "client.audio_ducking ducks an EXISTING music bed under an EXISTING voice track; nothing is generated. One of voice/voice_url, one of music/music_url. Voice may be audio or video (video returns a .mp4); music must be audio. Result: output_url, no stems.", "client.dubbing dubs one video into many languages in one async call; languages: en, zh_cn, ja, ko, pt, pt_br, es, es_419, de, fr, it, ru, th, ar, tr, vi, id, ta, ml, kn, gu, pa_in, sd_in, hi; default zh_cn,es,fr. Billed per language, no free trial.", "A DubbingResult has no audio/video/output_url. Its results live in result.outputs, a map of language code to dubbed .mp4 URL: use result.save(lang, path) or result.save_all(dir). dubbing's video_url must be https.", "client.video_analysis returns a creative BRIEF, not media: analyze() (not generate()), no save(). Music brief: result.segments (start/end/label/prompt) + result.variations[i].prompt; mode both (default) adds result.sfx_segments + result.sfx_prompt.", - "Pass a video_analysis variation's prompt straight to video_to_music / video_to_sfx / video_to_sound as their prompt. One of video/video_url plus optional prompt, variants_num (1-5, billed per brief), mode (both/music/sfx, same price); max 480s, 10s floor." + "Pass a video_analysis variation's prompt straight to video_to_music / video_to_sfx / video_to_sound as their prompt. One of video/video_url plus optional prompt, variants_num (1-5, billed per brief), mode (both/music/sfx, same price); max 480s, 10s floor.", + "client.proofread transcribes a video and translates the transcript into editable .srt files. One of video/video_url (https), optional languages (the same codes as dubbing) and source_language (a hint; omit it and the language is detected). Max 300s.", + "Omit languages (or send []) to get the source-language transcript alone. ProofreadResult.subtitles maps language to .srt URL and ALWAYS includes the detected source_language, so even a transcript-only request comes back with one file.", + "Proofread is the step before dubbing: save the scripts (result.save(lang, path) / save_all(dir)), correct the wording, then pass the files to client.dubbing as subtitles[] so the dub speaks exactly the approved lines.", + "proofread bills video seconds x the number of target languages ($0.001/sec; a transcript-only request counts as one), 2 free trial calls. result.warnings is language -> non-blocking issues (cue/code/severity + extras) and withholds nothing." ] } diff --git a/pyproject.toml b/pyproject.toml index 1cce918..2cc3412 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -4,7 +4,7 @@ build-backend = "hatchling.build" [project] name = "sonilo" -version = "0.19.0" +version = "0.20.0" description = "Official Python client for the Sonilo API" readme = "README.md" license = "MIT" diff --git a/sonilo-cli/README.md b/sonilo-cli/README.md index 2965996..5759123 100644 --- a/sonilo-cli/README.md +++ b/sonilo-cli/README.md @@ -84,6 +84,8 @@ production sign-in coexist without overwriting each other. # as a new .mp4 with the ducked mix muxed in sonilo video-analysis --video clip.mp4 --variants 2 # prints a creative brief as JSON; generates nothing + sonilo proofread --video clip.mp4 --languages ja,zh_cn --output scripts/clip.srt + # writes scripts/clip..srt, including the detected source language sonilo dubbing --video-url https://example.com/clip.mp4 --languages es,fr --output dubbed.mp4 # writes dubbed.es.mp4 and dubbed.fr.mp4 sonilo tasks get @@ -268,6 +270,47 @@ command that produces no media file — nothing is generated: sonilo video-analysis --video clip.mp4 --output brief.json sonilo video-to-music --video clip.mp4 --prompt "$(jq -r '.variations[0].prompt' brief.json)" +### Proofread + +`proofread` transcribes a video and translates the transcript into editable `.srt` files — one per +language, plus the source-language transcript. The video must have an audio track. Nothing is +dubbed: this is the step **before** `dubbing`, so the wording can be corrected before any voice is +rendered. + + sonilo proofread --video clip.mp4 --languages ja,zh_cn --output scripts/clip.srt + # writes scripts/clip.en.srt, scripts/clip.ja.srt, scripts/clip.zh_cn.srt + +- `--languages` is comma-separated and takes the same codes as `dubbing` (see [Dubbing](#dubbing) + below for the list), so a proofread script can go straight into a dub. Omit it for the + source-language transcript alone. +- `--source-language` tells transcription which language to expect, which helps on short, noisy or + mixed-language audio. Omit it to have the language detected; either way the detected code is + printed and names the source-language file. +- `--output` is a filename template, not a single destination, exactly as it is for `dubbing`: one + `.srt` is written per language with the code inserted before the extension, so + `--output scripts/clip.srt` writes `scripts/clip.en.srt`, `scripts/clip.fr.srt`, etc. Missing + directories are created. Default: `proofread.srt`. Every language is always written — the URLs on + the result are presigned and expire, and the files are the point. +- The source language is **always** returned alongside the requested targets, so a one-language + request writes two files. +- After the files, the command prints the detected source language, the cue count, and one line per + non-blocking warning (`Warning fr: high_text_speed (warning) at cue 33 — ...`). A warning never + withholds a file. +- Source videos may be at most 300 seconds long and 300 MB. Billing is per second of video + multiplied by the number of target languages at $0.001/second, a transcript-only request counts as + one, and billing has a 10-second floor; there are 2 free runs — see [Free trial](#free-trial) + below. +- `--timeout` defaults to 600 seconds, the usual default: a proofread job typically finishes in well + under a minute. If the wait does time out, the task keeps running server-side — resume it with + `sonilo tasks wait `. +- Edit the files, then feed them straight into `dubbing`, which makes the dub speak your exact + wording (drop the source-language file: `--subtitle` must match `--languages`): + + sonilo proofread --video clip.mp4 --languages es,fr --output scripts/clip.srt + # ... correct scripts/clip.es.srt and scripts/clip.fr.srt ... + sonilo dubbing --video clip.mp4 --languages es,fr \ + --subtitle es=scripts/clip.es.srt --subtitle fr=scripts/clip.fr.srt + ### Dubbing `dubbing` dubs a video into one or more target languages in a single async call: @@ -317,7 +360,7 @@ required: | Free runs | Endpoints | | --- | --- | -| 2 each | text-to-music, text-to-sfx, audio-ducking, video-analysis | +| 2 each | text-to-music, text-to-sfx, audio-ducking, video-analysis, proofread | | 1 each | video-to-music, video-to-sfx, video-to-video-music, video-to-video-sfx, video-to-sound, video-to-video-sound | | 0 | dubbing | diff --git a/sonilo-cli/pyproject.toml b/sonilo-cli/pyproject.toml index 09e3d5a..233a692 100644 --- a/sonilo-cli/pyproject.toml +++ b/sonilo-cli/pyproject.toml @@ -4,13 +4,13 @@ build-backend = "hatchling.build" [project] name = "sonilo-cli" -version = "0.18.0" +version = "0.19.0" description = "Command-line interface for the Sonilo API: generate music and sound effects from text or video" readme = "README.md" license = "MIT" requires-python = ">=3.9" authors = [{ name = "Sonilo AI" }] -dependencies = ["sonilo>=0.19.0,<0.20"] +dependencies = ["sonilo>=0.20.0,<0.21"] keywords = ["sonilo", "cli", "music", "sfx", "text-to-music", "video-to-music", "ai"] [project.urls] diff --git a/sonilo-cli/src/sonilo_cli/__init__.py b/sonilo-cli/src/sonilo_cli/__init__.py index 247c7f8..dbd8a6b 100644 --- a/sonilo-cli/src/sonilo_cli/__init__.py +++ b/sonilo-cli/src/sonilo_cli/__init__.py @@ -1,3 +1,3 @@ -__version__ = "0.18.0" +__version__ = "0.19.0" __all__ = ["__version__"] diff --git a/sonilo-cli/src/sonilo_cli/__main__.py b/sonilo-cli/src/sonilo_cli/__main__.py index c154b34..5c666df 100644 --- a/sonilo-cli/src/sonilo_cli/__main__.py +++ b/sonilo-cli/src/sonilo_cli/__main__.py @@ -629,13 +629,17 @@ def cmd_video_analysis(client: Sonilo, args: argparse.Namespace) -> None: DUBBING_WAIT_TIMEOUT = 7200.0 -def _language_path(out: str, language: str) -> str: +def _language_path(out: str, language: str, default_suffix: str = ".mp4") -> str: """Turn one --output value into a per-language path: `clip.mp4` + `es` becomes `clip.es.mp4`. A dubbing task returns one video per language, so a single literal destination cannot express the result. This is the same - transform _stem_path applies for --stem, so both flags read the same way.""" + transform _stem_path applies for --stem, so both flags read the same way. + + `default_suffix` only decides what an extension-less template gets, and + defaults to dubbing's `.mp4`; proofread passes `.srt` so `--output clip` + does not name its subtitles after a video container.""" base = Path(out) - return str(base.with_name(f"{base.stem}.{language}{base.suffix or '.mp4'}")) + return str(base.with_name(f"{base.stem}.{language}{base.suffix or default_suffix}")) def _subtitles(values: Optional[List[str]]) -> Optional[Dict[str, str]]: @@ -726,6 +730,53 @@ def cmd_dubbing(client: Sonilo, args: argparse.Namespace) -> None: print(_export_line(language, result.subtitle_export.get(language, {}))) +def _warning_line(language: str, issue: Any) -> str: + """One line per non-blocking issue. The measurement that came with the + code (`characters_per_second` on `high_text_speed`) is appended verbatim, + because the codes are server-owned and each brings its own.""" + where = f"cue {issue.cue}" if issue.cue is not None else "script" + line = f"Warning {language}: {issue.code} ({issue.severity}) at {where}" + if issue.extras: + details = ", ".join(f"{k}={v}" for k, v in sorted(issue.extras.items())) + line += f" — {details}" + return line + + +def cmd_proofread(client: Sonilo, args: argparse.Namespace) -> None: + """Proofread writes one .srt per language, the same way dubbing writes one + video per language and through the same --output template: the URLs on the + result are presigned and expire, and the whole point of the endpoint is the + files you then edit. The source language always comes back too, whether or + not any target languages were asked for.""" + out = args.output if args.output is not None else "proofread.srt" + languages = None + if args.languages is not None: + languages = [code.strip() for code in args.languages.split(",") if code.strip()] + if not languages: + _fail("--languages needs at least one language code, e.g. --languages es,fr") + result = client.proofread.generate( + video=args.video, + video_url=args.video_url, + languages=languages, + source_language=args.source_language, + timeout=args.timeout, + ) + if not result.subtitles: + _fail("task succeeded but returned no subtitle files") + # Unlike dubbing's, this template routinely names a directory of its own + # ("--output scripts/clip.srt"), so create it rather than fail on the write. + Path(out).parent.mkdir(parents=True, exist_ok=True) + for language in sorted(result.subtitles): + path = result.save(language, _language_path(out, language, ".srt")) + _wrote(path, path.stat().st_size) + print(f"Source language: {result.source_language or 'unknown'}") + if result.cue_count is not None: + print(f"Cues: {result.cue_count}") + for language in sorted(result.warnings): + for issue in result.warnings[language]: + print(_warning_line(language, issue)) + + def _identity(body: Any) -> Any: return body @@ -1154,6 +1205,38 @@ def build_parser() -> argparse.ArgumentParser: ) p_dub.set_defaults(func=cmd_dubbing) + p_pr = sub.add_parser( + "proofread", + help="Transcribe a video and translate the transcript into editable .srt files", + ) + _add_global(p_pr) + _add_video_source(p_pr) + p_pr.add_argument( + "--languages", default=None, + help="Comma-separated target languages to translate the transcript into. " + "Omit it for the source-language transcript alone. Same codes as " + "`sonilo dubbing`, so a proofread script can go straight into a dub.", + ) + p_pr.add_argument( + "--source-language", dest="source_language", default=None, + help="Tell transcription which language to expect, which helps on short, " + "noisy or mixed-language audio. One of the same codes. Omit it to " + "have the language detected.", + ) + p_pr.add_argument( + "--output", default=None, + help="Filename template, not a single destination: one .srt is written " + "per language with the code inserted before the extension " + "(scripts/clip.srt -> scripts/clip.en.srt). Missing directories are " + "created. Default: proofread.srt", + ) + p_pr.add_argument( + "--timeout", type=float, default=DEFAULT_WAIT_TIMEOUT, + help="Give up waiting after this many seconds. Default: 600. A timed-out " + "task may still finish — resume it with `sonilo tasks wait `.", + ) + p_pr.set_defaults(func=cmd_proofread) + p_tasks = sub.add_parser("tasks", help="Inspect async tasks") _add_global(p_tasks) tsub = p_tasks.add_subparsers(dest="tasks_command", metavar="") diff --git a/sonilo-cli/tests/test_cli.py b/sonilo-cli/tests/test_cli.py index 5439db3..0990d5e 100644 --- a/sonilo-cli/tests/test_cli.py +++ b/sonilo-cli/tests/test_cli.py @@ -963,6 +963,167 @@ def test_dubbing_rejects_a_subtitle_that_is_not_srt_or_vtt(tmp_path, capsys): assert not route.called +# --- proofread ------------------------------------------------------------- + +PROOFREAD_BODY = { + "task_id": "pr1", + "type": "proofread", + "status": "succeeded", + "duration_seconds": 206.32, + "source_language": "en", + "subtitles": {"en": "https://r2/en.srt", "fr": "https://r2/fr.srt"}, + "cue_count": 65, + "warnings": { + "fr": [ + { + "cue": 33, + "code": "high_text_speed", + "severity": "warning", + "characters_per_second": 26.92, + } + ] + }, +} + + +def _stub_proofread(body=None): + route = respx.post(f"{BASE}/v1/proofread").mock( + return_value=httpx.Response(202, json={"task_id": "pr1", "status": "processing"}) + ) + respx.get(f"{BASE}/v1/tasks/pr1").mock( + return_value=httpx.Response(200, json=body or PROOFREAD_BODY) + ) + for language in (body or PROOFREAD_BODY)["subtitles"]: + respx.get(f"https://r2/{language}.srt").mock( + return_value=httpx.Response(200, content=f"{language}-bytes".encode()) + ) + return route + + +@respx.mock +def test_proofread_writes_one_srt_per_language(tmp_path, capsys): + """--output is a template, the same one dubbing uses: the language code is + inserted before the extension, and the directory is created.""" + _stub_proofread() + run([ + "proofread", + "--video-url", "https://x/v.mp4", + "--languages", "fr", + "--output", str(tmp_path / "scripts" / "clip.srt"), + ]) + # The detected source language comes back alongside the requested target, + # so a one-language request still writes two files. + assert (tmp_path / "scripts" / "clip.en.srt").read_bytes() == b"en-bytes" + assert (tmp_path / "scripts" / "clip.fr.srt").read_bytes() == b"fr-bytes" + out = capsys.readouterr().out + assert "Source language: en" in out + assert "Cues: 65" in out + + +@respx.mock +def test_proofread_prints_the_warnings(tmp_path, capsys): + _stub_proofread() + run([ + "proofread", "--video-url", "https://x/v.mp4", + "--languages", "fr", "--output", str(tmp_path / "clip.srt"), + ]) + out = capsys.readouterr().out + assert "Warning fr: high_text_speed (warning) at cue 33" in out + # The code's own measurement is printed with it, not dropped. + assert "characters_per_second=26.92" in out + + +@respx.mock +def test_proofread_sends_languages_as_a_json_array(tmp_path): + route = _stub_proofread() + run([ + "proofread", "--video-url", "https://x/v.mp4", + "--languages", " fr , en ", "--output", str(tmp_path / "clip.srt"), + ]) + body = unquote_plus(route.calls.last.request.content.decode()) + assert '["fr", "en"]' in body + + +@respx.mock +def test_proofread_without_languages_omits_the_field(tmp_path): + route = _stub_proofread({ + "task_id": "pr1", "status": "succeeded", + "source_language": "en", + "subtitles": {"en": "https://r2/en.srt"}, + }) + run([ + "proofread", "--video-url", "https://x/v.mp4", + "--output", str(tmp_path / "clip.srt"), + ]) + # Omitting --languages asks for the transcript alone; sending the field at + # all would be a different request. + assert b"languages" not in route.calls.last.request.content + assert (tmp_path / "clip.en.srt").exists() + + +@respx.mock +def test_proofread_sends_source_language(tmp_path): + route = _stub_proofread() + run([ + "proofread", "--video-url", "https://x/v.mp4", + "--source-language", "en", "--output", str(tmp_path / "clip.srt"), + ]) + assert "source_language=en" in unquote_plus(route.calls.last.request.content.decode()) + + +@respx.mock +def test_proofread_output_without_an_extension_still_gets_srt(tmp_path): + """The shared template transform defaults an extension-less value to + dubbing's .mp4; proofread has to override that or name subtitles after a + video container.""" + _stub_proofread() + run([ + "proofread", "--video-url", "https://x/v.mp4", "--languages", "fr", + "--output", str(tmp_path / "interview"), + ]) + assert (tmp_path / "interview.fr.srt").exists() + assert not (tmp_path / "interview.fr.mp4").exists() + + +@respx.mock +def test_proofread_default_output_template(tmp_path, monkeypatch): + """With no --output the files land beside the caller as proofread..srt, + the way dubbing defaults to output.mp4.""" + _stub_proofread() + monkeypatch.chdir(tmp_path) + run(["proofread", "--video-url", "https://x/v.mp4", "--languages", "fr"]) + assert (tmp_path / "proofread.en.srt").exists() + assert (tmp_path / "proofread.fr.srt").exists() + + +def test_proofread_requires_a_video_source(capsys): + with pytest.raises(SystemExit) as exc: + main(["--api-key", "sk-test", "proofread"]) + assert exc.value.code == 1 + # Assert the message too: a bare exit code would also pass if argparse + # bailed out for some unrelated reason. + assert "--video" in capsys.readouterr().err + + +@respx.mock +def test_proofread_non_https_url_exits_1(capsys): + with pytest.raises(SystemExit) as exc: + main(["--api-key", "sk-test", "proofread", "--video-url", "http://x/v.mp4"]) + assert exc.value.code == 1 + assert "https" in capsys.readouterr().err + + +@respx.mock +def test_proofread_empty_languages_value_exits_1(capsys): + with pytest.raises(SystemExit) as exc: + main([ + "--api-key", "sk-test", "proofread", + "--video-url", "https://x/v.mp4", "--languages", " , ", + ]) + assert exc.value.code == 1 + assert "--languages" in capsys.readouterr().err + + # --- --segments ----------------------------------------------------------- # # The two shapes are not interchangeable: music segments are diff --git a/src/sonilo/__init__.py b/src/sonilo/__init__.py index d6c055a..b8f848f 100644 --- a/src/sonilo/__init__.py +++ b/src/sonilo/__init__.py @@ -23,6 +23,8 @@ MusicResult, MusicStems, MusicTitle, + ProofreadIssue, + ProofreadResult, Segment, SfxMedia, SfxResult, @@ -51,6 +53,8 @@ "MusicStems", "MusicTitle", "PaymentRequiredError", + "ProofreadIssue", + "ProofreadResult", "RateLimitError", "Segment", "SfxMedia", diff --git a/src/sonilo/_async_client.py b/src/sonilo/_async_client.py index be7fcc6..eb09df7 100644 --- a/src/sonilo/_async_client.py +++ b/src/sonilo/_async_client.py @@ -10,6 +10,7 @@ from sonilo.resources.account import AsyncAccount from sonilo.resources.audio_ducking import AsyncAudioDucking from sonilo.resources.dubbing import AsyncDubbing +from sonilo.resources.proofread import AsyncProofread from sonilo.resources.video_analysis import AsyncVideoAnalysis from sonilo.resources.tasks import AsyncTasks from sonilo.resources.text_to_music import AsyncTextToMusic @@ -54,6 +55,7 @@ def __init__( self.video_to_video_sound = AsyncVideoToVideoSound(self) self.audio_ducking = AsyncAudioDucking(self) self.dubbing = AsyncDubbing(self) + self.proofread = AsyncProofread(self) self.video_analysis = AsyncVideoAnalysis(self) self.account = AsyncAccount(self) self.tasks = AsyncTasks(self) diff --git a/src/sonilo/_client.py b/src/sonilo/_client.py index 614100a..0742d94 100644 --- a/src/sonilo/_client.py +++ b/src/sonilo/_client.py @@ -11,6 +11,7 @@ from sonilo.resources.account import Account from sonilo.resources.audio_ducking import AudioDucking from sonilo.resources.dubbing import Dubbing +from sonilo.resources.proofread import Proofread from sonilo.resources.tasks import Tasks from sonilo.resources.text_to_music import TextToMusic from sonilo.resources.text_to_sfx import TextToSfx @@ -84,6 +85,7 @@ def __init__( self.video_to_video_sound = VideoToVideoSound(self) self.audio_ducking = AudioDucking(self) self.dubbing = Dubbing(self) + self.proofread = Proofread(self) self.video_analysis = VideoAnalysis(self) self.account = Account(self) self.tasks = Tasks(self) diff --git a/src/sonilo/_requests.py b/src/sonilo/_requests.py index 2ec6a04..653bedd 100644 --- a/src/sonilo/_requests.py +++ b/src/sonilo/_requests.py @@ -299,6 +299,63 @@ def build_dubbing_parts( return data, files or None, MultiClose(opened) if opened else None +def build_proofread_parts( + video: Any, + video_url: Optional[str], + languages: Optional[List[str]] = None, + source_language: Optional[str] = None, +) -> Tuple[Dict[str, str], Optional[Dict[str, tuple]], bool]: + """Build the multipart parts for POST /v1/proofread. + + Proofread is dubbing's sibling — it transcribes the video and translates + the transcript so the `.srt` files can be corrected before they are sent + back to /v1/dubbing as `subtitles[]` — so the two builders agree + field for field wherever they overlap. `languages` travels as one opaque + form field holding a JSON array string, exactly as it does for dubbing, + and is omitted entirely when unset (which asks for the source-language + transcript alone). Unlike dubbing's, an empty list IS meaningful here and + is sent as `[]`: it is the explicit spelling of "transcript only". + + `source_language` is a hint for the transcription, sent only when given; + without it the language is detected and reported back on the finished + task. Language codes are deliberately NOT checked here — the backend owns + that list, exactly as it does for dubbing, and a hardcoded copy would make + this SDK reject codes added later. + + The https check on `video_url` is local for the same reason as dubbing's: + this pipeline fetches the source URL itself and requires https + specifically, so a plain-http URL is a guaranteed server-side 422. + + Only one file can ever be opened here (there are no subtitle uploads on + the way in), so this returns the plain `opened` bool the single-input + builders use rather than dubbing's `MultiClose`. + """ + if (video is None) == (video_url is None): + raise SoniloError("Provide exactly one of video or video_url") + + # Assemble data dict completely before opening any files + data: Dict[str, str] = {} + if video_url is not None: + if not video_url.lower().startswith("https://"): + raise SoniloError( + "video_url must use https — the proofread pipeline requires an https URL" + ) + data["video_url"] = video_url + if languages is not None: + data["languages"] = json.dumps(languages) + if source_language is not None: + data["source_language"] = source_language + + # Now open files (only after data is fully assembled) + files: Optional[Dict[str, tuple]] = None + opened = False + if video is not None: + filename, fileobj, opened = normalize_video(video) + files = {"video": (filename, fileobj, "video/mp4")} + + return data, files, opened + + def build_video_analysis_parts( video: Any, video_url: Optional[str], diff --git a/src/sonilo/_version.py b/src/sonilo/_version.py index 11ac8e1..5f4bb0b 100644 --- a/src/sonilo/_version.py +++ b/src/sonilo/_version.py @@ -1 +1 @@ -__version__ = "0.19.0" +__version__ = "0.20.0" diff --git a/src/sonilo/resources/proofread.py b/src/sonilo/resources/proofread.py new file mode 100644 index 0000000..77c15a0 --- /dev/null +++ b/src/sonilo/resources/proofread.py @@ -0,0 +1,129 @@ +from __future__ import annotations + +from typing import TYPE_CHECKING, Any, List, Optional + +from sonilo._requests import build_proofread_parts +from sonilo.resources.tasks import ( + DEFAULT_POLL_INTERVAL, + DEFAULT_WAIT_TIMEOUT, + parse_proofread_result, + parse_sfx_task, +) +from sonilo.types import ProofreadResult, SfxTask + +if TYPE_CHECKING: + from sonilo._async_client import AsyncSonilo + from sonilo._client import Sonilo + +PATH = "/v1/proofread" + + +class Proofread: + """Transcribe one video and translate the transcript into editable + subtitle files. Async only; the result carries a language → `.srt`-URL map + under `subtitles`. + + This is the step before `client.dubbing`, not a replacement for it: + proofread returns one `.srt` per language plus the source-language + transcript, you review or correct the wording, and the corrected files go + to `client.dubbing` as `subtitles[]` so the dub speaks exactly + the approved lines. The language codes are the same on both endpoints, so + a proofread script can go straight into a dub. + + Pass exactly one of `video` / `video_url` (`video_url` must be **https**), + plus optional `languages` — the target languages to translate into, sent + as a JSON array. Omit it, or pass `[]`, for the source-language transcript + alone. `source_language` is an optional hint telling transcription which + language to expect, which helps on short, noisy or mixed-language audio; + without it the language is detected, and either way the finished task + reports what the transcript is in. + + Billing is per second of video multiplied by the number of target + languages; a transcript-only request counts as one. + """ + + def __init__(self, client: "Sonilo") -> None: + self._client = client + + def submit( + self, + *, + video: Any = None, + video_url: Optional[str] = None, + languages: Optional[List[str]] = None, + source_language: Optional[str] = None, + ) -> SfxTask: + data, files, opened = build_proofread_parts( + video, video_url, languages, source_language + ) + close_after = files["video"][1] if files is not None and opened else None + return parse_sfx_task( + self._client._post_json(PATH, data=data, files=files, close_after=close_after) + ) + + def generate( + self, + *, + video: Any = None, + video_url: Optional[str] = None, + languages: Optional[List[str]] = None, + source_language: Optional[str] = None, + poll_interval: float = DEFAULT_POLL_INTERVAL, + timeout: float = DEFAULT_WAIT_TIMEOUT, + ) -> ProofreadResult: + task = self.submit( + video=video, video_url=video_url, languages=languages, + source_language=source_language, + ) + return self._client.tasks.wait( + task.task_id, + poll_interval=poll_interval, + timeout=timeout, + parser=parse_proofread_result, + ) + + +class AsyncProofread: + """Async twin of Proofread; same parameters and same result shape.""" + + def __init__(self, client: "AsyncSonilo") -> None: + self._client = client + + async def submit( + self, + *, + video: Any = None, + video_url: Optional[str] = None, + languages: Optional[List[str]] = None, + source_language: Optional[str] = None, + ) -> SfxTask: + data, files, opened = build_proofread_parts( + video, video_url, languages, source_language + ) + close_after = files["video"][1] if files is not None and opened else None + return parse_sfx_task( + await self._client._post_json( + PATH, data=data, files=files, close_after=close_after + ) + ) + + async def generate( + self, + *, + video: Any = None, + video_url: Optional[str] = None, + languages: Optional[List[str]] = None, + source_language: Optional[str] = None, + poll_interval: float = DEFAULT_POLL_INTERVAL, + timeout: float = DEFAULT_WAIT_TIMEOUT, + ) -> ProofreadResult: + task = await self.submit( + video=video, video_url=video_url, languages=languages, + source_language=source_language, + ) + return await self._client.tasks.wait( + task.task_id, + poll_interval=poll_interval, + timeout=timeout, + parser=parse_proofread_result, + ) diff --git a/src/sonilo/resources/tasks.py b/src/sonilo/resources/tasks.py index 7d1367d..6a93f2a 100644 --- a/src/sonilo/resources/tasks.py +++ b/src/sonilo/resources/tasks.py @@ -15,6 +15,8 @@ MusicResult, MusicStems, MusicTitle, + ProofreadIssue, + ProofreadResult, SfxMedia, SfxResult, SfxTask, @@ -286,6 +288,83 @@ def parse_dubbing_result(body: Dict[str, Any]) -> "DubbingResult": raise SoniloError(f"Malformed task response: missing {e.args[0]!r}") from e +_ISSUE_NAMED_FIELDS = ("cue", "code", "severity") + + +def _proofread_issue_from(data: Any) -> Optional[ProofreadIssue]: + """Coerce one warning entry. Everything the check reports beyond the three + named fields is kept in `extras` rather than dropped: the issue codes are + server-owned and each brings its own measurement (`high_text_speed` brings + `characters_per_second`), so a code added later must still arrive whole. + `cue` is read leniently — the pipeline's store can hand a number back as a + string — and reads as None when it is neither.""" + if not isinstance(data, dict): + return None + try: + cue: Optional[int] = int(data["cue"]) + except (KeyError, TypeError, ValueError): + cue = None + return ProofreadIssue( + code=str(data.get("code") or "unknown"), + severity=str(data.get("severity") or "warning"), + cue=cue, + extras={k: v for k, v in data.items() if k not in _ISSUE_NAMED_FIELDS}, + ) + + +def _proofread_warnings_from(data: Any) -> Dict[str, List[ProofreadIssue]]: + """Coerce the language → issue-list map, dropping malformed entries, for + the same reason _url_map_from coerces dubbing's outputs: a differently + shaped entry from a backend change should surface here and not as an + AttributeError deep inside the caller's loop.""" + if not isinstance(data, dict): + return {} + warnings: Dict[str, List[ProofreadIssue]] = {} + for language, issues in data.items(): + if not isinstance(issues, list): + continue + warnings[str(language)] = [ + issue for issue in map(_proofread_issue_from, issues) if issue is not None + ] + return warnings + + +def _int_or_none(value: Any) -> Optional[int]: + """`cue_count` may arrive as a number or, from a store that keeps numbers + as strings, as a string. Read either; anything else is None.""" + try: + return int(value) + except (TypeError, ValueError): + return None + + +def parse_proofread_result(body: Dict[str, Any]) -> "ProofreadResult": + """Map a GET /v1/tasks/{id} body for a proofread task to ProofreadResult; + unknown fields are ignored. + + `subtitles` is coerced exactly as dubbing's `outputs` is — see + _url_map_from — and always carries the detected source language alongside + the requested targets. `warnings` is `{}` when the scripts raised nothing, + which is the common case; it never gates delivery of a file. + """ + try: + return ProofreadResult( + task_id=body["task_id"], + status=body["status"], + type=body.get("type"), + subtitles=_url_map_from(body.get("subtitles")), + source_language=body.get("source_language"), + cue_count=_int_or_none(body.get("cue_count")), + warnings=_proofread_warnings_from(body.get("warnings")), + duration_seconds=body.get("duration_seconds"), + cost=body.get("cost"), + error=body.get("error"), + refunded=body.get("refunded"), + ) + except KeyError as e: + raise SoniloError(f"Malformed task response: missing {e.args[0]!r}") from e + + def _analysis_segment_from(data: Any) -> Optional[AnalysisSegment]: if not isinstance(data, dict): return None diff --git a/src/sonilo/types.py b/src/sonilo/types.py index 5dc5e35..f1bf098 100644 --- a/src/sonilo/types.py +++ b/src/sonilo/types.py @@ -777,6 +777,129 @@ async def asave_all_subtitles( } +@dataclass +class ProofreadIssue: + """One non-blocking validation issue on a proofread script. + + The fields every issue carries are named; anything else the check reports + is kept verbatim in `extras` rather than dropped, because the issue codes + are server-owned and grow — `high_text_speed` comes with + `characters_per_second`, and a code added later will come with fields this + SDK has never heard of. `cue` is the 1-based index of the subtitle cue the + issue is about, and is None only when the report omitted it. + """ + + code: str + severity: str + cue: Optional[int] = None + extras: Dict[str, Any] = field(default_factory=dict) + + def get(self, key: str, default: Any = None) -> Any: + """Read one pass-through field, e.g. `issue.get("characters_per_second")`.""" + return self.extras.get(key, default) + + +@dataclass +class ProofreadResult: + """State of a proofread task (`tasks.get`) or its final result + (`wait`/`generate`). + + Shaped like DubbingResult, because the two are two halves of one workflow: + proofread transcribes a video and translates the transcript, you correct + the wording, and the corrected files go back to `client.dubbing` as + `subtitles[]` so the dub speaks exactly what was approved. + + `subtitles` is a language → presigned `.srt`-URL map and always includes + the DETECTED source language (reported in `source_language`) alongside one + entry per requested target language, so a request with no `languages` at + all still comes back with one file. `save(language, path)` fetches one and + `save_all(dir)` fetches every one of them, mirroring DubbingResult. + + `cue_count` is the number of subtitle cues in the source script; every + language has the same count, since translation is cue-by-cue. `warnings` + maps a language to the non-blocking issues its script raised and is empty + when there are none — nothing in it fails the task or withholds a file. + """ + + task_id: str + status: str + type: Optional[str] = None + subtitles: Dict[str, str] = field(default_factory=dict) + source_language: Optional[str] = None + cue_count: Optional[int] = None + warnings: Dict[str, List[ProofreadIssue]] = field(default_factory=dict) + duration_seconds: Optional[float] = None + cost: Optional[float] = None + error: Optional[Dict[str, Any]] = None + refunded: Optional[bool] = None + + def _url(self, language: str) -> str: + if language not in self.subtitles: + available = ", ".join(sorted(self.subtitles)) or "none" + raise SoniloError( + f"No subtitle for language {language!r} on this result " + f"(status={self.status}; available: {available})" + ) + return self.subtitles[language] + + def save( + self, + language: str, + path: Union[str, Path], + *, + timeout: float = DOWNLOAD_TIMEOUT, + ) -> Path: + """Download one language's `.srt` to `path` and return it. The URL is + presigned — no API key is sent.""" + return _download_to(self._url(language), path, timeout) + + async def asave( + self, + language: str, + path: Union[str, Path], + *, + timeout: float = DOWNLOAD_TIMEOUT, + ) -> Path: + """Async variant of save().""" + return await _adownload_to(self._url(language), path, timeout) + + def save_all( + self, + directory: Union[str, Path], + *, + prefix: str = "proofread", + timeout: float = DOWNLOAD_TIMEOUT, + ) -> Dict[str, Path]: + """Download every language into `directory` as + `{prefix}.{language}.srt`, returning the language → path map. The + directory is created if it does not exist.""" + target = Path(directory) + target.mkdir(parents=True, exist_ok=True) + return { + language: self.save( + language, target / f"{prefix}.{language}.srt", timeout=timeout + ) + for language in sorted(self.subtitles) + } + + async def asave_all( + self, + directory: Union[str, Path], + *, + prefix: str = "proofread", + timeout: float = DOWNLOAD_TIMEOUT, + ) -> Dict[str, Path]: + """Async variant of save_all().""" + target = Path(directory) + target.mkdir(parents=True, exist_ok=True) + return { + language: await self.asave( + language, target / f"{prefix}.{language}.srt", timeout=timeout + ) + for language in sorted(self.subtitles) + } + + @dataclass class AnalysisSegment: """One time-aligned section of the analyzed video, with the creative diff --git a/tests/test_proofread.py b/tests/test_proofread.py new file mode 100644 index 0000000..b02bba1 --- /dev/null +++ b/tests/test_proofread.py @@ -0,0 +1,371 @@ +from pathlib import Path +from urllib.parse import unquote_plus + +import httpx +import pytest +import respx + +from sonilo import AsyncSonilo, ProofreadIssue, ProofreadResult, Sonilo +from sonilo._requests import build_proofread_parts +from sonilo.errors import SoniloError +from sonilo.types import SfxTask +from sonilo.resources.tasks import parse_proofread_result + +# The real prod body (task 4288764d..., 2026-09-18), URLs shortened. Keeping +# the actual shape means the parser is checked against what the API sends — +# including `warnings` carrying a code with its own measurement field. +SUCCESS_BODY = { + "task_id": "4288764d-0057-4a77-885e-03c0391c7c1d", + "type": "proofread", + "status": "succeeded", + "duration_seconds": 206.32, + "source_language": "en", + "subtitles": { + "en": "https://r2/en.srt", + "ko": "https://r2/ko.srt", + "fr": "https://r2/fr.srt", + "de": "https://r2/de.srt", + "ar": "https://r2/ar.srt", + "th": "https://r2/th.srt", + "ru": "https://r2/ru.srt", + }, + "cue_count": 65, + "warnings": { + "fr": [ + { + "cue": 33, + "code": "high_text_speed", + "severity": "warning", + "characters_per_second": 26.92, + } + ] + }, +} + +ACK = {"task_id": "pr1", "status": "processing"} + + +# --- result parsing -------------------------------------------------------- + + +def test_parse_proofread_result_reads_the_subtitles_map(): + result = parse_proofread_result(SUCCESS_BODY) + assert result.task_id == "4288764d-0057-4a77-885e-03c0391c7c1d" + assert result.status == "succeeded" + assert result.type == "proofread" + assert result.source_language == "en" + assert result.cue_count == 65 + assert result.duration_seconds == 206.32 + # The detected source language is always in the map, alongside the + # requested targets. + assert "en" in result.subtitles + assert result.subtitles["fr"] == "https://r2/fr.srt" + assert len(result.subtitles) == 7 + + +def test_parse_proofread_result_reads_the_warnings(): + result = parse_proofread_result(SUCCESS_BODY) + assert set(result.warnings) == {"fr"} + issue = result.warnings["fr"][0] + assert isinstance(issue, ProofreadIssue) + assert issue.cue == 33 + assert issue.code == "high_text_speed" + assert issue.severity == "warning" + # Everything past the three named fields is kept verbatim, because each + # server-owned code brings its own measurement. + assert issue.extras == {"characters_per_second": 26.92} + assert issue.get("characters_per_second") == 26.92 + assert issue.get("nothing_like_this") is None + + +def test_parse_proofread_result_defaults_warnings_to_empty(): + result = parse_proofread_result( + {"task_id": "pr1", "status": "succeeded", "subtitles": {"en": "https://r2/en.srt"}} + ) + assert result.warnings == {} + assert result.source_language is None + assert result.cue_count is None + + +def test_parse_proofread_result_defaults_subtitles_to_empty(): + result = parse_proofread_result({"task_id": "pr1", "status": "processing"}) + assert result.subtitles == {} + + +def test_parse_proofread_result_tolerates_stringy_numbers(): + """A store that keeps numbers as strings is what made dubbing's reports + arrive as strings; read either shape here rather than crash.""" + result = parse_proofread_result( + { + "task_id": "pr1", + "status": "succeeded", + "cue_count": "65", + "warnings": {"fr": [{"cue": "7", "code": "high_text_speed"}]}, + } + ) + assert result.cue_count == 65 + assert result.warnings["fr"][0].cue == 7 + # severity is defaulted rather than dropped, so callers can print it. + assert result.warnings["fr"][0].severity == "warning" + + +def test_parse_proofread_result_drops_malformed_warning_entries(): + result = parse_proofread_result( + { + "task_id": "pr1", + "status": "succeeded", + "subtitles": "nope", + "warnings": {"fr": "not-a-list", "de": ["not-an-issue", {"code": "x"}]}, + } + ) + assert result.subtitles == {} + assert "fr" not in result.warnings + assert [issue.code for issue in result.warnings["de"]] == ["x"] + + +def test_parse_proofread_result_rejects_a_body_without_a_task_id(): + with pytest.raises(SoniloError): + parse_proofread_result({"status": "succeeded"}) + + +# --- saving ---------------------------------------------------------------- + + +@respx.mock +def test_save_downloads_one_language(tmp_path): + respx.get("https://r2/fr.srt").mock( + return_value=httpx.Response(200, content=b"1\nbonjour\n") + ) + result = parse_proofread_result(SUCCESS_BODY) + path = result.save("fr", tmp_path / "clip.fr.srt") + assert path.read_bytes() == b"1\nbonjour\n" + + +def test_save_rejects_a_language_the_task_did_not_produce(tmp_path): + result = parse_proofread_result(SUCCESS_BODY) + with pytest.raises(SoniloError): + result.save("es", tmp_path / "clip.es.srt") + + +@respx.mock +def test_save_all_writes_one_srt_per_language(tmp_path): + for language in SUCCESS_BODY["subtitles"]: + respx.get(f"https://r2/{language}.srt").mock( + return_value=httpx.Response(200, content=f"{language}-bytes".encode()) + ) + result = parse_proofread_result(SUCCESS_BODY) + paths = result.save_all(tmp_path / "out") + assert set(paths) == set(SUCCESS_BODY["subtitles"]) + assert (tmp_path / "out" / "proofread.en.srt").read_bytes() == b"en-bytes" + assert (tmp_path / "out" / "proofread.ko.srt").read_bytes() == b"ko-bytes" + + +@respx.mock +async def test_asave_all_writes_one_srt_per_language(tmp_path): + for language in SUCCESS_BODY["subtitles"]: + respx.get(f"https://r2/{language}.srt").mock( + return_value=httpx.Response(200, content=f"{language}-bytes".encode()) + ) + result = parse_proofread_result(SUCCESS_BODY) + paths = await result.asave_all(tmp_path / "out", prefix="clip") + assert set(paths) == set(SUCCESS_BODY["subtitles"]) + assert (tmp_path / "out" / "clip.en.srt").read_bytes() == b"en-bytes" + assert (tmp_path / "out" / "clip.th.srt").read_bytes() == b"th-bytes" + + +@respx.mock +async def test_asave_downloads_one_language(tmp_path): + respx.get("https://r2/de.srt").mock( + return_value=httpx.Response(200, content=b"1\nhallo\n") + ) + result = parse_proofread_result(SUCCESS_BODY) + path = await result.asave("de", tmp_path / "clip.de.srt") + assert path.read_bytes() == b"1\nhallo\n" + + +# --- request shape --------------------------------------------------------- + + +@respx.mock +def test_submit_posts_to_v1_proofread(): + route = respx.post("https://api.sonilo.com/v1/proofread").mock( + return_value=httpx.Response(202, json=ACK) + ) + with Sonilo(api_key="sk-test") as client: + task = client.proofread.submit( + video_url="https://x/v.mp4", languages=["ja", "zh_cn"] + ) + assert isinstance(task, SfxTask) + assert task.task_id == "pr1" and task.status == "processing" + # video_url-only submissions have no file part, so httpx sends this as + # application/x-www-form-urlencoded (not multipart) and percent-escapes + # the JSON array; unquote before checking for the raw JSON substring. + sent = unquote_plus(route.calls.last.request.content.decode()) + assert "video_url=https://x/v.mp4" in sent + assert '["ja", "zh_cn"]' in sent + + +@respx.mock +def test_submit_omits_languages_when_unset(): + """Omitting the field is what asks for the transcript alone; sending + `languages=None` as a string would be a 422.""" + route = respx.post("https://api.sonilo.com/v1/proofread").mock( + return_value=httpx.Response(202, json=ACK) + ) + with Sonilo(api_key="sk-test") as client: + client.proofread.submit(video_url="https://x/v.mp4") + assert b"languages" not in route.calls.last.request.content + + +@respx.mock +def test_submit_sends_an_explicit_empty_language_list(): + """Unlike dubbing, `[]` is meaningful here — it is the explicit spelling + of "transcript only" — so it goes on the wire rather than being dropped.""" + route = respx.post("https://api.sonilo.com/v1/proofread").mock( + return_value=httpx.Response(202, json=ACK) + ) + with Sonilo(api_key="sk-test") as client: + client.proofread.submit(video_url="https://x/v.mp4", languages=[]) + sent = unquote_plus(route.calls.last.request.content.decode()) + assert "languages=[]" in sent + + +@respx.mock +def test_submit_sends_source_language_only_when_given(): + route = respx.post("https://api.sonilo.com/v1/proofread").mock( + return_value=httpx.Response(202, json=ACK) + ) + with Sonilo(api_key="sk-test") as client: + client.proofread.submit(video_url="https://x/v.mp4", source_language="en") + client.proofread.submit(video_url="https://x/v.mp4") + bodies = [unquote_plus(c.request.content.decode()) for c in route.calls] + assert "source_language=en" in bodies[0] + # Unset → omitted entirely so the server detects the language itself. + assert "source_language" not in bodies[1] + + +@respx.mock +def test_submit_uploads_a_local_video_as_a_file_part(tmp_path): + route = respx.post("https://api.sonilo.com/v1/proofread").mock( + return_value=httpx.Response(202, json=ACK) + ) + clip = tmp_path / "interview.mp4" + clip.write_bytes(b"fake-mp4") + with Sonilo(api_key="sk-test") as client: + client.proofread.submit(video=str(clip), languages=["ja"]) + body = route.calls.last.request.content.decode(errors="replace") + assert 'name="video"' in body + assert 'filename="interview.mp4"' in body + # multipart, so the JSON array travels as a plain field value. + assert '["ja"]' in body + + +@respx.mock +def test_submit_rejects_a_non_https_url_before_sending(): + route = respx.post("https://api.sonilo.com/v1/proofread") + with Sonilo(api_key="sk-test") as client: + with pytest.raises(SoniloError): + client.proofread.submit(video_url="http://x/v.mp4") + assert not route.called + + +@respx.mock +def test_submit_requires_exactly_one_video_input(): + route = respx.post("https://api.sonilo.com/v1/proofread") + with Sonilo(api_key="sk-test") as client: + with pytest.raises(SoniloError): + client.proofread.submit() + with pytest.raises(SoniloError): + client.proofread.submit(video=b"bytes", video_url="https://x/v.mp4") + assert not route.called + + +@respx.mock +def test_submit_closes_a_local_video_even_when_the_api_rejects_it(tmp_path, monkeypatch): + respx.post("https://api.sonilo.com/v1/proofread").mock( + return_value=httpx.Response(422, json={"error": {"code": "unprocessable_entity"}}) + ) + clip = tmp_path / "clip.mp4" + clip.write_bytes(b"fake-mp4") + opened = [] + real_open = Path.open + + def spy(self, *args, **kwargs): + handle = real_open(self, *args, **kwargs) + opened.append(handle) + return handle + + monkeypatch.setattr(Path, "open", spy) + with Sonilo(api_key="sk-test") as client: + with pytest.raises(SoniloError): + client.proofread.submit(video=str(clip)) + assert opened and all(handle.closed for handle in opened) + + +def test_build_proofread_parts_passes_unknown_codes_through(): + """Language codes are server-owned, exactly as on dubbing: a code this + SDK has never heard of must reach the API and earn its own 422.""" + data, _, _ = build_proofread_parts(None, "https://x/v.mp4", ["klingon"], "klingon") + assert data["languages"] == '["klingon"]' + assert data["source_language"] == "klingon" + + +# --- polling --------------------------------------------------------------- + + +@respx.mock +def test_generate_polls_to_a_proofread_result(): + respx.post("https://api.sonilo.com/v1/proofread").mock( + return_value=httpx.Response(202, json=ACK) + ) + respx.get("https://api.sonilo.com/v1/tasks/pr1").mock( + return_value=httpx.Response(200, json=dict(SUCCESS_BODY, task_id="pr1")) + ) + with Sonilo(api_key="sk-test") as client: + result = client.proofread.generate( + video_url="https://x/v.mp4", languages=["fr", "de"], poll_interval=0 + ) + assert isinstance(result, ProofreadResult) + assert result.source_language == "en" + assert result.subtitles["fr"] == "https://r2/fr.srt" + assert result.warnings["fr"][0].code == "high_text_speed" + + +@respx.mock +def test_tasks_get_parses_a_proofread_task(): + respx.get("https://api.sonilo.com/v1/tasks/pr1").mock( + return_value=httpx.Response(200, json=dict(SUCCESS_BODY, task_id="pr1")) + ) + with Sonilo(api_key="sk-test") as client: + result = client.tasks.get("pr1", parser=parse_proofread_result) + assert result.type == "proofread" + assert result.cue_count == 65 + + +@respx.mock +async def test_async_generate_polls_to_a_proofread_result(): + respx.post("https://api.sonilo.com/v1/proofread").mock( + return_value=httpx.Response(202, json=ACK) + ) + respx.get("https://api.sonilo.com/v1/tasks/pr1").mock( + return_value=httpx.Response(200, json=dict(SUCCESS_BODY, task_id="pr1")) + ) + async with AsyncSonilo(api_key="sk-test") as client: + result = await client.proofread.generate( + video_url="https://x/v.mp4", source_language="en", poll_interval=0 + ) + assert result.subtitles["en"] == "https://r2/en.srt" + + +@respx.mock +async def test_async_submit_sends_the_same_fields(): + route = respx.post("https://api.sonilo.com/v1/proofread").mock( + return_value=httpx.Response(202, json=ACK) + ) + async with AsyncSonilo(api_key="sk-test") as client: + await client.proofread.submit( + video_url="https://x/v.mp4", languages=["ja"], source_language="en" + ) + sent = unquote_plus(route.calls.last.request.content.decode()) + assert '["ja"]' in sent + assert "source_language=en" in sent diff --git a/uv.lock b/uv.lock index c936885..c9c254a 100644 --- a/uv.lock +++ b/uv.lock @@ -267,7 +267,7 @@ wheels = [ [[package]] name = "sonilo" -version = "0.19.0" +version = "0.20.0" source = { editable = "." } dependencies = [ { name = "httpx" },