Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 9 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -509,9 +509,14 @@ The result is the work order — a time-aligned `segments` plan plus one
counterparts.

Pass exactly one of `video` / `video_url`, plus optional `prompt` (guidance
for the analysis, at most 2000 characters) and `variants_num` (1-5, default
1 — billed per brief). Source videos may be at most 360 seconds long, and
billing has a 10-second floor, so a very short clip still costs the same as a
for the analysis, at most 2000 characters), `variants_num` (1-5, default
1 — billed per brief) and `mode`. `mode` defaults to `both`: the music brief
in `segments`/`variations` plus a sound-design brief in `sfx_segments`
(shot-sized sections) and `sfx_prompt` (one string, authored once regardless
of `variants_num`). Pass `mode="music"` or `mode="sfx"` for just one brief;
`mode="music"` reproduces the previous result shape. The price is the same
for all three. Source videos may be at most 480 seconds long, and billing
has a 10-second floor, so a very short clip still costs the same as a
10-second one.

```python
Expand All @@ -525,6 +530,7 @@ with Sonilo() as client:
)
for segment in brief.segments:
print(f"{segment.start}-{segment.end}s [{segment.label}] {segment.prompt}")
print(brief.sfx_prompt) # the sound-design brief (mode="both")

# Feed a variation's prompt straight into a generation call.
track = client.video_to_music.generate_async(
Expand Down
4 changes: 2 additions & 2 deletions context7.json
Original file line number Diff line number Diff line change
Expand Up @@ -32,7 +32,7 @@
"client.audio_ducking ducks an EXISTING music bed under an EXISTING voice track; nothing is generated. One of voice/voice_url, one of music/music_url. Voice may be audio or video (video returns a .mp4); music must be audio. Result: output_url, no stems.",
"client.dubbing dubs one video into many languages in one async call; languages: en, zh_cn, ja, ko, pt, pt_br, es, es_419, de, fr, it, ru, th, ar, tr, vi, id, ta, ml, kn, gu, pa_in, sd_in, hi; default zh_cn,es,fr. Billed per language, no free trial.",
"A DubbingResult has no audio/video/output_url. Its results live in result.outputs, a map of language code to dubbed .mp4 URL: use result.save(lang, path) or result.save_all(dir). dubbing's video_url must be https.",
"client.video_analysis returns a creative BRIEF, not media: the method is analyze() (not generate()) and VideoAnalysisResult has no save(). Read result.segments (start/end/label/prompt) and result.variations[i].prompt.",
"Pass a video_analysis variation's prompt straight to video_to_music / video_to_sfx / video_to_sound as their prompt. It takes one of video/video_url plus optional prompt and variants_num (1-5, billed per brief); max 360s, 10s billing floor."
"client.video_analysis returns a creative BRIEF, not media: analyze() (not generate()), no save(). Music brief: result.segments (start/end/label/prompt) + result.variations[i].prompt; mode both (default) adds result.sfx_segments + result.sfx_prompt.",
"Pass a video_analysis variation's prompt straight to video_to_music / video_to_sfx / video_to_sound as their prompt. One of video/video_url plus optional prompt, variants_num (1-5, billed per brief), mode (both/music/sfx, same price); max 480s, 10s floor."
]
}
2 changes: 1 addition & 1 deletion pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ build-backend = "hatchling.build"

[project]
name = "sonilo"
version = "0.18.2"
version = "0.19.0"
description = "Official Python client for the Sonilo API"
readme = "README.md"
license = "MIT"
Expand Down
5 changes: 4 additions & 1 deletion sonilo-cli/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -259,7 +259,10 @@ command that produces no media file — nothing is generated:
time-aligned section plan) and `variations` (one ready-to-use generation prompt each). Pass
`--output brief.json` to write it to a file instead.
- `--variants` is 1-5 (default 1) and is **billed per brief**.
- Source videos may be at most 360 seconds long, and billing has a 10-second floor.
- `--mode` is `both` (default), `music` or `sfx`. `both` adds a sound-design brief to the JSON:
`sfx_segments` (shot-sized sections) and `sfx_prompt` (one string). `music` reproduces the
previous output shape. Same price for all three.
- Source videos may be at most 480 seconds long, and billing has a 10-second floor.
- Feed a variation's prompt straight into the next command:

sonilo video-analysis --video clip.mp4 --output brief.json
Expand Down
4 changes: 2 additions & 2 deletions sonilo-cli/pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -4,13 +4,13 @@ build-backend = "hatchling.build"

[project]
name = "sonilo-cli"
version = "0.17.2"
version = "0.18.0"
description = "Command-line interface for the Sonilo API: generate music and sound effects from text or video"
readme = "README.md"
license = "MIT"
requires-python = ">=3.9"
authors = [{ name = "Sonilo AI" }]
dependencies = ["sonilo>=0.18.0,<0.19"]
dependencies = ["sonilo>=0.19.0,<0.20"]
keywords = ["sonilo", "cli", "music", "sfx", "text-to-music", "video-to-music", "ai"]

[project.urls]
Expand Down
2 changes: 1 addition & 1 deletion sonilo-cli/src/sonilo_cli/__init__.py
Original file line number Diff line number Diff line change
@@ -1,3 +1,3 @@
__version__ = "0.17.2"
__version__ = "0.18.0"

__all__ = ["__version__"]
40 changes: 30 additions & 10 deletions sonilo-cli/src/sonilo_cli/__main__.py
Original file line number Diff line number Diff line change
Expand Up @@ -570,30 +570,44 @@ def _analysis_payload(result: Any) -> dict:
this output is meant to be piped into another tool (or read by an agent),
and it should look identical to what GET /v1/tasks returned. None-valued
accounting fields are dropped so a brief stays readable.

`mode` is echoed when the API sent one. The sound-design half of a
`both` brief — `sfx_segments` and `sfx_prompt` — is emitted only when
present, so a `music` or `sfx` brief keeps the shape it always had.
"""
payload: dict = {
"task_id": result.task_id,
"status": result.status,
"segments": [
{"start": s.start, "end": s.end, "label": s.label, "prompt": s.prompt}
for s in result.segments
],
"variations": [{"prompt": v.prompt} for v in result.variations],
}
payload: dict = {"task_id": result.task_id, "status": result.status}
mode = getattr(result, "mode", None)
if mode is not None:
payload["mode"] = mode
payload["segments"] = _segments_payload(result.segments)
payload["variations"] = [{"prompt": v.prompt} for v in result.variations]
sfx_segments = getattr(result, "sfx_segments", None) or []
if sfx_segments or mode == "both":
payload["sfx_segments"] = _segments_payload(sfx_segments)
sfx_prompt = getattr(result, "sfx_prompt", None)
if sfx_prompt is not None:
payload["sfx_prompt"] = sfx_prompt
for key in ("variants_num", "duration_seconds", "cost"):
value = getattr(result, key, None)
if value is not None:
payload[key] = value
return payload


def _segments_payload(segments: Any) -> list:
return [
{"start": s.start, "end": s.end, "label": s.label, "prompt": s.prompt}
for s in segments
]


def cmd_video_analysis(client: Sonilo, args: argparse.Namespace) -> None:
"""video-analysis is the one command that produces no media file. The
brief goes to stdout so it can be piped straight into the next tool;
--output is the opt-in for keeping a copy on disk."""
result = client.video_analysis.analyze(
video=args.video, video_url=args.video_url,
prompt=args.prompt, variants_num=args.variants,
prompt=args.prompt, variants_num=args.variants, mode=args.mode,
timeout=args.timeout,
)
if not result.variations:
Expand Down Expand Up @@ -1063,6 +1077,12 @@ def build_parser() -> argparse.ArgumentParser:
help="How many independent briefs to author for the same video (1-5). "
"Billed per brief. Default: 1",
)
p_va.add_argument(
"--mode", default=None, choices=["both", "music", "sfx"],
help="Which brief to return: both (default) returns a music brief and a "
"sound-effects brief in one call; music or sfx returns only that "
"one. Same price for all three.",
)
p_va.add_argument(
"--output", default=None,
help="Write the brief to this .json file instead of printing it to stdout.",
Expand Down
74 changes: 72 additions & 2 deletions sonilo-cli/tests/test_cli.py
Original file line number Diff line number Diff line change
Expand Up @@ -1609,12 +1609,22 @@ def test_empty_env_var_falls_through_to_the_credential(monkeypatch):
}


def _mock_analysis():
ANALYSIS_BOTH_BODY = {
**ANALYSIS_BODY,
"mode": "both",
"sfx_segments": [
{"start": 0, "end": 4, "label": "none", "prompt": "wind, distant traffic"},
],
"sfx_prompt": "naturalistic exterior ambience",
}


def _mock_analysis(body=ANALYSIS_BODY):
respx.post(f"{BASE}/v1/video-analysis").mock(
return_value=httpx.Response(202, json=ANALYSIS_ACK)
)
respx.get(f"{BASE}/v1/tasks/va1").mock(
return_value=httpx.Response(200, json=ANALYSIS_BODY)
return_value=httpx.Response(200, json=body)
)


Expand Down Expand Up @@ -1658,6 +1668,66 @@ def test_video_analysis_omits_unset_optionals():
body = unquote_plus(route.calls.last.request.content.decode())
assert "prompt" not in body
assert "variants_num" not in body
assert "mode" not in body


@respx.mock
def test_video_analysis_sends_mode():
_mock_analysis()
route = respx.routes[0]
run(["video-analysis", "--video-url", "https://x/v.mp4", "--mode", "sfx"])
body = unquote_plus(route.calls.last.request.content.decode())
assert "mode=sfx" in body


def test_video_analysis_rejects_an_unknown_mode(capsys):
"""argparse owns the choice list, so a typo fails before any request
(the CLI maps argparse errors to exit 1, like every other usage error)."""
with pytest.raises(SystemExit) as exc:
main(["--api-key", "sk-test", "video-analysis", "--video-url",
"https://x/v.mp4", "--mode", "ambience"])
assert exc.value.code == 1
assert "invalid choice" in capsys.readouterr().err


@respx.mock
def test_video_analysis_prints_the_sound_design_brief_in_both_mode(capsys):
"""A `both` brief carries mode, sfx_segments and sfx_prompt in the wire
shape, in a readable key order."""
_mock_analysis(ANALYSIS_BOTH_BODY)
run(["video-analysis", "--video-url", "https://x/v.mp4"])
out = json.loads(capsys.readouterr().out)
assert out["mode"] == "both"
assert out["sfx_segments"] == [
{"start": 0, "end": 4, "label": "none", "prompt": "wind, distant traffic"}
]
assert out["sfx_prompt"] == "naturalistic exterior ambience"
assert list(out)[:6] == [
"task_id", "status", "mode", "segments", "variations", "sfx_segments"
]
assert list(out)[6] == "sfx_prompt"


@respx.mock
def test_video_analysis_omits_sound_design_keys_when_absent(capsys):
"""A music/sfx brief (and a pre-`mode` body) keeps the old shape: no
mode, sfx_segments or sfx_prompt key at all, not empty placeholders."""
_mock_analysis()
run(["video-analysis", "--video-url", "https://x/v.mp4"])
out = json.loads(capsys.readouterr().out)
assert "mode" not in out
assert "sfx_segments" not in out
assert "sfx_prompt" not in out


@respx.mock
def test_video_analysis_music_mode_echo_only(capsys):
_mock_analysis({**ANALYSIS_BODY, "mode": "music"})
run(["video-analysis", "--video-url", "https://x/v.mp4", "--mode", "music"])
out = json.loads(capsys.readouterr().out)
assert out["mode"] == "music"
assert "sfx_segments" not in out
assert "sfx_prompt" not in out


@respx.mock
Expand Down
10 changes: 7 additions & 3 deletions src/sonilo/_requests.py
Original file line number Diff line number Diff line change
Expand Up @@ -304,12 +304,14 @@ def build_video_analysis_parts(
video_url: Optional[str],
prompt: Optional[str],
variants_num: Optional[int],
mode: Optional[str] = None,
) -> Tuple[Dict[str, str], Optional[Dict[str, tuple]], bool]:
"""Build the multipart parts for POST /v1/video-analysis.

Both optionals are omitted when unset so the server's own defaults apply
(no prompt, one variation). The 1-5 bound on variants_num and the 2000-char
bound on prompt are deliberately NOT checked here — the backend owns them,
Every optional is omitted when unset so the server's own defaults apply
(no prompt, one variation, mode "both"). The 1-5 bound on variants_num,
the 2000-char bound on prompt and the allowed values of mode ("both",
"music", "sfx") are deliberately NOT checked here — the backend owns them,
and a hardcoded copy would make this SDK reject values a later API widens.
"""
if (video is None) == (video_url is None):
Expand All @@ -323,6 +325,8 @@ def build_video_analysis_parts(
data["prompt"] = prompt
if variants_num is not None:
data["variants_num"] = str(variants_num)
if mode is not None:
data["mode"] = mode

# Now open files (only after data is fully assembled)
files: Optional[Dict[str, tuple]] = None
Expand Down
2 changes: 1 addition & 1 deletion src/sonilo/_version.py
Original file line number Diff line number Diff line change
@@ -1 +1 @@
__version__ = "0.18.2"
__version__ = "0.19.0"
31 changes: 23 additions & 8 deletions src/sonilo/resources/tasks.py
Original file line number Diff line number Diff line change
Expand Up @@ -303,6 +303,14 @@ def _analysis_segment_from(data: Any) -> Optional[AnalysisSegment]:
return None


def _analysis_segments_from(raw: Any) -> List[AnalysisSegment]:
"""Coerce a raw segment list (`segments` or `sfx_segments`), dropping
malformed entries; anything that is not a list reads as empty."""
if not isinstance(raw, list):
return []
return [s for s in map(_analysis_segment_from, raw) if s is not None]


def _analysis_variation_from(data: Any) -> Optional[AnalysisVariation]:
if not isinstance(data, dict):
return None
Expand All @@ -316,23 +324,27 @@ def parse_video_analysis_result(body: Dict[str, Any]) -> "VideoAnalysisResult":
"""Map a GET /v1/tasks/{id} body for a video-analysis task to
VideoAnalysisResult; unknown fields are ignored.

Both lists are coerced entry-by-entry and malformed entries are dropped,
for the same reason parse_dubbing_result coerces `outputs`: a
All three lists are coerced entry-by-entry and malformed entries are
dropped, for the same reason parse_dubbing_result coerces `outputs`: a
differently-shaped entry from a backend change should not surface as an
AttributeError deep inside the caller's loop, long after the parse.
`sfx_segments`/`sfx_prompt` only exist in mode "both"; a blank or
non-string `sfx_prompt` reads as None so callers can test it for truth.
"""
raw_segments = body.get("segments")
segments = (
[s for s in map(_analysis_segment_from, raw_segments) if s is not None]
if isinstance(raw_segments, list)
else []
)
segments = _analysis_segments_from(body.get("segments"))
sfx_segments = _analysis_segments_from(body.get("sfx_segments"))
raw_variations = body.get("variations")
variations = (
[v for v in map(_analysis_variation_from, raw_variations) if v is not None]
if isinstance(raw_variations, list)
else []
)
raw_sfx_prompt = body.get("sfx_prompt")
sfx_prompt = (
raw_sfx_prompt.strip()
if isinstance(raw_sfx_prompt, str) and raw_sfx_prompt.strip()
else None
)
try:
return VideoAnalysisResult(
task_id=body["task_id"],
Expand All @@ -345,6 +357,9 @@ def parse_video_analysis_result(body: Dict[str, Any]) -> "VideoAnalysisResult":
error=body.get("error"),
refunded=body.get("refunded"),
variants_num=body.get("variants_num"),
mode=body.get("mode"),
sfx_segments=sfx_segments,
sfx_prompt=sfx_prompt,
)
except KeyError as e:
raise SoniloError(f"Malformed task response: missing {e.args[0]!r}") from e
Expand Down
Loading
Loading