Turn raw podcast mic recordings into a polished, broadcast-ready episode with one command.
podcare host.wav guest.wav -o episode.mp3Input: one or more WAV/MP3/FLAC files (one per mic/recorder; anything ffmpeg reads). Output: one WAV/MP3/FLAC/M4A — aligned, gap-filled, declipped, de-thumped, de-hummed, de-plosived, either optionally restored from severe speech-codec damage (Sidon neural resynthesis) or conventionally denoised (DeepFilterNet3) and dereverbed, then tonally balanced, de-clicked, de-essed, resonance-tamed, clarity-enhanced, crosstalk-gated, breath-controlled, filler-words removed, pauses tightened, loudness-leveled, multiband-compressed, presence-excited, loudness-normalized and true-peak limited using default delivery targets of −16 LUFS / −1.5 dBTP. Output is 44.1 kHz (16-bit + dither for WAV/FLAC); processing runs internally at 48 kHz float32. Defaults favor quality over speed.
Requires uv, ffmpeg on PATH, and Python ≥ 3.11. From a source checkout:
git clone https://github.com/voho/podcare.git
cd podcare
uv sync # installs torch, DeepFilterNet, faster-whisper, nara-wpe, …
uv run podcare --helpThe standard pipeline runs on CPU. The first run that uses filler removal downloads the
Whisper model (large-v3 is ~3 GB; pick a smaller --whisper-model to skip
that) plus the wav2vec2 forced-alignment model (~360 MB), both cached afterward.
The first denoise likewise fetches the small (~9 MB) DeepFilterNet3 model from
GitHub and caches it. Cached models can be reused without downloading them again.
The optional --restore-compressed stage lazily downloads only the selected
CPU or CUDA pair from the official
Sidon v0.1 model,
about 1.0 GB total, and caches it too. Its upstream CPU performance is not
benchmarked and is likely slow for long episodes; an NVIDIA CUDA GPU is strongly
preferred. --restore-device auto chooses CUDA when available and otherwise
falls back to CPU; the published bundles do not support Apple MPS.
Install its supported runtime before first use:
uv sync --extra restore
# packaged install: pip install 'podcare[restore]'# Two mics, full polish, MP3 out
uv run podcare host.wav guest.wav -o episode.mp3
# Gentler overall treatment
uv run podcare host.wav guest.wav -o episode.mp3 --strength 0.4
# Single track to 16-bit/44.1k WAV, force-aggressive filler removal
uv run podcare interview.mp3 -o clean.wav --filler-sensitivity 0.9
# Pin the language to avoid language auto-detection errors. Unsupported
# alignment languages leave the filler pass unedited; see the Filler section.
uv run podcare cz-show.flac -o out.mp3 --language cs
# Turn off the stages you don't want
uv run podcare raw.wav -o out.wav --no-dereverb --no-tighten
# Opt in to generative restoration for heavily compressed speech. CUDA is
# strongly preferred; reduce --restore-chunk if GPU memory is tight.
uv run podcare crushed.mp3 -o restored.wav --restore-compressed \
--restore-device cuda --restore-chunk 25
# Clean audio for a video edit without timeline cuts. WAV plus the source rate
# preserves the frame count; lossy outputs may add codec delay or padding.
uv run podcare scene-audio.wav -o scene-clean.wav --nocut --out-sr 48000
# Practical CPU preview: skip multi-GB ASR and the expensive WPE pass
uv run podcare a.flac b.flac -o draft.mp3 --no-fillers --no-dereverb --strength 0.5
# Debug: write every stage's intermediate audio (and the final master) so you
# can A/B them
uv run podcare a.wav b.wav -o out.wav --keep-stems stems/
# Save the delivery QC measurement of the finished file as JSON
uv run podcare a.wav b.wav -o episode.mp3 --report episode-qc.jsonEvery run ends with the delivered file measured as a file — not as the arrays the pipeline worked on — so the loudness and true-peak targets are verified rather than assumed:
✓ episode.mp3 (62.4 min in → 58.9 min out)
-19.3 LUFS file / -16.3 LUFS dual-mono (target -16.0), -2.9 dBTP (ceiling -1.5), LRA 1.7 LU
Two loudness numbers because the output is mono: a standard meter reads the file
~3 dB below what it sounds like on stereo playback, which is what --lufs
targets. See delivery QC.
Each input file is treated as one speaker's mic. Give Podcare the separate
recorder/mic tracks, not a pre-mixed file, so it can align them, gate crosstalk,
balance levels, and detect each speaker's fillers on a clean isolated voice
before summing. A single pre-mixed file works too — the multi-track-only stages
are unnecessary, alignment skips a lone track, raw (--no-master) mixdown still
enforces headroom, and filler detection runs on that track.
There is no information-theoretic way to “decompress” an MP3: discarded detail is gone. Sidon instead predicts clean speech features and resynthesizes a perceptually plausible voice. That can greatly reduce codec warble and restore bandwidth, but it is not the original waveform and can alter identity, prosody, or fine pronunciation. Keep the source and A/B the result before publishing.
Podcare ships an optional Model Context Protocol
server (src/podcare/mcp_server.py) — a thin layer that publishes every
pipeline stage, plus the full pipeline, as MCP tools, so an assistant (Claude
Desktop, an IDE MCP client, the Agent SDK, …) can clean audio for you.
The server lives behind an optional extra so the core install stays lean:
uv sync --extra mcp # adds the `mcp` SDK on top of the audio deps
uv run podcare-mcp # serve over stdio(With pip: pip install 'podcare[mcp]', then podcare-mcp.) ffmpeg must be on
PATH, exactly as for the CLI.
Point any MCP client at the podcare-mcp command over stdio. For example, in
Claude Desktop's claude_desktop_config.json:
{
"mcpServers": {
"podcare": {
"command": "uv",
"args": ["run", "--directory", "/path/to/podcare", "podcare-mcp"]
}
}
}There are two ways to use it: the one-shot process tool (the whole
pipeline, equivalent to the CLI), or the per-stage tools to run a single
stage in isolation and chain them yourself.
| Tool | What it does | Key params (defaults) |
|---|---|---|
process |
Full pipeline → one delivery file (returns the delivery QC measurement under delivery) |
strength=0.8, nocut=false, lufs=-16, out_sr=44100, bitrate=192k, whisper_model=large-v3, language, filler_sensitivity, restore_compressed=false, restore_device=auto, restore_chunk_s=25, intro_sound, outro_sound, disable=[…] |
dropouts |
Packet-loss gap fill (two-sided LPC) | strength=0.8 |
repair |
Declick + declip + rumble HPF | declip=true, hpf_hz=80 |
dethump |
LF handling-noise gate (mic bumps, table knocks) | strength=0.8 |
dehum |
Mains-hum harmonic notching | strength=0.8 |
align |
Inter-track offset + polarity (≥ 2 tracks) | window_s=300, min_confidence=12 |
plosives |
Plosive ("p-pop") ducking | strength=0.8, max_hz=150 |
codec_restore |
Sidon generative restoration for severely compressed speech | device=auto, chunk_s=25 |
denoise |
DeepFilterNet3 neural denoise | strength=0.8, dry_db=-15 |
dereverb |
WPE dereverberation | strength=0.8, chunk_s=15, wpe_delay=3 |
tonebalance |
LTAS → broadcast-voice EQ | strength=0.8 |
declick |
Mouth-click / de-crackle | strength=0.8 |
deess |
Sibilance control | strength=0.8, lo_hz=4500, hi_hz=9500 |
resonance |
Dynamic resonance / harshness taming | strength=0.8 |
clarity |
Consonant-clarity / transient presence enhancer | strength=0.8 |
gate |
Crosstalk gate + level match | strength=0.8, level_target_dbfs=-20 |
breath |
Breath ducking | strength=0.8 |
fillers |
Filler-word removal (ASR + align) | strength=0.8, sensitivity, whisper_model, language, pad_s |
mixdown |
Sum cleaned tracks → mono program | — |
tighten |
Pause / dead-air tightening | strength=0.8, max_pause_s, target_pause_s, lead_trail_s=0.5 |
leveler |
Slow segment-loudness ride | strength=0.8 |
master |
MB-comp + exciter + loudnorm + TP-limit + bookends + encode | strength=0.8, compress=true, exciter=true, exciter_amount, exciter_drive=4, lufs=-16, true_peak_db=-1.5, out_sr=44100, bitrate=192k, intro_sound, outro_sound |
Every param defaults to the same value as the CLI (sourced from
src/podcare/config.py); strength is the one universal
intensity knob (see below).
Inputs/outputs. Each stage tool takes one or more input files and an
output_dir, applies just that stage, and writes the result as 48 kHz float
WAV(s) (loss-free, so stages chain cleanly) — one output per input, named after
the source. master and process instead take an output_path and produce a
real delivery file, choosing the container from its extension
(.wav/.mp3/.flac/.m4a/.aac) and resampling once at the end. The process
tool's disable list turns stages off by name (e.g. ["dereverb", "tighten"]),
mirroring the CLI's --no-* flags. Codec restoration is opt-in instead:
set restore_compressed=true on process, or call codec_restore directly.
In the full pipeline, enabling it automatically bypasses DeepFilterNet denoise
and WPE dereverb so the resynthesized voice is not destructively cleaned twice.
Podcare has one universal intensity dial, --strength (0–1, default
0.8). Every stage maps it to its own notion of "how hard to work" — more
noise removed, deeper de-essing, tighter pauses, firmer compression, and so on.
At --strength 0 strength-driven enhancement stages are skipped: the pipeline
becomes just align → mixdown → loudness-normalize → true-peak limit → encode
(so the output still targets your --lufs setting and remains peak-limited, but
its tone and dynamics are otherwise untreated — useful as an A/B baseline).
1 is the most aggressive. The exact per-stage mapping is listed in each
pipeline section below, and every value is derived from
src/podcare/config.py (the single source of truth).
align ignores strength on purpose: it's a correctness operation (fix
timing/polarity), not a matter of degree, so it still runs at --strength 0.
repair scales only its detection thresholds (declick/declip are otherwise
on/off and the high-pass cutoff is fixed), and like every enhancement stage it is
skipped entirely at --strength 0, so the raw baseline carries no
declicking/declipping.
Loudness normalization and the true-peak limiter are absolute delivery settings
and also run regardless of strength.
codec restoration is another deliberate exception: it is an opt-in generative resynthesis, not a wet/dry effect, so nonzero strength does not make it “more restored.” At strength 0 it is skipped with the other enhancement stages.
You can still override individual stages — --filler-sensitivity, --max-pause,
--target-pause, --lufs, --language — and an explicit value always wins over
what --strength would have chosen. Overrides cannot resurrect a stage at
--strength 0, though: --filler-sensitivity 0.9 --strength 0 still performs no
filler cuts, because the no-op contract outranks the override.
Overriding only one of --max-pause / --target-pause is fine — the
strength-derived other side yields to it (a lone --max-pause 0.5 keeps a
0.25 s beat rather than the 0.56 s strength would have chosen). Only an
explicitly contradictory pair (--target-pause ≥ --max-pause) is an error.
| Flag | Default | Description |
|---|---|---|
AUDIO… (positional) |
— | One or more input files, one per mic/recorder. Any format ffmpeg can decode. All are resampled to 48 kHz mono float internally. |
-o, --output PATH |
required | Output file. The extension picks the container: .wav, .mp3, .flac, .m4a/.aac. Parent dirs are created automatically. |
-v, --verbose |
off | Debug-level logging (per-stage decisions, offsets, gains, frame counts). Downgrades --progress auto to plain; an explicit --progress value still wins. |
--progress {auto,rich,plain,none} |
auto |
Progress display. auto: a live bar (overall stage progress + ETA, plus a sub-bar showing chunk/transcription progress inside the long stages) when stderr is a terminal, otherwise plain log lines. rich: force the live bar. plain: log lines only. none: warnings + the final ✓ summary only. |
--nocut |
off | Keep the original edit timeline. Skips alignment, filler-word cuts, and pause/silence tightening. Every length-preserving cleanup stage (denoise, EQ, de-ess, gate, leveler, master, …) still runs. For an exact output frame count, use WAV/FLAC and set --out-sr to the source rate; lossy codecs may add delay or padding. |
--version |
— | Print version and exit. |
| Flag | Default | Description |
|---|---|---|
--strength 0..1 |
0.8 |
Universal processing intensity (see above). Scales every strength-driven parameter in the pipeline. |
--filler-sensitivity 0..1 |
follows --strength |
Override how aggressively non-lexical fillers ("um", "uh", "ehm", "hmm", …) are cut. 0 disables the stage. When unset, defaults to 0.7 × strength. |
--whisper-model NAME |
large-v3 |
faster-whisper model used to transcribe before forced alignment. Smaller models (medium/small/base/tiny) trade accuracy for speed and download size. |
--language CODE |
auto | Force the spoken language for filler detection (e.g. en, cs, de). When unset, the language is auto-detected per track. |
--max-pause SECONDS |
follows --strength |
Override: silences longer than this get shortened. When unset, lerp(4.0 → 1.0) over strength. Must exceed --target-pause. |
--target-pause SECONDS |
follows --strength |
Override: the length an over-long pause is shortened to. When unset, lerp(1.2 → 0.4) over strength. |
--lufs DB |
-16 |
Output integrated-loudness target (ITU-R BS.1770 / EBU R128), measured as mono intended for playback through two speakers (dual_mono). A standard meter pointed at the mono file reads ~3 dB lower — -16 here delivers the familiar -19 LUFS mono; delivery QC prints both numbers. Validated to -40 … -5. |
--out-sr HZ |
44100 |
Output sample rate (validated 8000 … 192000). Resampling happens exactly once, at the end. |
--bitrate RATE |
192k |
Bitrate for lossy outputs (MP3/AAC). Ignored for WAV/FLAC. |
--keep-stems DIR |
off | Write each stage's intermediate audio and the final master into DIR as numbered files. |
--report JSON |
off | Write the delivery QC measurement of the finished file (loudness, true peak, loudness range) to a JSON file. Rejected up front if it would overwrite an input, the output, or a bookend. The same figures print at the end of every run regardless. |
--intro-sound AUDIO |
off | Sound placed before the finished program (anything ffmpeg reads; downmixed to mono, loudness-matched to the output target, joined with a 100 ms equal-power crossfade). Ignored with --nocut. |
--outro-sound AUDIO |
off | Sound placed after the finished program (same treatment as --intro-sound). Ignored with --nocut. |
| Flag | Default | Description |
|---|---|---|
--restore-compressed |
off | Run Sidon v0.1 generative speech restoration after plosive control. Intended for severe codec artifacts and bandwidth loss; it is full resynthesis, not exact MP3 decompression. Automatically bypasses DeepFilterNet denoise and WPE dereverb. |
--restore-device {auto,cpu,cuda} |
auto |
Inference device. auto prefers NVIDIA CUDA when available, otherwise CPU. CPU works but has no published upstream benchmark and is likely slow for long-form audio; the published bundles do not support Apple MPS. |
--restore-chunk SECONDS |
25 |
Maximum neural-restoration chunk length. The conservative default follows Sidon's official dataset-cleansing example; raise it only when memory allows. Must be finite and at least 0.1. |
Advanced/library callers can set the corresponding Config fields:
codec_restore, codec_restore_device, and codec_restore_chunk_s.
Each switch disables one stage; everything else still runs.
| Flag | Disables |
|---|---|
--no-dropouts |
Dropout / short-gap restoration |
--no-declip |
Distortion repair (declick + declip + rumble high-pass) |
--no-dethump |
LF handling-noise gate (mic bumps, table knocks, footsteps) |
--no-dehum |
Mains-hum (50/60 Hz) harmonic removal |
--no-align |
Inter-track time-offset and polarity correction |
--no-denoise |
Noise reduction (DeepFilterNet3) |
--no-dereverb |
WPE dereverberation |
--no-tonebalance |
Tonal-balance / LTAS corrective EQ |
--no-declick |
Mouth-click / de-crackle removal |
--no-plosives |
Plosive ("p-pop") ducking |
--no-deess |
De-essing (sibilance control) |
--no-resonance |
Dynamic resonance / harshness suppression |
--no-clarity |
Consonant-clarity / transient presence enhancer |
--no-gate |
Crosstalk/room gate and per-track level matching |
--no-breath |
Breath ducking |
--no-fillers |
Filler-word removal |
--no-tighten |
Pause tightening / dead-air trimming |
--no-leveler |
Slow segment-loudness leveling |
--no-exciter |
Harmonic presence exciter |
--no-master |
Multiband compression + loudness normalization + limiting |
Stages run in this fixed order. Track-level stages process each mic independently; session-level stages see all tracks at once. The ordering is deliberate: conservative restoration comes first; exceptional low-frequency plosive bursts are removed before the optional generative or classic learned/statistical enhancers; corrective tone and dynamics follow; and all per-mic cleanup — including filler detection — finishes before the tracks are summed. The only post-sum edits operate on the single mono program, so tracks cannot drift apart. When codec restoration is active, the classic denoise and dereverb stages are bypassed automatically.
┌──────────────────────────────── per track ────────────────────────────────┐
decode ─▶ dropouts ─ repair ─ dethump ─ dehum ─ align ─ plosives
(load (LPC gap (declick (LF (mains (offset (p-pop
48k) fill) declip thumps) hum + pol. ducking)
HPF) notch) fix) │
├─ codec restore (Sidon, opt-in) ─┐
│ │
└─ denoise (DFN3) ─ dereverb (WPE) ┤
▼
tonebalance ─▶ declick ─▶ deess ─▶ resonance
(LTAS→voice) (mouth (sibilance (dynamic
curve EQ) clicks) control) notching)
│
▼
clarity ─▶ gate ─▶ breath ─▶ fillers ─┐
(consonant (level (duck (per-mic │
presence) match) inhale) ASR cuts)│
▼
mixdown ─▶ tighten ─▶ leveler ─▶ master ──────────────▶ encode ──▶ QC
(sum→mono) (shorten (slow (mb-comp + exciter + (44.1k/16 (measure
dead air) loudness loudnorm + TP-limiter) + dither) the file)
ride)
The Sidon branch replaces the DFN3/WPE branch when `--restore-compressed` is set.
QC measures the delivered file and reports/verifies it; it never alters the audio.
Every stage section below follows the same layout:
- Fixes — the podcast defect it removes.
- How it works — the algorithm, in plain terms.
- Strength — how
--strengthdrives it (or why it ignores strength). - Parameters — every knob and how it's controlled:
Strength(scaled by--strength), a CLI flag, aConfigfield (advanced, edit src/podcare/config.py), orHardcoded. All strength-derived numbers show the0 → 1range and the value at the default0.8.
Fixes. Format/rate/channel-count differences between source files.
How it works. Every input is decoded through ffmpeg to 48 kHz mono float32. 48 kHz is DeepFilterNet3's native rate, so the entire chain runs at one rate and resamples just once at the very end. Stereo inputs are downmixed to mono. A file that decodes to zero samples is rejected immediately.
Strength. Not applicable — this is I/O.
| Parameter | Value | Controlled by |
|---|---|---|
| Internal sample rate | 48000 Hz | Config.sr (do not change — DFN3 requires 48 k) |
| Channels | mono | Hardcoded |
Fixes. Brief signal holes — 3–50 ms of missing audio on remote-guest tracks (VoIP, double-enders over flaky links) where packets were lost, audible as tiny stutters.
A hole in the stream is distinguished from ordinary articulation by depth: it collapses to near-digital silence, at or below the track's own room tone. The voiceless closure of a "p", "t" or "k" is also a short, speech-flanked drop of 25 dB or more, but it still carries room tone and sits above the noise floor. Without that test the stage rewrote every stop closure with fabricated LPC audio — measured, 60 "dropouts" and 591 ms of synthesis in a clean 30 s take.
How it works. Each hole is refilled by linear prediction: the speech is extrapolated forward from the audio before the gap and backward from the audio after it (using Levinson-Durbin LPC on an 80 ms context window), and the two estimates are crossfade-blended across the gap with complementary raised-cosine weights. The weights sum to one, avoiding the correlated-signal gain bump of an equal-power blend, while each predictor anchors the real waveform at its seam.
Strict caps keep it honest. Only gaps up to a strength-scaled maximum (0 ms at strength 0 → 50 ms at 1) are filled; only when the surrounding context is speech-active (a quiet moment inside a real pause is not a dropout); and no more than ~1.2 s per minute — beyond that the track is corrupt, not packet-lossy, and filling more would fabricate content.
Runs first so every later stage (including declick and denoise) sees gap-free audio.
Strength. Scales the longest fillable gap: 0 → 50 ms (40 ms @ 0.8).
At strength 0 the maximum gap is 0 ms — a true no-op (and the stage is
skipped).
| Parameter | Value | Controlled by |
|---|---|---|
| Stage enabled | on | CLI --no-dropouts |
| Max fillable gap | 0 → 50 ms (40 ms @ 0.8) |
Strength |
| Fill budget | ~1.2 s per minute | Hardcoded |
| Min gap | 3 ms (1 RMS block) | Hardcoded |
| Gap depth | within 6 dB of the track's noise floor and below −60 dBFS | Hardcoded |
| LPC context | 80 ms (both sides) | Hardcoded |
| LPC order | 96 | Hardcoded |
Fixes. Clicks, glitches, hard clipping (recorded too hot), subsonic rumble.
How it works. Three ffmpeg filters in series: adeclick interpolates over
impulsive clicks; adeclip reconstructs samples driven past full-scale; and a
2-pole high-pass removes rumble, desk thumps, and HVAC roar below the voice
fundamental. Distortion repaired here can't be denoised away later, so it's first.
Both restoration filters run in method=save, which patches only the intervals
the detector flags. This matters more than it sounds: in ffmpeg's default
method=add mode they resynthesize the whole signal by overlap-adding every
window through an autoregressive model, rewriting clean samples along with damaged
ones. Measured through this chain, that cost ~28 dB of fidelity on clean material
regardless of threshold (a pure 1 kHz sine came back with a −28.6 dB residual;
white noise −7.6 dB) — enough that lightly clicky or lightly clipped speech came
out further from the clean reference than it went in. In method=save the stage
measures as a bit-exact no-op on clean tone, noise and speech, while repairing
clicks ~19 dB better and mild clipping ~10 dB better than the old default.
Strength. Scales detection sensitivity (lower threshold = more sensitive).
Both thresholds stay at or above 6 at every strength: below that, adeclick's
detector starts firing on ordinary noise-like signal — which is exactly what
speech fricatives and breath are.
| Parameter | Value | Controlled by |
|---|---|---|
| Declick + declip enabled | on | CLI --no-declip |
| Reconstruction mode | method=save |
Hardcoded |
| Declick threshold | 12 → 6 | --strength |
| Declip threshold | 14 → 6 | --strength |
| High-pass cutoff | 80 Hz | Config.hpf_hz |
| High-pass slope | 2-pole (−12 dB/oct) | Hardcoded |
Fixes. Transient low-frequency thumps — mic bumps, table knocks, footsteps, chair creaks, wind gusts — that the fixed 80 Hz high-pass barely dents and that the master's low band would otherwise pump and loudnorm would lift into an audible "whump". One of the most common "amateur on a cheap desk" tells, and one nothing else in the chain targets (the plosive ducker needs the burst to dominate the spectrum, which a knock under speech does not; the gate only acts in pauses).
How it works. A split-band gate on a narrow ~50–160 Hz band (the sub-voice
shoulder the high-pass leaves), extracted zero-phase in single precision exactly
like the de-esser, so full = lf_band + rest holds and only the low band is
touched. A thump is a transient spike in that band — well above its own ~300 ms
median — that is not followed by sustained voicing: a spoken word's onset also
spikes the low band (its fundamental appears) but is followed by sustained mid/high
formant energy, whereas a knock decays into nothing. Confirming on what comes
after the spike (rather than what coincides with it — a sharp knock is itself
broadband) is what protects a low male voice fundamental, which lives in the same
band, during speech. Flagged spans duck the low band only, with a fast attack and
gentle release. Runs after repair (the high-pass is done) and before de-hum (the
stationary hum comb is not yet notched).
That "not followed by voicing" test is measured two ways at once, and both must hold: the HF energy after the spike must fall relative to the spike and land near the track's own HF noise floor. The relative half alone is self-defeating — a knock is itself broadband, so half of its own HF peak is a bar that loud continuing speech clears easily. Measured on a clean 30 s take, the one-sided test fired on 83 blocks, 96 % of them inside loud speech.
Strict by design. Cut-only, low-band-only, and bounded — a clean track triggers nothing (no spike → no-op) and the duck depth is capped modest so even a misfire only dips 50–160 Hz briefly. The case it deliberately leaves alone is a knock fully buried under continuous loud speech: it is largely masked there, and removing it would risk the voice.
Strength. Tightens the detection margin (the spike threshold) and deepens the duck. At strength 0 the duck depth is 0 dB — a no-op (and the stage is skipped).
| Parameter | Value | Controlled by |
|---|---|---|
| Stage enabled | on | CLI --no-dethump |
| Duck band | 50–160 Hz | Hardcoded |
| Detection margin (spike over local median) | 18 → 6 dB (8.4 dB @ 0.8) |
Strength |
| Duck depth | 0 → 12 dB (9.6 dB @ 0.8) |
Strength |
| Local-median window | ~300 ms | Hardcoded |
| Sustained-voicing check | 100–250 ms after the spike, relative and within ~18 dB of the HF noise floor | Hardcoded |
| Attack / release | ~5 ms / ~60 ms | Hardcoded |
Fixes. Steady 50/60 Hz mains hum and its harmonics (100/120, 150/180, …) and ground-loop/USB buzz — one of the most instantly-noticeable amateur tells, which the 80 Hz high-pass only dents and the broadband neural denoiser doesn't reliably kill as a tonal comb.
How it works. The mains fundamental is found from the long-term spectrum
(Welch PSD): the sharpest prominent peak within ±3 Hz of 50 or 60 Hz is refined
by parabolic interpolation and retained at its measured frequency. This
matters on drifting grids and imperfect recorder clocks; snapping a Q≈30 notch
to exactly 50/60 Hz can miss a 49/59 Hz hum. Narrow zero-phase notches
(iirnotch via sosfiltfilt) are then placed at each harmonic that actually
protrudes above the local floor — so a clean track gets zero notches and the
voice between harmonics is untouched. Runs early so the stationary tones don't
bias GCC-PHAT alignment, get smeared by WPE, or feed the denoiser's noise estimate.
Strength. Scales how many harmonics are removed and how readily hum is detected. Two no-op guards at strength 0: zero harmonics and an unreachable detection margin (plus the stage is skipped).
| Parameter | Value | Controlled by |
|---|---|---|
| Stage enabled | on | CLI --no-dehum |
| Max harmonics | 0 → 12 (10 @ 0.8) |
Strength |
| Detection margin over floor | 20 → 6 dB (≈9 dB @ 0.8) |
Strength |
| Notch Q | 30 (a few Hz wide) | Hardcoded |
| Highest harmonic | 1500 Hz | Hardcoded |
Fixes. Recorders started at different instants (offset tracks smear crosstalk into echo) and miswired/inverted mics (phase cancellation when summed).
How it works. Both are fixed against the first track as reference. A coarse GCC-PHAT cross-correlation (on an 8 kHz downsample of the opening window) estimates the offset, then a direct sample-domain waveform correlation confirms it (|r| ≥ 0.08) — so genuinely uncorrelated remote recordings are left untouched. A negative confirming correlation flips the track's polarity. Tracks are then zero-padded to equal length.
Strength. Not strength-scaled — a correctness fix.
| Parameter | Value | Controlled by |
|---|---|---|
| Stage enabled | on | CLI --no-align |
| Offset search window | first 300 s | Config.align_window_s |
| Peak confidence gate | z ≥ 12 | Config.align_min_confidence |
| Waveform confirm threshold | |r| ≥ 0.08 | Hardcoded |
| Polarity-flip threshold | r < −0.05 | Hardcoded |
Fixes. "P"/"B" pops — the burst of low-frequency energy a plosive blasts into the mic.
How it works. Per-track STFT (in 60 s chunks so the spectrogram can never OOM a multi-hour render). Frames are flagged where energy below the plosive ceiling is both abnormally high (≫ the whole-track LF-energy median — computed once in a bounded-memory pre-pass, so an identical plosive near a chunk boundary is ducked the same on both sides) and dominates the frame's spectrum; just those LF bins are ducked back toward the typical level, feathered ±1 frame. It runs before DeepFilterNet and WPE so an exceptional LF burst cannot bias either enhancer's estimate.
Strength. Lowers the detection thresholds and deepens the duck.
| Parameter | Value | Controlled by |
|---|---|---|
| Stage enabled | on | CLI --no-plosives |
| Plosive band ceiling | 150 Hz | Config.plosive_max_hz |
| Burst threshold (×median) | 24 → 4 (8.0 @ 0.8) |
Strength |
| Spectral-dominance threshold | 0.80 → 0.40 (0.48 @ 0.8) |
Strength |
| Duck target (×median) | 8 → 3 (4.0 @ 0.8) |
Strength |
Fixes. Severe speech-codec warble, metallic pre-echo, smeared transients, and missing bandwidth that ordinary EQ or denoising cannot reconstruct.
How it works. Sidon v0.1 is an MIT-licensed, multilingual speech-restoration model. A w2v-BERT 2.0 feature predictor estimates clean speech features from a 16 kHz representation of the degraded recording, then a neural decoder generates a new 48 kHz waveform. Its published training degradations explicitly include MP3 at 65–245 kbps, alongside noise, reverb, band limiting, clipping, and packet loss; see the paper.
This is generative/perceptual restoration, not true decompression. MP3 discarded information cannot be recovered exactly, so the model supplies plausible detail and may change voice identity, prosody, or fine pronunciation. The published MP3 training range also does not promise recovery below 65 kbps. Use it only on speech that needs it, keep the original, and audition the result.
The implementation safely peak-normalizes and high-passes model input, processes
long tracks in bounded chunks with a 1 s overlap/crossfade, and pins the result
exactly to the input timeline. A short final chunk gets the 1.5 s right-tail
padding used by Sidon's
hosted demo,
then is trimmed; it is not inflated to a
full 25 s inference job. The output is RMS-matched to the source and peak-capped
at 0.98 before the later pipeline stages. The stage uses the
official TorchScript feature-extractor/decoder pair from the
model release.
The release is pinned to revision
b3b02d8bbd55fdbc410e6e46e76ef95ace4fbf52 for reproducibility. Only the
chosen *_cpu.pt or *_cuda.pt pair is downloaded lazily (about 1.0 GB total)
and then reused from the Hugging Face cache. auto prefers CUDA and otherwise
uses CPU; upstream publishes no CPU benchmark, so CUDA is strongly preferred
for long audio.
In the full pipeline this stage runs after plosive control and automatically
bypasses DeepFilterNet denoise and WPE dereverb. Sidon already performs
general restoration; running both classic passes afterward can damage or
over-dry its resynthesized voice. The dedicated MCP codec_restore tool does
only this stage, so callers chaining MCP tools should avoid adding those two
passes themselves.
Strength. All-or-nothing: a dry/wet mix of the resynthesized and original
waveforms would risk phase artifacts. The stage is explicitly enabled by
--restore-compressed; --strength 0 still skips it under the global no-op
contract.
| Parameter | Value | Controlled by |
|---|---|---|
| Stage enabled | off | CLI --restore-compressed / Config.codec_restore |
| Model | Sidon v0.1, multilingual, 48 kHz output, MIT | Pinned official release |
| Device | auto (prefer CUDA, else CPU) |
CLI --restore-device / Config.codec_restore_device |
| Maximum chunk length | 25 s | CLI --restore-chunk / Config.codec_restore_chunk_s |
| Chunk overlap / crossfade | up to 1 s | Hardcoded |
| Output level guard | source RMS match, then 0.98 peak cap | Hardcoded |
| Model download | selected CPU or CUDA pair, ~1.0 GB | Lazy, cached |
| DeepFilterNet / WPE | bypassed while active | Automatic |
Fixes. Broadband noise — room tone, hiss, fans, distant traffic, breath noise.
How it works. DeepFilterNet3, a full-band 48 kHz neural speech-enhancement model that separates voice from noise far more cleanly than classical methods. The ~9 MB model is downloaded from GitHub on first use and cached, so only the very first denoise needs a network. Processed in 60 s chunks with a 1 s crossfade so memory stays bounded on hour-long files. A fixed share of the dry signal is mixed back ("ambience preservation"): it bounds the worst-case suppression near 15 dB, so marginal quiet words and room tone are softened rather than erased — full removal sounds unnaturally dead and can swallow soft speech.
Strength. Sets the attenuation ceiling — a continuous, finite dB value; 60 dB at the top is already effectively full suppression for speech.
| Parameter | Value | Controlled by |
|---|---|---|
| Stage enabled | on | CLI --no-denoise |
| Model | DeepFilterNet3 (48 kHz full-band) | Hardcoded |
| Attenuation ceiling | 0 → 60 dB (48 dB @ 0.8) |
Strength |
| Dry mix (suppression floor) | −15 dB | Config.denoise_dry_db |
| Chunk / crossfade | 60 s / 1 s | Hardcoded |
Note: DeepFilterNet 0.5.x imports
torchaudio.backend.common, which newer torchaudio removed. A tiny compatibility shim (src/podcare/_compat.py) handles it — no action needed.
Fixes. The room — the late-reverberation "tail" that makes voices sound distant or boxy.
How it works. WPE (Weighted Prediction Error) linear-prediction dereverb: it estimates a per-frequency filter that predicts the reverb tail from the recent past and subtracts it. Complementary to the neural denoiser. Run in 15 s chunks with crossfade — short chunks keep WPE's large per-chunk transients from contending for RAM right after the neural denoiser, and a static room needs no more context than this.
Strength. Lengthens the prediction filter and adds one refinement pass. The practical ceiling is 10 taps / 2 iterations: it captures most of WPE's spoken-word benefit while avoiding the steep solve-time and conditioning cost of longer filters. Silent chunks and numerically unstable results bypass safely.
| Parameter | Value | Controlled by |
|---|---|---|
| Stage enabled | on | CLI --no-dereverb |
| WPE taps (filter length) | 6 → 10 (9 @ 0.8) |
Strength |
| WPE iterations | 1 → 2 (2 @ 0.8) |
Strength |
| WPE prediction delay | 3 | Config.wpe_delay |
| Chunk length | 15 s | Config.dereverb_chunk_s |
Fixes. The chain has no spectral shaping otherwise, so a dull lavalier and a bright condenser stay timbrally mismatched after all the cleanup; gross tilt, proximity-effect low-mid boom, and dull/harsh tops go uncorrected.
How it works. Each track's long-term average spectrum (LTAS) is measured over
speech-active, contiguous FFT frames only, normalized to the speech body,
and compared to a fixed produced broadcast-voice target curve. Speech is a
gate rather than a splice: non-adjacent fragments are never concatenated, which
would create false broadband energy at every edit boundary. The deviation is
realized as four broad minimal-phase RBJ biquads (low/high shelf + low-mid and
presence bells) via sosfilt — length-preserving, zero added latency. Only the
spectral shape is matched rather than applying an intentional broadband gain;
the final master sets delivery loudness. Deliberately gentle: broad filters,
boosts clamped tighter than cuts (a noisy band is never lifted), and only a
fraction of the deviation applied (strength / 3), so it stays subtle even at
full strength.
Analysis bands above the source's bandwidth limit are dropped from the
averages. A track that came from an MP3 has nothing above ~15 kHz, so its 12.5 and
16 kHz points read 37 and 48 dB below the speech body — which the raw deviation
reads as "far too dull here, boost it". Averaged into the air band those two
dead points outvoted the three live ones and flipped the high shelf from a
−0.9 dB cut to a +1.1 dB boost, lifting the hiss region on exactly the lossy
inputs the tool advertises support for. Only the dead top of the spectrum is
excluded — sub-80 Hz points sit low because repair's rumble high-pass put them
there, not for want of bandwidth, so they keep their vote.
Strength. Scales the fraction of the measured deviation applied — strength / 3
(0 at strength 0).
| Parameter | Value | Controlled by |
|---|---|---|
| Stage enabled | on | CLI --no-tonebalance |
| Correction fraction | strength / 3 (0.27 @ 0.8) |
Strength |
| Boost clamp | +3 dB | Hardcoded |
| Cut clamp | −6 dB | Hardcoded |
| Dead-band exclusion | grid points >35 dB below the speech body, above the highest live one | Hardcoded |
| Filters | low shelf (120 Hz), "air" high shelf (7 kHz), low-mid (300 Hz) + presence (3 kHz) bells | Hardcoded |
| Target curve | produced broadcast-voice LTAS | Hardcoded |
Fixes. Wet mouth clicks, lip smacks, tongue clicks and saliva crackle that
adeclick (vinyl/digital impulses) and the neural denoiser leave intact — and
which become more audible as compression and loudnorm lift the quiet inter-word
detail.
How it works. STFT transient detection (≈1.3 ms hop). A click is a short mid-band (1.5–6 kHz) energy spike with a high crest factor over a robust local median that is also a local peak, sits in a quiet neighbourhood (the median-filtered broadband level there is well below the speech reference), and is not followed by speech within ~60 ms. That last gate is what separates a click from a consonant: a "t"/"k"/"p" burst is also an isolated high-crest mid-band spike in a momentarily quiet gap, but it is the beginning of a word, so speech follows it — whereas a click decays back into the gap. Together these gates protect real consonants. Flagged frames have their 1.5–9 kHz bins (a click's real reach) ducked toward the baseline, feathered ±1 frame — the air/brilliance region above 9 kHz, which the master exciter restores, is left untouched, so de-clicking never hollows the highs. The detection "quiet neighbourhood" reference is the whole-track p90 speech level (not each 60 s chunk's own), so an identical click is judged the same on either side of a chunk boundary. Chunked like the other spectral stages.
Strength. Lowers the crest threshold (catch progressively subtler clicks). Huge at strength 0 (nothing triggers).
| Parameter | Value | Controlled by |
|---|---|---|
| Stage enabled | on | CLI --no-declick |
| Crest threshold | 40 → 8× (14.4× @ 0.8) |
Strength |
| Detection band | 1500–6000 Hz | Hardcoded |
| Duck band | 1500–9000 Hz (air above untouched) | Hardcoded |
| Quiet-neighbourhood gate | < 10% of the whole-track speech reference (energy, ≈ −10 dB) | Hardcoded |
| Nothing-follows gate | next ~60 ms also below that level | Hardcoded |
| Local-median window | ~268 ms (201 frames) | Hardcoded |
Fixes. Harsh "S"/"SH"/"T" sibilance.
How it works. A zero-phase split-band design (in single precision, so a
long episode never spawns a multi-GB float64 copy) guarantees
full = sibilant_band + rest exactly: whenever the band's short-time energy
exceeds a fraction of the full-band energy, the band is attenuated and recombined.
Fast attack / slow release keeps it transparent. The audibility gate is relative
to the track's own speech level (≈30 dB below it, floored at −55 dBFS), so it
engages correctly on a quiet mic.
Strength. Lowers the trigger ratio and raises the maximum reduction.
| Parameter | Value | Controlled by |
|---|---|---|
| Stage enabled | on | CLI --no-deess |
| Sibilance band | 4500–9500 Hz | Config.deess_lo_hz / deess_hi_hz |
| Trigger ratio (band/full RMS) | 0.90 → 0.30 (0.42 @ 0.8) |
Strength |
| Max reduction | 0 → 14 dB (11.2 dB @ 0.8) |
Strength |
| Attack / release | ~3 ms / ~30 ms | Hardcoded |
Fixes. Resonant peaks that come and go with the voice — ringy room modes, nasal honk, 2–5 kHz harshness spikes — which static tonal-balance EQ cannot catch because they are dynamic. A major cause of earbud fatigue on dense podcast mixes.
How it works. A "Soothe-lite" dynamic resonance tamer. Per STFT frame, a median filter across frequency bins estimates the broad spectral envelope (~420 Hz window at 48 kHz); any bin that protrudes more than a margin above its own envelope is pulled back by the excess, capped at a maximum cut, with asymmetric one-pole attack/release smoothing so notches fade in fast and out gently — cut engages in ~5 ms, releases over ~80 ms. Cut-only, and only between 800 Hz and 9 kHz (leaving the voice fundamental and the air region untouched).
The excess must also be at least ~140 Hz wide (3 bins). A resonance is a broad hump; a single protruding bin is a harmonic of the speaker's own pitch. At 46.9 Hz bins the envelope window sits in the valleys between the harmonics of any voice above ~150 Hz, so without the width requirement every harmonic read as an excess and the stage quietly shaved the speech band — a long-term spectral shift it must not cause (measured −0.62 dB across 800–1500 Hz on a real take, now −0.02 dB).
Strength. Lowers the detection margin (more sensitive) and raises the maximum cut. At strength 0 the cut cap is 0 dB — a bitwise no-op (and the stage is skipped).
| Parameter | Value | Controlled by |
|---|---|---|
| Stage enabled | on | CLI --no-resonance |
| Detection margin | 18 → 6 dB (8.4 dB @ 0.8) |
Strength |
| Max cut | 0 → 10 dB (8 dB @ 0.8) |
Strength |
| Active band | 800–9000 Hz | Hardcoded |
| Attack / release | ~5 ms / ~80 ms | Hardcoded |
| Spectral median window | ~420 Hz (~9 bins @ 48k/1024) | Hardcoded |
Fixes. Mushy, distant-sounding consonants. The rest of the chain is subtractive for intelligibility — tonal balance is gentle, de-ess cuts sibilance, denoise/dereverb soften consonant attacks, and the master exciter only adds >8 kHz "air" (not the 1.5–4 kHz band that carries word recognition) — so a clean, heavily denoised voice can end up clear but lifeless. This is the chain's one additive intelligibility stage.
How it works. Per STFT frame, the positive spectral flux (the rising edge of a consonant attack) is measured in a 1.5–4 kHz band — kept strictly below the de-ess band (4.5 kHz) so it can never re-excite the sibilance de-ess just tamed. That band gets a short, transient-gated boost only during onsets, smoothed with a fast attack and gentle release and gated to speech-level frames so the noise floor between words is never lifted.
"Onset" needs an explicit threshold, because running speech has a persistent positive-flux floor — every frame of a vowel shows some rising bins. The flux is therefore referenced to its own distribution over speech frames: the 90th percentile counts as "no onset" and the 99th as a full one, so the median frame gets no lift at all. Feeding the raw flux ratio into the gain instead made almost every speech frame lift a little — measured, a near-static +0.97 dB tilt across 1.5–4 kHz, precisely the static EQ this stage promises not to be (now +0.28 dB). Because only onsets are touched, steady-state vowels and silence pass through unchanged — the long-term spectrum (tonal balance) is preserved: this restores definition/bite, not a static EQ tilt. Runs after resonance (harshness already tamed, so no resonant peak is amplified) and before the gate.
Strength. Sets the onset-boost amount; 0 dB at strength 0 (identity, stage skipped). Deliberately gentle — over-boosting consonant onsets sounds "spitty".
| Parameter | Value | Controlled by |
|---|---|---|
| Stage enabled | on | CLI --no-clarity |
| Consonant band | 1500–4000 Hz (strictly below de-ess) | Hardcoded |
| Onset boost | 0 → 3 dB (2.4 dB @ 0.8) |
Strength |
| Onset measure | positive spectral flux, speech-gated, referenced to its p90–p99 over speech frames | Hardcoded |
| Attack / release | ~3 ms / ~20 ms | Hardcoded |
Fixes. Crosstalk/bleed and room tone between phrases; mismatched speaker levels.
How it works. A downward expander (gate): below an adaptive speech/noise threshold the track is pushed down (2:1, floored at the gate depth), suppressing the other speaker's bleed. Standard gate timing: it opens fast (~10 ms) so a word arriving right after a ducked pause keeps its onset — a slow-opening gate audibly "swallows" quiet first words — and closes slow (~160 ms hold) so word tails ring out and the gate never chatters. Level matching: each track's speech-active RMS is normalized toward a target so a quiet guest and a loud host arrive balanced. Only blocks within 15 dB of the track's own loud speech count toward that measurement — on a two-mic recording the other person's bleed also clears the gate threshold, and averaging it in dragged the estimate down and over-boosted the track by an amount that depended on how leaky the room was. The correction is clamped to ±12 dB, so a track containing nothing but crosstalk cannot be promoted to full program level.
If the adaptive threshold lands above the entire level envelope — possible on a
take with very little dynamic range, e.g. a very noisy source with
--no-denoise — there is no pause to gate and the track is passed through
untouched rather than expanded down as a whole.
Strength. Sets how deep the gate cuts.
| Parameter | Value | Controlled by |
|---|---|---|
| Stage enabled | on | CLI --no-gate |
| Gate depth (max attenuation) | 0 → 24 dB (19.2 dB @ 0.8) |
Strength |
| Speech-level target | −20 dBFS | Config.level_target_dbfs |
| Level-match window | blocks within 15 dB of the track's own p90 | Hardcoded |
| Level-match clamp | ±12 dB | Hardcoded |
| Expansion ratio | 2:1 | Hardcoded |
| Open / close (hold) time | ~10 ms / ~160 ms | Hardcoded |
Fixes. Audible inhale breaths between phrases. The gate only acts below its threshold and breaths usually sit above it, so they survive; after compression and loudnorm push the quiet detail forward, they become a prominent earbud tell.
How it works. Breaths are detected by what they are, not level alone: short segments that are unvoiced (high zero-crossing rate, no low fundamental), audible (above the noise floor), below speech level, of breath-like duration, and confirmed by spectral shape.
The level test does not protect fricatives — a real "s", "f" or "sh" sits 10–20 dB below the vowel it follows, in exactly the same band as an inhale. Two rejections do that work: any voiced block within ±40 ms marks the run as part of a word rather than an isolated breath (regardless of how loud that neighbour is), and a segment with more energy above 4.5 kHz than in the mid band is sibilance, not an inhale. Without them, word-final fricatives of normally spoken words were ducked by 6–9 dB.
Flagged spans are ducked (never muted) by a capped amount, ducking in over ~40 ms and recovering fast so the ramp finishes inside the breath instead of fading in the word that follows.
Strength. Sets the duck depth — capped at 14 dB (a duck, never a mute).
| Parameter | Value | Controlled by |
|---|---|---|
| Stage enabled | on | CLI --no-breath |
| Duck depth | 0 → 14 dB (11.2 dB @ 0.8) |
Strength |
| Breath duration window | 0.08–0.7 s | Hardcoded |
| Voiced/unvoiced split | zero-crossing rate | Hardcoded |
| Level gate | below speech, above 2× noise floor | Hardcoded |
| Word guard | any voiced block within ±40 ms | Hardcoded |
| Sibilance rejection | >4.5 kHz energy above mid-band energy | Hardcoded |
| Attack / release | ~40 ms / ~10 ms | Hardcoded |
17. Filler-word removal — --no-fillers, --filler-sensitivity, --whisper-model, --language (session-level, per-track detection, before mixdown)
Fixes. Non-lexical fillers — "um", "uh", "ehm", "er", "hmm", "mm", … — located by transcription, not by listening for a sound.
How it works. Runs before mixdown, per track, so the ASR and aligner always see one clean isolated voice rather than the summed mix. For each mic: faster-whisper transcribes it (biased verbatim with a filler-laden prompt), then WhisperX force-aligns the transcript with a wav2vec2 CTC model for phoneme-tight word boundaries. To keep every track frame-aligned, a candidate is cut only when every other track is silent during it; the surviving intervals are removed identically from all tracks. Models are loaded once and reused; any ASR/alignment failure (unsupported language, bad model, download error) degrades to a logged no-op rather than aborting the render.
Strength. Sets the alignment-score floor 0.9 − 0.6·sens, minimum duration
0.24 − 0.2·sens s, and (below 0.7) an isolation requirement. Effective
sensitivity follows strength conservatively (0.7 × strength).
| Parameter | Value | Controlled by |
|---|---|---|
| Stage enabled | on | CLI --no-fillers |
| Sensitivity | 0.7 × strength (0.56 @ 0.8) |
Strength / CLI --filler-sensitivity |
| Cross-track safety | cut only where all other tracks are silent | Hardcoded |
| Transcription model | large-v3 (faster-whisper) | CLI --whisper-model |
| Language | auto-detect per track | CLI --language |
| Forced aligner | wav2vec2 CTC (WhisperX) | Hardcoded |
| Filler lexicon | um/uh/ehm/er/hmm/mm/… (affirmative "mhm" excluded) | Hardcoded |
Fixes. Many cleaned mics → one coherent program bus.
How it works. All cleaned, gated, level-matched, filler-trimmed tracks are
summed to one mono floating-point program. On the normal mastered path the
sum is deliberately not peak-normalized: one stray transient must not turn down
the entire episode before the fixed-threshold compressor. Loudness normalization
and the final limiter establish delivery level and headroom later. A raw
--no-master delivery is instead scaled only when needed to remain below
−1 dBFS before integer or lossy encoding. From here on there is exactly one
timeline.
Strength. Not applicable.
| Parameter | Value | Controlled by |
|---|---|---|
| Mastered bus | unnormalized float32 sum | Hardcoded |
Raw --no-master safety headroom |
−1 dBFS | Hardcoded |
Fixes. Dead air — pacing.
How it works. Block-RMS detection finds silent runs on the mixed program; a run longer than the max-pause is shortened to the target-pause with a crossfade, and lead/trail silence is trimmed. The threshold sits well below speech (and below the post-gate noise floor) so breaths, beats and quiet reactions survive.
Strength. Shortens both the trigger and the kept beat.
| Parameter | Value | Controlled by |
|---|---|---|
| Stage enabled | on | CLI --no-tighten |
| Max pause (trigger) | 4.0 → 1.0 s (1.60 s @ 0.8) |
Strength / CLI --max-pause |
| Target pause (kept) | 1.2 → 0.4 s (0.56 s @ 0.8) |
Strength / CLI --target-pause |
| Lead/tail trim | 0.5 s | Config.lead_trail_s |
| Edit crossfade | 30 ms | Hardcoded |
Fixes. Minutes-scale loudness drift that the rest of the chain ignores: a guest who fades over a segment, a host who leans back, the gap between an intro and a tired late take. Integrated loudnorm fixes only the whole-file average and the master compressor reacts far too fast.
How it works. A slow short-term loudness envelope is computed over ~3 s windows counting only speech blocks (so pauses are neither pulled down nor boosted), each window is pulled toward the program's median speech loudness with a tightly clamped gain, and the result is applied heavily smoothed at multi-second time constants — inaudible as processing but very audible in the result. Runs on the mono bus before the master compressor so it sees consistent macro-dynamics.
The "boosts belong to speech, not room tone" clamp is applied to the raw gain curve, before the smoothing, so the multi-second time constant governs the transition too. Clamping the already-smoothed curve wrote a hard per-block decision back over it, and the envelope slammed from the full boost to unity inside one 100 ms block at every pause edge (measured 5.22 dB, against 0.18 dB now) — under-gaining the first phrase after every short pause.
Strength. Sets the maximum ride range (±dB); 0 at strength 0.
| Parameter | Value | Controlled by |
|---|---|---|
| Stage enabled | on | CLI --no-leveler |
| Ride range | ±0 → ±8 dB (±6.4 dB @ 0.8) |
Strength |
| Short-term window | 3 s | Hardcoded |
| Target | program median speech loudness | Hardcoded |
| Gating | speech-only (pauses never boosted) | Hardcoded |
Fixes. Band-specific dynamics, inconsistent loudness, and inter-sample peaks — the finishing chain.
How it works. Four steps. Multiband compression splits the bus into three
phase-coherent bands (acrossover, LR4 at 250 Hz / 4 kHz) and compresses each
independently (acompressor per band) so a boomy low-mid or a sibilant peak no
longer ducks the whole program — denser, more consistent loudness than a single
broadband compressor (off at --strength 0, where every band is 1:1). The bands
also differ in timing: the low band rides slow (25 ms attack / 350 ms release,
soft knee) so it controls proximity boom without modulating the voice fundamental —
a fast attack on a sub-250 Hz band is shorter than one cycle of a 60–120 Hz
fundamental and reads as low-end pumping — while the high band stays fast to catch
harsh transients and the mid sits between them. Then a
harmonic exciter (--no-exciter) synthesizes a touch of new harmonics from the
4–8 kHz consonant/sibilance band (ffmpeg aexciter, harmonics land at 8–16 kHz)
to restore the "air" heavy denoise/dereverb removes and to cut through tiny
speakers — rather than boosting (possibly noisy) existing highs. Runs before the
loudness measurement so the added energy is counted in the loudness math. Strength
maps amount 0 → 1.0 at a smooth drive of 4 — deliberately gentle: a touch of
air (~+3 dB in the presence band), not a sheen. Over-driving it is the easiest way
to make a mix sound harsh and noisy, so the ceiling is low; Config.exciter_amount
overrides it directly (a fine-tuning knob) and the body/punch of the mix lives in
the low-mids and loudness below, not here. Then two-pass EBU R128
loudness normalization (loudnorm, always on): the first pass measures, the
second applies a linear correction toward the target LUFS. Peak limiting,
resampling, and lossy encoding can move the delivered measurement slightly.
Mono is measured with dual_mono=true, meaning the −16 LUFS default describes
the intended two-speaker playback of the mono program rather than under-reading
it by 3 LU. Finally a lookahead delivery limiter (alimiter, level=false
so it never fights the loudness target) catches sample overs left by loudnorm's
single linear gain. It runs at 48 kHz, where it sees sample rather than
reconstruction peaks, so its ceiling is set a small margin
(Config.tp_limiter_margin_db, 1 dB) below the loudnorm true-peak target.
That reserve covers inter-sample peaks that can surface after the 44.1 kHz
resample or codec reconstruction, keeping the delivered file under −1.5 dBTP on
a consumer DAC. A
silent/near-silent program (below loudnorm's −70 LUFS gate) skips normalization
rather than erroring. Lookahead latency is compensated, so enabling the limiter
does not shift the program or truncate its tail.
Strength. Firms up the per-band compression (higher ratios, lower thresholds) and raises the exciter amount; loudness, true-peak target, and the limiter are absolute delivery settings, not strength-scaled (the limiter runs at every strength, even 0).
| Parameter | Value | Controlled by |
|---|---|---|
| Stage enabled | on | CLI --no-master |
| Multiband compression | 3-band LR4 @ 250 Hz / 4 kHz (off at strength 0) | Config.compress |
| Mid-band ratio | 1.0 → 3.5 (3.0 @ 0.8); low ×1.1, high ×0.8 |
Strength |
| Mid-band threshold | 0.30 → 0.10 amplitude (0.14 @ 0.8) |
Strength |
| Per-band timing | low 25 ms/350 ms (knee 6), mid 10 ms/200 ms, high 5 ms/120 ms | Hardcoded |
| Exciter enabled | on | CLI --no-exciter |
| Exciter source band | 4–8 kHz (freq=4000, ceil=16000) | Hardcoded |
| Exciter amount | 0 → 1.0 (0.8 @ 0.8) |
Strength / Config.exciter_amount |
| Exciter drive | 4.0 (smoother harmonics) | Config.exciter_drive |
| Integrated loudness target | −16 LUFS | CLI --lufs |
| True-peak ceiling | −1.5 dBTP | Config.true_peak_db |
| Normalization | two-pass, linear, dual-mono measurement | Hardcoded |
| Limiter ISP margin | 1.0 dB below the ceiling (reserves inter-sample headroom) | Config.tp_limiter_margin_db |
| Delivery limiter | lookahead brickwall at ceiling − margin (alimiter, no makeup, latency compensated) |
Hardcoded |
Fixes. Delivery format and the single, clean rate conversion.
How it works. A single resample to the output rate with soxr (VHQ) closes the chain. WAV and FLAC are written 16-bit with triangular-HP dither; MP3/AAC are encoded from float. The container is chosen from the file extension. Encoding happens through a same-directory temporary file that atomically replaces the destination only after ffmpeg succeeds, so an interrupted or failed render cannot destroy the previous delivery. The final log line confirms the delivered duration, rate, bit depth, loudness target, and size.
Strength. Not applicable — delivery.
| Parameter | Value | Controlled by |
|---|---|---|
| Output sample rate | 44100 Hz | CLI --out-sr |
| WAV/FLAC bit depth | 16-bit + triangular-HP dither | Hardcoded |
| Resampler | soxr VHQ | Hardcoded |
| Lossy bitrate | 192k | CLI --bitrate |
Fixes. The gap between what mastering aimed at and what the file actually is.
How it works. Everything upstream of the encoder works on arrays. The
delivered file is not an array — it has been through a resample, a bit-depth
reduction (or a lossy codec) and a container, and ffmpeg's loudnorm can quietly
fall back to dynamic normalization when its linear gain would breach the peak
ceiling. So the finished file is measured as a file: integrated loudness,
loudness range and true peak, compared against the delivery targets. The figures
print at the end of every run; --report out.json also writes them as JSON, and
the MCP process tool returns them under delivery.
Anything outside tolerance (±1 LU on loudness, any true-peak over, LRA above 15 LU) is reported as a warning rather than silently accepted. A render that finishes is still a success — if the measurement pass itself cannot run, that is logged and the delivery stands.
Why two loudness numbers. Podcare delivers mono and masters with
dual_mono=true, which is what makes --lufs -16 mean "sounds like −16 LUFS on
a listener's stereo playback". A standard meter pointed at the mono file itself
reports ~3 dB lower, because BS.1770 weights the one channel once while playback
feeds it to both speakers — the same reason the podcast specs say "−16 LUFS
stereo / −19 LUFS mono". Both numbers are correct and both are reported:
qc: episode.mp3 — on target · -19.3 LUFS file / -16.3 LUFS dual-mono (target -16.0),
-2.9 dBTP (ceiling -1.5), LRA 1.7 LU
Reading the raw file number against --lufs and concluding the output is 3 dB
quiet is the easiest mistake to make with this tool; printing both is the fix.
Strength. Not applicable — measurement.
| Parameter | Value | Controlled by |
|---|---|---|
| Loudness tolerance | ±1.0 LU | Hardcoded |
| True-peak tolerance | 0.1 dB | Hardcoded |
| Wide-LRA warning | > 15 LU | Hardcoded |
| JSON report | off | CLI --report |
--intro-sound and --outro-sound attach a sting or theme around the finished
program after all processing: each file (anything ffmpeg reads; downmixed to
mono) is loudness-matched to the program — so a hot music sting can't blast ears
relative to speech — then joined with a 100 ms equal-power crossfade (clamped to
half the bookend's length so a short sting is never consumed by its own fade). A
final true-peak limiter pass runs over the joins to handle any momentary overlap
peaks from the crossfade. Not available with --nocut, whose purpose is to
preserve the original edit timeline.
Two details make the match hold in the cases that matter most:
- Short stings are measured by tiling. ffmpeg's
loudnormneeds a few seconds of material before its gated integrated measurement is defined; below that it reports-inf. A 250 ms sting is exactly what people use as an intro, and measuring it raw reported "silent", skipping the level match and delivering it at full scale. The clip is repeated up to the measurement minimum first — tiling cannot change the loudness of the material — and the resulting linear gain is applied directly, becauseloudnormalso leaves clips shorter than ~0.5 s only partly scaled. - The target is the program, not an absolute number. On the mastered path the
program has just been normalized to
--lufs, so the two are the same thing. With--no-masternothing normalizes the program, and matching the bookend to an absolute--lufsproduced an intro measured 24 dB louder than the show.
- Python 3.11 + uv. One linear pipeline over a
Session(a list ofTracks- sample rate). Each stage is a small function
(Session|Track, Config) → Session|Track; the stage list lives in src/podcare/pipeline.py, every tunable (and the strength→stage mapping) in src/podcare/config.py.
- sample rate). Each stage is a small function
- ffmpeg for decode/encode and the repair + master filters (incl. the multiband compressor and true-peak limiter); numpy/scipy for the hand-written DSP (dropout restoration, de-thump, de-hum, tonal-balance EQ, de-click, plosives, de-ess, resonance, clarity, gate, breath, leveler, tighten, align); DeepFilterNet/torch, optional Sidon v0.1 TorchScript fetched through huggingface-hub, nara-wpe, and faster-whisper + WhisperX for the ML/heavy stages; soxr for the single final resample.
- Robustness by design. Heavy stages (codec restoration, denoise, dereverb, clarity, resonance, plosives, de-click) are chunked so working memory stays bounded on multi-hour episodes; the chunk accumulator and the block→sample gain interpolator work in float32 and in slices rather than materialising whole-episode float64 temporaries, which is what actually makes that bound hold. Codec restoration loads only one device-specific model pair and bypasses DeepFilterNet/WPE; on the classic path the neural denoiser's model is freed at the denoise→dereverb boundary so it never contends with WPE's large working set, and the Whisper/wav2vec2 models are freed once the filler pass ends rather than staying resident through mixdown and mastering. The optional ML filler pass degrades to a no-op (with a warning) rather than aborting a render; silent programs, bad CLI inputs, output/input collisions, and missing intro/outro files fail fast and cleanly; delivery writes are atomic; Ctrl+C exits with a clean "interrupted" line instead of a traceback.
- Tests use synthetic fixtures with known ground truth (recover a known offset, reduce injected sibilance/hum/clicks, even out a loudness drift, hold the true-peak ceiling, final loudness within ±1.5 LU of target) plus the strength-mapping invariants, the cross-track filler-safety logic, and CLI validation.
uv run pytest -m "not slow" # fast suite; skips real DFN/WPE
uv run pytest -m slow # real DeepFilterNet/WPE checks; may take minutes
uv run pytest # all tests available in this environment
uv sync --extra mcp
uv run pytest tests/test_mcp.py -m "not slow" # optional MCP layer
uv run pytest tests/test_mcp.py # MCP layer including slow model checksThe first slow DeepFilterNet run may download its model. A real
--restore-compressed run separately downloads and caches the selected ~1.0 GB
Sidon pair; ordinary test and processing paths do not opt in to it. MCP tests
skip cleanly unless the mcp extra is installed.
