Skip to content
vohoPublic

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

logo

Podcare

Turn raw podcast mic recordings into a polished, broadcast-ready episode with one command.

podcare host.wav guest.wav -o episode.mp3

Input: one or more WAV/MP3/FLAC files (one per mic/recorder; anything ffmpeg reads). Output: one WAV/MP3/FLAC/M4A — aligned, gap-filled, declipped, de-thumped, de-hummed, de-plosived, either optionally restored from severe speech-codec damage (Sidon neural resynthesis) or conventionally denoised (DeepFilterNet3) and dereverbed, then tonally balanced, de-clicked, de-essed, resonance-tamed, clarity-enhanced, crosstalk-gated, breath-controlled, filler-words removed, pauses tightened, loudness-leveled, multiband-compressed, presence-excited, loudness-normalized and true-peak limited using default delivery targets of −16 LUFS / −1.5 dBTP. Output is 44.1 kHz (16-bit + dither for WAV/FLAC); processing runs internally at 48 kHz float32. Defaults favor quality over speed.


Install

Requires uv, ffmpeg on PATH, and Python ≥ 3.11. From a source checkout:

git clone https://github.com/voho/podcare.git
cd podcare
uv sync            # installs torch, DeepFilterNet, faster-whisper, nara-wpe, …
uv run podcare --help

The standard pipeline runs on CPU. The first run that uses filler removal downloads the Whisper model (large-v3 is ~3 GB; pick a smaller --whisper-model to skip that) plus the wav2vec2 forced-alignment model (~360 MB), both cached afterward. The first denoise likewise fetches the small (~9 MB) DeepFilterNet3 model from GitHub and caches it. Cached models can be reused without downloading them again.

The optional --restore-compressed stage lazily downloads only the selected CPU or CUDA pair from the official Sidon v0.1 model, about 1.0 GB total, and caches it too. Its upstream CPU performance is not benchmarked and is likely slow for long episodes; an NVIDIA CUDA GPU is strongly preferred. --restore-device auto chooses CUDA when available and otherwise falls back to CPU; the published bundles do not support Apple MPS. Install its supported runtime before first use:

uv sync --extra restore
# packaged install: pip install 'podcare[restore]'

Usage

# Two mics, full polish, MP3 out
uv run podcare host.wav guest.wav -o episode.mp3

# Gentler overall treatment
uv run podcare host.wav guest.wav -o episode.mp3 --strength 0.4

# Single track to 16-bit/44.1k WAV, force-aggressive filler removal
uv run podcare interview.mp3 -o clean.wav --filler-sensitivity 0.9

# Pin the language to avoid language auto-detection errors. Unsupported
# alignment languages leave the filler pass unedited; see the Filler section.
uv run podcare cz-show.flac -o out.mp3 --language cs

# Turn off the stages you don't want
uv run podcare raw.wav -o out.wav --no-dereverb --no-tighten

# Opt in to generative restoration for heavily compressed speech. CUDA is
# strongly preferred; reduce --restore-chunk if GPU memory is tight.
uv run podcare crushed.mp3 -o restored.wav --restore-compressed \
  --restore-device cuda --restore-chunk 25

# Clean audio for a video edit without timeline cuts. WAV plus the source rate
# preserves the frame count; lossy outputs may add codec delay or padding.
uv run podcare scene-audio.wav -o scene-clean.wav --nocut --out-sr 48000

# Practical CPU preview: skip multi-GB ASR and the expensive WPE pass
uv run podcare a.flac b.flac -o draft.mp3 --no-fillers --no-dereverb --strength 0.5

# Debug: write every stage's intermediate audio (and the final master) so you
# can A/B them
uv run podcare a.wav b.wav -o out.wav --keep-stems stems/

# Save the delivery QC measurement of the finished file as JSON
uv run podcare a.wav b.wav -o episode.mp3 --report episode-qc.json

Every run ends with the delivered file measured as a file — not as the arrays the pipeline worked on — so the loudness and true-peak targets are verified rather than assumed:

✓ episode.mp3  (62.4 min in → 58.9 min out)
  -19.3 LUFS file / -16.3 LUFS dual-mono (target -16.0), -2.9 dBTP (ceiling -1.5), LRA 1.7 LU

Two loudness numbers because the output is mono: a standard meter reads the file ~3 dB below what it sounds like on stereo playback, which is what --lufs targets. See delivery QC.

Each input file is treated as one speaker's mic. Give Podcare the separate recorder/mic tracks, not a pre-mixed file, so it can align them, gate crosstalk, balance levels, and detect each speaker's fillers on a clean isolated voice before summing. A single pre-mixed file works too — the multi-track-only stages are unnecessary, alignment skips a lone track, raw (--no-master) mixdown still enforces headroom, and filler detection runs on that track.

There is no information-theoretic way to “decompress” an MP3: discarded detail is gone. Sidon instead predicts clean speech features and resynthesizes a perceptually plausible voice. That can greatly reduce codec warble and restore bandwidth, but it is not the original waveform and can alter identity, prosody, or fine pronunciation. Keep the source and A/B the result before publishing.


MCP server

Podcare ships an optional Model Context Protocol server (src/podcare/mcp_server.py) — a thin layer that publishes every pipeline stage, plus the full pipeline, as MCP tools, so an assistant (Claude Desktop, an IDE MCP client, the Agent SDK, …) can clean audio for you.

Install

The server lives behind an optional extra so the core install stays lean:

uv sync --extra mcp        # adds the `mcp` SDK on top of the audio deps
uv run podcare-mcp         # serve over stdio

(With pip: pip install 'podcare[mcp]', then podcare-mcp.) ffmpeg must be on PATH, exactly as for the CLI.

Register it with a client

Point any MCP client at the podcare-mcp command over stdio. For example, in Claude Desktop's claude_desktop_config.json:

{
  "mcpServers": {
    "podcare": {
      "command": "uv",
      "args": ["run", "--directory", "/path/to/podcare", "podcare-mcp"]
    }
  }
}

Tools

There are two ways to use it: the one-shot process tool (the whole pipeline, equivalent to the CLI), or the per-stage tools to run a single stage in isolation and chain them yourself.

Tool What it does Key params (defaults)
process Full pipeline → one delivery file (returns the delivery QC measurement under delivery) strength=0.8, nocut=false, lufs=-16, out_sr=44100, bitrate=192k, whisper_model=large-v3, language, filler_sensitivity, restore_compressed=false, restore_device=auto, restore_chunk_s=25, intro_sound, outro_sound, disable=[…]
dropouts Packet-loss gap fill (two-sided LPC) strength=0.8
repair Declick + declip + rumble HPF declip=true, hpf_hz=80
dethump LF handling-noise gate (mic bumps, table knocks) strength=0.8
dehum Mains-hum harmonic notching strength=0.8
align Inter-track offset + polarity (≥ 2 tracks) window_s=300, min_confidence=12
plosives Plosive ("p-pop") ducking strength=0.8, max_hz=150
codec_restore Sidon generative restoration for severely compressed speech device=auto, chunk_s=25
denoise DeepFilterNet3 neural denoise strength=0.8, dry_db=-15
dereverb WPE dereverberation strength=0.8, chunk_s=15, wpe_delay=3
tonebalance LTAS → broadcast-voice EQ strength=0.8
declick Mouth-click / de-crackle strength=0.8
deess Sibilance control strength=0.8, lo_hz=4500, hi_hz=9500
resonance Dynamic resonance / harshness taming strength=0.8
clarity Consonant-clarity / transient presence enhancer strength=0.8
gate Crosstalk gate + level match strength=0.8, level_target_dbfs=-20
breath Breath ducking strength=0.8
fillers Filler-word removal (ASR + align) strength=0.8, sensitivity, whisper_model, language, pad_s
mixdown Sum cleaned tracks → mono program —
tighten Pause / dead-air tightening strength=0.8, max_pause_s, target_pause_s, lead_trail_s=0.5
leveler Slow segment-loudness ride strength=0.8
master MB-comp + exciter + loudnorm + TP-limit + bookends + encode strength=0.8, compress=true, exciter=true, exciter_amount, exciter_drive=4, lufs=-16, true_peak_db=-1.5, out_sr=44100, bitrate=192k, intro_sound, outro_sound

Every param defaults to the same value as the CLI (sourced from src/podcare/config.py); strength is the one universal intensity knob (see below).

Inputs/outputs. Each stage tool takes one or more input files and an output_dir, applies just that stage, and writes the result as 48 kHz float WAV(s) (loss-free, so stages chain cleanly) — one output per input, named after the source. master and process instead take an output_path and produce a real delivery file, choosing the container from its extension (.wav/.mp3/.flac/.m4a/.aac) and resampling once at the end. The process tool's disable list turns stages off by name (e.g. ["dereverb", "tighten"]), mirroring the CLI's --no-* flags. Codec restoration is opt-in instead: set restore_compressed=true on process, or call codec_restore directly. In the full pipeline, enabling it automatically bypasses DeepFilterNet denoise and WPE dereverb so the resynthesized voice is not destructively cleaned twice.


The --strength knob

Podcare has one universal intensity dial, --strength (0–1, default 0.8). Every stage maps it to its own notion of "how hard to work" — more noise removed, deeper de-essing, tighter pauses, firmer compression, and so on. At --strength 0 strength-driven enhancement stages are skipped: the pipeline becomes just align → mixdown → loudness-normalize → true-peak limit → encode (so the output still targets your --lufs setting and remains peak-limited, but its tone and dynamics are otherwise untreated — useful as an A/B baseline). 1 is the most aggressive. The exact per-stage mapping is listed in each pipeline section below, and every value is derived from src/podcare/config.py (the single source of truth).

align ignores strength on purpose: it's a correctness operation (fix timing/polarity), not a matter of degree, so it still runs at --strength 0. repair scales only its detection thresholds (declick/declip are otherwise on/off and the high-pass cutoff is fixed), and like every enhancement stage it is skipped entirely at --strength 0, so the raw baseline carries no declicking/declipping. Loudness normalization and the true-peak limiter are absolute delivery settings and also run regardless of strength.

codec restoration is another deliberate exception: it is an opt-in generative resynthesis, not a wet/dry effect, so nonzero strength does not make it “more restored.” At strength 0 it is skipped with the other enhancement stages.

You can still override individual stages — --filler-sensitivity, --max-pause, --target-pause, --lufs, --language — and an explicit value always wins over what --strength would have chosen. Overrides cannot resurrect a stage at --strength 0, though: --filler-sensitivity 0.9 --strength 0 still performs no filler cuts, because the no-op contract outranks the override.

Overriding only one of --max-pause / --target-pause is fine — the strength-derived other side yields to it (a lone --max-pause 0.5 keeps a 0.25 s beat rather than the 0.56 s strength would have chosen). Only an explicitly contradictory pair (--target-pause ≥ --max-pause) is an error.


Command-line options

General

Flag Default Description
AUDIO… (positional) — One or more input files, one per mic/recorder. Any format ffmpeg can decode. All are resampled to 48 kHz mono float internally.
-o, --output PATH required Output file. The extension picks the container: .wav, .mp3, .flac, .m4a/.aac. Parent dirs are created automatically.
-v, --verbose off Debug-level logging (per-stage decisions, offsets, gains, frame counts). Downgrades --progress auto to plain; an explicit --progress value still wins.
--progress {auto,rich,plain,none} auto Progress display. auto: a live bar (overall stage progress + ETA, plus a sub-bar showing chunk/transcription progress inside the long stages) when stderr is a terminal, otherwise plain log lines. rich: force the live bar. plain: log lines only. none: warnings + the final ✓ summary only.
--nocut off Keep the original edit timeline. Skips alignment, filler-word cuts, and pause/silence tightening. Every length-preserving cleanup stage (denoise, EQ, de-ess, gate, leveler, master, …) still runs. For an exact output frame count, use WAV/FLAC and set --out-sr to the source rate; lossy codecs may add delay or padding.
--version — Print version and exit.

Tuning

Flag Default Description
--strength 0..1 0.8 Universal processing intensity (see above). Scales every strength-driven parameter in the pipeline.
--filler-sensitivity 0..1 follows --strength Override how aggressively non-lexical fillers ("um", "uh", "ehm", "hmm", …) are cut. 0 disables the stage. When unset, defaults to 0.7 × strength.
--whisper-model NAME large-v3 faster-whisper model used to transcribe before forced alignment. Smaller models (medium/small/base/tiny) trade accuracy for speed and download size.
--language CODE auto Force the spoken language for filler detection (e.g. en, cs, de). When unset, the language is auto-detected per track.
--max-pause SECONDS follows --strength Override: silences longer than this get shortened. When unset, lerp(4.0 → 1.0) over strength. Must exceed --target-pause.
--target-pause SECONDS follows --strength Override: the length an over-long pause is shortened to. When unset, lerp(1.2 → 0.4) over strength.
--lufs DB -16 Output integrated-loudness target (ITU-R BS.1770 / EBU R128), measured as mono intended for playback through two speakers (dual_mono). A standard meter pointed at the mono file reads ~3 dB lower — -16 here delivers the familiar -19 LUFS mono; delivery QC prints both numbers. Validated to -40 … -5.
--out-sr HZ 44100 Output sample rate (validated 8000 … 192000). Resampling happens exactly once, at the end.
--bitrate RATE 192k Bitrate for lossy outputs (MP3/AAC). Ignored for WAV/FLAC.
--keep-stems DIR off Write each stage's intermediate audio and the final master into DIR as numbered files.
--report JSON off Write the delivery QC measurement of the finished file (loudness, true peak, loudness range) to a JSON file. Rejected up front if it would overwrite an input, the output, or a bookend. The same figures print at the end of every run regardless.
--intro-sound AUDIO off Sound placed before the finished program (anything ffmpeg reads; downmixed to mono, loudness-matched to the output target, joined with a 100 ms equal-power crossfade). Ignored with --nocut.
--outro-sound AUDIO off Sound placed after the finished program (same treatment as --intro-sound). Ignored with --nocut.

Compressed-speech restoration (opt-in)

Flag Default Description
--restore-compressed off Run Sidon v0.1 generative speech restoration after plosive control. Intended for severe codec artifacts and bandwidth loss; it is full resynthesis, not exact MP3 decompression. Automatically bypasses DeepFilterNet denoise and WPE dereverb.
--restore-device {auto,cpu,cuda} auto Inference device. auto prefers NVIDIA CUDA when available, otherwise CPU. CPU works but has no published upstream benchmark and is likely slow for long-form audio; the published bundles do not support Apple MPS.
--restore-chunk SECONDS 25 Maximum neural-restoration chunk length. The conservative default follows Sidon's official dataset-cleansing example; raise it only when memory allows. Must be finite and at least 0.1.

Advanced/library callers can set the corresponding Config fields: codec_restore, codec_restore_device, and codec_restore_chunk_s.

Stage toggles

Each switch disables one stage; everything else still runs.

Flag Disables
--no-dropouts Dropout / short-gap restoration
--no-declip Distortion repair (declick + declip + rumble high-pass)
--no-dethump LF handling-noise gate (mic bumps, table knocks, footsteps)
--no-dehum Mains-hum (50/60 Hz) harmonic removal
--no-align Inter-track time-offset and polarity correction
--no-denoise Noise reduction (DeepFilterNet3)
--no-dereverb WPE dereverberation
--no-tonebalance Tonal-balance / LTAS corrective EQ
--no-declick Mouth-click / de-crackle removal
--no-plosives Plosive ("p-pop") ducking
--no-deess De-essing (sibilance control)
--no-resonance Dynamic resonance / harshness suppression
--no-clarity Consonant-clarity / transient presence enhancer
--no-gate Crosstalk/room gate and per-track level matching
--no-breath Breath ducking
--no-fillers Filler-word removal
--no-tighten Pause tightening / dead-air trimming
--no-leveler Slow segment-loudness leveling
--no-exciter Harmonic presence exciter
--no-master Multiband compression + loudness normalization + limiting

The pipeline

Stages run in this fixed order. Track-level stages process each mic independently; session-level stages see all tracks at once. The ordering is deliberate: conservative restoration comes first; exceptional low-frequency plosive bursts are removed before the optional generative or classic learned/statistical enhancers; corrective tone and dynamics follow; and all per-mic cleanup — including filler detection — finishes before the tracks are summed. The only post-sum edits operate on the single mono program, so tracks cannot drift apart. When codec restoration is active, the classic denoise and dereverb stages are bypassed automatically.

          ┌──────────────────────────────── per track ────────────────────────────────┐
decode ─▶ dropouts ─ repair ─ dethump ─ dehum ─ align ─ plosives
 (load    (LPC gap  (declick (LF      (mains  (offset (p-pop
  48k)     fill)     declip   thumps)  hum     + pol.  ducking)
                     HPF)              notch)  fix)        │
                                                         ├─ codec restore (Sidon, opt-in) ─┐
                                                         │                                  │
                                                         └─ denoise (DFN3) ─ dereverb (WPE) ┤
                                                                                            ▼
 tonebalance ─▶ declick ─▶ deess ─▶ resonance
 (LTAS→voice)  (mouth     (sibilance (dynamic
  curve EQ)     clicks)    control)   notching)
                                        │
                                        ▼
      clarity ─▶ gate ─▶ breath ─▶ fillers ─┐
    (consonant  (level   (duck     (per-mic │
     presence)   match)   inhale)   ASR cuts)│
                                            ▼
     mixdown ─▶ tighten ─▶ leveler ─▶ master ──────────────▶ encode ──▶ QC
    (sum→mono) (shorten   (slow      (mb-comp + exciter +   (44.1k/16   (measure
                dead air)  loudness   loudnorm + TP-limiter)  + dither)   the file)
                          ride)

The Sidon branch replaces the DFN3/WPE branch when `--restore-compressed` is set.
QC measures the delivered file and reports/verifies it; it never alters the audio.

Every stage section below follows the same layout:

  • Fixes — the podcast defect it removes.
  • How it works — the algorithm, in plain terms.
  • Strength — how --strength drives it (or why it ignores strength).
  • Parameters — every knob and how it's controlled: Strength (scaled by --strength), a CLI flag, a Config field (advanced, edit src/podcare/config.py), or Hardcoded. All strength-derived numbers show the 0 → 1 range and the value at the default 0.8.

0. Decode

Fixes. Format/rate/channel-count differences between source files.

How it works. Every input is decoded through ffmpeg to 48 kHz mono float32. 48 kHz is DeepFilterNet3's native rate, so the entire chain runs at one rate and resamples just once at the very end. Stereo inputs are downmixed to mono. A file that decodes to zero samples is rejected immediately.

Strength. Not applicable — this is I/O.

Parameter Value Controlled by
Internal sample rate 48000 Hz Config.sr (do not change — DFN3 requires 48 k)
Channels mono Hardcoded

1. Dropout restoration — --no-dropouts (per track)

Fixes. Brief signal holes — 3–50 ms of missing audio on remote-guest tracks (VoIP, double-enders over flaky links) where packets were lost, audible as tiny stutters.

A hole in the stream is distinguished from ordinary articulation by depth: it collapses to near-digital silence, at or below the track's own room tone. The voiceless closure of a "p", "t" or "k" is also a short, speech-flanked drop of 25 dB or more, but it still carries room tone and sits above the noise floor. Without that test the stage rewrote every stop closure with fabricated LPC audio — measured, 60 "dropouts" and 591 ms of synthesis in a clean 30 s take.

How it works. Each hole is refilled by linear prediction: the speech is extrapolated forward from the audio before the gap and backward from the audio after it (using Levinson-Durbin LPC on an 80 ms context window), and the two estimates are crossfade-blended across the gap with complementary raised-cosine weights. The weights sum to one, avoiding the correlated-signal gain bump of an equal-power blend, while each predictor anchors the real waveform at its seam.

Strict caps keep it honest. Only gaps up to a strength-scaled maximum (0 ms at strength 0 → 50 ms at 1) are filled; only when the surrounding context is speech-active (a quiet moment inside a real pause is not a dropout); and no more than ~1.2 s per minute — beyond that the track is corrupt, not packet-lossy, and filling more would fabricate content.

Runs first so every later stage (including declick and denoise) sees gap-free audio.

Strength. Scales the longest fillable gap: 0 → 50 ms (40 ms @ 0.8). At strength 0 the maximum gap is 0 ms — a true no-op (and the stage is skipped).

Parameter Value Controlled by
Stage enabled on CLI --no-dropouts
Max fillable gap 0 → 50 ms (40 ms @ 0.8) Strength
Fill budget ~1.2 s per minute Hardcoded
Min gap 3 ms (1 RMS block) Hardcoded
Gap depth within 6 dB of the track's noise floor and below −60 dBFS Hardcoded
LPC context 80 ms (both sides) Hardcoded
LPC order 96 Hardcoded

2. Repair — --no-declip (per track)

Fixes. Clicks, glitches, hard clipping (recorded too hot), subsonic rumble.

How it works. Three ffmpeg filters in series: adeclick interpolates over impulsive clicks; adeclip reconstructs samples driven past full-scale; and a 2-pole high-pass removes rumble, desk thumps, and HVAC roar below the voice fundamental. Distortion repaired here can't be denoised away later, so it's first.

Both restoration filters run in method=save, which patches only the intervals the detector flags. This matters more than it sounds: in ffmpeg's default method=add mode they resynthesize the whole signal by overlap-adding every window through an autoregressive model, rewriting clean samples along with damaged ones. Measured through this chain, that cost ~28 dB of fidelity on clean material regardless of threshold (a pure 1 kHz sine came back with a −28.6 dB residual; white noise −7.6 dB) — enough that lightly clicky or lightly clipped speech came out further from the clean reference than it went in. In method=save the stage measures as a bit-exact no-op on clean tone, noise and speech, while repairing clicks ~19 dB better and mild clipping ~10 dB better than the old default.

Strength. Scales detection sensitivity (lower threshold = more sensitive). Both thresholds stay at or above 6 at every strength: below that, adeclick's detector starts firing on ordinary noise-like signal — which is exactly what speech fricatives and breath are.

Parameter Value Controlled by
Declick + declip enabled on CLI --no-declip
Reconstruction mode method=save Hardcoded
Declick threshold 12 → 6 --strength
Declip threshold 14 → 6 --strength
High-pass cutoff 80 Hz Config.hpf_hz
High-pass slope 2-pole (−12 dB/oct) Hardcoded

3. LF de-thump — --no-dethump (per track)

Fixes. Transient low-frequency thumps — mic bumps, table knocks, footsteps, chair creaks, wind gusts — that the fixed 80 Hz high-pass barely dents and that the master's low band would otherwise pump and loudnorm would lift into an audible "whump". One of the most common "amateur on a cheap desk" tells, and one nothing else in the chain targets (the plosive ducker needs the burst to dominate the spectrum, which a knock under speech does not; the gate only acts in pauses).

How it works. A split-band gate on a narrow ~50–160 Hz band (the sub-voice shoulder the high-pass leaves), extracted zero-phase in single precision exactly like the de-esser, so full = lf_band + rest holds and only the low band is touched. A thump is a transient spike in that band — well above its own ~300 ms median — that is not followed by sustained voicing: a spoken word's onset also spikes the low band (its fundamental appears) but is followed by sustained mid/high formant energy, whereas a knock decays into nothing. Confirming on what comes after the spike (rather than what coincides with it — a sharp knock is itself broadband) is what protects a low male voice fundamental, which lives in the same band, during speech. Flagged spans duck the low band only, with a fast attack and gentle release. Runs after repair (the high-pass is done) and before de-hum (the stationary hum comb is not yet notched).

That "not followed by voicing" test is measured two ways at once, and both must hold: the HF energy after the spike must fall relative to the spike and land near the track's own HF noise floor. The relative half alone is self-defeating — a knock is itself broadband, so half of its own HF peak is a bar that loud continuing speech clears easily. Measured on a clean 30 s take, the one-sided test fired on 83 blocks, 96 % of them inside loud speech.

Strict by design. Cut-only, low-band-only, and bounded — a clean track triggers nothing (no spike → no-op) and the duck depth is capped modest so even a misfire only dips 50–160 Hz briefly. The case it deliberately leaves alone is a knock fully buried under continuous loud speech: it is largely masked there, and removing it would risk the voice.

Strength. Tightens the detection margin (the spike threshold) and deepens the duck. At strength 0 the duck depth is 0 dB — a no-op (and the stage is skipped).

Parameter Value Controlled by
Stage enabled on CLI --no-dethump
Duck band 50–160 Hz Hardcoded
Detection margin (spike over local median) 18 → 6 dB (8.4 dB @ 0.8) Strength
Duck depth 0 → 12 dB (9.6 dB @ 0.8) Strength
Local-median window ~300 ms Hardcoded
Sustained-voicing check 100–250 ms after the spike, relative and within ~18 dB of the HF noise floor Hardcoded
Attack / release ~5 ms / ~60 ms Hardcoded

4. De-hum — --no-dehum (per track)

Fixes. Steady 50/60 Hz mains hum and its harmonics (100/120, 150/180, …) and ground-loop/USB buzz — one of the most instantly-noticeable amateur tells, which the 80 Hz high-pass only dents and the broadband neural denoiser doesn't reliably kill as a tonal comb.

How it works. The mains fundamental is found from the long-term spectrum (Welch PSD): the sharpest prominent peak within ±3 Hz of 50 or 60 Hz is refined by parabolic interpolation and retained at its measured frequency. This matters on drifting grids and imperfect recorder clocks; snapping a Q≈30 notch to exactly 50/60 Hz can miss a 49/59 Hz hum. Narrow zero-phase notches (iirnotch via sosfiltfilt) are then placed at each harmonic that actually protrudes above the local floor — so a clean track gets zero notches and the voice between harmonics is untouched. Runs early so the stationary tones don't bias GCC-PHAT alignment, get smeared by WPE, or feed the denoiser's noise estimate.

Strength. Scales how many harmonics are removed and how readily hum is detected. Two no-op guards at strength 0: zero harmonics and an unreachable detection margin (plus the stage is skipped).

Parameter Value Controlled by
Stage enabled on CLI --no-dehum
Max harmonics 0 → 12 (10 @ 0.8) Strength
Detection margin over floor 20 → 6 dB (≈9 dB @ 0.8) Strength
Notch Q 30 (a few Hz wide) Hardcoded
Highest harmonic 1500 Hz Hardcoded

5. Align + polarity — --no-align (session-level, ≥ 2 tracks)

Fixes. Recorders started at different instants (offset tracks smear crosstalk into echo) and miswired/inverted mics (phase cancellation when summed).

How it works. Both are fixed against the first track as reference. A coarse GCC-PHAT cross-correlation (on an 8 kHz downsample of the opening window) estimates the offset, then a direct sample-domain waveform correlation confirms it (|r| ≥ 0.08) — so genuinely uncorrelated remote recordings are left untouched. A negative confirming correlation flips the track's polarity. Tracks are then zero-padded to equal length.

Strength. Not strength-scaled — a correctness fix.

Parameter Value Controlled by
Stage enabled on CLI --no-align
Offset search window first 300 s Config.align_window_s
Peak confidence gate z ≥ 12 Config.align_min_confidence
Waveform confirm threshold |r| ≥ 0.08 Hardcoded
Polarity-flip threshold r < −0.05 Hardcoded

6. Plosive ducking — --no-plosives (per track)

Fixes. "P"/"B" pops — the burst of low-frequency energy a plosive blasts into the mic.

How it works. Per-track STFT (in 60 s chunks so the spectrogram can never OOM a multi-hour render). Frames are flagged where energy below the plosive ceiling is both abnormally high (≫ the whole-track LF-energy median — computed once in a bounded-memory pre-pass, so an identical plosive near a chunk boundary is ducked the same on both sides) and dominates the frame's spectrum; just those LF bins are ducked back toward the typical level, feathered ±1 frame. It runs before DeepFilterNet and WPE so an exceptional LF burst cannot bias either enhancer's estimate.

Strength. Lowers the detection thresholds and deepens the duck.

Parameter Value Controlled by
Stage enabled on CLI --no-plosives
Plosive band ceiling 150 Hz Config.plosive_max_hz
Burst threshold (×median) 24 → 4 (8.0 @ 0.8) Strength
Spectral-dominance threshold 0.80 → 0.40 (0.48 @ 0.8) Strength
Duck target (×median) 8 → 3 (4.0 @ 0.8) Strength

7. Compressed-speech restoration — --restore-compressed (per track, opt-in)

Fixes. Severe speech-codec warble, metallic pre-echo, smeared transients, and missing bandwidth that ordinary EQ or denoising cannot reconstruct.

How it works. Sidon v0.1 is an MIT-licensed, multilingual speech-restoration model. A w2v-BERT 2.0 feature predictor estimates clean speech features from a 16 kHz representation of the degraded recording, then a neural decoder generates a new 48 kHz waveform. Its published training degradations explicitly include MP3 at 65–245 kbps, alongside noise, reverb, band limiting, clipping, and packet loss; see the paper.

This is generative/perceptual restoration, not true decompression. MP3 discarded information cannot be recovered exactly, so the model supplies plausible detail and may change voice identity, prosody, or fine pronunciation. The published MP3 training range also does not promise recovery below 65 kbps. Use it only on speech that needs it, keep the original, and audition the result.

The implementation safely peak-normalizes and high-passes model input, processes long tracks in bounded chunks with a 1 s overlap/crossfade, and pins the result exactly to the input timeline. A short final chunk gets the 1.5 s right-tail padding used by Sidon's hosted demo, then is trimmed; it is not inflated to a full 25 s inference job. The output is RMS-matched to the source and peak-capped at 0.98 before the later pipeline stages. The stage uses the official TorchScript feature-extractor/decoder pair from the model release. The release is pinned to revision b3b02d8bbd55fdbc410e6e46e76ef95ace4fbf52 for reproducibility. Only the chosen *_cpu.pt or *_cuda.pt pair is downloaded lazily (about 1.0 GB total) and then reused from the Hugging Face cache. auto prefers CUDA and otherwise uses CPU; upstream publishes no CPU benchmark, so CUDA is strongly preferred for long audio.

In the full pipeline this stage runs after plosive control and automatically bypasses DeepFilterNet denoise and WPE dereverb. Sidon already performs general restoration; running both classic passes afterward can damage or over-dry its resynthesized voice. The dedicated MCP codec_restore tool does only this stage, so callers chaining MCP tools should avoid adding those two passes themselves.

Strength. All-or-nothing: a dry/wet mix of the resynthesized and original waveforms would risk phase artifacts. The stage is explicitly enabled by --restore-compressed; --strength 0 still skips it under the global no-op contract.

Parameter Value Controlled by
Stage enabled off CLI --restore-compressed / Config.codec_restore
Model Sidon v0.1, multilingual, 48 kHz output, MIT Pinned official release
Device auto (prefer CUDA, else CPU) CLI --restore-device / Config.codec_restore_device
Maximum chunk length 25 s CLI --restore-chunk / Config.codec_restore_chunk_s
Chunk overlap / crossfade up to 1 s Hardcoded
Output level guard source RMS match, then 0.98 peak cap Hardcoded
Model download selected CPU or CUDA pair, ~1.0 GB Lazy, cached
DeepFilterNet / WPE bypassed while active Automatic

8. Denoise — --no-denoise (per track)

Fixes. Broadband noise — room tone, hiss, fans, distant traffic, breath noise.

How it works. DeepFilterNet3, a full-band 48 kHz neural speech-enhancement model that separates voice from noise far more cleanly than classical methods. The ~9 MB model is downloaded from GitHub on first use and cached, so only the very first denoise needs a network. Processed in 60 s chunks with a 1 s crossfade so memory stays bounded on hour-long files. A fixed share of the dry signal is mixed back ("ambience preservation"): it bounds the worst-case suppression near 15 dB, so marginal quiet words and room tone are softened rather than erased — full removal sounds unnaturally dead and can swallow soft speech.

Strength. Sets the attenuation ceiling — a continuous, finite dB value; 60 dB at the top is already effectively full suppression for speech.

Parameter Value Controlled by
Stage enabled on CLI --no-denoise
Model DeepFilterNet3 (48 kHz full-band) Hardcoded
Attenuation ceiling 0 → 60 dB (48 dB @ 0.8) Strength
Dry mix (suppression floor) −15 dB Config.denoise_dry_db
Chunk / crossfade 60 s / 1 s Hardcoded

Note: DeepFilterNet 0.5.x imports torchaudio.backend.common, which newer torchaudio removed. A tiny compatibility shim (src/podcare/_compat.py) handles it — no action needed.


9. Dereverb — --no-dereverb (per track, WPE)

Fixes. The room — the late-reverberation "tail" that makes voices sound distant or boxy.

How it works. WPE (Weighted Prediction Error) linear-prediction dereverb: it estimates a per-frequency filter that predicts the reverb tail from the recent past and subtracts it. Complementary to the neural denoiser. Run in 15 s chunks with crossfade — short chunks keep WPE's large per-chunk transients from contending for RAM right after the neural denoiser, and a static room needs no more context than this.

Strength. Lengthens the prediction filter and adds one refinement pass. The practical ceiling is 10 taps / 2 iterations: it captures most of WPE's spoken-word benefit while avoiding the steep solve-time and conditioning cost of longer filters. Silent chunks and numerically unstable results bypass safely.

Parameter Value Controlled by
Stage enabled on CLI --no-dereverb
WPE taps (filter length) 6 → 10 (9 @ 0.8) Strength
WPE iterations 1 → 2 (2 @ 0.8) Strength
WPE prediction delay 3 Config.wpe_delay
Chunk length 15 s Config.dereverb_chunk_s

10. Tonal balance — --no-tonebalance (per track)

Fixes. The chain has no spectral shaping otherwise, so a dull lavalier and a bright condenser stay timbrally mismatched after all the cleanup; gross tilt, proximity-effect low-mid boom, and dull/harsh tops go uncorrected.

How it works. Each track's long-term average spectrum (LTAS) is measured over speech-active, contiguous FFT frames only, normalized to the speech body, and compared to a fixed produced broadcast-voice target curve. Speech is a gate rather than a splice: non-adjacent fragments are never concatenated, which would create false broadband energy at every edit boundary. The deviation is realized as four broad minimal-phase RBJ biquads (low/high shelf + low-mid and presence bells) via sosfilt — length-preserving, zero added latency. Only the spectral shape is matched rather than applying an intentional broadband gain; the final master sets delivery loudness. Deliberately gentle: broad filters, boosts clamped tighter than cuts (a noisy band is never lifted), and only a fraction of the deviation applied (strength / 3), so it stays subtle even at full strength.

Analysis bands above the source's bandwidth limit are dropped from the averages. A track that came from an MP3 has nothing above ~15 kHz, so its 12.5 and 16 kHz points read 37 and 48 dB below the speech body — which the raw deviation reads as "far too dull here, boost it". Averaged into the air band those two dead points outvoted the three live ones and flipped the high shelf from a −0.9 dB cut to a +1.1 dB boost, lifting the hiss region on exactly the lossy inputs the tool advertises support for. Only the dead top of the spectrum is excluded — sub-80 Hz points sit low because repair's rumble high-pass put them there, not for want of bandwidth, so they keep their vote.

Strength. Scales the fraction of the measured deviation applied — strength / 3 (0 at strength 0).

Parameter Value Controlled by
Stage enabled on CLI --no-tonebalance
Correction fraction strength / 3 (0.27 @ 0.8) Strength
Boost clamp +3 dB Hardcoded
Cut clamp −6 dB Hardcoded
Dead-band exclusion grid points >35 dB below the speech body, above the highest live one Hardcoded
Filters low shelf (120 Hz), "air" high shelf (7 kHz), low-mid (300 Hz) + presence (3 kHz) bells Hardcoded
Target curve produced broadcast-voice LTAS Hardcoded

11. Mouth-click / de-crackle — --no-declick (per track)

Fixes. Wet mouth clicks, lip smacks, tongue clicks and saliva crackle that adeclick (vinyl/digital impulses) and the neural denoiser leave intact — and which become more audible as compression and loudnorm lift the quiet inter-word detail.

How it works. STFT transient detection (≈1.3 ms hop). A click is a short mid-band (1.5–6 kHz) energy spike with a high crest factor over a robust local median that is also a local peak, sits in a quiet neighbourhood (the median-filtered broadband level there is well below the speech reference), and is not followed by speech within ~60 ms. That last gate is what separates a click from a consonant: a "t"/"k"/"p" burst is also an isolated high-crest mid-band spike in a momentarily quiet gap, but it is the beginning of a word, so speech follows it — whereas a click decays back into the gap. Together these gates protect real consonants. Flagged frames have their 1.5–9 kHz bins (a click's real reach) ducked toward the baseline, feathered ±1 frame — the air/brilliance region above 9 kHz, which the master exciter restores, is left untouched, so de-clicking never hollows the highs. The detection "quiet neighbourhood" reference is the whole-track p90 speech level (not each 60 s chunk's own), so an identical click is judged the same on either side of a chunk boundary. Chunked like the other spectral stages.

Strength. Lowers the crest threshold (catch progressively subtler clicks). Huge at strength 0 (nothing triggers).

Parameter Value Controlled by
Stage enabled on CLI --no-declick
Crest threshold 40 → 8× (14.4× @ 0.8) Strength
Detection band 1500–6000 Hz Hardcoded
Duck band 1500–9000 Hz (air above untouched) Hardcoded
Quiet-neighbourhood gate < 10% of the whole-track speech reference (energy, ≈ −10 dB) Hardcoded
Nothing-follows gate next ~60 ms also below that level Hardcoded
Local-median window ~268 ms (201 frames) Hardcoded

12. De-ess — --no-deess (per track)

Fixes. Harsh "S"/"SH"/"T" sibilance.

How it works. A zero-phase split-band design (in single precision, so a long episode never spawns a multi-GB float64 copy) guarantees full = sibilant_band + rest exactly: whenever the band's short-time energy exceeds a fraction of the full-band energy, the band is attenuated and recombined. Fast attack / slow release keeps it transparent. The audibility gate is relative to the track's own speech level (≈30 dB below it, floored at −55 dBFS), so it engages correctly on a quiet mic.

Strength. Lowers the trigger ratio and raises the maximum reduction.

Parameter Value Controlled by
Stage enabled on CLI --no-deess
Sibilance band 4500–9500 Hz Config.deess_lo_hz / deess_hi_hz
Trigger ratio (band/full RMS) 0.90 → 0.30 (0.42 @ 0.8) Strength
Max reduction 0 → 14 dB (11.2 dB @ 0.8) Strength
Attack / release ~3 ms / ~30 ms Hardcoded

13. Resonance suppression — --no-resonance (per track)

Fixes. Resonant peaks that come and go with the voice — ringy room modes, nasal honk, 2–5 kHz harshness spikes — which static tonal-balance EQ cannot catch because they are dynamic. A major cause of earbud fatigue on dense podcast mixes.

How it works. A "Soothe-lite" dynamic resonance tamer. Per STFT frame, a median filter across frequency bins estimates the broad spectral envelope (~420 Hz window at 48 kHz); any bin that protrudes more than a margin above its own envelope is pulled back by the excess, capped at a maximum cut, with asymmetric one-pole attack/release smoothing so notches fade in fast and out gently — cut engages in ~5 ms, releases over ~80 ms. Cut-only, and only between 800 Hz and 9 kHz (leaving the voice fundamental and the air region untouched).

The excess must also be at least ~140 Hz wide (3 bins). A resonance is a broad hump; a single protruding bin is a harmonic of the speaker's own pitch. At 46.9 Hz bins the envelope window sits in the valleys between the harmonics of any voice above ~150 Hz, so without the width requirement every harmonic read as an excess and the stage quietly shaved the speech band — a long-term spectral shift it must not cause (measured −0.62 dB across 800–1500 Hz on a real take, now −0.02 dB).

Strength. Lowers the detection margin (more sensitive) and raises the maximum cut. At strength 0 the cut cap is 0 dB — a bitwise no-op (and the stage is skipped).

Parameter Value Controlled by
Stage enabled on CLI --no-resonance
Detection margin 18 → 6 dB (8.4 dB @ 0.8) Strength
Max cut 0 → 10 dB (8 dB @ 0.8) Strength
Active band 800–9000 Hz Hardcoded
Attack / release ~5 ms / ~80 ms Hardcoded
Spectral median window ~420 Hz (~9 bins @ 48k/1024) Hardcoded

14. Consonant clarity — --no-clarity (per track)

Fixes. Mushy, distant-sounding consonants. The rest of the chain is subtractive for intelligibility — tonal balance is gentle, de-ess cuts sibilance, denoise/dereverb soften consonant attacks, and the master exciter only adds >8 kHz "air" (not the 1.5–4 kHz band that carries word recognition) — so a clean, heavily denoised voice can end up clear but lifeless. This is the chain's one additive intelligibility stage.

How it works. Per STFT frame, the positive spectral flux (the rising edge of a consonant attack) is measured in a 1.5–4 kHz band — kept strictly below the de-ess band (4.5 kHz) so it can never re-excite the sibilance de-ess just tamed. That band gets a short, transient-gated boost only during onsets, smoothed with a fast attack and gentle release and gated to speech-level frames so the noise floor between words is never lifted.

"Onset" needs an explicit threshold, because running speech has a persistent positive-flux floor — every frame of a vowel shows some rising bins. The flux is therefore referenced to its own distribution over speech frames: the 90th percentile counts as "no onset" and the 99th as a full one, so the median frame gets no lift at all. Feeding the raw flux ratio into the gain instead made almost every speech frame lift a little — measured, a near-static +0.97 dB tilt across 1.5–4 kHz, precisely the static EQ this stage promises not to be (now +0.28 dB). Because only onsets are touched, steady-state vowels and silence pass through unchanged — the long-term spectrum (tonal balance) is preserved: this restores definition/bite, not a static EQ tilt. Runs after resonance (harshness already tamed, so no resonant peak is amplified) and before the gate.

Strength. Sets the onset-boost amount; 0 dB at strength 0 (identity, stage skipped). Deliberately gentle — over-boosting consonant onsets sounds "spitty".

Parameter Value Controlled by
Stage enabled on CLI --no-clarity
Consonant band 1500–4000 Hz (strictly below de-ess) Hardcoded
Onset boost 0 → 3 dB (2.4 dB @ 0.8) Strength
Onset measure positive spectral flux, speech-gated, referenced to its p90–p99 over speech frames Hardcoded
Attack / release ~3 ms / ~20 ms Hardcoded

15. Gate + level match — --no-gate (per track, before mixdown)

Fixes. Crosstalk/bleed and room tone between phrases; mismatched speaker levels.

How it works. A downward expander (gate): below an adaptive speech/noise threshold the track is pushed down (2:1, floored at the gate depth), suppressing the other speaker's bleed. Standard gate timing: it opens fast (~10 ms) so a word arriving right after a ducked pause keeps its onset — a slow-opening gate audibly "swallows" quiet first words — and closes slow (~160 ms hold) so word tails ring out and the gate never chatters. Level matching: each track's speech-active RMS is normalized toward a target so a quiet guest and a loud host arrive balanced. Only blocks within 15 dB of the track's own loud speech count toward that measurement — on a two-mic recording the other person's bleed also clears the gate threshold, and averaging it in dragged the estimate down and over-boosted the track by an amount that depended on how leaky the room was. The correction is clamped to ±12 dB, so a track containing nothing but crosstalk cannot be promoted to full program level.

If the adaptive threshold lands above the entire level envelope — possible on a take with very little dynamic range, e.g. a very noisy source with --no-denoise — there is no pause to gate and the track is passed through untouched rather than expanded down as a whole.

Strength. Sets how deep the gate cuts.

Parameter Value Controlled by
Stage enabled on CLI --no-gate
Gate depth (max attenuation) 0 → 24 dB (19.2 dB @ 0.8) Strength
Speech-level target −20 dBFS Config.level_target_dbfs
Level-match window blocks within 15 dB of the track's own p90 Hardcoded
Level-match clamp ±12 dB Hardcoded
Expansion ratio 2:1 Hardcoded
Open / close (hold) time ~10 ms / ~160 ms Hardcoded

16. Breath control — --no-breath (per track)

Fixes. Audible inhale breaths between phrases. The gate only acts below its threshold and breaths usually sit above it, so they survive; after compression and loudnorm push the quiet detail forward, they become a prominent earbud tell.

How it works. Breaths are detected by what they are, not level alone: short segments that are unvoiced (high zero-crossing rate, no low fundamental), audible (above the noise floor), below speech level, of breath-like duration, and confirmed by spectral shape.

The level test does not protect fricatives — a real "s", "f" or "sh" sits 10–20 dB below the vowel it follows, in exactly the same band as an inhale. Two rejections do that work: any voiced block within ±40 ms marks the run as part of a word rather than an isolated breath (regardless of how loud that neighbour is), and a segment with more energy above 4.5 kHz than in the mid band is sibilance, not an inhale. Without them, word-final fricatives of normally spoken words were ducked by 6–9 dB.

Flagged spans are ducked (never muted) by a capped amount, ducking in over ~40 ms and recovering fast so the ramp finishes inside the breath instead of fading in the word that follows.

Strength. Sets the duck depth — capped at 14 dB (a duck, never a mute).

Parameter Value Controlled by
Stage enabled on CLI --no-breath
Duck depth 0 → 14 dB (11.2 dB @ 0.8) Strength
Breath duration window 0.08–0.7 s Hardcoded
Voiced/unvoiced split zero-crossing rate Hardcoded
Level gate below speech, above 2× noise floor Hardcoded
Word guard any voiced block within ±40 ms Hardcoded
Sibilance rejection >4.5 kHz energy above mid-band energy Hardcoded
Attack / release ~40 ms / ~10 ms Hardcoded

17. Filler-word removal — --no-fillers, --filler-sensitivity, --whisper-model, --language (session-level, per-track detection, before mixdown)

Fixes. Non-lexical fillers — "um", "uh", "ehm", "er", "hmm", "mm", … — located by transcription, not by listening for a sound.

How it works. Runs before mixdown, per track, so the ASR and aligner always see one clean isolated voice rather than the summed mix. For each mic: faster-whisper transcribes it (biased verbatim with a filler-laden prompt), then WhisperX force-aligns the transcript with a wav2vec2 CTC model for phoneme-tight word boundaries. To keep every track frame-aligned, a candidate is cut only when every other track is silent during it; the surviving intervals are removed identically from all tracks. Models are loaded once and reused; any ASR/alignment failure (unsupported language, bad model, download error) degrades to a logged no-op rather than aborting the render.

Strength. Sets the alignment-score floor 0.9 − 0.6·sens, minimum duration 0.24 − 0.2·sens s, and (below 0.7) an isolation requirement. Effective sensitivity follows strength conservatively (0.7 × strength).

Parameter Value Controlled by
Stage enabled on CLI --no-fillers
Sensitivity 0.7 × strength (0.56 @ 0.8) Strength / CLI --filler-sensitivity
Cross-track safety cut only where all other tracks are silent Hardcoded
Transcription model large-v3 (faster-whisper) CLI --whisper-model
Language auto-detect per track CLI --language
Forced aligner wav2vec2 CTC (WhisperX) Hardcoded
Filler lexicon um/uh/ehm/er/hmm/mm/… (affirmative "mhm" excluded) Hardcoded

18. Mixdown (session-level)

Fixes. Many cleaned mics → one coherent program bus.

How it works. All cleaned, gated, level-matched, filler-trimmed tracks are summed to one mono floating-point program. On the normal mastered path the sum is deliberately not peak-normalized: one stray transient must not turn down the entire episode before the fixed-threshold compressor. Loudness normalization and the final limiter establish delivery level and headroom later. A raw --no-master delivery is instead scaled only when needed to remain below −1 dBFS before integer or lossy encoding. From here on there is exactly one timeline.

Strength. Not applicable.

Parameter Value Controlled by
Mastered bus unnormalized float32 sum Hardcoded
Raw --no-master safety headroom −1 dBFS Hardcoded

19. Pause tightening — --no-tighten, --max-pause, --target-pause (mono program)

Fixes. Dead air — pacing.

How it works. Block-RMS detection finds silent runs on the mixed program; a run longer than the max-pause is shortened to the target-pause with a crossfade, and lead/trail silence is trimmed. The threshold sits well below speech (and below the post-gate noise floor) so breaths, beats and quiet reactions survive.

Strength. Shortens both the trigger and the kept beat.

Parameter Value Controlled by
Stage enabled on CLI --no-tighten
Max pause (trigger) 4.0 → 1.0 s (1.60 s @ 0.8) Strength / CLI --max-pause
Target pause (kept) 1.2 → 0.4 s (0.56 s @ 0.8) Strength / CLI --target-pause
Lead/tail trim 0.5 s Config.lead_trail_s
Edit crossfade 30 ms Hardcoded

20. Segment loudness leveler — --no-leveler (mono program)

Fixes. Minutes-scale loudness drift that the rest of the chain ignores: a guest who fades over a segment, a host who leans back, the gap between an intro and a tired late take. Integrated loudnorm fixes only the whole-file average and the master compressor reacts far too fast.

How it works. A slow short-term loudness envelope is computed over ~3 s windows counting only speech blocks (so pauses are neither pulled down nor boosted), each window is pulled toward the program's median speech loudness with a tightly clamped gain, and the result is applied heavily smoothed at multi-second time constants — inaudible as processing but very audible in the result. Runs on the mono bus before the master compressor so it sees consistent macro-dynamics.

The "boosts belong to speech, not room tone" clamp is applied to the raw gain curve, before the smoothing, so the multi-second time constant governs the transition too. Clamping the already-smoothed curve wrote a hard per-block decision back over it, and the envelope slammed from the full boost to unity inside one 100 ms block at every pause edge (measured 5.22 dB, against 0.18 dB now) — under-gaining the first phrase after every short pause.

Strength. Sets the maximum ride range (±dB); 0 at strength 0.

Parameter Value Controlled by
Stage enabled on CLI --no-leveler
Ride range ±0 → ±8 dB (±6.4 dB @ 0.8) Strength
Short-term window 3 s Hardcoded
Target program median speech loudness Hardcoded
Gating speech-only (pauses never boosted) Hardcoded

21. Master — --no-master, --lufs, --bitrate

Fixes. Band-specific dynamics, inconsistent loudness, and inter-sample peaks — the finishing chain.

How it works. Four steps. Multiband compression splits the bus into three phase-coherent bands (acrossover, LR4 at 250 Hz / 4 kHz) and compresses each independently (acompressor per band) so a boomy low-mid or a sibilant peak no longer ducks the whole program — denser, more consistent loudness than a single broadband compressor (off at --strength 0, where every band is 1:1). The bands also differ in timing: the low band rides slow (25 ms attack / 350 ms release, soft knee) so it controls proximity boom without modulating the voice fundamental — a fast attack on a sub-250 Hz band is shorter than one cycle of a 60–120 Hz fundamental and reads as low-end pumping — while the high band stays fast to catch harsh transients and the mid sits between them. Then a harmonic exciter (--no-exciter) synthesizes a touch of new harmonics from the 4–8 kHz consonant/sibilance band (ffmpeg aexciter, harmonics land at 8–16 kHz) to restore the "air" heavy denoise/dereverb removes and to cut through tiny speakers — rather than boosting (possibly noisy) existing highs. Runs before the loudness measurement so the added energy is counted in the loudness math. Strength maps amount 0 → 1.0 at a smooth drive of 4 — deliberately gentle: a touch of air (~+3 dB in the presence band), not a sheen. Over-driving it is the easiest way to make a mix sound harsh and noisy, so the ceiling is low; Config.exciter_amount overrides it directly (a fine-tuning knob) and the body/punch of the mix lives in the low-mids and loudness below, not here. Then two-pass EBU R128 loudness normalization (loudnorm, always on): the first pass measures, the second applies a linear correction toward the target LUFS. Peak limiting, resampling, and lossy encoding can move the delivered measurement slightly. Mono is measured with dual_mono=true, meaning the −16 LUFS default describes the intended two-speaker playback of the mono program rather than under-reading it by 3 LU. Finally a lookahead delivery limiter (alimiter, level=false so it never fights the loudness target) catches sample overs left by loudnorm's single linear gain. It runs at 48 kHz, where it sees sample rather than reconstruction peaks, so its ceiling is set a small margin (Config.tp_limiter_margin_db, 1 dB) below the loudnorm true-peak target. That reserve covers inter-sample peaks that can surface after the 44.1 kHz resample or codec reconstruction, keeping the delivered file under −1.5 dBTP on a consumer DAC. A silent/near-silent program (below loudnorm's −70 LUFS gate) skips normalization rather than erroring. Lookahead latency is compensated, so enabling the limiter does not shift the program or truncate its tail.

Strength. Firms up the per-band compression (higher ratios, lower thresholds) and raises the exciter amount; loudness, true-peak target, and the limiter are absolute delivery settings, not strength-scaled (the limiter runs at every strength, even 0).

Parameter Value Controlled by
Stage enabled on CLI --no-master
Multiband compression 3-band LR4 @ 250 Hz / 4 kHz (off at strength 0) Config.compress
Mid-band ratio 1.0 → 3.5 (3.0 @ 0.8); low ×1.1, high ×0.8 Strength
Mid-band threshold 0.30 → 0.10 amplitude (0.14 @ 0.8) Strength
Per-band timing low 25 ms/350 ms (knee 6), mid 10 ms/200 ms, high 5 ms/120 ms Hardcoded
Exciter enabled on CLI --no-exciter
Exciter source band 4–8 kHz (freq=4000, ceil=16000) Hardcoded
Exciter amount 0 → 1.0 (0.8 @ 0.8) Strength / Config.exciter_amount
Exciter drive 4.0 (smoother harmonics) Config.exciter_drive
Integrated loudness target −16 LUFS CLI --lufs
True-peak ceiling −1.5 dBTP Config.true_peak_db
Normalization two-pass, linear, dual-mono measurement Hardcoded
Limiter ISP margin 1.0 dB below the ceiling (reserves inter-sample headroom) Config.tp_limiter_margin_db
Delivery limiter lookahead brickwall at ceiling − margin (alimiter, no makeup, latency compensated) Hardcoded

22. Encode + resample

Fixes. Delivery format and the single, clean rate conversion.

How it works. A single resample to the output rate with soxr (VHQ) closes the chain. WAV and FLAC are written 16-bit with triangular-HP dither; MP3/AAC are encoded from float. The container is chosen from the file extension. Encoding happens through a same-directory temporary file that atomically replaces the destination only after ffmpeg succeeds, so an interrupted or failed render cannot destroy the previous delivery. The final log line confirms the delivered duration, rate, bit depth, loudness target, and size.

Strength. Not applicable — delivery.

Parameter Value Controlled by
Output sample rate 44100 Hz CLI --out-sr
WAV/FLAC bit depth 16-bit + triangular-HP dither Hardcoded
Resampler soxr VHQ Hardcoded
Lossy bitrate 192k CLI --bitrate

23. Delivery QC — --report

Fixes. The gap between what mastering aimed at and what the file actually is.

How it works. Everything upstream of the encoder works on arrays. The delivered file is not an array — it has been through a resample, a bit-depth reduction (or a lossy codec) and a container, and ffmpeg's loudnorm can quietly fall back to dynamic normalization when its linear gain would breach the peak ceiling. So the finished file is measured as a file: integrated loudness, loudness range and true peak, compared against the delivery targets. The figures print at the end of every run; --report out.json also writes them as JSON, and the MCP process tool returns them under delivery.

Anything outside tolerance (±1 LU on loudness, any true-peak over, LRA above 15 LU) is reported as a warning rather than silently accepted. A render that finishes is still a success — if the measurement pass itself cannot run, that is logged and the delivery stands.

Why two loudness numbers. Podcare delivers mono and masters with dual_mono=true, which is what makes --lufs -16 mean "sounds like −16 LUFS on a listener's stereo playback". A standard meter pointed at the mono file itself reports ~3 dB lower, because BS.1770 weights the one channel once while playback feeds it to both speakers — the same reason the podcast specs say "−16 LUFS stereo / −19 LUFS mono". Both numbers are correct and both are reported:

qc: episode.mp3 — on target · -19.3 LUFS file / -16.3 LUFS dual-mono (target -16.0),
    -2.9 dBTP (ceiling -1.5), LRA 1.7 LU

Reading the raw file number against --lufs and concluding the output is 3 dB quiet is the easiest mistake to make with this tool; printing both is the fix.

Strength. Not applicable — measurement.

Parameter Value Controlled by
Loudness tolerance ±1.0 LU Hardcoded
True-peak tolerance 0.1 dB Hardcoded
Wide-LRA warning > 15 LU Hardcoded
JSON report off CLI --report

Intro / outro

--intro-sound and --outro-sound attach a sting or theme around the finished program after all processing: each file (anything ffmpeg reads; downmixed to mono) is loudness-matched to the program — so a hot music sting can't blast ears relative to speech — then joined with a 100 ms equal-power crossfade (clamped to half the bookend's length so a short sting is never consumed by its own fade). A final true-peak limiter pass runs over the joins to handle any momentary overlap peaks from the crossfade. Not available with --nocut, whose purpose is to preserve the original edit timeline.

Two details make the match hold in the cases that matter most:

  • Short stings are measured by tiling. ffmpeg's loudnorm needs a few seconds of material before its gated integrated measurement is defined; below that it reports -inf. A 250 ms sting is exactly what people use as an intro, and measuring it raw reported "silent", skipping the level match and delivering it at full scale. The clip is repeated up to the measurement minimum first — tiling cannot change the loudness of the material — and the resulting linear gain is applied directly, because loudnorm also leaves clips shorter than ~0.5 s only partly scaled.
  • The target is the program, not an absolute number. On the mastered path the program has just been normalized to --lufs, so the two are the same thing. With --no-master nothing normalizes the program, and matching the bookend to an absolute --lufs produced an intro measured 24 dB louder than the show.

How it's built

  • Python 3.11 + uv. One linear pipeline over a Session (a list of Tracks
  • ffmpeg for decode/encode and the repair + master filters (incl. the multiband compressor and true-peak limiter); numpy/scipy for the hand-written DSP (dropout restoration, de-thump, de-hum, tonal-balance EQ, de-click, plosives, de-ess, resonance, clarity, gate, breath, leveler, tighten, align); DeepFilterNet/torch, optional Sidon v0.1 TorchScript fetched through huggingface-hub, nara-wpe, and faster-whisper + WhisperX for the ML/heavy stages; soxr for the single final resample.
  • Robustness by design. Heavy stages (codec restoration, denoise, dereverb, clarity, resonance, plosives, de-click) are chunked so working memory stays bounded on multi-hour episodes; the chunk accumulator and the block→sample gain interpolator work in float32 and in slices rather than materialising whole-episode float64 temporaries, which is what actually makes that bound hold. Codec restoration loads only one device-specific model pair and bypasses DeepFilterNet/WPE; on the classic path the neural denoiser's model is freed at the denoise→dereverb boundary so it never contends with WPE's large working set, and the Whisper/wav2vec2 models are freed once the filler pass ends rather than staying resident through mixdown and mastering. The optional ML filler pass degrades to a no-op (with a warning) rather than aborting a render; silent programs, bad CLI inputs, output/input collisions, and missing intro/outro files fail fast and cleanly; delivery writes are atomic; Ctrl+C exits with a clean "interrupted" line instead of a traceback.
  • Tests use synthetic fixtures with known ground truth (recover a known offset, reduce injected sibilance/hum/clicks, even out a loudness drift, hold the true-peak ceiling, final loudness within ±1.5 LU of target) plus the strength-mapping invariants, the cross-track filler-safety logic, and CLI validation.

Tests

uv run pytest -m "not slow"             # fast suite; skips real DFN/WPE
uv run pytest -m slow                   # real DeepFilterNet/WPE checks; may take minutes
uv run pytest                           # all tests available in this environment

uv sync --extra mcp
uv run pytest tests/test_mcp.py -m "not slow"  # optional MCP layer
uv run pytest tests/test_mcp.py                 # MCP layer including slow model checks

The first slow DeepFilterNet run may download its model. A real --restore-compressed run separately downloads and caches the selected ~1.0 GB Sidon pair; ordinary test and processing paths do not opt in to it. MCP tests skip cleanly unless the mcp extra is installed.

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages