Fix/affricates - #19
arunasrivastava wants to merge 7 commits into
Conversation
|
Drafted by dotty, an AI assistant, for human review. Reviewed head P2: Preserve IPA_g2p_experiments.ipynb:379 explicitly selects only Minimal fix: combined_ds = combine_datasets(
datasets, seed=RANDOM_SEED, columns=["ipa", "ipa_tokens", "audio", "g2p"]
)Checks and limits 26 focused assertions passed using actual PR function bodies with datasets 4.4.1 / transformers 4.57.1: split versus true affricates, normalization, empty/missing-token fallback, mixed concatenation/interleaving, vocabulary, label selection, and synthetic exporter fixtures. The notebook omission above was reproduced separately. Tests used an isolated harness with real tokenizer/dataset operations but synthetic corpus inputs and audio/G2P test doubles; Python 3.12.14 / NumPy 2.5.3 differed from the full repository environment. Full corpus export, training, and listening review remain unverified. Please commit regression tests covering both callers and these boundary cases. Existing validation/rollout caveats
The metrics and cache behavior are pre-existing limitations, not new regressions introduced by this PR. |
Fixes L2Arctic affricate label handling by preserving phoneme-token boundaries when available.
Examples for review:
data/ExamplesWithComments/l2arctic_affricates/*.wavEach wav has a matching
.txtshowing the corrected token transcription vs the old joined parse. To my very biased ears, the fixes sound correct but feel free to verify. specifically verify you agree the "and jargon" one should be ʒ "zsh" not dʒ "dzh"Updated:
ipa_tokensfrom TextGrid phone boundaries.ipa_tokensdirectly when present, sot + ʃ/d + ʒare not accidentally merged intotʃ/dʒ.ipa_tokens.After this PR:
ipa_tokensin a separate PR.