Skip to content

Fix/affricates - #19

Open
arunasrivastava wants to merge 7 commits into
mainfrom
fix/affricates
Open

arunasrivastava wants to merge 7 commits into
mainfrom
fix/affricates

Conversation

@arunasrivastava

Copy link
Copy Markdown
Collaborator

Fixes L2Arctic affricate label handling by preserving phoneme-token boundaries when available.

Examples for review:
data/ExamplesWithComments/l2arctic_affricates/*.wav
Each wav has a matching .txt showing the corrected token transcription vs the old joined parse. To my very biased ears, the fixes sound correct but feel free to verify. specifically verify you agree the "and jargon" one should be ʒ "zsh" not dʒ "dzh"

Updated:

  • L2Arctic HF export now includes ipa_tokens from TextGrid phone boundaries.
  • Training uses ipa_tokens directly when present, so t + ʃ / d + ʒ are not accidentally merged into tʃ / dʒ.
  • Joined IPA strings still fall back to the old tokenizer path for datasets without ipa_tokens.

After this PR:

  • Regenerate and push only the L2Arctic HF datasets.
  • Later, migrate other datasets to expose ipa_tokens in a separate PR.

Copy link
Copy Markdown
Member

Drafted by dotty, an AI assistant, for human review. Reviewed head 7317d5c21e91367ebcf8bd4094afbfd1b70b08e0.

P2: Preserve ipa_tokens in the mixed-supervision notebook

IPA_g2p_experiments.ipynb:379 explicitly selects only ipa, audio, and g2p. This drops the new token column before preprocessing. For experiments with HUMAN_PROPORTION > 0, human L2Arctic labels still take the joined-string fallback: ["t", "ʃ"] becomes one tʃ target, and ["d", "ʒ"] becomes dʒ. This is an existing caller missed by the fix; the main fine-tuning notebook uses the new defaults correctly.

Minimal fix:

combined_ds = combine_datasets(
    datasets, seed=RANDOM_SEED, columns=["ipa", "ipa_tokens", "audio", "g2p"]
)

Checks and limits

26 focused assertions passed using actual PR function bodies with datasets 4.4.1 / transformers 4.57.1: split versus true affricates, normalization, empty/missing-token fallback, mixed concatenation/interleaving, vocabulary, label selection, and synthetic exporter fixtures. The notebook omission above was reproduced separately. Tests used an isolated harness with real tokenizer/dataset operations but synthetic corpus inputs and audio/G2P test doubles; Python 3.12.14 / NumPy 2.5.3 differed from the full repository environment. Full corpus export, training, and listening review remain unverified. Please commit regression tests covering both callers and these boundary cases.

Existing validation/rollout caveats

  • Metrics decode labels to joined strings, so ["t", "ʃ"] and ["tʃ"] both become "tʃ". Existing PER/FER cannot measure this distinction; add token-sequence assertions or a boundary-aware evaluation slice.
  • Existing processed-data directories are loaded unconditionally. After regenerating L2Arctic data, use a fresh/versioned processed directory and rerun vocabulary/model preparation and preprocessing together, keeping tokenizer and target IDs matched.
  • The “and jargon” ʒ versus dʒ judgment still needs human listening review.

The metrics and cache behavior are pre-existing limitations, not new regressions introduced by this PR.

This branch was successfully deployed

1 active deployment
security — 7317d5c2 Deployed Aug 9, 2026 by arunasrivastava via gitleaks #372
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants