Skip to content
View NassimaOULDOUALI's full-sized avatar
🎯
Focusing
🎯
Focusing

Organizations

@hi-paris

Block or report NassimaOULDOUALI

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
NassimaOULDOUALI/README.md

Header

👋 Hi, I'm Nassima OULD OUALI

Research Engineer in Speech @ Hi! PARIS / École Polytechnique

TTS · Prosody Control · Voice Conversion · NLP

LinkedIn Email GitHub Google Scholar Hugging Face


About me

I am a Research Engineer in Speech within the Hi! PARIS / École Polytechnique ecosystem, working at the intersection of speech generation, prosody control, voice conversion, and NLP.

My work combines scientific rigor with production-grade engineering — from controllable French TTS and WavLM-based resynthesis to zero-shot multilingual voice conversion and large-scale, reproducible experimentation on HPC (Jean Zay / A100). I collaborate closely with senior researchers including Éric Moulines and Reda Dehak.


🔬 Research focus

  • 🔊 Controllable French TTS with explicit prosody planning
  • 🧠 SSML-based modeling for pauses, rhythm, emphasis, and timing
  • 🧪 WavLM → audio resynthesis with adversarial training and layer ablations
  • 🚀 Zero-shot voice conversion in learned speech-representation spaces
  • 🌍 Multilingual TTS adaptation for European low-resource languages
  • 📊 Objective & perceptual evaluation for reproducible speech research

🛠️ Tech stack

Python PyTorch Hugging Face CUDA Slurm Docker Azure LaTeX

Specialties: DDP / AMP distributed training · GAN & flow-matching / diffusion generation · LoRA / PEFT fine-tuning · speech representation learning · RAG & LLM orchestration · reproducible evaluation pipelines


🚀 Featured projects

🎧 Prosody-Control French TTS (SSML)

Controllable French speech synthesis with explicit SSML planning for pauses, timing, and emphasis — symbolic pause planning, break prediction, SSML generation, and reproducible training/inference.

Repository · Demo · Paper — ICNLSP 2025 · HF: text2breaks · break2ssml

🧪 WavLM Vocoder for French

Waveform resynthesis from frozen WavLM representations: adversarial reconstruction (HiFi-GAN + MPD/MSD), learned weighted layer fusion over WavLM-Base+, chunked overlap-add inference, and layer-ablation studies. Trained on 238h of cleaned French speech (SIWIS, M-AILABS, Common Voice).

Repository · Demo · HF Models · Paper — JEP 2026

🌍 CosyVoice2-EU — Multilingual Zero-Shot TTS

Component-level adaptation of CosyVoice2 for European languages (French & German): language-specific fine-tuning of the text encoder, flow-matching module, and vocoder, evaluated across seen/unseen speakers.

Repository · Demo · Paper under submission — EUSIPCO 2026

🧭 Compass — Agentic RAG Assistant (HEC Paris)

Production-deployed retrieval-augmented assistant for HEC Paris, combining document retrieval with an agentic LLM pipeline, served as a full web application on Azure. Focus on robust orchestration, controlled tool-calling, and cost-efficient inference.

Live app (HEC login required)
Stack: RAG · LLM orchestration · vector retrieval · Flask/web app · Azure deployment

👁️ AIKON — Paleographic Analysis Module (École des Ponts / ENPC)

Developing a new computer-vision module for the aikon-demo platform that makes fine-grained, interpretable paleographic (historical-script) analysis accessible to historians without ML expertise. End-to-end pipeline (milestone 1): a new dataset object pairing text-line images with their transcriptions, a training/inference API that fine-tunes networks for morphological script-metrological analysis (following Vlachou-Efstathiou et al., ICDAR 2026), and an interactive front-end to explore results — character prototypes, morphological differences, distances, and diachronic evolution. Built for non-expert testing, open-source release, and documentation; later extensible to typographers. Developed with Mathieu Aubry's group at École des Ponts (ENPC).

Platform · Repository · Stack: Computer Vision · Python · PyTorch · web platform (front-end + API)


📚 Publications

Venue Title Status
ICNLSP 2025 Improving French Synthetic Speech Quality via SSML Prosody Control Published — ACL Anthology
JEP 2026 WavLM-Vocoder-French: Neural Waveform Resynthesis from Frozen WavLM Representations Published — ISCA Archive
EUSIPCO 2026 Europeanizing Modular Zero-Shot TTS: A Component-Level Adaptation Framework for French and German Under submission

Released models: ssml-text2breaks-fr-lora · ssml-break2ssml-fr-lora Demos: WavLM2Audio · Prosody-Control TTS · CosyVoice2-EU


🎓 Education

Degree Institution
Master's degree Université Gustave Eiffel
Master's degree UVSQ (Versailles Saint-Quentin-en-Yvelines)
Bachelor's degree Université Paris Descartes

🤝 Collaboration

Open to research collaborations and industry partnerships in controllable TTS, prosody modeling, French speech technology, multilingual voice conversion, and reproducible evaluation for speech systems — especially where scientific rigor, data confidentiality, and engineering quality matter.

If you find a repository useful, a ⭐ helps visibility and supports continued maintenance.

Footer

Popular repositories Loading

  1. Prosody-Control-French-TTS Prosody-Control-French-TTS Public archive

    An End-to-End Pipeline for Enhanced French Text-to-Speech with SSML Prosody Control

    Python 19

  2. CosyVoice2-EU-original CosyVoice2-EU-original Public

    Jupyter Notebook 3

  3. ManimDemoTTS ManimDemoTTS Public

    Python 2

  4. NassimaOULDOUALI NassimaOULDOUALI Public

    1

  5. wavlm-vocoder-french wavlm-vocoder-french Public archive

    Python 1

  6. Summer_school2025 Summer_school2025 Public

    Python