TTS · Prosody Control · Voice Conversion · NLP
I am a Research Engineer in Speech within the Hi! PARIS / École Polytechnique ecosystem, working at the intersection of speech generation, prosody control, voice conversion, and NLP.
My work combines scientific rigor with production-grade engineering — from controllable French TTS and WavLM-based resynthesis to zero-shot multilingual voice conversion and large-scale, reproducible experimentation on HPC (Jean Zay / A100). I collaborate closely with senior researchers including Éric Moulines and Reda Dehak.
- 🔊 Controllable French TTS with explicit prosody planning
- 🧠 SSML-based modeling for pauses, rhythm, emphasis, and timing
- 🧪 WavLM → audio resynthesis with adversarial training and layer ablations
- 🚀 Zero-shot voice conversion in learned speech-representation spaces
- 🌍 Multilingual TTS adaptation for European low-resource languages
- 📊 Objective & perceptual evaluation for reproducible speech research
Specialties: DDP / AMP distributed training · GAN & flow-matching / diffusion generation · LoRA / PEFT fine-tuning · speech representation learning · RAG & LLM orchestration · reproducible evaluation pipelines
Controllable French speech synthesis with explicit SSML planning for pauses, timing, and emphasis — symbolic pause planning, break prediction, SSML generation, and reproducible training/inference.
Repository · Demo · Paper — ICNLSP 2025 · HF: text2breaks · break2ssml
Waveform resynthesis from frozen WavLM representations: adversarial reconstruction (HiFi-GAN + MPD/MSD), learned weighted layer fusion over WavLM-Base+, chunked overlap-add inference, and layer-ablation studies. Trained on 238h of cleaned French speech (SIWIS, M-AILABS, Common Voice).
Repository · Demo · HF Models · Paper — JEP 2026
Component-level adaptation of CosyVoice2 for European languages (French & German): language-specific fine-tuning of the text encoder, flow-matching module, and vocoder, evaluated across seen/unseen speakers.
Repository · Demo · Paper under submission — EUSIPCO 2026
Production-deployed retrieval-augmented assistant for HEC Paris, combining document retrieval with an agentic LLM pipeline, served as a full web application on Azure. Focus on robust orchestration, controlled tool-calling, and cost-efficient inference.
Live app (HEC login required)
Stack: RAG · LLM orchestration · vector retrieval · Flask/web app · Azure deployment
Developing a new computer-vision module for the aikon-demo platform that makes fine-grained, interpretable paleographic (historical-script) analysis accessible to historians without ML expertise. End-to-end pipeline (milestone 1): a new dataset object pairing text-line images with their transcriptions, a training/inference API that fine-tunes networks for morphological script-metrological analysis (following Vlachou-Efstathiou et al., ICDAR 2026), and an interactive front-end to explore results — character prototypes, morphological differences, distances, and diachronic evolution. Built for non-expert testing, open-source release, and documentation; later extensible to typographers. Developed with Mathieu Aubry's group at École des Ponts (ENPC).
Platform · Repository · Stack: Computer Vision · Python · PyTorch · web platform (front-end + API)
| Venue | Title | Status |
|---|---|---|
| ICNLSP 2025 | Improving French Synthetic Speech Quality via SSML Prosody Control | Published — ACL Anthology |
| JEP 2026 | WavLM-Vocoder-French: Neural Waveform Resynthesis from Frozen WavLM Representations | Published — ISCA Archive |
| EUSIPCO 2026 | Europeanizing Modular Zero-Shot TTS: A Component-Level Adaptation Framework for French and German | Under submission |
Released models: ssml-text2breaks-fr-lora · ssml-break2ssml-fr-lora Demos: WavLM2Audio · Prosody-Control TTS · CosyVoice2-EU
| Degree | Institution |
|---|---|
| Master's degree | Université Gustave Eiffel |
| Master's degree | UVSQ (Versailles Saint-Quentin-en-Yvelines) |
| Bachelor's degree | Université Paris Descartes |
Open to research collaborations and industry partnerships in controllable TTS, prosody modeling, French speech technology, multilingual voice conversion, and reproducible evaluation for speech systems — especially where scientific rigor, data confidentiality, and engineering quality matter.
If you find a repository useful, a ⭐ helps visibility and supports continued maintenance.



