so-vits-svc fork with realtime support, improved interface and more features.
-
Updated
Sep 11, 2026 - Python
so-vits-svc fork with realtime support, improved interface and more features.
Self-Supervised Speech Pre-training and Representation Learning Toolkit
[ICASSP 2023] Mingling or Misalignment? Temporal Shift for Speech Emotion Recognition with Pre-trained Representations
This repo contains the source code of the first deep learning-base singing voice beat tracking system. It leverages WavLM and DistilHuBERT pre-trained speech models to create vocal embeddings and trains linear multi-head self-attention layers on top of them to extract vocal beat activations. Then, it uses HMM decoder to infer signing beats and t…
The code for the MAPSS measures for source separation evaluation (ICLR, 2026)
🦜 Synthetic Voice Detection
unsupervised spoken utterances scoring
Rethinking Leveraging Pre-Trained Multi-Layer Representations for Speaker Verification, ISCA Interspeech 2025
Layer-aware TDNN: Speaker Recognition Using Multi-Layer Features from Pre-Trained Models, to appear in ICAIIC 2026
code for our paper DistilALHuBERT: A Distilled Parameter Sharing Audio Representation Model
Unofficial PyTorch implementation of Higgs Audio V2 Tokenizer with HuBERT semantic features. Complete training pipeline for semantic-acoustic audio tokenization with 960x downsampling and 8-layer RVQ.
Speech Keyword detection using Wav2Vec Model
Functionality for speech data processing including time alignment, encoding with speech encoders (tokenizers) and data preprocessing of common datasets
Patches to train RVC v2 voice models for free on modern Google Colab (Python 3.12 / numpy 2.x / torch 2.x). Removes fairseq, fixes PyAV+matplotlib breakage.
Comparing text-only (RoBERTa), audio-only (wav2vec2/WavLM), and multimodal transfer learning approaches for Speech Emotion Recognition on RAVDESS, MELD, and IEMOCAP. PyTorch Lightning + HuggingFace + Gradio demo.
Partial Spoofed Speech Detection using HuBERT, Wav2Vec2, Mixture of Experts and ECAPA-TDNN-MSTA.
Disentangling prosodic representations from adaptation capacity in frozen HuBERT ASR.
A unified interface to extract hidden representations from various speech foundation models
Cross-corpus Speech Emotion Recognition using hybrid SSL + MFCC features, domain alignment, adaptive blending, and multiple classifiers.
Voice Cloning Bark HuBERT - Enables voice cloning from personalized audio samples by processing model's outputs into semantic tokens compatible with text-to-audio system.
To associate your repository with the hubert topic, visit your repo's landing page and select "manage topics."