GenTSE: Enhancing Target Speaker Extraction via a Coarse-to-Fine Generative Language Model

ByAdmin

Dec 26, 2025

arXiv:2512.20978v1 Announce Type: cross
Abstract: Language Model (LM)-based generative modeling has emerged as a promising direction for TSE, offering potential for improved generalization and high-fidelity speech. We present GenTSE, a two-stage decoder-only generative LM approach for TSE: Stage-1 predicts coarse semantic tokens, and Stage-2 generates fine acoustic tokens. Separating semantics and acoustics stabilizes decoding and yields more faithful, content-aligned target speech. Both stages use continuous SSL or codec embeddings, offering richer context than discretized-prompt methods. To reduce exposure bias, we employ a Frozen-LM Conditioning training strategy that conditions the LMs on predicted tokens from earlier checkpoints to reduce the gap between teacher-forcing training and autoregressive inference. We further employ DPO to better align outputs with human perceptual preferences. Experiments on Libri2Mix show that GenTSE surpasses previous LM-based systems in speech quality, intelligibility, and speaker consistency.

GenTSE: Enhancing Target Speaker Extraction via a Coarse-to-Fine Generative Language Model

ByAdmin

By Admin

Related Post

Reward Learning through Ranking Mean Squared Error

Introduction to optimization methods for training SciML models

Who Owns the Text? Design Patterns for Preserving Authorship in AI-Assisted Writing

Leave a Reply Cancel reply

You missed

AI Survival Stories: a Taxonomic Analysis of AI Existential Risk

Bridging Semantic Understanding and Popularity Bias with LLMs

Who Owns the Text? Design Patterns for Preserving Authorship in AI-Assisted Writing

Introduction to optimization methods for training SciML models

THE AI TODAY