Probabilistic Nonlinear Spatial Filtering for Multi-Talker Speech Enhancement

Multi-Source Speech Presence Probability Estimation and DNN-Guided MVDR Beamforming in Multichannel Reverberant Environments

Master Thesis (2026)
Author(s)

Y. Liu (TU Delft - Electrical Engineering, Mathematics and Computer Science)

Contributor(s)

R.C. Hendriks – Mentor (TU Delft - Electrical Engineering, Mathematics and Computer Science)

Jorge Abraham Martinez Castaneda – Graduation committee member (TU Delft - Electrical Engineering, Mathematics and Computer Science)

Faculty
Electrical Engineering, Mathematics and Computer Science
More Info
expand_more
Publication Year
2026
Language
English
Graduation Date
22-09-2026
Awarding Institution
Delft University of Technology
Programme
Electrical Engineering, Signals and Systems
Faculty
Electrical Engineering, Mathematics and Computer Science
Reuse Rights

Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.

Abstract

A linear MVDR beamformer followed by a single-channel post-filter is MMSE-optimal only when the interference is Gaussian. That assumption fails when competing talkers are present, and the optimal estimator is then nonlinear. This thesis derives that estimator from a composite hypothesis structure which splits the per-bin signal-activity state into two questions: whether the target is present, and which interferer dominates. In closed form the estimator is a set of target-directed MVDR–Wiener filters, one per interference hypothesis, each weighted by an interferer posterior, with the weighted sum multiplied by a target posterior. These two posteriors are the multi-source speech presence probabilities, and the chain rule supplies the product structure exactly, without assuming that they factorise.

The estimator is shown to be structurally equivalent to the nonlinear spatial filter of Tesch and Gerkmann. Two sub-networks with no shared parameters estimate only these posteriors; all filtering is performed by the model-based filters. On a simulated five-talker reverberant corpus captured with a three-microphone array, the system improves SI-SDR by 11.26 dB and PESQ by 0.53, exceeding an end-to-end complex-mask baseline of comparable cost and more than doubling the gain of an oracle MVDR. Because the network outputs posteriors rather than a mask, the estimated quantities can be inspected directly and are shown to recover the intended signal-activity decomposition.

Files

TU_Delft_Report_Thesis.pdf
(pdf | 4.89 Mb)
License info not available