YL

Y. Liu

info

Please Note

1 records found

Multi-Source Speech Presence Probability Estimation and DNN-Guided MVDR Beamforming in Multichannel Reverberant Environments

A linear MVDR beamformer followed by a single-channel post-filter is MMSE-optimal only when the interference is Gaussian. That assumption fails when competing talkers are present, and the optimal estimator is then nonlinear. This thesis derives that estimator from a composite hypothesis structure which splits the per-bin signal-activity state into two questions: whether the target is present, and which interferer dominates. In closed form the estimator is a set of target-directed MVDR–Wiener filters, one per interference hypothesis, each weighted by an interferer posterior, with the weighted sum multiplied by a target posterior. These two posteriors are the multi-source speech presence probabilities, and the chain rule supplies the product structure exactly, without assuming that they factorise.

The estimator is shown to be structurally equivalent to the nonlinear spatial filter of Tesch and Gerkmann. Two sub-networks with no shared parameters estimate only these posteriors; all filtering is performed by the model-based filters. On a simulated five-talker reverberant corpus captured with a three-microphone array, the system improves SI-SDR by 11.26 dB and PESQ by 0.53, exceeding an end-to-end complex-mask baseline of comparable cost and more than doubling the gain of an oracle MVDR. Because the network outputs posteriors rather than a mask, the estimated quantities can be inspected directly and are shown to recover the intended signal-activity decomposition. ...