Jesper Jensen
Please Note
14 records found
1
Wireless acoustic sensor networks (WASNs) can be used for centralized multi-microphone noise reduction, where the processing is done in a fusion center (FC). To perform the noise reduction, the data needs to be transmitted to the FC. Considering the limited battery life of the devices in a WASN, the total data rate at which the FC can communicate with the different network devices should be constrained. In this article, we propose a rate-constrained multi-microphone noise reduction algorithm, which jointly finds the best rate allocation and estimation weights for the microphones across all frequencies. The optimal linear estimators are found to be the quantized Wiener filters, and the rates are the solutions to a filter-dependent reverse water-filling problem. The performance of the proposed framework is evaluated using simulations in terms of mean square error and predicted speech intelligibility. The results show that the proposed method is very close in performance to that of the existing optimal method based on discrete optimization. However, the proposed approach can do this at a much lower complexity, while the existing optimal reference method needs a non-tractable exhaustive search to find the best rate allocation across microphones.
Compared to monaural hearing aids (HAs), binaural hearing aid systems, in which there is a communication link between the two devices, have improved noise reduction capabilities and the ability to preserve binaural spatial information. However, the limited HA battery lifetime puts constraints on the amount of information that can be shared between the two devices. In other words, the rate of transmission between the devices is an important constraint that needs to be considered, while preserving the spatial information. In this article, a linearly constrained noise reduction problem is proposed, which jointly finds the optimal rate allocation and the optimal estimation (beamforming) weights across all sensors and frequencies, while preserving the binaural spatial cues of point sources. The proposed method considers a rate constraint together with linear constraints to preserve the binaural spatial cues of point sources. Minimizing the mean square error on the estimated target speech at the left and the right side beamformers, the optimal weights are found to be rate-constrained linearly constrained minimum variance (LCMV) filters, and the optimal rates are found to be the solutions to a set of reverse water filling problems. The performance of the proposed method is evaluated using the averaged binaural signal-to-noise ratio (SNR), the interaural level difference (ILD) error and the interaural time difference (ITD) error. The results show that the proposed method outperforms spatially correct noise reduction approaches that use naive/random rate allocation strategies.
Factor analysis is a popular tool in multivariate statistics, applied in several areas of study such as psychology, economics, chemistry and signal processing. Given a set of observed random variables, factor analysis aims at explaining and analyzing the correlation between these random variables. This is done by finding a meaningful structural model representation for the correlation matrix of the observed random variables, and subsequently estimating the underlying model parameters. In this paper, we focus on factor analysis methods applied to a commonly used signal model for sensor arrays applications and use it to jointly estimate the underlying model parameters. In addition we discuss practical considerations of these methods.
One of the biggest challenges in multimicrophone applications is the estimation of the parameters of the signal model, such as the power spectral densities (PSDs) of the sources, the early (relative) acoustic transfer functions of the sources with respect to the microphones, the PSD of late reverberation, and the PSDs of microphone-self noise. Typically, existing methods estimate subsets of the aforementioned parameters and assume some of the other parameters to be known a priori. This may result in inconsistencies and inaccurately estimated parameters and potential performance degradation in the applications using these estimated parameters. So far, there is no method to jointly estimate all the aforementioned parameters. In this paper, we propose a robust method for jointly estimating all the aforementioned parameters using confirmatory factor analysis. The estimation accuracy of the signal-model parameters thus obtained outperforms existing methods in most cases. We experimentally show significant performance gains in several multimicrophone applications over state-of-the-art methods.
The recently proposed relaxed binaural beamforming (RBB) optimization problem provides a flexible tradeoff between noise suppression and binaural-cue preservation of the sound sources in the acoustic scene. It minimizes the output noise power, under the constraints, which guarantee that the target remains unchanged after processing and the binaural-cue distortions of the acoustic sources will be less than a user-defined threshold. However, the RBB problem is a computationally demanding non convex optimization problem. The only existing suboptimal method which approximately solves the RBB is a successive convex optimization (SCO) method which, typically, requires to solve multiple convex optimization problems per frequency bin, in order to converge. Convergence is achieved when all constraints of the RBB optimization problem are satisfied. In this paper, we propose a semidefinite convex relaxation (SDCR) of the RBB optimization problem. The proposed suboptimal SDCR method solves a single convex optimization problem per frequency bin, resulting in a much lower computational complexity than the SCO method. Unlike the SCO method, the SDCR method does not guarantee user-controlled upper-bounded binaural-cue distortions. To tackle this problem, we also propose a suboptimal hybrid method that combines the SDCR and SCO methods. Instrumental measures combined with a listening test show that the SDCR and hybrid methods achieve significantly lower computational complexity than the SCO method, and in most cases better tradeoff between predicted intelligibility and binaural-cue preservation than the SCO method.
Binaural hearing aids (HAs) can potentially perform advanced noise reduction algorithms, leading to an improvement over monaural/bilateral HAs. Due to the limited transmission capacities between the HAs and given knowledge of the complete joint noisy signal statistics, the optimal rate-constrained beamforming strategy is known from the literature. However, as these joint statistics are unknown in practice, sub-optimal strategies have been presented. In this paper, we present a unified framework to study the performance of these existing optimal and sub-optimal rate-constrained beamforming methods for binaural HAs. Moreover, we propose to use an asymmetric sequential coding scheme to estimate the joint statistics between the microphones in the two HAs. We show that under certain assumptions, this leads to sub-optimal performance in one HA but allows to obtain the truly optimal performance in the second HA. Based on the mean square error distortion measure, we evaluate the performance improvement between monaural beamforming (no communication) and the proposed scheme, as well as the optimal and the existing sub-optimal strategies in terms of the information bit-rate. The results show that the proposed method outperforms existing practical approaches in most scenarios, especially at middle rates and high rates, without having the prior knowledge of the joint statistics.
While the majority of binaural beamformers aim to minimize the output noise power while (approximately) preserving the binaural cues of the sources using constraints, we propose in this paper to minimize the binaural-cue distortions of the sources in the acoustic scene, such that the output noise power is below a predefined threshold. This new problem formulation is a convex QCQP problem, which leads to an efficient trade-off between noise reduction, binaural-cue preservation and complexity. In particular, the proposed beamformer provides a better trade-off between noise reduction and binaural-cue preservation (in terms of interaural level and phase differences) compared to the well-known binaural minimum variance distortionless response-η beamformer.
In this paper, we perceptually evaluate two recently proposed binaural multi-microphone speech enhancement methods in terms of intelligibility improvement and binaural-cue preservation. We compare these two methods with the well-known binaural minimum variance distortionless response (BMVDR) method. More specifically, we measure the 50% speech reception threshold, and the localization error of all dominant point sources in three different acoustic scenes. The listening tests are divided into a parameter selection phase and a testing phase. The parameter selection phase is used to select the algorithms' parameters based on one acoustic scene. In the testing phase, the two methods are evaluated in two other acoustic scenes in order to examine their robustness. Both methods achieve significantly better intelligiblity compared to the unprocessed scene, and slightly worse intelligibility than the BMVDR method. However, unlike the BMVDR method which severely distorts the binaural cues of all interferers, the new methods achieve localization errors which are not significantly different compared to those of the unprocessed scene.
Modern binaural hearing aids (HAs) can collaborate wirelessly with each other as well as with other assistive (wireless) devices. This enables multi-microphone noise reduction over small wireless acoustic sensor networks (WASNs) to increase the intelligibility under adverse conditions. In this work, we assume one of the HAs to serve as the fusion center (FC). The optimal beamforming strategy for processing the received data at the FC depends on the acoustic scene and physical constraints (e.g., the bit-rate for transmission to the FC), and might be frequency dependent. Selection of the optimal beamforming strategy, while satisfying rate constraints on the communication between the different devices is an important challenge in such setups. In this paper, we propose an operational rate-constrained beamforming system for optimal rate allocation and strategy selection across frequency. We show an example of the proposed framework, where both the algorithm selection as well as the required rates to transmit the necessary microphone signals are optimized using uniform quantizers, while minimizing the mean-square error (MSE) distortion measure. In contrast to a well-known (theoretically optimal) reference method based on remote source coding for two devices, the presented algorithm is practically implementable and only requires knowledge of joint signal statistics at the FC. Evaluations (based on simulation experiments) show clear improvement over other practically implementable strategies.