CL

C. Li

info

Please Note

9 records found

Journal article (2026) - Sen Yuan, Changheng Li, Francesco Fioranelli
This study addresses the combined challenges of reduced maximum unambiguous Doppler velocity and DOA estimation blurring inherent in time-division multiplexing (TDM) multi-input–multi-output (MIMO) frequency-modulated continuous wave (FMCW) radar. We propose a novel joint de-aliasing framework that integrates peak searching via matched filtering with a compensated direction-of-arrival (DOA) processing scheme. A key advantage of the proposed method is its versatility: it can work with general MIMO configurations without the requirement of overlapped virtual elements, whereas existing state-of-the-art methods typically necessitate specific virtual array configurations. The framework efficacy is validated through numerical simulations to establish theoretical performance, experimental data from a radar target simulator (RTS) utilizing an irregular MIMO topology based on Texas instruments (TIs) cascaded radar boards, and real-world driving scenarios to demonstrate practical robustness. Results indicate that the proposed method achieves an estimation error comparable to state-of-the-art benchmarks—within a 1% margin in RTS experiments—while providing reliable target localization in complex dynamic environments. ...
Doctoral thesis (2025) - C. Li, A.J. van der Veen, R.C. Hendriks
Many modern devices, such as mobile phones, hearing aids and (hands-free) acoustic human-machine interfaces are equipped with microphone arrays that can be used for various applications. These applications include source separation, audio quality enhancement, speech intelligibility improvement and source localization. In an ideal anechoic chamber, the signals received by ideal microphones are just attenuated and delayed version of the original sound. However, in practice, obstacles such as the floor, the ceiling and the surrounding walls will reflect the sound to the microphones. Also, the microphone itself will generate noise, distorting the recorded signals. Lastly, it is possible that multiple point sources are active simultaneously. When we consider one point source as the target signal, the other sources could be considered interfering signals. These distortions make it difficult to get access to the target signal. Therefore, spatial filtering is often applied to the microphone signals.

To achieve satisfying performance, these spatial filters typically need to be adaptive to the (changing) scene. Specifically, the filter coefficients depend on the acoustic-scene related parameters that model the microphone signals. These parameters, such as the relative transfer functions (RTFs) of the sources, the power spectral densities (PSDs) of the sources, the late reverberation and the ambient noise, are typically unknown in practice. Therefore, estimation of these parameters is crucial and thus the main focus of the dissertation. While it is relatively straightforward to estimate these parameters in less complex acoustic scenes, these algorithms are usually not applicable and not extendable to more complex acoustic scenes. Therefore, the complexity of the estimation methods needed depends on the complexity of the acoustic scene.

In Chapter 3, we consider the simplest acoustic scene in this dissertation, where there is only a single source in a reverberant and noiseless environment. The parameters that we aim to estimate are the RTFs, the PSDs of the target signal and the PSDs of the late reverberation. A joint estimator using a single time frame is first proposed, having a closed form. Then, a joint estimator using multiple time frames having the same RTF is proposed, where the solution for each iteration step is in closed form. The parameter estimation accuracy and the additional performance of noise reduction, speech quality and speech intelligibility of the proposed method are compared to various state-of-the-art reference methods. The proposed method reduces computational costs and improves performance as demonstrated by the experiments.

Next, we extend the noiseless signal model in Chapter 3 to the noisy model in Chapters 4 and 5. In Chapter 4, we focus on RTF estimation and propose an estimator that is robust to the late reverberation and noise PSD errors. This is achieved by using only off-diagonal elements of a simplified covariance matrix. The experiments demonstrate the effectiveness of the proposed method. In Chapter 5, a joint estimator of the RTFs, the PSDs of the source, the PSDs of the late reverberation, and the PSDs of the ambient noise is proposed when using a single time frame as well as when using multiple time frames that share the same RTF.

Beyond the acoustic scene of a single point source, in Chapter 6 and 7, we consider the scenario of multiple point sources. In Chapter 6, we first consider the case where the environment is close to non-reverberant and noiseless. Under this assumption, we propose a method to estimate the RTFs. We obtain satisfying estimates by averaging covariance matrices for as many time frames as possible without suffering too much from model mismatch errors caused by distortion signals. This method is based on a comparison of several estimates from different averaged covariance matrices, which is somewhat heuristically motivated and not satisfying for reverberant and noisy environments. Therefore, in Chapter 7, we propose a robust method that works in reverberant and noisy environment and estimates not only the RTFs but also the PSDs of the sources and the late reverberation.

As in most of the works we have introduced, we use the prior information that several consecutive time frames share the same RTF. However, this is only possible if the source stays at the same position during these time frames. In Chapter 8, we therefore propose a method to adaptively segment signals into segments where the source is considered static. The proposed method is combined with the estimator we proposed in Chapter 3 for estimating the parameters of a single non-static source. It is shown in the experiments that with our proposed adaptive time segmentation, the estimation performance is improved over the use of a fixed time segmentation.
...
Journal article (2024) - Zhonggang Li, Changheng Li, Raj Thilak Rajan
Consensus control of multiagent systems arises in various applications such as rendezvous and formation control. The input to these algorithms, e.g., the (relative) positions of neighboring agents need to be measured using various sensors. Recent works aim to reconstruct these positions, i.e., achieve localization using Euclidean distance measurements instead of displacements, for cost efficiency and scalability. However, this approach inherently introduces ambiguities, such as a rotation or a reflection, which can cause stability issues in practice without corrections by some anchors. In this letter, we conduct a thorough analysis of the stability of consensus control in the presence of localization-induced rotational ambiguities, in several scenarios including, e.g., proper and improper rotation, and the homogeneity of rotations. We give stability criteria and stability margin on the rotations, which are numerically verified with two traditional examples of consensus control. ...
Conference paper (2023) - Changheng Li, Richard C.Hendriks
Estimating the parameters that describe the acoustic scene is very important for many microphone array applications. For example, consider the power spectral densities (PSDs) or relative acoustic transfer functions (RTFs) that are required when estimating a particular sound source using multi-microphone noise reduction. State-of-the-art algorithms estimate the pa-rameters per segment, where each segment consists of a fixed number of time frames. These algorithms exploit the assumption that PSDs are constant per time frame, and RTFs are constant per segment. However, in practice, sound sources will move relative to the microphone array. Improved per-formance is therefore expected when the actual time frames that are used to form the segments are adapted such that time frames all share the same (unknown) RTF. In this paper, we therefore present an algorithm to obtain an optimal adaptive time segmentation and combine this with our previously pub-lished joint maximum likelihood estimator (JMLE) for jointly estimating the RTF, source PSD and late reverberation PSD of a single source in a reverberant environment. ...
Journal article (2023) - Changheng Li, Richard C. Hendriks
Acoustic-scene-related parameters such as relative transfer functions (RTFs) and power spectral densities (PSDs) of the target source, late reverberation and ambient noise are essential for microphone array signal processing and are challenging to estimate. Existing methods typically only estimate a subset of the parameters by assuming the other parameters are known. This can lead to unmatched scenarios and reduced estimation performance on the parameters of interest. Moreover, many methods process time frames independently, despite they share common information such as the same RTF. In this work, we consider a noisy scenario by modelling the noise component as a spatially homogeneous sound field with a time-invariant spatial coherence matrix and time-varying PSD. We first modify an existing alternating least squares (ALS) method to obtain more accurate estimates using a single time frame. Then, we extend the method to use multiple time frames that share the same RTF. Furthermore, we propose more robust constraints on the PSDs to avoid large estimation errors. We compare our proposed methods to the state-of-the-art simultaneously confirmatory factor analysis (SCFA) method, a joint maximum likelihood estimation (JMLE) method and an existing ALS-based method. The experimental results in terms of estimation accuracy, noise reduction performance, predicted speech quality, and predicted speech intelligibility demonstrate that our proposed methods achieve similar performance compared to the state-of-the-art SCFA method, which outperforms the existing ALS method in all scenarios and outperforms the JMLE method particularly in low SNR scenarios. Moreover, our proposed methods have significantly lower computational complexity than SCFA. ...
Conference paper (2023) - Changheng Li, Richard Hendriks
Spatial filtering techniques typically rely on estimates of the target relative transfer function (RTF). However, the target speech signal is typically corrupted by late reverberation and ambient noise, which complicates RTF estimation. Existing methods subtract the noise covariance matrix to obtain the target plus late reverberation covariance matrix, from where the RTF is estimated. However, the noise covariance matrix is typically unknown. More specifically, the noise power spectral density (PSD) is typically unknown, while the spatial coherence matrix can be assumed known as it might remain time-invariant for a longer time. Using the spatial coherence matrices we simplify the signal model such that the off-diagonal elements are not affected by the PSDs of the late reverberation and the ambient noise. Then we use these elements to estimate the target covariance matrix, from where the RTF can be obtained. Hence, the resulting estimate of the RTF is insensitive to the noise PSD. Experiments demonstrate the estimation performance of our proposed method. ...
Estimation of the acoustic-scene related parameters such as relative transfer functions (RTFs) from source to microphones, source power spectral densities (PSDs) and PSDs of the late reverberation is essential and also challenging. Existing maximum likelihood estimators typically consider only subsets of these parameters and use each time frame separately. In this paper we explicitly focus on the single source scenario and first propose a joint maximum likelihood estimator (MLE) to estimate all parameters jointly using a single time frame. Since the RTFs are typically invariant for a number of consecutive time frames we also propose a joint maximum likelihood estimator (MLE) using multiple time frames which has similar estimation performance compared to a recently proposed reference algorithm called simultaneously confirmatory factor analysis (SCFA), but at a much lower complexity. Moreover, we present experimental results which demonstrate that the estimation accuracy, together with the performance of noise reduction, speech quality and speech intelligibility, of our proposed joint MLE outperform those of existing MLE based approaches that use only a single time frame. ...
Conference paper (2022) - Changheng Li, Jorge Martinez, Richard C. Hendriks
Many multi-microphone algorithms depend on knowing the relative acoustic transfer functions (RTFs) of the individual sound sources in the acoustic scene. However, accurate joint RTF estimation for multiple sources is a challenging problem. Existing methods to jointly estimate the RTF for multiple sources have either no satisfying performance, or, suffer from a very large computational complexity. In this paper, we propose a method for robust estimation of the individual RTFs in a multi-source acoustic scenario. The presented algorithm is based on linear algebraic concepts and therefore of lower computational complexity compared to a recently presented state-of-the-art algorithm, while having a similar performance. Experimental results are presented to demonstrate the RTF estimation performance as well as the noise reduction performance when combining the estimated RTFs with a beamformer. ...
Journal article (2021) - Jie Zhang, Changheng Li
For hearing-impaired listeners, both ambient noise suppression and binaural cues preservation of directional sources are required, such that a complete spatial awareness of the acoustic scene can be obtained. It was shown that the binaural multichannel Wiener filter (MWF) with partial noise preservation can achieve joint noise reduction and binaural cues preservation and incorporating an external microphone signal improves the performance of binaural MWFs. Motivated by this, we propose a binaural MWF incorporating external wireless devices in this paper. First, we theoretically analyze the performance of the MWF in terms of output signal-to-noise ratio (SNR) and binaural cues preservation errors. As in practice the external devices are power driven with a limited amount of battery resource and the power consumption heavily depends on the transmission rate, given an expected noise reduction performance we then optimize the bit-rate for a single external microphone. Further, we consider to minimize the total power consumption over multiple external devices under a constraint on the output SNR, which turns out to be a rate distribution problem. The proposed rate-distributed binaural MWF is evaluated using a hearing-aid setup with various dynamics. It is shown that the proposed method can obtain a desired SNR at a much lower bit-rate, and an expected trade-off between SNR gain and binaural cues preservation accuracy can be obtained by optimizing the bit-rates. Increasing the bit-rates improves both instrumental speech quality and speech intelligibility. ...