<p>This page displays the records of the person named above and is not linked to a unique person identifier. This record may need to be merged to a profile.</p>
The efficient coding hypothesis posits that biological sensory systems maximize information transfer to the brain while minimizing neural resources. Although extensively studied in humans, its role in non-human auditory perception remains relatively unexplored. Here, we apply sparse coding to bat echolocation calls to test whether their vocalizations are intrinsically optimized for efficient representation. Unlike prior bat studies using black-box models, our approach examines how acoustic selectivity can emerge in early auditory structures from call structure alone, independent of higher-level neural processing. The learned kernel representations are compact, sparse, and functionally specialized, with distinct activation profiles encoding specific call shapes. These findings suggest that bat auditory systems are tuned to conspecific vocalizations and underscore the advantages of sparse coding over traditional signal representations. They also improve the interpretability of animal auditory processing and provide a computational basis for modeling animal signals, supporting future research in interspecies communication and decoding animal vocalizations.
...
The efficient coding hypothesis posits that biological sensory systems maximize information transfer to the brain while minimizing neural resources. Although extensively studied in humans, its role in non-human auditory perception remains relatively unexplored. Here, we apply sparse coding to bat echolocation calls to test whether their vocalizations are intrinsically optimized for efficient representation. Unlike prior bat studies using black-box models, our approach examines how acoustic selectivity can emerge in early auditory structures from call structure alone, independent of higher-level neural processing. The learned kernel representations are compact, sparse, and functionally specialized, with distinct activation profiles encoding specific call shapes. These findings suggest that bat auditory systems are tuned to conspecific vocalizations and underscore the advantages of sparse coding over traditional signal representations. They also improve the interpretability of animal auditory processing and provide a computational basis for modeling animal signals, supporting future research in interspecies communication and decoding animal vocalizations.
Dysarthric speech poses significant challenges for automatic speech recognition (ASR) systems due to its high variability and reduced intelligibility. In this work we explore the use of diffusion models for dysarthric speech enhancement, which is based on the hypothesis that using diffusion-based speech enhancement moves the distribution of dysarthric speech closer to that of typical speech, which could potentially improve dysarthric speech recognition performance. We assess the effect of two diffusion-based and one signal-processing-based speech enhancement algorithms on intelligibility and speech quality of two English dysarthric speech corpora. We applied speech enhancement to both typical and dysarthric speech and evaluate the ASR performance using Whisper-Turbo, and the subjective and objective speech quality of the original and enhanced dysarthric speech. We also fine-tuned Whisper-Turbo on the enhanced speech to assess its impact on recognition performance.
...
Dysarthric speech poses significant challenges for automatic speech recognition (ASR) systems due to its high variability and reduced intelligibility. In this work we explore the use of diffusion models for dysarthric speech enhancement, which is based on the hypothesis that using diffusion-based speech enhancement moves the distribution of dysarthric speech closer to that of typical speech, which could potentially improve dysarthric speech recognition performance. We assess the effect of two diffusion-based and one signal-processing-based speech enhancement algorithms on intelligibility and speech quality of two English dysarthric speech corpora. We applied speech enhancement to both typical and dysarthric speech and evaluate the ASR performance using Whisper-Turbo, and the subjective and objective speech quality of the original and enhanced dysarthric speech. We also fine-tuned Whisper-Turbo on the enhanced speech to assess its impact on recognition performance.
In this paper we propose a robust loudspeaker beamforming algorithm which is used to enhance the performance of voice driven applications in scenarios where the loudspeakers introduce the majority of the noise, e.g. when music is playing loudly. The loudspeaker beamformer modifies the loudspeaker playback signals to create a low-acoustic-energy region around the device that implements automatic speech recognition for a voice driven application (VDA). The algorithm utilises a distortion measure based on human auditory perception to limit the distortion perceived by human listeners. Simulations and real-world experiments show that the proposed loudspeaker beamformer improves the speech recognition performance in all tested scenarios. Moreover, the algorithm allows to further reduce the acoustic energy around the VDA device at the expense of reduced objective audio quality at the listener’s location
...
In this paper we propose a robust loudspeaker beamforming algorithm which is used to enhance the performance of voice driven applications in scenarios where the loudspeakers introduce the majority of the noise, e.g. when music is playing loudly. The loudspeaker beamformer modifies the loudspeaker playback signals to create a low-acoustic-energy region around the device that implements automatic speech recognition for a voice driven application (VDA). The algorithm utilises a distortion measure based on human auditory perception to limit the distortion perceived by human listeners. Simulations and real-world experiments show that the proposed loudspeaker beamformer improves the speech recognition performance in all tested scenarios. Moreover, the algorithm allows to further reduce the acoustic energy around the VDA device at the expense of reduced objective audio quality at the listener’s location
Cookie settings
We use necessary cookies to make the TU Delft Repository work.
Help us improve the Repository
With your permission, we use privacy-friendly Matomo analytics to understand how people use the
Repository — for example, which features are used and where we can improve the search experience. The analytics are managed by TU Delft and are not used for advertising or commercial tracking. Your IP
address is anonymized, and analytics data is not shared with third parties.
Choosing “Accept all” helps the Library improve the Repository for researchers, students, and other
users. You can change your choice at any time using the cookie settings icon in the footer. For more information, read our
privacy statement.