DM
D.B. Meszka
info
Please Note
<p>This page displays the records of the person named above and is not linked to a unique person identifier. This record may need to be merged to a profile.</p>
1 records found
1
Dysarthric speech remains challenging for state-of-the-art automatic speech recognition (ASR) systems, despite their high accuracy on typical speech. Personalized dysarthric ASR offers a way to address the substantial differences in speech characteristics between individuals, but it is constrained by the limited amount of speaker-specific dysarthric data available for training. Speech produced by the same target speaker in another language provides an additional source of labeled data while preserving speaker-specific and dysarthric speech characteristics.
This study investigates bilingual phoneme-level contrastive learning (CL) for personalized dysarthric ASR in a case study involving a single speaker with severe dysarthria producing speech in Dutch and English. Phoneme-level modeling is used as phonemes provide a natural unit for identifying relationships between speech sounds across languages. English and Dutch speech are jointly modeled by constructing contrastive pairs from phonemes that are either shared across the two languages or manually identified as phonetically equivalent. Because these explicitly cross-lingual relationships constitute only a subset of the possible positive pairs, three positive sampling strategies and corresponding loss-weighting variants are evaluated to investigate whether giving them greater influence during training improves recognition or more strongly shapes the learned representation space. Two approaches for incorporating additional typical speech are also evaluated to investigate the effect of training-data composition.
Evaluation considers both phoneme error rate (PER) and the structure of the learned phoneme embedding space using cosine-distance, nearest-neighbor, and silhouette-based analyses. To the best of our knowledge, this is the first study to investigate bilingual phoneme-level contrastive learning for personalized dysarthric ASR.
Results show that phoneme-level CL improves recognition over bilingual fine-tuning using only Connectionist Temporal Classification (CTC), but explicitly prioritizing cross-lingual phoneme relationships does not provide a statistically significant improvement over the simpler same-symbol contrastive learning. At the representation level, bilingual and weighted objectives can bring cross-lingual phoneme representations closer together and produce better-separated phoneme clusters, but these changes do not consistently result in lower PER. Replacing dysarthric speech with typical speech significantly worsens recognition when the amount of training data is kept approximately constant, whereas adding typical speech on top of the full dysarthric training set significantly improves recognition; however, the latter condition also contains more training data, so the improvement cannot be attributed specifically to typical speech. Overall, the results show a benefit from phoneme-level contrastive learning for personalized dysarthric ASR, but no clear additional benefit from explicitly encoding canonical Dutch--English phoneme relationships. This suggests that the usefulness of bilingual contrastive learning may depend on defining cross-lingual relationships that better reflect the target speaker's actual speech patterns. ...
This study investigates bilingual phoneme-level contrastive learning (CL) for personalized dysarthric ASR in a case study involving a single speaker with severe dysarthria producing speech in Dutch and English. Phoneme-level modeling is used as phonemes provide a natural unit for identifying relationships between speech sounds across languages. English and Dutch speech are jointly modeled by constructing contrastive pairs from phonemes that are either shared across the two languages or manually identified as phonetically equivalent. Because these explicitly cross-lingual relationships constitute only a subset of the possible positive pairs, three positive sampling strategies and corresponding loss-weighting variants are evaluated to investigate whether giving them greater influence during training improves recognition or more strongly shapes the learned representation space. Two approaches for incorporating additional typical speech are also evaluated to investigate the effect of training-data composition.
Evaluation considers both phoneme error rate (PER) and the structure of the learned phoneme embedding space using cosine-distance, nearest-neighbor, and silhouette-based analyses. To the best of our knowledge, this is the first study to investigate bilingual phoneme-level contrastive learning for personalized dysarthric ASR.
Results show that phoneme-level CL improves recognition over bilingual fine-tuning using only Connectionist Temporal Classification (CTC), but explicitly prioritizing cross-lingual phoneme relationships does not provide a statistically significant improvement over the simpler same-symbol contrastive learning. At the representation level, bilingual and weighted objectives can bring cross-lingual phoneme representations closer together and produce better-separated phoneme clusters, but these changes do not consistently result in lower PER. Replacing dysarthric speech with typical speech significantly worsens recognition when the amount of training data is kept approximately constant, whereas adding typical speech on top of the full dysarthric training set significantly improves recognition; however, the latter condition also contains more training data, so the improvement cannot be attributed specifically to typical speech. Overall, the results show a benefit from phoneme-level contrastive learning for personalized dysarthric ASR, but no clear additional benefit from explicitly encoding canonical Dutch--English phoneme relationships. This suggests that the usefulness of bilingual contrastive learning may depend on defining cross-lingual relationships that better reflect the target speaker's actual speech patterns. ...
Dysarthric speech remains challenging for state-of-the-art automatic speech recognition (ASR) systems, despite their high accuracy on typical speech. Personalized dysarthric ASR offers a way to address the substantial differences in speech characteristics between individuals, but it is constrained by the limited amount of speaker-specific dysarthric data available for training. Speech produced by the same target speaker in another language provides an additional source of labeled data while preserving speaker-specific and dysarthric speech characteristics.
This study investigates bilingual phoneme-level contrastive learning (CL) for personalized dysarthric ASR in a case study involving a single speaker with severe dysarthria producing speech in Dutch and English. Phoneme-level modeling is used as phonemes provide a natural unit for identifying relationships between speech sounds across languages. English and Dutch speech are jointly modeled by constructing contrastive pairs from phonemes that are either shared across the two languages or manually identified as phonetically equivalent. Because these explicitly cross-lingual relationships constitute only a subset of the possible positive pairs, three positive sampling strategies and corresponding loss-weighting variants are evaluated to investigate whether giving them greater influence during training improves recognition or more strongly shapes the learned representation space. Two approaches for incorporating additional typical speech are also evaluated to investigate the effect of training-data composition.
Evaluation considers both phoneme error rate (PER) and the structure of the learned phoneme embedding space using cosine-distance, nearest-neighbor, and silhouette-based analyses. To the best of our knowledge, this is the first study to investigate bilingual phoneme-level contrastive learning for personalized dysarthric ASR.
Results show that phoneme-level CL improves recognition over bilingual fine-tuning using only Connectionist Temporal Classification (CTC), but explicitly prioritizing cross-lingual phoneme relationships does not provide a statistically significant improvement over the simpler same-symbol contrastive learning. At the representation level, bilingual and weighted objectives can bring cross-lingual phoneme representations closer together and produce better-separated phoneme clusters, but these changes do not consistently result in lower PER. Replacing dysarthric speech with typical speech significantly worsens recognition when the amount of training data is kept approximately constant, whereas adding typical speech on top of the full dysarthric training set significantly improves recognition; however, the latter condition also contains more training data, so the improvement cannot be attributed specifically to typical speech. Overall, the results show a benefit from phoneme-level contrastive learning for personalized dysarthric ASR, but no clear additional benefit from explicitly encoding canonical Dutch--English phoneme relationships. This suggests that the usefulness of bilingual contrastive learning may depend on defining cross-lingual relationships that better reflect the target speaker's actual speech patterns.
This study investigates bilingual phoneme-level contrastive learning (CL) for personalized dysarthric ASR in a case study involving a single speaker with severe dysarthria producing speech in Dutch and English. Phoneme-level modeling is used as phonemes provide a natural unit for identifying relationships between speech sounds across languages. English and Dutch speech are jointly modeled by constructing contrastive pairs from phonemes that are either shared across the two languages or manually identified as phonetically equivalent. Because these explicitly cross-lingual relationships constitute only a subset of the possible positive pairs, three positive sampling strategies and corresponding loss-weighting variants are evaluated to investigate whether giving them greater influence during training improves recognition or more strongly shapes the learned representation space. Two approaches for incorporating additional typical speech are also evaluated to investigate the effect of training-data composition.
Evaluation considers both phoneme error rate (PER) and the structure of the learned phoneme embedding space using cosine-distance, nearest-neighbor, and silhouette-based analyses. To the best of our knowledge, this is the first study to investigate bilingual phoneme-level contrastive learning for personalized dysarthric ASR.
Results show that phoneme-level CL improves recognition over bilingual fine-tuning using only Connectionist Temporal Classification (CTC), but explicitly prioritizing cross-lingual phoneme relationships does not provide a statistically significant improvement over the simpler same-symbol contrastive learning. At the representation level, bilingual and weighted objectives can bring cross-lingual phoneme representations closer together and produce better-separated phoneme clusters, but these changes do not consistently result in lower PER. Replacing dysarthric speech with typical speech significantly worsens recognition when the amount of training data is kept approximately constant, whereas adding typical speech on top of the full dysarthric training set significantly improves recognition; however, the latter condition also contains more training data, so the improvement cannot be attributed specifically to typical speech. Overall, the results show a benefit from phoneme-level contrastive learning for personalized dysarthric ASR, but no clear additional benefit from explicitly encoding canonical Dutch--English phoneme relationships. This suggests that the usefulness of bilingual contrastive learning may depend on defining cross-lingual relationships that better reflect the target speaker's actual speech patterns.