Detecting and analysing spontaneous oral cancer speech in the wild

Conference Paper (2020)
Author(s)

Bence Mark Halpern (Nederlands Kanker Instituut - Antoni van Leeuwenhoek ziekenhuis, Universiteit van Amsterdam, TU Delft - Electrical Engineering, Mathematics and Computer Science)

Rob van Son (Universiteit van Amsterdam, Nederlands Kanker Instituut - Antoni van Leeuwenhoek ziekenhuis)

Michiel W.M. van den Brekel (Universiteit van Amsterdam, Nederlands Kanker Instituut - Antoni van Leeuwenhoek ziekenhuis)

Odette Scharenborg (TU Delft - Electrical Engineering, Mathematics and Computer Science)

Research Group
Multimedia Computing
DOI related publication
https://doi.org/10.21437/Interspeech.2020-1598 Final published version
More Info
expand_more
Publication Year
2020
Language
English
Research Group
Multimedia Computing
Pages (from-to)
4826 - 4830
Event
INTERSPEECH 2020 (2020-10-25 - 2020-10-29), Shanghai, China
Page Views
259

Abstract

Oral cancer speech is a disease which impacts more than half a million people worldwide every year. Analysis of oral cancer speech has so far focused on read speech. In this paper, we 1) present and 2) analyse a three-hour long spontaneous oral cancer speech dataset collected from YouTube. 3) We set baselines for an oral cancer speech detection task on this dataset. The analysis of these explainable machine learning baselines shows that sibilants and stop consonants are the most important indicators for spontaneous oral cancer speech detection.