EK

Evgeny Krivosheev

info

Please Note

3 records found

Journal article (2021) - Burcu Sayin, Evgeny Krivosheev, Jie Yang, Andrea Passerini, Fabio Casati
Training data creation is increasingly a key bottleneck for developing machine learning, especially for deep learning systems. Active learning provides a cost-effective means for creating training data by selecting the most informative instances for labeling. Labels in real applications are often collected from crowdsourcing, which engages online crowds for data labeling at scale. Despite the importance of using crowdsourced data in the active learning process, an analysis of how the existing active learning approaches behave over crowdsourced data is currently missing. This paper aims to fill this gap by reviewing the existing active learning approaches and then testing a set of benchmarking ones on crowdsourced datasets. We provide a comprehensive and systematic survey of the recent research on active learning in the hybrid human–machine classification setting, where crowd workers contribute labels (often noisy) to either directly classify data instances or to train machine learning models. We identify three categories of state of the art active learning methods according to whether and how predefined queries employed for data sampling, namely fixed-strategy approaches, dynamic-strategy approaches, and strategy-free approaches. We then conduct an empirical study on their cost-effectiveness, showing that the performance of the existing active learning approaches is affected by many factors in hybrid classification contexts, such as the noise level of data, label fusion technique used, and the specific characteristics of the task. Finally, we discuss challenges and identify potential directions to design active learning strategies for hybrid classification problems. ...
Conference paper (2021) - Burcu Sayin, Evgeny Krivosheev, Jorge Ramirez, Fabio Casati, Ekaterina Taran, Veronika Malanina, Jie Yang
Hybrid classification services are online services that combine machine learning (ML) and humans - either crowd workers or experts - to achieve a classification objective, from relatively simple ones such as deriving the sentiment of a text to more complex ones such as medical diagnoses. This paper takes the first steps toward a science for hybrid classification services, discussing key concepts, challenges, and architectures, and then focusing on a central aspect, that of ML calibration and how it can be achieved with crowdsourced labels. ...
Journal article (2020) - Evgeny Krivosheev, Burcu Sayin, Alessandro Bozzon, Zoltán Szlávik
In this paper, we explore how to efficiently combine crowdsourcing and machine intelligence for the problem of document screening, where we need to screen documents with a set of machine-learning filters. Specifically, we focus on building a set of machine learning classifiers that evaluate documents, and then screen them efficiently. It is a challenging task since the budget is limited and there are countless number of ways to spend the given budget on the problem. We propose a multi-label active learning screening specific sampling technique -objective-aware samplingfor querying unlabelled documents for annotating. Our algorithm takes a decision on which machine filter need more training data and how to choose unlabeled items to annotate in order to minimize the risk of overall classification errors rather than minimizing a single filter error. We demonstrate that objective-aware sampling significantly outperforms the state of the art active learning sampling strategies. ...