IK

I. Kondyurin

info

Please Note

8 records found

How has the portrayal of female characters in fanfiction evolved in response to the #MeToo movement and fourth-wave feminism, as analyzed with the help of NLP techniques?

Bachelor thesis (2025) - I. Marinescu, H.S. Hung, E. Eisemann, C. Hao, I. Kondyurin
This paper explores how the portrayal of female characters in fanfiction evolved in response to the #MeToo movement and fourth-wave feminism, with the aim of assessing whether the impact of the awareness of the campaign was broad enough to visibly alter how the average author portrays women in narrative contexts. To analyze these trends, fanfiction data from Archive of Our Own (AO3) spanning 2015–2019 was parsed, and two Natural Language Processing (NLP) pipelines — Word2Vec and GloVe, and BERT — were developed. The study finds that bias scores, aggregated through formulas created to compare gendered associations, show a stronger stereotypization of women before 2017 compared to after. Furthermore, a similar trend is discovered in the representation of women in fanfiction. While the BERT pipeline proved most effective for capturing contextual nuances, it is significantly limited by its reliance on binary labels and computational intensity. This further indicates the need for more inclusive and sustainable methods, making the Word2Vec/GloVe models more appropriate for this task. The paper concludes with recommendations for future work, including broader representation, longer-term analysis, and enhanced detection of evolving language patterns. ...

Fine Tuning a BERT-based Pre-Trained Language Model for Named Entity Extraction within the Domain of Fanfiction

Bachelor thesis (2025) - N.P.A. Kindt, H.S. Hung, C. Hao, I. Kondyurin, E. Eisemann
The introduction of Pretrained Language Models (PLMs) has revolutionised the field of Natural Language Processing (NLP) and paved the way for many new, exciting large-scale studies for various areas of research. One such field presents itself in the emerging digital literary corpus that is fanfiction, providing research opportunities within the fields of (NLP), Computational (Socio-) Linguistics, the Social Sciences and Digital Humanities. However, because of the unique linguistic characteristics of this literary domain many modern NLP solutions utilizing PLMs encounter difficulties when applied on fanfiction texts. This paper aims to indicate that the performance of various NLP tasks performed by PLMs on fanfiction texts can be improved by applying Domain Adaptive Pre-Training (DAPT) to PLMs. A case-study is performed to show that the performance of a BERT-based PLM can be improved for the downstream NLP task of Named Entity Recognition (NER) by applying supervised domain specific fine-tuning. While we gain a 6% increase in F1 score performance, we are sceptical about these results due to the limited amount of annotated data available leading to the model overfitting and show a lack of capacity to generalize to unseen data from the CoNLL NER dataset. ...

A computational analysis of linear correlations between emotional behavior and popularity

Fanfiction writers always look for ways to make their stories more engaging. Analyzing what influences the popularity of fanfiction provides insights into readers' preferences and allows writers to tailor to these. This paper attempts to find linear correlations between fanfiction stories and the emotional journey of their characters. It does so by computationally extracting these journeys from 319 Good Omens fanfiction stories, defining and extracting several features from them and using simple linear regression to determine their correlation to fanfiction popularity. Five features were found to have a significant influence on fanfiction popularity. It was also determined that readers prefer characters that have low emotional fluctuations in their behavior. ...

How does fan-fiction differ in style to its original canon and does it affect its success?

Bachelor thesis (2025) - R.C. Lambert, H.S. Hung, I. Kondyurin, E. Eisemann
Natural language processing, specifically style, is not explored significantly in the context of fan-fiction. By using function word frequency analysis, this paper explores the similarity in style between original works and fan-fictions derived from them as well as the impact of those stylistic similarities on the fan-fictions' successes. Investigating the works of Worm and Narnia and utilising a control set, this hypothesis is examined. Similarity in style seems to be variable, but it is suggested that certain stylistic features lead to success more than other regardless of the original work's style. More research would be needed to confirm those numbers. ...
Bachelor thesis (2025) - J.Q.Q. Ye, E. Eisemann, H.S. Hung, C. Hao, I. Kondyurin
This study investigates how genre preferences and sentiment influence fanfiction popularity across multiple languages, focusing on English, Mandarin, Russian, and Spanish datasets. Leveraging advanced natural language processing techniques, including multilingual sentiment analysis, genre classification, and topic modeling, this research explores the interplay between cultural and linguistic factors in storytelling. Preprocessing steps, such as translation and named entity recognition, ensured consistency and reduced noise across the multilingual dataset. Key findings reveal cross-linguistic patterns, such as the popularity of genres like Alter- nate Universe and Romance, alongside cultural distinctions in sentiment and engagement. This work contributes to computational fan studies by demonstrating how linguistic and cultural factors influence storytelling trends and audience preferences in fanfiction. ...

Understanding the meaning of gestures in densely crowded social settings

Recent studies have shown that gesture annotation schemes should account for the multidimensional nature of gestures and define their meaning in terms of referentiality and pragmatic meaning. However, accurately annotating gesture meaning in densely crowded social settings using such a coding scheme remains to be accomplished. This study uses the MultiModal MultiDimensional (M3D) labelling scheme and the EUDICO Linguistic Annotator (ELAN) tool to annotate video data from the Conference Living Lab (ConfLab) dataset. The ConfLab dataset contains 8 video recordings of standing conversations at a conference, captured from an overhead perspective, and low-frequency audio recordings of the conversations. A total of 1119 clips of individual gesture instances are generated. This data is then fed into a VideoMAE model pre-trained on the UCF101 dataset. The model achieves an overall accuracy score of 49% on the test set but shows a significant bias towards one class due to the imbalanced dataset. Due to the small size of the dataset and the similarities between gestures with different meanings, the model cannot identify different gesture types. The results demonstrate that high-frequency audio or transcripts of the conversations are vital to avoid strong and potentially incorrect assumptions when annotating gesture meaning. Further investigation is required into the annotation and classification of pragmatic meanings and Machine Learning solutions for multi-class, multi-label video classification problems. ...

Employing gesture coding schemes and machine learning to predict physical features of hand gestures in video footage from a crowded social setting

Bachelor thesis (2024) - F.J. Latała, I. Kondyurin, Z. Li, H.S. Hung, M.A. Neerincx
Researching hand gestures in real-world social interactions requires very careful analysis. While gesture coding schemes were created with that purpose in mind, they are not widely utilised in research. Moreover, studies on gesture classification rarely focus on the physical nature of movements involved in gesturing, despite the fact that being able to quantify the motion could reveal useful patterns and correlations. To address those points, this research proposes the following approach: using machine learning models to automatically classify physical features of hand gestures, according to a coding scheme. Two such classifiers were created, for the left and right hand respectively. Overall, the results are quite promising - despite a small and imbalanced training set and complex features both models achieved an accuracy of roughly 60%. Moreover, the results indicate that by avoiding some of the simplifications that this research makes, and by using more balanced training data, the accuracy could be significantly increased. This is concrete evidence that machine learning models can indeed be used to classify the physical aspects of hand gestures, as defined
by a coding scheme, in social interactions in the wild.
...

Classification of gesture phases in a crowded social setting recorded from top-view angle

Bachelor thesis (2024) - A. Grigore, H.S. Hung, I. Kondyurin, Z. Li, M.A. Neerincx
Hand gestures play a crucial role in communication, especially in social interactions. This research investigates the viability of using coding schemes to describe hand gestures and how accurately they can be classified in crowded environments by using fine-tuned visual transformers such as VideoMAE. The dataset used during training is based on the Conflab dataset and contains top-view video recordings of social interactions in a crowded social setting. The videos are manually annotated for gesture phases (preparation, hold, stroke, recovery) and gesture units. The two classifiers obtain high accuracies after fine tuning, with an overall accuracy of 95% for the gesture phase classification and 93% for classifying whether a clip is a gesture unit or not. These findings suggest that the proposed approach is effective in crowded environments and can be adapted for real-time applications. ...