H.S. Hung
Please Note
79 records found
1
Hold Your Cup
Evaluating Smart Cup IMU Cues for Movement Events in Conversation
The results show that smart cup motion contains useful movement information, but reliable temporal localization remains difficult under natural interaction conditions. Window based baselines reveal local class-discriminative information but are limited by fixed boundaries and severe class imbalance. Supervised temporal action localization provides the most stable sample-level performance, whereas few-shot LLM prompting can sometimes identify useful candidate intervals but remains sensitive to reference construction and recording variation.
The findings support a cautious role for smart cup IMU data as an auxiliary cue for locating movement-related intervals in long multimodal recordings, rather than as independent evidence of conversational meaning. Future work should refine the behaviour category system, expand annotations, and integrate cup motion with video, speech, body keypoints, and interaction context. ...
The results show that smart cup motion contains useful movement information, but reliable temporal localization remains difficult under natural interaction conditions. Window based baselines reveal local class-discriminative information but are limited by fixed boundaries and severe class imbalance. Supervised temporal action localization provides the most stable sample-level performance, whereas few-shot LLM prompting can sometimes identify useful candidate intervals but remains sensitive to reference construction and recording variation.
The findings support a cautious role for smart cup IMU data as an auxiliary cue for locating movement-related intervals in long multimodal recordings, rather than as independent evidence of conversational meaning. Future work should refine the behaviour category system, expand annotations, and integrate cup motion with video, speech, body keypoints, and interaction context.
A Gaussian Mixture Model fit on four room-level features (adult word count, auditory overlap, displacement, and teacher distance) and selected by a stability-aware rule (lowest mean BIC among K values whose cluster assignments reproduce across random initialisations, mean pairwise Adjusted Rand Index >= 0.80) over 2 to 20 components recovers six latent activity contexts, each with a distinct sensor profile. Each context is then given a post-hoc descriptive label drawn from activity types familiar in inclusive preschool classrooms: dispersed transition, peer-driven activity, independent / parallel work, adult-scaffolded peer activity, seated guided work, and whole-class instruction / read-aloud. These labels describe the recovered clusters and are not validated against an external ground truth.
Within each context, a linear mixed model with a per-child random intercept compares HL and TH children on three sensor-derivable behavioural markers of inclusion: peer co-presence (time in spatial groups, where a spatial group is a set of children simultaneously co-present and mutually body-oriented; §5.2.3), vocal participation rate (utterances per minute), and peer affiliation patterns (time in same-diagnosis vs. mixed-diagnosis spatial groups).
The HL/TH signal distributes differently across contexts for each marker. Peer co-presence shows no substantial HL/TH difference in three of six contexts. HL children exceed TH children in independent / parallel work (+7.8% of the minute, q < 0.0001, d = +1.59) and fall below TH children in peer-driven activity (-4.3% of the minute, q = 0.02) and seated guided work (-5.3% of the minute, q = 0.002). Vocal participation rate shows an HL deficit concentrated in the contexts that combine high peer communicative demand with reduced adult scaffolding (peer-driven activity q = 0.03 and adult-scaffolded peer activity q = 0.05, both significant after false-discovery-rate correction; dispersed transition in the same direction but only marginal, q = 0.07), but not in the three contexts where high adult word count or very low auditory overlap slows the pace of exchange. Peer affiliation is the most consistent asymmetry: TH children concentrate grouped time in TH-only spatial groups more than HL children concentrate theirs in HL-only spatial groups in five of six contexts (significantly in four), and HL children spend more grouped time in mixed spatial groups than TH children do across all six (significantly in four). The 6:7 cohort composition produces a baseline TH-above-HL gap of approximately 0.08 on the homophily index under random affiliation; the four significant clusters exceed this baseline (observed gap 0.12-0.19, of which 0.04-0.11 is preference beyond availability), while the two non-significant clusters are within the magnitude expected from cohort composition alone. ...
A Gaussian Mixture Model fit on four room-level features (adult word count, auditory overlap, displacement, and teacher distance) and selected by a stability-aware rule (lowest mean BIC among K values whose cluster assignments reproduce across random initialisations, mean pairwise Adjusted Rand Index >= 0.80) over 2 to 20 components recovers six latent activity contexts, each with a distinct sensor profile. Each context is then given a post-hoc descriptive label drawn from activity types familiar in inclusive preschool classrooms: dispersed transition, peer-driven activity, independent / parallel work, adult-scaffolded peer activity, seated guided work, and whole-class instruction / read-aloud. These labels describe the recovered clusters and are not validated against an external ground truth.
Within each context, a linear mixed model with a per-child random intercept compares HL and TH children on three sensor-derivable behavioural markers of inclusion: peer co-presence (time in spatial groups, where a spatial group is a set of children simultaneously co-present and mutually body-oriented; §5.2.3), vocal participation rate (utterances per minute), and peer affiliation patterns (time in same-diagnosis vs. mixed-diagnosis spatial groups).
The HL/TH signal distributes differently across contexts for each marker. Peer co-presence shows no substantial HL/TH difference in three of six contexts. HL children exceed TH children in independent / parallel work (+7.8% of the minute, q < 0.0001, d = +1.59) and fall below TH children in peer-driven activity (-4.3% of the minute, q = 0.02) and seated guided work (-5.3% of the minute, q = 0.002). Vocal participation rate shows an HL deficit concentrated in the contexts that combine high peer communicative demand with reduced adult scaffolding (peer-driven activity q = 0.03 and adult-scaffolded peer activity q = 0.05, both significant after false-discovery-rate correction; dispersed transition in the same direction but only marginal, q = 0.07), but not in the three contexts where high adult word count or very low auditory overlap slows the pace of exchange. Peer affiliation is the most consistent asymmetry: TH children concentrate grouped time in TH-only spatial groups more than HL children concentrate theirs in HL-only spatial groups in five of six contexts (significantly in four), and HL children spend more grouped time in mixed spatial groups than TH children do across all six (significantly in four). The 6:7 cohort composition produces a baseline TH-above-HL gap of approximately 0.08 on the homophily index under random affiliation; the four significant clusters exceed this baseline (observed gap 0.12-0.19, of which 0.04-0.11 is preference beyond availability), while the two non-significant clusters are within the magnitude expected from cohort composition alone.
Analysis of Results in the ML Research Field
How well can an LLM decide the reproducibility of a paper?
Analysis of results in the ML research field
Investigating the Efficacy of LLMs in Extracting Stated Research Limitations
fication and evaluation of responsible research checklists impose a significant burden on
reviewers. This study investigates the ability of Large Language Models (LLMs) to au-
tomatically classify research papers as empirical, theoretical, or hybrid, and to extract
checklist compliance data. Using a dataset of publicly available NeurIPS papers, we
designed an automated pipeline and evaluated its outputs against a human-annotated
ground truth. Our results demonstrate that the LLM achieves high accuracy in the
core classification task, reliably distinguishing the papers core methodology by iden-
tifying clear structural indicators like mathematical proofs and benchmark datasets.
Furthermore, the model excels at extracting objective checklist elements, performing
well on close-ended extraction tasks that rely on clear structural indicators. However,
performance noticeably decreased on structurally scattered or subjective criteria, such
as broader impacts and the declaration of AI usage. This drop highlights a limitation in
the model’s broader reading comprehension, as it struggles to merge contextual infor-
mation without explicit headers. Notably, this automated failure closely mirrors human
task ambiguity, as these exact subjective items also generated the lower inter-annotator
agreement among human annotators. Conclusively, while LLMs provide a highly con-
sistent baseline for classifying paper typologies and extracting explicit methodological
data, their reliance on structural cues indicates they should serve as assistive screening
tools rather than autonomous evaluators in academic peer review. ...
fication and evaluation of responsible research checklists impose a significant burden on
reviewers. This study investigates the ability of Large Language Models (LLMs) to au-
tomatically classify research papers as empirical, theoretical, or hybrid, and to extract
checklist compliance data. Using a dataset of publicly available NeurIPS papers, we
designed an automated pipeline and evaluated its outputs against a human-annotated
ground truth. Our results demonstrate that the LLM achieves high accuracy in the
core classification task, reliably distinguishing the papers core methodology by iden-
tifying clear structural indicators like mathematical proofs and benchmark datasets.
Furthermore, the model excels at extracting objective checklist elements, performing
well on close-ended extraction tasks that rely on clear structural indicators. However,
performance noticeably decreased on structurally scattered or subjective criteria, such
as broader impacts and the declaration of AI usage. This drop highlights a limitation in
the model’s broader reading comprehension, as it struggles to merge contextual infor-
mation without explicit headers. Notably, this automated failure closely mirrors human
task ambiguity, as these exact subjective items also generated the lower inter-annotator
agreement among human annotators. Conclusively, while LLMs provide a highly con-
sistent baseline for classifying paper typologies and extracting explicit methodological
data, their reliance on structural cues indicates they should serve as assistive screening
tools rather than autonomous evaluators in academic peer review.
Large Language Models for Reviewing Research Papers
Evaluating Claim-Level Completeness in Machine Learning Research
Investigating Narratives of Social Intention in Restaurant Interactions
Researching scenarios for Intention Prediction
This thesis presents the design, implementation, and evaluation of a low-cost, modular smart carpet for human behaviour tracking. The proposed system uses a resistive matrix sensing approach based on Velostat and copper electrodes, combined with a tile-based architecture that enables flexible deployment over larger areas. A custom-designed printed circuit board centralises signal routing and data acquisition, reducing wiring complexity while avoiding the use of active electronics within individual tiles. To support scalability, the system separates real-time data acquisition from offline signal processing, where cross-talk mitigation, filtering, and visualisation are performed.
The system is evaluated through a series of qualitative and quantitative experiments that examine its ability to capture pressure distributions, detect multiple contact regions, and represent dynamic interactions such as weight shifting and walking. The results show that the system reliably captures coarse spatial interaction patterns, but that limitations in spatial resolution and material behaviour reduce its effectiveness for distinguishing subtle pressure differences. These findings indicate that, while the system provides useful information for activity and behaviour analysis in shared spaces, its sensing capabilities are less expressive than initially anticipated for fine-grained interpretation.
Overall, this work explores a design point that prioritises affordability, modularity, and privacy, and demonstrates how these priorities shape the trade-offs observed in large-area, pressure-based human sensing systems. ...
This thesis presents the design, implementation, and evaluation of a low-cost, modular smart carpet for human behaviour tracking. The proposed system uses a resistive matrix sensing approach based on Velostat and copper electrodes, combined with a tile-based architecture that enables flexible deployment over larger areas. A custom-designed printed circuit board centralises signal routing and data acquisition, reducing wiring complexity while avoiding the use of active electronics within individual tiles. To support scalability, the system separates real-time data acquisition from offline signal processing, where cross-talk mitigation, filtering, and visualisation are performed.
The system is evaluated through a series of qualitative and quantitative experiments that examine its ability to capture pressure distributions, detect multiple contact regions, and represent dynamic interactions such as weight shifting and walking. The results show that the system reliably captures coarse spatial interaction patterns, but that limitations in spatial resolution and material behaviour reduce its effectiveness for distinguishing subtle pressure differences. These findings indicate that, while the system provides useful information for activity and behaviour analysis in shared spaces, its sensing capabilities are less expressive than initially anticipated for fine-grained interpretation.
Overall, this work explores a design point that prioritises affordability, modularity, and privacy, and demonstrates how these priorities shape the trade-offs observed in large-area, pressure-based human sensing systems.
Laughter Accelerometer-Based Detection in Natural Social Interactions
Investigating segmentation and inter-modality annotation strategies for wearable laughter detection
Prediction-based Anomaly Detection in Multivariate Time-Series Data
Improving Wahoo Fitness Cycling Data Quality by Addressing Sensor Errors
Exploring Genre Preferences and Audience Engagement in Multilingual Fanfiction
A Study of Popularity and Preferences
What if fan-fiction, but also coding
How does fan-fiction differ in style to its original canon and does it affect its success?
What if fanfiction, but also coding: Investigating cultural differences in fanfiction writing and reviewing with machine learning methods
Fine Tuning a BERT-based Pre-Trained Language Model for Named Entity Extraction within the Domain of Fanfiction
What if fanfiction, but also coding: Investigating cultural differences in fanfiction writing and reviewing with machine learning methods
How has the portrayal of female characters in fanfiction evolved in response to the #MeToo movement and fourth-wave feminism, as analyzed with the help of NLP techniques?
The impact of emotional journeys on fanfiction popularity
A computational analysis of linear correlations between emotional behavior and popularity
Laughter in Motion: Pose-Based Detection Across Annotation Modalities in Natural Social Interactions
Investigating modality annotation impact for detecting laughter in the wild