S. Tan
Please Note
30 records found
1
This thesis asks whether a Transformer can be trained using only the forward-pass operations a photonic substrate provides, and characterizes what such training costs. The proposed method combines a layer-wise Forward-Forward prototype-based objective, directional-derivative gradient estimation (with no backward pass and no automatic differentiation), and a softmax-free Spherical attention adapted from the Kramers-Kronig kernel, with the four attention projections trained one at a time in a round-robin schedule; six training variants are compared across seven vision and sequence tasks to isolate the gradient estimator and the update schedule. The answer is affirmative: the fully forward-only variant trains a Transformer to a useful operating point on six of the seven tasks. Locality, not the forward-only gradient alone, is what makes this possible, and the depth-resilience of local learning is shown to extend, conditionally, to self-attention. The remaining gap to backpropagation has three sources, two inherent to forward-only learning (local credit assignment and gradient-estimation variance) and one architectural (the softmax-free attention cannot form the sharp selection that content-addressed retrieval requires). Because every operation reduces to a forward pass and a measured scalar loss, the method is a candidate for in-situ training on a photonic chip, the validation step this work points to. ...
This thesis asks whether a Transformer can be trained using only the forward-pass operations a photonic substrate provides, and characterizes what such training costs. The proposed method combines a layer-wise Forward-Forward prototype-based objective, directional-derivative gradient estimation (with no backward pass and no automatic differentiation), and a softmax-free Spherical attention adapted from the Kramers-Kronig kernel, with the four attention projections trained one at a time in a round-robin schedule; six training variants are compared across seven vision and sequence tasks to isolate the gradient estimator and the update schedule. The answer is affirmative: the fully forward-only variant trains a Transformer to a useful operating point on six of the seven tasks. Locality, not the forward-only gradient alone, is what makes this possible, and the depth-resilience of local learning is shown to extend, conditionally, to self-attention. The remaining gap to backpropagation has three sources, two inherent to forward-only learning (local credit assignment and gradient-estimation variance) and one architectural (the softmax-free attention cannot form the sharp selection that content-addressed retrieval requires). Because every operation reduces to a forward pass and a measured scalar loss, the method is a candidate for in-situ training on a photonic chip, the validation step this work points to.
A Gaussian Mixture Model fit on four room-level features (adult word count, auditory overlap, displacement, and teacher distance) and selected by a stability-aware rule (lowest mean BIC among K values whose cluster assignments reproduce across random initialisations, mean pairwise Adjusted Rand Index >= 0.80) over 2 to 20 components recovers six latent activity contexts, each with a distinct sensor profile. Each context is then given a post-hoc descriptive label drawn from activity types familiar in inclusive preschool classrooms: dispersed transition, peer-driven activity, independent / parallel work, adult-scaffolded peer activity, seated guided work, and whole-class instruction / read-aloud. These labels describe the recovered clusters and are not validated against an external ground truth.
Within each context, a linear mixed model with a per-child random intercept compares HL and TH children on three sensor-derivable behavioural markers of inclusion: peer co-presence (time in spatial groups, where a spatial group is a set of children simultaneously co-present and mutually body-oriented; §5.2.3), vocal participation rate (utterances per minute), and peer affiliation patterns (time in same-diagnosis vs. mixed-diagnosis spatial groups).
The HL/TH signal distributes differently across contexts for each marker. Peer co-presence shows no substantial HL/TH difference in three of six contexts. HL children exceed TH children in independent / parallel work (+7.8% of the minute, q < 0.0001, d = +1.59) and fall below TH children in peer-driven activity (-4.3% of the minute, q = 0.02) and seated guided work (-5.3% of the minute, q = 0.002). Vocal participation rate shows an HL deficit concentrated in the contexts that combine high peer communicative demand with reduced adult scaffolding (peer-driven activity q = 0.03 and adult-scaffolded peer activity q = 0.05, both significant after false-discovery-rate correction; dispersed transition in the same direction but only marginal, q = 0.07), but not in the three contexts where high adult word count or very low auditory overlap slows the pace of exchange. Peer affiliation is the most consistent asymmetry: TH children concentrate grouped time in TH-only spatial groups more than HL children concentrate theirs in HL-only spatial groups in five of six contexts (significantly in four), and HL children spend more grouped time in mixed spatial groups than TH children do across all six (significantly in four). The 6:7 cohort composition produces a baseline TH-above-HL gap of approximately 0.08 on the homophily index under random affiliation; the four significant clusters exceed this baseline (observed gap 0.12-0.19, of which 0.04-0.11 is preference beyond availability), while the two non-significant clusters are within the magnitude expected from cohort composition alone. ...
A Gaussian Mixture Model fit on four room-level features (adult word count, auditory overlap, displacement, and teacher distance) and selected by a stability-aware rule (lowest mean BIC among K values whose cluster assignments reproduce across random initialisations, mean pairwise Adjusted Rand Index >= 0.80) over 2 to 20 components recovers six latent activity contexts, each with a distinct sensor profile. Each context is then given a post-hoc descriptive label drawn from activity types familiar in inclusive preschool classrooms: dispersed transition, peer-driven activity, independent / parallel work, adult-scaffolded peer activity, seated guided work, and whole-class instruction / read-aloud. These labels describe the recovered clusters and are not validated against an external ground truth.
Within each context, a linear mixed model with a per-child random intercept compares HL and TH children on three sensor-derivable behavioural markers of inclusion: peer co-presence (time in spatial groups, where a spatial group is a set of children simultaneously co-present and mutually body-oriented; §5.2.3), vocal participation rate (utterances per minute), and peer affiliation patterns (time in same-diagnosis vs. mixed-diagnosis spatial groups).
The HL/TH signal distributes differently across contexts for each marker. Peer co-presence shows no substantial HL/TH difference in three of six contexts. HL children exceed TH children in independent / parallel work (+7.8% of the minute, q < 0.0001, d = +1.59) and fall below TH children in peer-driven activity (-4.3% of the minute, q = 0.02) and seated guided work (-5.3% of the minute, q = 0.002). Vocal participation rate shows an HL deficit concentrated in the contexts that combine high peer communicative demand with reduced adult scaffolding (peer-driven activity q = 0.03 and adult-scaffolded peer activity q = 0.05, both significant after false-discovery-rate correction; dispersed transition in the same direction but only marginal, q = 0.07), but not in the three contexts where high adult word count or very low auditory overlap slows the pace of exchange. Peer affiliation is the most consistent asymmetry: TH children concentrate grouped time in TH-only spatial groups more than HL children concentrate theirs in HL-only spatial groups in five of six contexts (significantly in four), and HL children spend more grouped time in mixed spatial groups than TH children do across all six (significantly in four). The 6:7 cohort composition produces a baseline TH-above-HL gap of approximately 0.08 on the homophily index under random affiliation; the four significant clusters exceed this baseline (observed gap 0.12-0.19, of which 0.04-0.11 is preference beyond availability), while the two non-significant clusters are within the magnitude expected from cohort composition alone.
Our evaluation, utilizing geometric metrics such as Centered Kernel Alignment (CKA) and Principal Component Analysis (PCA), showcases how PEPITA exhibits higher similarity to BP, but is more geometrically aligned with the unsupervised SoftHebb. Furthermore, empirical hardware profiling exposes a significant implementation paradox: despite the theoretical efficiency of BP-free methods, high-level framework bottlenecks currently make algorithms like PEPITA computationally expensive on traditional digital architectures. ...
Our evaluation, utilizing geometric metrics such as Centered Kernel Alignment (CKA) and Principal Component Analysis (PCA), showcases how PEPITA exhibits higher similarity to BP, but is more geometrically aligned with the unsupervised SoftHebb. Furthermore, empirical hardware profiling exposes a significant implementation paradox: despite the theoretical efficiency of BP-free methods, high-level framework bottlenecks currently make algorithms like PEPITA computationally expensive on traditional digital architectures.
Automating Key-Point Analysis with LLMs
Argument-to-Key-Point Mapping
Decoding Legislative Discourse: Transformer-Based Topic Modeling of U.S. Congressional Hearings
A Comparative Analysis of Standard and Zero-Shot BERTopic
Laughter in Motion: Pose-Based Detection Across Annotation Modalities in Natural Social Interactions
Investigating modality annotation impact for detecting laughter in the wild
Evaluating modern computer vision techniques for Shape Language classification in meetings
Automatic understanding of meetings and negotiations
Breaking down negotiations
Analyzing negotiations using the Coloured Trails Game & the NegotiAct
In the experiment a total of 15 participants took part and it was run a total of 20 times, resulting in 3 hours and 10 minutes of recordings. Encoding them with the NegotiAct resulted in the discerning of a total of 87 offers made, 24 offers accepted, 26 offers rejected, and 16 requests for offer modifications. Based on this coherent mapping it can be concluded that the Colored Trails Game is a suitable choice for the workflow of gathering data to be put through the analysis of the NegotiAct.
...
In the experiment a total of 15 participants took part and it was run a total of 20 times, resulting in 3 hours and 10 minutes of recordings. Encoding them with the NegotiAct resulted in the discerning of a total of 87 offers made, 24 offers accepted, 26 offers rejected, and 16 requests for offer modifications. Based on this coherent mapping it can be concluded that the Colored Trails Game is a suitable choice for the workflow of gathering data to be put through the analysis of the NegotiAct.
Gesture Recognition for Enhanced Meeting Analysis
Segmenting and Tracking Hand Movements During Human Interaction
Identifying Speaking and Drinking Events Within Audio Recordings for Multiactivity Analysis
Rethinking Ubiquitous Smart Sensing of Social Behaviour in the Wild
Personalized Gesture Range Detection Using Transductive Parameter Transfer
Rethinking Ubiquitous Smart Sensing of Social Behaviour In The Wild