ST

S. Tan

info

Please Note

30 records found

Master thesis (2026) - D. Krylov, S. Tan, D.M.J. Tax, R.A. Norte
Training state-of-the-art AI models is growing far more costly than digital hardware can sustain, motivating analog photonic substrates that perform matrix multiplication directly in the propagation of light. Inference on such hardware is established, but training is not: backpropagation requires a backward pass and exact gradients that an analog chip cannot natively provide. Forward-only, backpropagation-free learning offers a way around this, yet it has been demonstrated only for multilayer perceptrons and convolutional networks, never for the Transformer, the architecture that now dominates AI and whose self-attention couples every position across the sequence, resisting purely local objectives.

This thesis asks whether a Transformer can be trained using only the forward-pass operations a photonic substrate provides, and characterizes what such training costs. The proposed method combines a layer-wise Forward-Forward prototype-based objective, directional-derivative gradient estimation (with no backward pass and no automatic differentiation), and a softmax-free Spherical attention adapted from the Kramers-Kronig kernel, with the four attention projections trained one at a time in a round-robin schedule; six training variants are compared across seven vision and sequence tasks to isolate the gradient estimator and the update schedule. The answer is affirmative: the fully forward-only variant trains a Transformer to a useful operating point on six of the seven tasks. Locality, not the forward-only gradient alone, is what makes this possible, and the depth-resilience of local learning is shown to extend, conditionally, to self-attention. The remaining gap to backpropagation has three sources, two inherent to forward-only learning (local credit assignment and gradient-estimation variance) and one architectural (the softmax-free attention cannot form the sharp selection that content-addressed retrieval requires). Because every operation reduces to a forward pass and a measured scalar loss, the method is a candidate for in-situ training on a photonic chip, the validation step this work points to. ...
Master thesis (2026) - S.A. Sebastian, H.S. Hung, S. Tan
In inclusive classrooms, the physical presence of children with hearing loss (HL) does not guarantee social inclusion with their typically hearing (TH) peers. Self-report is biased and manual observation is too labour-intensive to cover a full school day at fine temporal resolution. This study adapts a multimodal sensing pipeline that combines ultra-wideband (UWB) spatial tracking with Language Environment Analysis (LENA) audio recordings, applied in a single inclusive preschool with 13 children (6 HL, 7 TH).

A Gaussian Mixture Model fit on four room-level features (adult word count, auditory overlap, displacement, and teacher distance) and selected by a stability-aware rule (lowest mean BIC among K values whose cluster assignments reproduce across random initialisations, mean pairwise Adjusted Rand Index >= 0.80) over 2 to 20 components recovers six latent activity contexts, each with a distinct sensor profile. Each context is then given a post-hoc descriptive label drawn from activity types familiar in inclusive preschool classrooms: dispersed transition, peer-driven activity, independent / parallel work, adult-scaffolded peer activity, seated guided work, and whole-class instruction / read-aloud. These labels describe the recovered clusters and are not validated against an external ground truth.

Within each context, a linear mixed model with a per-child random intercept compares HL and TH children on three sensor-derivable behavioural markers of inclusion: peer co-presence (time in spatial groups, where a spatial group is a set of children simultaneously co-present and mutually body-oriented; §5.2.3), vocal participation rate (utterances per minute), and peer affiliation patterns (time in same-diagnosis vs. mixed-diagnosis spatial groups).

The HL/TH signal distributes differently across contexts for each marker. Peer co-presence shows no substantial HL/TH difference in three of six contexts. HL children exceed TH children in independent / parallel work (+7.8% of the minute, q < 0.0001, d = +1.59) and fall below TH children in peer-driven activity (-4.3% of the minute, q = 0.02) and seated guided work (-5.3% of the minute, q = 0.002). Vocal participation rate shows an HL deficit concentrated in the contexts that combine high peer communicative demand with reduced adult scaffolding (peer-driven activity q = 0.03 and adult-scaffolded peer activity q = 0.05, both significant after false-discovery-rate correction; dispersed transition in the same direction but only marginal, q = 0.07), but not in the three contexts where high adult word count or very low auditory overlap slows the pace of exchange. Peer affiliation is the most consistent asymmetry: TH children concentrate grouped time in TH-only spatial groups more than HL children concentrate theirs in HL-only spatial groups in five of six contexts (significantly in four), and HL children spend more grouped time in mixed spatial groups than TH children do across all six (significantly in four). The 6:7 cohort composition produces a baseline TH-above-HL gap of approximately 0.08 on the homophily index under random affiliation; the four significant clusters exceed this baseline (observed gap 0.12-0.19, of which 0.04-0.11 is preference beyond availability), while the two non-significant clusters are within the magnitude expected from cohort composition alone. ...
Bachelor thesis (2026) - O. Argherie, S. Tan, Y. Guo, R.L. Lagendijk
Backpropagation-free learning rules depict an affinity towards neuromorphic and energy constrained hardware, yet the final representations that they learn remain not well understood. We dive deep on two local Hebbian rules that appear to compute distinct objectives: (i) Oja’s rule computes the first principal component; (ii) SoftHebb extends it to a soft winner-take-all network whose fixed points are normalized component means. In the batch setting, Ding and He (2004) have shown that K-means and PCA are strongly related, that is, the subspace spanned by the cluster centroids coincides with the span of the first K − 1 principal directions of the data covariance. We analyze if the same correspondence survives sample by sample in a streaming setting, where updates are noisy and the weight vectors are renormalized. As such, we first provide a self contained fixed-point analysis, which we are going to use it as the common lens for both rules. Second, on controlled two dimensional Gaussian data, we assess some geometric conditions under the rules agree or disagree, yielding an actionable criterion for predicting, on a given dataset, whether the rules converge to the same representation. Third, we show the disagreement is not as the naive picture suggests, that is, an expected divergence does not hold and is replaced with a quantitative account depicted by a ratio of the cluster width to the inter cluster offset. ...
Bachelor thesis (2026) - Ștefan Stoian, S. Tan, Y. Guo, R.L. Lagendijk
Equilibrium Propagation (EP) is a backpropagation-free learning algorithm for energy-based networks; its standard estimator computes the gradient by comparing the equilibrium states reached in a free phase and a single nudged phase, but carries a bias that limits how closely EP can match backpropagation. The centered estimator reduces this bias and improves accuracy, but adds a second nudged phase per update, raising the training cost. To balance accuracy against compute, we introduce hybrid EP, a family of estimators that mix the standard and centered updates on a per-batch basis, and show analytically that the mixing probability controls this bias, so that annealing it interpolates between the two regimes. We evaluate three hybrid schedules - a cosine anneal, its inverse, and a fixed stochastic mix - against standard and centered EP on MNIST, Fashion-MNIST, and CIFAR-10, in order of increasing complexity. On the easier tasks the hybrids match centered EP at lower compute. On CIFAR-10 standard EP collapses to near-chance accuracy, and the cosine and inverse schedules collapse with it: each concentrates its biased updates into one long stretch, whereas only the stochastic mix, which spreads the same biased updates evenly across batches, trains stably. The compute savings of hybrid EP are therefore real but task-dependent: they are realized most cleanly when standard EP is itself viable, and training stability is governed not by the number of biased updates but by their distribution over the course of training. ...
The aim of this paper is to explore the potential of adapting the Mono-Forward algorithm with Zeroth-Order Optimization for backpropagation (BP) and automatic-differentiation(AD)-free image classification, assessing its feasibility in scenarios where exact gradients are unavailable. The Mono-Forward method introduces a novel approach to training neural networks without the need for backpropagation or multiple forward passes typically required in forward-forward algorithms; however it still relies on AD for local training of model layers when implemented with modern deep learning frameworks. This work proposes MF+DD, which replaces AD in Mono-Forward with zeroth-order gradient estimation via directional derivatives, resulting in a training algorithm that is free of AD and global BP. This paper also introduces a random projection based modification to adress the limitation of Mono-Forward in architectures with large intermediate activation tensors, for increased computational efficiency. Experiments on MNIST, FashionMNIST, CIFAR-10, and CIFAR-100 with both MLP and CNN architectures show that MF+DD achieves comparable accuracy to MF with AD on simpler datasets, while the accuracy gap widens on more complex benchmarks, suggesting that the noise introduced by the directional derivative estimator becomes more impactful as task difficulty increases. Results further show that increasing the number of perturbation directions P improves both accuracy and training stability with a downside of increased computational cost. ...
Bachelor thesis (2026) - D. Mustata, S. Tan, Y. Guo, R.L. Lagendijk
The backpropagation (BP) algorithm, though fundamental to modern deep learning, faces severe biological, computational, and physical limitations that hinder its applicability on energy-efficient neuromorphic systems. This has motivated the search for BP-free learning paradigms, with Hebbian-based algorithms being a notable alternative. Existing approaches range from purely local, unsupervised schemes such as SoftHebb-which naturally clusters data based on structural variance-to supervised, error-modulated approaches like PEPITA, which provides task-specific feedback through a second forward pass. However, the tradeoffs between these extremes remain largely unexplored. This study systematically compares the internal representations and hardware efficiencies of three Multi-Layer Perceptrons (MLP) trained with Backpropagation, PEPITA, and SoftHebb.

Our evaluation, utilizing geometric metrics such as Centered Kernel Alignment (CKA) and Principal Component Analysis (PCA), showcases how PEPITA exhibits higher similarity to BP, but is more geometrically aligned with the unsupervised SoftHebb. Furthermore, empirical hardware profiling exposes a significant implementation paradox: despite the theoretical efficiency of BP-free methods, high-level framework bottlenecks currently make algorithms like PEPITA computationally expensive on traditional digital architectures. ...
Bachelor thesis (2026) - A. Radu, S. Tan, Y. Guo, R.L. Lagendijk
Recent interest in biologically plausible alternatives to backpropagation has renewed attention on Spiking Neural Networks and the Forward-Forward algorithm, where learning is driven by local layer-wise goodness functions rather than global error gradients. In most Forward-Forward learning implementations of Spiking Neural Networks, goodness is defined as spike-count activity, leaving temporal properties of neural activity unused. This work investigates whether temporal spike stability, measured using the inter-spike interval coefficient of variation (ISI-CV), can improve Forward-Forward learning in fully connected leaky integrate-and-fire spiking neural networks. Using MNIST as a benchmark, we evaluate several ISI-CV-based extensions, including direct temporal penalties, contrastive gap losses, plasticity based approaches, and candidate scoring. Directly optimizing for temporal regularity conflicts with the Forward-Forward goodness margin and destabilizes training. The use of ISI-CV as a plasticity control signal, that reduces updates to temporally stable neurons, can be used to fine-tune the model. ISI-CV-based candidate scoring performs above chance, indicating that spike timing contains class-related information, but remains weaker than standard goodness-based classification. ...
Master thesis (2026) - S. Vacanas, H.S. Hung, S. Tan, M. Kok
Results show that combined dual-sensor features consistently outperform single-sensor variants, and that coarser grids yield higher exact accuracy while mean physical error remains stable across resolutions at approximately 70–80 cm. At 0.5 m resolution, the best configuration places 71% of predictions within 50 cm of the true location, approaching the accuracy of a UWB baseline system while requiring no installed infrastructure. Zone merging and hexagonal tessellation do not provide consistent improvements over the plain square grid, suggesting that magnetic ambiguity rather than data imbalance or cell geometry is the dominant source of error. The findings demonstrate that infrastructure-free magnetic fingerprinting is a practically viable approach for coarse spatial awareness in socially dynamic indoor environments. ...

Argument-to-Key-Point Mapping

A well-functioning democracy depends on an informed population. To help informing citizens, summaries of arguments in political transcripts can be made. An approach to argument summarization is the creation of summaries through distillation of the arguments into higher-level key points. In this approach, mapping arguments to key points is an important subtask. This study examines how model selection, prompting strategy, choice of domain, and input batching influence the performance of large language models (LLMs) in matching arguments to key points. We introduce a self-annotated dataset from U.S. Congress committee transcripts and evaluate both generative and embedding-based models on this task. Generative LLMs (GPT-3.5-turbo, o4-mini) outperform both untuned and fine-tuned RoBERTa in zero-shot argument-to-keypoint mapping (up to 0.880 macro-F1), while sparse two-shot prompting yields no gains. Moderate batching (n=32) boosts throughput without losing accuracy. These results show that a fully automated KPA pipeline—argument extraction, key-point generation, and mapping—is achievable with current LLMs. ...
Congressional hearings are at the center of legislation, yet their analysis is hindered by the volume and complexity of the transcripts. While recent advances in Natural Language Processing (NLP) have enabled political discourse analysis using automated tools, conventional topic modeling methods often struggle to produce semantically coherent topics due to their reliance on context-free word frequencies. This paper evaluates the performance of a new transformer-based topic modeling technique, focusing on its application to policy discussions through a detailed case study. Two variants of BERTopic are considered: (1) a parameter-tuned model and (2) a zero-shot variant, evaluated on U.S. congressional hearing transcripts from 2021 to 2024. The results demonstrate that the zero-shot version achieves competitive coherence with increased interpretability and stability, making it a useful resource for policymakers and researchers alike. This paper establishes a foundational methodological framework for automated legislative text analysis. It also outlines the trade-offs between unsupervised and semi-supervised topic modeling in political usage. ...
Bachelor thesis (2025) - A. Nikolaidis, S. Tan, E. Salas Gironés
U.S. congressional hearing transcripts offer a valuable window into national policy discourse, but they are prohibitively large for manual analysis. This study explores the use of large language models (LLMs) for multi-speaker, multi-target stance detection, a task that involves identifying each speaker's position on multiple topics within a single hearing. To this end, a novel annotation framework is introduced to produce stance labels for a small corpus of hearings from the House Oversight and Government Reform Committee. The study then evaluates the classification performance of zero- and few-shot prompting and investigates how chain-of-thought reasoning influences the results. The evaluation is conducted using OpenAI's GPT-4o and o3 models. Initial experimental results indicate that combining chain-of-thought with few-shot prompting yields the highest performance, suggesting a promising direction for automating stance analysis using LLMs in complex political discourse. ...

Investigating modality annotation impact for detecting laughter in the wild

Bachelor thesis (2025) - V. Guenov, H.S. Hung, L. Li, S. Tan
Laughter is a complex multimodal behavior and one of the most essential aspects of social interactions. Although previous research has used both auditory and facial cues for laughter detection, these approaches are commonly afflicted with difficulties in noisy, occluded, and privacy-sensitive settings. This paper explores the potential of using body posture alone—captured through 2D keypoint estimation as a robust signal for automatic laughter detection in naturalistic settings. We create a machine learning pipeline using the ConfLab dataset, which segments pose data, extracts motion-based features, and trains Random Forest classifiers on various annotation modalities (audio-only, video-only, and audiovisual) and segmentation methods (fixed and variable length). We show that, while variable-length segmentation yields optimal performance, it leads to overfitting. On the other hand, fixed-duration segmentation with three-second windows and audiovisual annotations achieves a pragmatic compromise and reaches F1-scores (65\%) comparable to earlier efforts in ideal environments. Upper-body movement, especially head and arm motion, is seen to be salient cues to laughter via feature importance analysis. Annotation modality is also found to significantly affect both classification performance and relative pose feature importance. These findings demonstrate the viability of pose-based laughter detection and reveal how annotation choices shape model behavior, offering insights for affective computing in the wild. ...
Bachelor thesis (2025) - S.A. Stan, S. Tan, E. Salas Gironés, M.S. Pera
Meetings represent a key component of collabora- tion in the workplace, serving purposes like brain- storming, discussion, and negotiation. Despite their importance, reaching a consensus among partici- pants can frequently be difficult because different people can leave the debate with different perspec- tives. In order to promote efficient communication and decision-making in organisational contexts, the use of the Shape Language is proposed. The Shape Language consists of shapes that people can use in meetings in order to represent abstract ideas, that would be difficult to represent by only words. In order to track how people interact with these ob- jects, computer vision tools can be used. This study aims to explore the current existing computer vi- sion tools for segmenting and classifying objects in meetings, aiming to find limitations in how well these models are able to recognize objects in the context of meetings and negotiations. Results of this study show that after fine-tuning four models on the custom dataset, they can recognize the three shapes provided as classes in most of the cases, but still make mistakes when assigning classes, or miss objects that they should classify all together, which show limitations of these modern tools. ...
Bachelor thesis (2025) - E. Milinović, S. Tan, E. Salas Gironés, M.S. Pera
Meetings are a vital part of discussions and negotiations. Unfortunately, individuals often leave with a vague understanding of the topics covered during the meeting and tend to forget even more of what transpired as time goes on. Driven by previous research that attempts to solve the issue by using architectural shapes as a way of removing ambiguity along with recent advancements in Automatic Speech Recognition (ASR) and Natural Language Processing (NLP) this research attempts to improve user understanding of key topics discussed in meetings by combining ASR models with NLP tools to create a visual summary that would improve user understanding of key topics covered during meetings. To achieve this the research utilizes the speech-to-text transcription and speaker identification capabilities of the WhisperX model with noun phrase extraction features provided by Spacy and key topic recognition functionality of Microsoft's DeBERTa model. Finally, the data is presented as a node-based graph utilizing the D3.js library. The results show that the system is able to identify between 33% - 58% of meeting key topics. This shows the potential of combining ASR models with NLP tools for creating concise meeting summaries but also raises new questions such as why some topics were missed, how the system performance can be improved, and how to design an optimal user interface for such a task. ...

Analyzing negotiations using the Coloured Trails Game & the NegotiAct

Bachelor thesis (2025) - A.B. Kichukov, S. Tan, E. Salas Gironés, M.S. Pera
Within the field of negotiations, a recent publication is a paper called the NegotiAct[9], which analyzed existing coding schemes of negotiations and introduced an improvement on them that promises a viable way to analyze negotiations in depth. In this research, my goal is to develop a workflow for gathering information and analyzing it with the NegotiAct. To this end the Colored Trails Game [7] is used to design an experiment that simulates real-life negotiations. The Colored Trails Game is a game where each player, through negotiating tries to maximize their own score under limited resources. The game at its core offers the opportunity for both cooperation and competitiveness.
In the experiment a total of 15 participants took part and it was run a total of 20 times, resulting in 3 hours and 10 minutes of recordings. Encoding them with the NegotiAct resulted in the discerning of a total of 87 offers made, 24 offers accepted, 26 offers rejected, and 16 requests for offer modifications. Based on this coherent mapping it can be concluded that the Colored Trails Game is a suitable choice for the workflow of gathering data to be put through the analysis of the NegotiAct.
...

Segmenting and Tracking Hand Movements During Human Interaction

Bachelor thesis (2025) - A.M. Semov, S. Tan

Rethinking Ubiquitous Smart Sensing of Social Behaviour in the Wild

Multiactivity analysis investigates one's coordination of actions within a social context, such as gestures and speech, usually using video recordings of the social activity, to further understand the rules of human behaviour. This paper focuses specifically on the coordination between speaking and drinking activities within a social setting, and explores the possibility of automatically identifying these events using audio captured from a drinking glass. As social interactions occur in vastly different contexts, this paper also investigates the effect that background noise might have on the accuracy of identifying these events. Different parameters and audio features were compared. Linear classification models LR and SVM with a linear kernel were able to achieve 100% accuracy for all sample lengths between 2 and 8 seconds using the first 20 PCA components from 60 audio features. The best performing feature in identifying speaking and drinking events was MFCCs, achieving an F1 score of 99.4% on average across models with a training sample length of 3 seconds. Background noise had different effects on classification accuracy depending on the type, with music lowering the F1 score to 74.3%, noisy room audio to 64.7%, and podcast audio simulating the presence of other speakers to 59.6% using MFCCs and a 3-second sample length ...

Rethinking Ubiquitous Smart Sensing of Social Behaviour In The Wild

This research investigates the detection of gestures using a torso-worn accelerometer sensor. Using the Conflab dataset, we focus on gestures during conversations in mingling scenarios. Due to significant variability in gesture styles among individuals, traditional methods face challenges in building personalized models. Our experiments demonstrate that Transductive Parameter Transfer (TPT), an adaptive transfer learning method, can more effectively model these individual differences in gesturing. To gain insights into individual expressiveness, we classify gestures into three classes: 'no gesture,' 'normal,' and 'large' gestures. TPT performed an average AUC score of 0.84 in binary classification and 0.77 in multiclass classification. These findings highlight the potential of using a single torso-worn accelerometer to understand social behavior in naturalistic settings. ...
Human activity recognition plays an interesting and important role nowadays as there are a variety of use cases. It is utilized in health monitoring, in the development of human-computer interaction system and in security monitoring. However current methods involve usage of privacy sensitive data and impractical sensors for everyday usage. To tackle this problem, we aim to answer the research question "How to maximize the capabilities of in-mouth sensors for human activity recognition?". The main contributions of this paper are the classification of different gestures using an in-mouth device, implementation of a classifier directly onto a microcontroller and the evaluation whether the models can generalize to multiple people. To investigate this, we experimented with popular classical machine learning classifiers: Decision Tree, K-Nearest Neighbors, Support Vector Machine, Logistic Regression and Random Forest classifiers. The results shows that the F1-score of all classification problems are above 80% using the various classifiers along with different parameters. ...