VC

V.P. Chatalbasheva

info

Please Note

2 records found

Master thesis (2025) - V.P. Chatalbasheva, H. Jamali-Rad, S. Rastegar, E. Isufi, Holger Caesar, Hamid Palangi
Text-to-image (T2I) diffusion models have achieved remarkable image quality but still struggle to produce images that align with the compositional information from the input text prompt, especially when it comes to spatial cues. We attribute this limitation to two key factors: the lack of clear fine-grained spatial supervision in common training datasets, and the inability of the CLIP text encoder, used in the pretraining of stable diffusion models, to represent spatial semantics. While recent work has addressed object omission and attribute mismatches, accurately generating objects in the spatial locations defined in the text prompt remains an open challenge. Prior solutions typically rely on fine-tuning, which introduces computational overhead and risks degrading the pretrained model’s generative prior on other tasks unrelated to spatial reasoning. In this paper, we introduce InfSplign, a simple and training-free method that improves spatial understanding in T2I diffusion models. InfSplign leverages attention maps and a centroid-based loss to guide object placement during sampling at inference time without modifying the pretrained model. Our approach is modular, lightweight and compatible with any pretrained diffusion model. InfSplign achieves strong performance on spatial benchmarks such as VISOR, T2I-CompBench and GenEval, outperforming baselines in many scenarios. ...
Sedentary activity recognition is an important research field due to its various positive implications in people’s life. This study builds upon previous research which is based on low level features extracted from the gaze signals using a fixation filter and uses a dataset of 24 participants performing 8 different sedentary activities. The main research question are related to extracting features from the raw data and selecting the most relevant ones which improve the classification accuracy. The novelty of this paper is using dynamic thresholds in the fixation filter to ensure the fixation-specific measurements reported by literature as well as contributing to the human activity recognition (HAR) field by developing an additional low-level gaze feature in combination with the fixation dispersion area. The machine learning (ML) models, Random Forest, k-NN (k-Nearest Neighbour) and SVM (Support Vector Machine), used for the classification task are evaluated using the within dataset evaluation protocol, with cross validation and hyperparameter tuning. The overall recognition accuracy of the Random Forest model is 0.94 (f1-score). ...