QS

Q. Song

info

Please Note

13 records found

Human activity recognition plays an interesting and important role nowadays as there are a variety of use cases. It is utilized in health monitoring, in the development of human-computer interaction system and in security monitoring. However current methods involve usage of privacy sensitive data and impractical sensors for everyday usage. To tackle this problem, we aim to answer the research question "How to maximize the capabilities of in-mouth sensors for human activity recognition?". The main contributions of this paper are the classification of different gestures using an in-mouth device, implementation of a classifier directly onto a microcontroller and the evaluation whether the models can generalize to multiple people. To investigate this, we experimented with popular classical machine learning classifiers: Decision Tree, K-Nearest Neighbors, Support Vector Machine, Logistic Regression and Random Forest classifiers. The results shows that the F1-score of all classification problems are above 80% using the various classifiers along with different parameters. ...

Rethinking Ubiquitous Smart Sensing of Social Behaviour in the Wild

Multiactivity analysis investigates one's coordination of actions within a social context, such as gestures and speech, usually using video recordings of the social activity, to further understand the rules of human behaviour. This paper focuses specifically on the coordination between speaking and drinking activities within a social setting, and explores the possibility of automatically identifying these events using audio captured from a drinking glass. As social interactions occur in vastly different contexts, this paper also investigates the effect that background noise might have on the accuracy of identifying these events. Different parameters and audio features were compared. Linear classification models LR and SVM with a linear kernel were able to achieve 100% accuracy for all sample lengths between 2 and 8 seconds using the first 20 PCA components from 60 audio features. The best performing feature in identifying speaking and drinking events was MFCCs, achieving an F1 score of 99.4% on average across models with a training sample length of 3 seconds. Background noise had different effects on classification accuracy depending on the type, with music lowering the F1 score to 74.3%, noisy room audio to 64.7%, and podcast audio simulating the presence of other speakers to 59.6% using MFCCs and a 3-second sample length ...

Rethinking Ubiquitous Smart Sensing of Social Behaviour In The Wild

This research investigates the detection of gestures using a torso-worn accelerometer sensor. Using the Conflab dataset, we focus on gestures during conversations in mingling scenarios. Due to significant variability in gesture styles among individuals, traditional methods face challenges in building personalized models. Our experiments demonstrate that Transductive Parameter Transfer (TPT), an adaptive transfer learning method, can more effectively model these individual differences in gesturing. To gain insights into individual expressiveness, we classify gestures into three classes: 'no gesture,' 'normal,' and 'large' gestures. TPT performed an average AUC score of 0.84 in binary classification and 0.77 in multiclass classification. These findings highlight the potential of using a single torso-worn accelerometer to understand social behavior in naturalistic settings. ...

Predicting force feedback of liquids for haptic bilateral teleoperation

This paper presents a novel approach to simulating liquid interactions for haptic bilateral teleoperation without complex fluid dynamics simulations. We propose a model combining a simplified drag equation with a "position tail" mechanism to approximate force feedback for viscous and turbulent liquids. The model was evaluated through theoretical analysis, manual testing, and a user study. Results of the latter demonstrate the ability to convincingly simulate different liquid types, with participants correctly identifying viscous and turbulent liquid types. The study also revealed that visual feedback delays up to 125 ms did not significantly impact user experience. While promising, the system remains sensitive to noisy input data. This research contributes to the development of efficient and believable haptic simulations for teleoperation, paving the way for future developments. ...
Haptic bilateral teleoperation aims to create systems, where human precision and dexterity can be applied over long distances with the aid of robotic devices. These systems could make remote surgery or spacecraft repair a reality. The main challenge for such systems is latency. TU Delft's Networked Systems group is developing a teleoperation system, which combines predicted force feedback with delayed vision feedback. This project investigates the force feedback prediction of cutting interactions by implementing a simple physics based model to estimate the forces of cutting objects with a knife in a haptic simulation. The quality of the interaction and impact of visual delay is evaluated in a user study. It is found that a simple physics based cutting model can create convincing cutting interactions, which users find to be sufficiently controllable up to a delay of 75ms. ...
Bachelor thesis (2024) - D.L.A. de Klein, Kees Kroep, R.R. Venkatesha Prasad, Q. Song
Being able to do something while not being in the same room could be benefit more people than you know. For this, the improvements made in the world of robotics and network systems world lead to using haptic bilateral teleoperation to being a vi- able option to help with this. However when con- trolling a robot, delays make the experience a lot more difficult and with very precise tasks, like sur- geon trying to make incisions, this could be unac- ceptable when going wrong as there is possibly no way back. A way to mitigate this is by using predic- tive force feedback to aid in this. For this paper, we look into if we can approximate hard physical tran- sitions, like puncturing, to gain satisfactory force feedback for haptic bilateral teleoperation. The basis to getting physical state transition was to model a needle puncturing a piece of paper. For this to be modelled, we have opted to go for a thresh- old mechanism that separates the states from not puncturing to puncturing. To get a first guess on what this threshold should be, physical testing was performed with a real needle and paper. For this actions of applying force with a needle on paper, how been recorded and from it a threshold force has been measured. This results is used for the thresh- old used in the models creation in Bullet physics engine provided by the Networked Systems group. The threshold mechanism, a friction mechanism and mechanism to apply delay on the model simu- lating the delay of the robot have been implemented for this research. This model was then tested for its behaviour over time, which resulted in it performing the way it was expected with respects to what was implemented. Then the model was evaluated in an user study on which different delays have been applied on the model, to gain insight in how immersive the expe- rience is with these delays applied which simulates the reality more when it could be used in combina- tion with a robot arm. The evaluation resulted in it being immersive, while the difficulty of the task of puncturing a piece of paper does increase the more the delay is. With hard physical transitions being able to be approximated with a the model behaving like expected and the force feedback resulting for it being satisfying enough for the user for the model to be immersive enough, makes this hard physical transitions worth continuing researching about in the future. ...
Natural disasters often present significant challenges for rescue operations due to the complex and hazardous environments they create. Teleoperation technology, particularly haptic bilateral teleoperation, offers promising solutions to enhance the efficiency, safety, and effectiveness of disaster response. This study explores the feasibility of approximating dynamic object movement to achieve satisfactory force feedback for haptic bilateral teleoperation systems, focusing on minimizing the impact of network delays on user experience. We conducted a series of experiments and user studies to assess the system’s performance under various delay conditions. Our findings indicate that dynamic object movement can be successfully approximated to provide realistic force feedback. The user study revealed that network delays greater than 75 ms significantly impact task difficulty, and delays beyond 125 ms affect system usability, although the system remains functional up to a total delay of 200 ms. The acceptable total delay, including system latency of approximately 135 ms, is found to be around 185 ms for optimal user experience. This research contributes to the development of teleoperation systems by providing insights into the
acceptable levels of network delay and the practical implementation of force feedback mechanisms, paving the way for future advancements in the field. ...
Bachelor thesis (2024) - S. Stoicescu, R.R. Venkatesha Prasad, Kees Kroep, Q. Song
In this paper, we investigate the use of heightmap-based models to predict deformations in granular materials during teleoperation tasks. Our primary objective is to provide real-time cutting feedback and contact predictions for haptic bilateral teleoperation applications, thereby enhancing operator control in remote manipulation scenarios. We developed a simulation model that utilizes heightmaps to represent deformable materials and implemented a cutting algorithm to simulate complex interactions with granular soils.

Through a series of experiments, we demonstrated that higher resolution heightmaps significantly improve the accuracy and stability of cutting feedback. Our findings suggest that this approach is a viable method for keeping track of changes in a deformable object's shape and can be further expanded to include a wider range of teleoperation tasks, potentially improving the safety and effectiveness of remote operations in hazardous environments. ...

The Use of Rich Metadata and Graph Learning to Estimate Task Transferability

Master thesis (2024) - H.J. van der Wilk, R. Hai, Z. Li, A. Anand, Q. Song
The democratization of machine learning through public repositories, often known as model zoos, has significantly increased the availability of pre-trained models for practitioners. However, this abundance can make it difficult to choose the most suitable pre-trained model for fine-tuning on new tasks. Although various methods have been proposed in the field of transferability estimation to address this issue, these methods can take hours to execute and may still fail to find the optimal pre-trained model for fine-tuning. By exploring a new graph learning-based approach to transferability estimation, we outperform state-of-the-art methods such as LogME, improving the accuracy of the best-predicted model by up to 31.5\% in less than 5 minutes. ...
Master thesis (2024) - J. Liu, O.E. Scharenborg, Q. Song, Z. Yue
Dysarthric speech, characterized by articulation problems and a slower speech rate, shows lower automatic speech recognition (ASR) performance compared to normal speech. To improve performance, researchers often try to enhance dysarthric speech to be more like normal speech before passing it through an ASR trained on normal speech. In this project, we compare different signal processing and voice conversion techniques for dysarthric-to-normal speech enhancement. The resulting enhanced speech is objectively evaluated using an ASR system trained on normal speech. Also, the naturalness and intelligibility of the enhanced dysarthric speech are evaluated through listening experiments. Finally, the correlation between subjective and objective evaluations was analyzed. We found that among the techniques investigated, time-stretching demonstrated superior performance in objective evaluation experiments, surpassing state-of-the-art voice conversion methods. Across all methods, improvements in naturalness and intelligibility were positively correlated with improvements in automatic speech recognition (ASR) performance. However, this correlation was significant for some methods but not for others. ...
Master thesis (2024) - B. Visser, R.R. Venkatesha Prasad, Q. Song, G.M. Willemsen
In response to the growing demand for sustainable transportation solutions, bike-sharing systems have gained prominence in supporting an eco-friendly means of commuting. Within this landscape, Skopei, a forward-thinking company specializing in innovative sharing propositions, has developed a smart bike lock capable of autonomous rentals and returns. This master’s thesis delves into the application of radio frequency-based distance measurements to create an indoor positioning system for the efficient management of bike storage. Specifically, the system is designed to determine whether bikes have been correctly parked at docking locations, enabling users to conclude their rentals autonomously. The architecture of this system uses a network of anchor nodes (known-location routers) that should ascertain the positions of mobile nodes (bikes) within the bicycle storage area. Notably, the solution developed in this thesis employs sophisticated distance measurement techniques, including frequency hopping and phase shift analysis. By finding the amplitude and phase shift over multiple frequencies, we can find the channel impulse response and estimate the distance using machine learning. We employ a novel Multi-layer Perceptron neural network regressor to improve the accuracy in the presence of complex environmental factors in bike storage environments. In the bike storage test case, we achieved a mean absolute error in position estimation of 1.68m compared to 3.80m of a naive approach. We improved the parking state classification from 75.99% of a naive approach to 98.09% with our machine-learning-based approach. This thesis underscores the importance of cutting-edge distance measurement methods and real world field studies in advancing indoor positioning systems, specifically for smart bike storage management. By bridging the gap between technology and sustainable transportation, this work aims to make urban bike-sharing systems more scalable, efficient, user-friendly, and environmentally conscious. ...
In the last decades our usage of the Radio Frequency (RF) spectrum has intensified a lot due to the increasing number of mobile devices. The spectrum is getting crowded and new communication alternatives might become necessary in the future. One such alternative is Visible Light Communication (VLC),which uses the visible light spectrum instead of RF. Passive VLC, in particular, combines this property with a low power usage, suited for battery powered devices. However, high data rates are lacking.

This thesis proposes a multi-channel passive VLC system, based on light dispersion principles, as a novel concept to try and enhance these data rates. Results of the system so far only show low data rates however, at a maximum of 4 bits per second. The low data rate in particular is caused by choice of transmitter and receiver, thereby limiting the bandwidth and the number of channels that can be created. Yet, the concept itself is still considered a valid and valuable approach, and higher data rates can be expected from future works.

In this thesis we will see exactly how we can go from light dispersion to creating a multi-channel passive VLC system. As we will see this requires an optical design, a transmitter and a receiver, all of which will be working together. Besides an evaluation of the performance it will give insight in challenges this concept faces and, also, what improvements might increase the system performance. ...
Recently non-image-based data capturing methods such as sensors like RF, ultrasonic or radars, Wi-Fi, Bluetooth, etc for People Counting (PC) applications have gained momentum as an alternative to camera-based systems due to the preference for privacy preservation. Among them mm-wave radars are a strong choice for data capture since they consume less power, they are robust to the extremes of weather and environment. Further, they record data in the form of a collection of points making them computationally lightweight. However, such type of data becomes a challenge when used outdoors since those environments are complex. It becomes extremely difficult to separately identify individuals from the collection of points. This makes PC outcomes inaccurate when using radar data.
Therefore, in our project, we aimed to address these limitations of using mm-wave radar data. We treated groups of people as a unit, distinguishing them based on the number of people in the group. In this way, we were able to count the number of people by detecting the group size. To achieve this, we processed the raw radar output data made available by the AMS Institute into a dataset to suit our chosen deep learning model. Through a literature survey, we found that computer vision algorithms based on Deep Neural Networks (DNN), detect objects while maintaining a balance between speed and accuracy. Moreover, DNNs are capable of extracting features more effectively than traditional Machine Learning models. We selected the YOLOv8 model by Ultralytics after concluding our literature survey, as the most suitable model for our research problem. However, the selected model accepts three-channel images and videos to detect objects. Therefore, we customized the YOLOv8 model as well as formatted our radar data into a four-channel tensor input which would be acceptable by the model. We were successful in detecting groups of up to four people outdoors with a detection accuracy of 79.21%.
...