Z. Yue
Please Note
14 records found
1
Backdoor Attacks in Active Learning
An Extensive Analysis of Backdoor Injection in Active Learning-Trained Computer Vision Models
Here, the vulnerabilities of active learning against backdoor attacks in computer vision models were analyzed. Various configurations, datasets and deep learning models were used to evaluate their effectiveness. Backdoor attacks managed to hit ASR values over 95% with just 1% of the data being poisoned on simple datasets like MNIST, particularly when using certainty-based sampling and CNNs. More complex datasets like CIFAR-10 and models like ResNet proved to be more resilient. Furthermore, different attack techniques were explored, such as progressive parameter adjustment, sub-trigger division and clean label attacks on advanced backdoor triggers like LIRA and WaNet. The analysis revealed that although global LIRA triggers were the most effective, sub-trigger and progressive poisoning methods offered promising alternatives, especially because they allow poisoning smaller parts of images across training cycles. Additionally, it is revealed that attack success in clean-label scenarios was highly dependent on the number of poisoned samples per cycle, due to the post-query constraint that only allows poisoning already-selected samples, often limiting impact when the target label appears infrequently. Finally, different poisoning timings were compared. From this, post-query poisoning consistently outperformed pre-query methods in terms of ASR, even at low poison rates. However, it also pointed out that this approach has its limitations in real-world scenarios, where attackers usually do not have control over the samples being queried. Clean accuracy remained unaltered to a large extent, demonstrating the stealth and hidden threat of backdoor attacks in active learning settings.
...
Here, the vulnerabilities of active learning against backdoor attacks in computer vision models were analyzed. Various configurations, datasets and deep learning models were used to evaluate their effectiveness. Backdoor attacks managed to hit ASR values over 95% with just 1% of the data being poisoned on simple datasets like MNIST, particularly when using certainty-based sampling and CNNs. More complex datasets like CIFAR-10 and models like ResNet proved to be more resilient. Furthermore, different attack techniques were explored, such as progressive parameter adjustment, sub-trigger division and clean label attacks on advanced backdoor triggers like LIRA and WaNet. The analysis revealed that although global LIRA triggers were the most effective, sub-trigger and progressive poisoning methods offered promising alternatives, especially because they allow poisoning smaller parts of images across training cycles. Additionally, it is revealed that attack success in clean-label scenarios was highly dependent on the number of poisoned samples per cycle, due to the post-query constraint that only allows poisoning already-selected samples, often limiting impact when the target label appears infrequently. Finally, different poisoning timings were compared. From this, post-query poisoning consistently outperformed pre-query methods in terms of ASR, even at low poison rates. However, it also pointed out that this approach has its limitations in real-world scenarios, where attackers usually do not have control over the samples being queried. Clean accuracy remained unaltered to a large extent, demonstrating the stealth and hidden threat of backdoor attacks in active learning settings.
Our work proposes an innovative way to bridge the gap between vulnerability data (CVEs) and security alert data originating from multiple security tools that protect servers using MITRE ATT&CK tactics. That would provide more context to the alerts which would be useful in their classification as attacks or false positives. We use DeBERTa (Decoding-enhanced BERT with Disentangled Attention), a deeplearning state-of-the-art model, to map CVE descriptions to MITRE ATT&CK tactics. Then, we map security alerts to MITRE ATT&CK tactics, which will be used as input to a context-enriched machinelearning model (by CVEs and tactics). That machine-learning model is used to classify security alerts as malicious or benign. We tested our approach using over 5.5 million security alert data combined with red-team exercise attacks and incident response labelling from the company, a large international organization with over 60,000 employees. Our CVE+tactic model (without hyperparameter tuning) detects 64% more true positives than the machine-learning model without that information. In addition, the SOC needs to investigate less than 1400 alerts to catch the red-team attacks in our test set, compared to more than 5500 generated by the model without CVE and tactics. Moreover, assuming a standard response time of 8 minutes per alert, this improved model would save the SOC team up to 550 person hours. That yields a model that catches red-team attacks without overwhelming the SOC with too many false positives. ...
Our work proposes an innovative way to bridge the gap between vulnerability data (CVEs) and security alert data originating from multiple security tools that protect servers using MITRE ATT&CK tactics. That would provide more context to the alerts which would be useful in their classification as attacks or false positives. We use DeBERTa (Decoding-enhanced BERT with Disentangled Attention), a deeplearning state-of-the-art model, to map CVE descriptions to MITRE ATT&CK tactics. Then, we map security alerts to MITRE ATT&CK tactics, which will be used as input to a context-enriched machinelearning model (by CVEs and tactics). That machine-learning model is used to classify security alerts as malicious or benign. We tested our approach using over 5.5 million security alert data combined with red-team exercise attacks and incident response labelling from the company, a large international organization with over 60,000 employees. Our CVE+tactic model (without hyperparameter tuning) detects 64% more true positives than the machine-learning model without that information. In addition, the SOC needs to investigate less than 1400 alerts to catch the red-team attacks in our test set, compared to more than 5500 generated by the model without CVE and tactics. Moreover, assuming a standard response time of 8 minutes per alert, this improved model would save the SOC team up to 550 person hours. That yields a model that catches red-team attacks without overwhelming the SOC with too many false positives.
Objective. Our study aims to identify users’ preferences on ethical principles that a virtual coach should follow to decide when to allocate human feedback to individuals preparing to quit smoking.
Methods. Our research was based on pre-gathered data, that included participants’ responses to open and closed questions regarding feedback allocation principles. Thematic analysis was conducted on these responses. Triangulation was performed using a qualitative literature review and quantitative data analysis.
Results. Four main themes were identified: (1) Struggling the Most (63.75%), (2) Increasing Chances of Success the Most (13.75%), (3) Equal Treatment (11.25%), and (4) Appreciating the Most (11.25%). Participants prioritized support for those experiencing the greatest difficulty in smoking cessation. The triangulation supported the validity of these themes.
Conclusions. Our study highlights the importance of integrating user-preferred ethical principles in virtual coaching systems for smoking cessation. Prioritization of users who struggle the most can increase the effectiveness and fairness of such systems, potentially increasing success rates. Future research should explore additional ethical principles, combining several principles into systems, and real-world application of these findings to further refine virtual coaching in healthcare. ...
Objective. Our study aims to identify users’ preferences on ethical principles that a virtual coach should follow to decide when to allocate human feedback to individuals preparing to quit smoking.
Methods. Our research was based on pre-gathered data, that included participants’ responses to open and closed questions regarding feedback allocation principles. Thematic analysis was conducted on these responses. Triangulation was performed using a qualitative literature review and quantitative data analysis.
Results. Four main themes were identified: (1) Struggling the Most (63.75%), (2) Increasing Chances of Success the Most (13.75%), (3) Equal Treatment (11.25%), and (4) Appreciating the Most (11.25%). Participants prioritized support for those experiencing the greatest difficulty in smoking cessation. The triangulation supported the validity of these themes.
Conclusions. Our study highlights the importance of integrating user-preferred ethical principles in virtual coaching systems for smoking cessation. Prioritization of users who struggle the most can increase the effectiveness and fairness of such systems, potentially increasing success rates. Future research should explore additional ethical principles, combining several principles into systems, and real-world application of these findings to further refine virtual coaching in healthcare.
We perform exploratory research to investigate the ability of the screen compendium to show network structures that reflect known biological processes. We find that with stringent requirements on interactions the screen compendium shows enrichment for a wide range of biological processes and known protein-protein interactions. We further conclude that the experimental design biases network behaviour and needs to be accounted for when constructing networks. We recommended a mutual k-nearest neighbor network construction approach, which yielded networks with the most biological relevance.
We compare the CELL-seq screens using findings from the approaches to the DepMap dataset, a well-known collection of synthetic lethality CRISPR screens, and find that the behaviour of these datasets is in many ways mirrored. We conclude that this is both due to the biology they represent and the differences in the number of screens in each dataset. Finally, we compare the coverage of biological processes between the HAP1 compendium and DepMap, and show large overlap in their coverage. Nonetheless, the differences they do show leads us to bring forward two hypotheses for gene-gene interactions that score strongly uniquely in the CELL-seq networks which are biologically plausible but are not found in DepMap or curated literature, warranting future investigations.
All code pertaining to the methods and figures in this work are hosted on GitLab by the High Performance Computing Facility of the Netherlands Cancer Institute. As such the code can be viewed by supervisors, but further details could be shared upon request. ...
We perform exploratory research to investigate the ability of the screen compendium to show network structures that reflect known biological processes. We find that with stringent requirements on interactions the screen compendium shows enrichment for a wide range of biological processes and known protein-protein interactions. We further conclude that the experimental design biases network behaviour and needs to be accounted for when constructing networks. We recommended a mutual k-nearest neighbor network construction approach, which yielded networks with the most biological relevance.
We compare the CELL-seq screens using findings from the approaches to the DepMap dataset, a well-known collection of synthetic lethality CRISPR screens, and find that the behaviour of these datasets is in many ways mirrored. We conclude that this is both due to the biology they represent and the differences in the number of screens in each dataset. Finally, we compare the coverage of biological processes between the HAP1 compendium and DepMap, and show large overlap in their coverage. Nonetheless, the differences they do show leads us to bring forward two hypotheses for gene-gene interactions that score strongly uniquely in the CELL-seq networks which are biologically plausible but are not found in DepMap or curated literature, warranting future investigations.
All code pertaining to the methods and figures in this work are hosted on GitLab by the High Performance Computing Facility of the Netherlands Cancer Institute. As such the code can be viewed by supervisors, but further details could be shared upon request.
A Comparative Analysis of Learning Curve Models and their Applicability in Different Scenarios
Finding datasets patterns which lead to certain parametric curve model
Non-Monotonicity in Empirical Learning Curves
Identifying non-monotonicity through slope approximations on discrete points
Empirical Investigation of Learning Curves
Assessing Convexity Characteristics
”How Much Data is Enough?” Learning curves for machine learning
Investigating alternatives to the Levenberg-Marquardt algorithm for learning curve extrapolation
The paper specifically explores the Learning Curve Database (LCDB) and investigates alternative fitting algorithms to the employed Levenberg-Marquardt (LM). These algorithms are Gradient Descent and BFGS, and the paper aims to determine whether they are more suitable for fitting learning curves than LM.
The algorithms were implemented, both in their default and optimised states, and the results were compared to LM. The results measured mean-squared error (MSE), L1 Loss, individual parametric model performance, and computation time.
The findings showed that Gradient Descent is not a suitable alternative to LM; however, BFGS proved to be competitive, as it is practically identical in performance while being significantly faster than LM. The results answered the proposed aim of the paper and generated new questions that need answering.
Further exploration of the BFGS algorithm and its application on learning curve fitting is recommended. Comparisons between the MSE distribution of LM and BFGS can be further explored, as well as comparisons on new parametric models, learners, and datasets. ...
The paper specifically explores the Learning Curve Database (LCDB) and investigates alternative fitting algorithms to the employed Levenberg-Marquardt (LM). These algorithms are Gradient Descent and BFGS, and the paper aims to determine whether they are more suitable for fitting learning curves than LM.
The algorithms were implemented, both in their default and optimised states, and the results were compared to LM. The results measured mean-squared error (MSE), L1 Loss, individual parametric model performance, and computation time.
The findings showed that Gradient Descent is not a suitable alternative to LM; however, BFGS proved to be competitive, as it is practically identical in performance while being significantly faster than LM. The results answered the proposed aim of the paper and generated new questions that need answering.
Further exploration of the BFGS algorithm and its application on learning curve fitting is recommended. Comparisons between the MSE distribution of LM and BFGS can be further explored, as well as comparisons on new parametric models, learners, and datasets.