JO

J.T. O'Dwyer Wha Binda

info

Please Note

2 records found

Transferable & Parameter Efficient LLM Fine Tuning

With the increasing popularity of Large Language Models (LLMs), fine-tuning them has become increasingly computationally expensive. Parameter Efficient Fine-Tuning (PEFT) methods like LoRA and Adapters, introduced by Microsoft and Google, respectively, aim to reduce the number of trainable parameters, with the current state-of-the-art combining both methods as LoRA Adapters. This paper introduces Transformer Modules as a PEFT method. These modules utilize Modular Transformer Blocks (MTBs) inserted into a frozen pre-trained model, achieving competitive performance while significantly reducing computation costs. Compared to the current state-of-the-art using GPT-2, BERT, and T5, Transformer Modules further reduced compute time by 39.7\% and training memory by 72.7\%, with a performance cost of 4.5±2.51\% on the GLUE benchmark. Additionally, the paper presents the Transformer Bridge, a continuous vector transformer designed to transfer Transformer Modules across different models. This could enable cross-model fine-tuning, allowing model-agnostic modules, such as an ethics or medical module, to be used across various LLMs without retraining or access to the original dataset. Although the current implementation of the Transformer Bridge did not fully succeed in mapping embedding spaces, analysis of the results suggests that further refinements using traditional model distillation techniques could lead to success in future iterations. ...
In recent years there has been an increase in the number of patients for issues relating to mental illness. To this effect to help with this increase, schema mode assessment through a conversation agent is being used to conduct schema therapy, a form of psychological treatment. To train such an agent, training data labeled by humans is necessary but can be very expensive to conduct. The question being researched is through the use of Active Machine learning is it possible to reduce the amount of required labeled data to do such classification. Three experiments on the use of active learning with currently available classifiers were performed where the active learner attempted to train the classifiers to an accuracy within +/- 3% of the same classifier trained with traditional machine learning on the full data set. The experimental results found that in all cases the use of active learning drastically decreased the number of necessary labeled data the classifier needed to achieve a similar accuracy. Consistently reducing the number by 98% and above answering the initial question. Though possible limitations of the data set and classifiers for such texts may be positively influencing the magnitude of the reduction. ...