CT
C. Teodorescu
info
Please Note
<p>This page displays the records of the person named above and is not linked to a unique person identifier. This record may need to be merged to a profile.</p>
1 records found
1
Malware-Domain Continued Pre-Training for Binary Malware Classification
A Leakage-Aware Study of Code Models on the SBAN Corpus
Continued pre-training can adapt language models to a domain, but for malware classification it is un- clear whether gains come from malware-specific information or from additional training on code. We study this question on a leakage-controlled binary benchmark derived from the SBAN cor- pus, using strict BENIGN/MALWARE labels, exact duplicate removal, and matched compar- isons between TF-IDF baselines, CodeBERT, and Qwen2.5-Coder variants. Across the tested model sizes and pre-training data budgets, continued pre- training changes model behaviour but does not pro- duce a reliable downstream classification improve- ment. Even in the best malware-related setting, the margin over an equally trained general-code control is very small. Under these constraints, the overall trend is that malware-related continued pre-training does not improve binary classification in a reliable way.
...
Continued pre-training can adapt language models to a domain, but for malware classification it is un- clear whether gains come from malware-specific information or from additional training on code. We study this question on a leakage-controlled binary benchmark derived from the SBAN cor- pus, using strict BENIGN/MALWARE labels, exact duplicate removal, and matched compar- isons between TF-IDF baselines, CodeBERT, and Qwen2.5-Coder variants. Across the tested model sizes and pre-training data budgets, continued pre- training changes model behaviour but does not pro- duce a reliable downstream classification improve- ment. Even in the best malware-related setting, the margin over an equally trained general-code control is very small. Under these constraints, the overall trend is that malware-related continued pre-training does not improve binary classification in a reliable way.