How Much Do Code Language Models Remember? An Investigation on Data Extraction Attacks Before and After Fine-tuning

Conference Paper (2025)
Author(s)

Fabio Salerno (TU Delft - Electrical Engineering, Mathematics and Computer Science)

Ali Al-Kaswan (TU Delft - Electrical Engineering, Mathematics and Computer Science)

Maliheh Izadi (TU Delft - Electrical Engineering, Mathematics and Computer Science)

Research Group
Software Engineering
DOI related publication
https://doi.org/10.1109/MSR66628.2025.00080 Final published version
More Info
expand_more
Publication Year
2025
Language
English
Research Group
Software Engineering
Pages (from-to)
465-477
Publisher
IEEE
ISBN (electronic)
9798331501839
Event
22nd IEEE/ACM International Conference on Mining Software Repositories, MSR 2025 (2025-04-28 - 2025-04-29), Ottawa, Canada
Page Views
50
Reuse Rights

Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.

Abstract

Code language models, while widely popular, are often trained on unsanitized source code gathered from across the Internet. Previous work revealed that pre-trained models can remember the content of their training data and regurgitate them through data extraction attacks. Due to the large size of current models, only a few entities have the resources for pre-training such models. However, fine-tuning requires fewer resources and is increasingly used by both small and large entities for its effectiveness on specialized data. Such small curated data for finetuning might contain sensitive information or proprietary assets. In this study, we attack both pre-trained and fine-tuned code language models to investigate the extent of data extractability. We first develop a custom benchmark to assess the vulnerability of both pre-training and fine-tuning samples to extraction attacks. Our findings reveal that 54.9% of extractable pre-training data could be retrieved from StarCoder2-15B, whereas this number decreased to 23.5% after fine-tuning. This indicates that finetuning reduces the extractability of pre-training data. However, compared to larger models, fine-tuning smaller models increases their vulnerability to data extraction attacks on fine-tuning data. Given the potential sensitivity of fine-tuning data, this can lead to more severe consequences. Lastly, we also manually analyzed 2000 extractable samples before and after fine-tuning. We also found that data carriers and licensing information are the most likely data categories to be memorized from pre-trained and finetuned models, while the latter is the most likely to be forgotten after fine-tuning.

Files

How_Much_Do_Code_Language_Mode... (pdf)
(pdf | 0.697 Mb)
- Embargo expired in 01-11-2025
– Personal use only – Dutch Copyright Act (Article 25fa)