Metamorphic-Based Many-Objective Distillation of LLMs for Code-related Tasks

None, None

Metamorphic-Based Many-Objective Distillation of LLMs for Code-related Tasks

Conference Paper (2025)

Author(s)

A. Panichella (TU Delft - Software Engineering)

Research Group

Software Engineering

DOI related publication

https://doi.org/10.1109/ICSE55347.2025.00230

Large Language Models Sustainability Knowledge Distillation Search-based Software Engineering Green AI Many-Objective Optimisation Metamorphic Testing AI for SE

To reference this document use:

https://resolver.tudelft.nl/uuid:482693bf-be60-4ab9-b3e9-938d8cb178a9

More Info

expand_more

Publication Year

2025

Language

English

Research Group

Software Engineering

Bibliographical Note

Green Open Access added to TU Delft Institutional Repository as part of the Taverne amendment. More information about this copyright law amendment can be found at https://www.openaccess.nl. Otherwise as indicated in the copyright section: the publisher is the copyright holder of this work and the author uses the Dutch legislation to make this work public. @en

Pages (from-to)

1001-1013

Publisher

IEEE / ACM

ISBN (electronic)

9798331505691

Reuse Rights

Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.

Abstract

Knowledge distillation compresses large language models (LLMs) into more compact and efficient versions that achieve similar accuracy on code-related tasks. However, as we demonstrate in this study, compressed models are four times less robust than the original LLMs when evaluated with metamorphic code. They exhibit a 440% higher probability of misclassifying code clones due to minor changes in the code fragment under analysis, such as replacing parameter names with synonyms. To address this issue, we propose Morph, a novel method that combines metamorphic testing with many-objective optimization for a robust distillation of LLMs for code. Morph efficiently explores the models' configuration space and generates Paretooptimal models that effectively balance accuracy, efficiency, and robustness to metamorphic code. Metamorphic testing measures robustness as the number of code fragments for which a model incorrectly makes different predictions between the original and their equivalent metamorphic variants (prediction flips). We evaluate Morph on two tasks-code clone and vulnerability detection-targeting CodeBERT and GraphCodeBERT for distillation. Our comparison includes Morph, the state-of-theart distillation method AVATAR, and the fine-tuned non-distilled LLMs. Compared to Avatar, Morph produces compressed models that are (i) 47% more robust, (ii) 25% more efficient (fewer floating-point operations), while maintaining (iii) equal or higher accuracy (up to +6%), and (iv) similar model size.

Files

Main.pdf

(pdf | 0.511 Mb)

Metamorphic-Based_Many-Objecti... (pdf)

(pdf | 0.821 Mb)

- Embargo expired in 05-01-2026

License info not available