Chain-of-Thought LLM-Based Code Translation
Using Chain-of-Thought to Improve LLM-Based Translation of C++ to Java
B.I. Gunev (TU Delft - Electrical Engineering, Mathematics and Computer Science)
S.S. Chakraborty – Mentor (TU Delft - Electrical Engineering, Mathematics and Computer Science)
P. Pawelczak – Mentor (TU Delft - Electrical Engineering, Mathematics and Computer Science)
A. van Deursen – Graduation committee member (TU Delft - Electrical Engineering, Mathematics and Computer Science)
More Info
expand_more
Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.
Abstract
Large language models (LLMs) have recently gained significant traction for their ability to assist with day-to-day tasks, especially coding-related tasks. At the same time, cybersecurity threats are becoming increasingly complex, sometimes requiring thorough analysis by multiple experts in the field. Therefore, it is valuable to assess the ability of LLMs to aid in malware inspection.
This research focuses on the use of LLMs as a code translation tool for beginner malware researchers. Novices can use such a tool to translate malware source code from C++ to Java, helping them understand its functionality by making the code more readable. Two zero-shot prompting frameworks are presented and evaluated for their effectiveness, achieving syntactically correct outputs at rates between 62.2% and 63.9%. Of these outputs, between 39.0% and 42.9% produce functionally preserved translations.