Not All Local LLMs Are Equal
A Benchmark of Energy and Performance
Simão Cunha (University of Minho)
Francisco Ribeiro (New York University Abu Dhabi)
Luís Cruz (TU Delft - Electrical Engineering, Mathematics and Computer Science)
João Saraiva (University of Minho)
More Info
expand_more
Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.
Abstract
The rapid adoption of Large Language Models (LLMs) is transforming research, education, software development and everyday life. As their use grows, so does the diversity of available models, from general-purpose to code-oriented LLMs that can run both in data centers and on edge devices. Several benchmarks have emerged to evaluate their performance in code generation and completion tasks, yet their energy and time efficiency remain underexplored. This paper evaluates five local LLMs on HumanEval-X and MBPP+ to analyze their accuracy, runtime and energy consumption under CPU-only inference, reflecting realistic on-device deployment scenarios where GPUs are unavailable. The results reveal clear trade-offs between effectiveness and efficiency: while some models achieve higher accuracy, others deliver comparable results with substantially lower energy use. In particular, 3-shot prompting consistently improves runtime and energy efficiency compared to 0-shot, without sacrificing code quality. These findings emphasize that prompt design and model selection must be considered together when deploying LLMs for coding tasks and call for the creation of practical prompt-efficiency guidelines to support more sustainable and efficient use of local LLMs.