Not All Local LLMs Are Equal

A Benchmark of Energy and Performance

Conference Paper (2026)
Author(s)

Simão Cunha (University of Minho)

Francisco Ribeiro (New York University Abu Dhabi)

Luís Cruz (TU Delft - Electrical Engineering, Mathematics and Computer Science)

João Saraiva (University of Minho)

Research Group
Software Engineering
DOI related publication
https://doi.org/10.1145/3786148.3788630 Final published version
More Info
expand_more
Publication Year
2026
Language
English
Research Group
Software Engineering
Pages (from-to)
99-106
Publisher
ACM
ISBN (electronic)
9798400723810
Event
10th International Workshop on Green and Sustainable Software, GREENS 2026 (2026-04-12 - 2026-04-18), Rio de Janeiro, Brazil
Page Views
42
Reuse Rights

Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.

Abstract

The rapid adoption of Large Language Models (LLMs) is transforming research, education, software development and everyday life. As their use grows, so does the diversity of available models, from general-purpose to code-oriented LLMs that can run both in data centers and on edge devices. Several benchmarks have emerged to evaluate their performance in code generation and completion tasks, yet their energy and time efficiency remain underexplored. This paper evaluates five local LLMs on HumanEval-X and MBPP+ to analyze their accuracy, runtime and energy consumption under CPU-only inference, reflecting realistic on-device deployment scenarios where GPUs are unavailable. The results reveal clear trade-offs between effectiveness and efficiency: while some models achieve higher accuracy, others deliver comparable results with substantially lower energy use. In particular, 3-shot prompting consistently improves runtime and energy efficiency compared to 0-shot, without sacrificing code quality. These findings emphasize that prompt design and model selection must be considered together when deploying LLMs for coding tasks and call for the creation of practical prompt-efficiency guidelines to support more sustainable and efficient use of local LLMs.