Introns and Templates Matter

Rethinking Linkage in GP-GOMEA

Conference Paper (2026)
Author(s)

Johannes Koch (TU Delft - Electrical Engineering, Mathematics and Computer Science, Centrum Wiskunde & Informatica (CWI))

Tanja Alderliesten (Leiden University Medical Center)

Peter Bosman (Centrum Wiskunde & Informatica (CWI), TU Delft - Electrical Engineering, Mathematics and Computer Science)

Research Group
Algorithmics
DOI related publication
https://doi.org/10.1145/3795095.3805139 Final published version
More Info
expand_more
Publication Year
2026
Language
English
Research Group
Algorithmics
Pages (from-to)
780-789
Publisher
ACM
ISBN (electronic)
9798400724879
Event
Genetic and Evolutionary Computation Conference, GECCO 2026 (2026-07-13 - 2026-07-17), San Jose, Costa Rica
Page Views
44
Reuse Rights

Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.

Abstract

GP-GOMEA is among the state-of-the-art for symbolic regression, especially when it comes to finding small and potentially interpretable solutions. A key mechanism employed in any GOMEA variant is the exploitation of linkage, the dependencies between variables, to ensure efficient evolution. In GP-GOMEA, mutual information between node positions in GP trees has so far been used to learn linkage. For this, a fixed expression template is used. This, however, leads to introns for expressions smaller than the full template. As introns have no impact on fitness, their occurrences are not directly linked to selection. Consequently, introns can adversely affect the extent to which mutual information captures dependencies between tree nodes. To overcome this, we propose two new measures for linkage learning, one that explicitly considers introns in mutual information estimates, and one that revisits linkage learning in GP-GOMEA from a grey-box perspective, yielding a measure that needs not to be learned from the population but is derived directly from the template. Across five standard symbolic regression problems, GP-GOMEA achieves substantial improvements using both measures. We also find that the newly learned linkage structure closely reflects the template linkage structure, and that explicitly using the template structure yields the best performance overall.