Loading...
Improving selection methods and data representation towards increasing explainability in evolutionary algorithms
Citations
Altmetric:
Date
2025-12
Abstract
While black-box Machine Learning (ML) models, such as deep neural networks, achieve remarkable performance across a wide range of applications, their lack of transparency has raised concerns on their application in critical areas such as healthcare, law and manufacturing industries, where interpretability has become an essential component. Genetic Programming (GP) offers a more suitable alternative. Unlike black-box models, the solution of GP presents usually a more human-readable representation, such as a mathematical expression or a tree-based structure, allowing the interpretation of the relationships between input features and predicted outputs. While an acceptable metric for interpretability is still an open issue, it is reasonable to assume that, at least for symbolic expression-based ML models, the shorter the solution, the more interpretable it is. However, when evolving solutions in GP, a common phenomenon is known as bloat, which can be defined as the increase in the size of the solutions without an equivalent improvement in fitness. In this way, handling this issue is essential to enhance the interpretability of the solutions. In addition, converting input data into a more meaningful representation can make the features and consequent solutions more interpretable. In this work, we apply lexicographic parsimony pressure to Lexicase selection by including a tie-breaking step based on size, to handle the bloat issue. Furthermore, we explore representing input data using fuzzy logic by separating the features into specific parts of their domain and associating each with a descriptive term. In this direction, we use GP to evolve Fuzzy Pattern Trees, tree-based structures in which the internal nodes are fuzzy operators, and the leaves are fuzzy features, while applying bloat control methods based on Lexicase selection. We assess our proposal using Boolean, classification, and regression problems. The results show that our approaches significantly reduced the size while maintaining similar accuracy. Moreover, we also conduct experiments using different metrics, including the area under the receiver operating characteristic curve, an important metric for evaluating and comparing different approaches in medical classification tasks, which we assess in the breast cancer data domain.
Supervisor
Description
Peer-reviewed
Publisher
University of Limerick
Citation
Files
ULRR Identifiers
Funding code
Funding Information
Sustainable Development Goals
External Link
License
Attribution-NonCommercial-ShareAlike 4.0 International
