Harris Hawks Optimizer for High-Dimensional Biomarker Discovery: A Statistically Rigorous Comparative Study of Eleven Metaheuristic Algorithms on Leukemia Gene Expression Data
Abstract
The rapid accumulation of high-throughput omics data has transformed disease diagnosis, yet the extreme dimensionality of gene expression profiles continues to impose severe computational and statistical obstacles for biomarker discovery. Although many nature-inspired metaheuristic algorithms have been proposed for feature selection, the literature lacks rigorous, large-scale, statistically validated comparative studies on noisy, epistatic biological landscapes. This study presents an exhaustive evaluation of the Harris Hawks Optimizer (HHO) against ten state-of-the-art metaheuristics for binary feature selection on the benchmark Leukemia (ALL versus AML) microarray dataset. A binary HHO wrapper built around a Support Vector Machine with radial basis function kernel is benchmarked against ten competitors under strictly unified budgets of 30 agents, 100 iterations, 30 independent runs, and stratified 10-fold cross-validation. Algorithms are assessed on accuracy, F1-score, Matthews Correlation Coefficient, area under the curve, subset size, and runtime, with significance verified by Wilcoxon signed-rank and Friedman tests. HHO attained 98.61% accuracy, an AUC of 0.992, and an MCC of 0.971 using only 12.4 genes on average, significantly outperforming every competitor at p < 0.05 and ranking first under the Friedman test with a mean rank of 1.00. Biological enrichment analysis confirmed that the selected genes are established leukemia drivers. The dynamic escape energy mechanism of HHO provides an exceptional exploration-exploitation balance for high-dimensional biological feature selection, offering a reliable, interpretable, and clinically translatable biomarker discovery tool.