Hybrid Metaheuristic Algorithms for Deep Neural Network Hyperparameter Optimization: A Comprehensive Review and Framework
Abstract
Deep Neural Networks (DNNs) have achieved state-of-the-art performance across numerous domains, yet their efficacy is intrinsically tied to the selection of optimal hyperparameters, a process known as Hyperparameter Optimization (HPO). Traditional exhaustive search methods, such as Grid Search (GS) and Random Search (RS), suffer from prohibitive computational costs and poor scalability, particularly as model complexity and dataset size increase. This review addresses the limitations of conventional HPO techniques by comprehensively analyzing the application of Metaheuristic Algorithms (MAs) in this domain. We formulate the HPO problem as a black-box global optimization task, clearly defining the objective function and the constrained search space. Furthermore, this paper introduces a structured taxonomy for Hybrid Metaheuristic Algorithms (HMAs), classifying them into sequential, parallel, and integrated architectures. We provide detailed reviews of foundational algorithms Genetic Algorithm (GA), Particle Swarm Optimization (PSO), Ant Colony Optimization (ACO), and Differential Evolution (DE) and subsequently analyze their synergistic combinations, focusing on frameworks like GA-PSO and Memetic Algorithms. To demonstrate practical efficacy, a case study involving the hyperparameter tuning of a Convolutional Neural Network (CNN) on the Canadian Intitute for Advanced Research-10 (CIFAR-10) benchmark is presented, illustrating how HMAs balance global exploration with fine-grained exploitation. This comprehensive review establishes a foundational framework for designing robust and computationally efficient HPO strategies tailored for modern Deep Learning (DL) systems, serving as a critical resource for future research.
Keywords:
Hyperparameter optimization, Deep neural networks, Hybrid metaheuristics, Genetic algorithm, Particle swarm optimizationReferences
- [1] LeCun, Y., Bengio, Y., & Hinton, G. (2015). Deep learning. Nature, 521(7553), 436–444. https://doi.org/10.1038/nature14539
- [2] Goodfellow, I. (2016). Deep learning. MIT Press. http://www.deeplearningbook.org
- [3] Smith, L. N. (2017). Cyclical learning rates for training neural networks. 2017 IEEE winter conference on applications of computer vision (WACV) (pp. 464-472). IEEE. https://doi.org/10.1109/WACV.2017.58
- [4] Feurer, M., & Hutter, F. (2019). Automated machine learning. Cham: Springer. https://doi.org/10.1007/978-3-030-05318-5
- [5] Liashchynskyi, P., & Liashchynskyi, P. (2019). Grid search, random search, genetic algorithm: A big comparison for NAS. https://doi.org/10.48550/arXiv.1912.06059
- [6] Bergstra, J., & Bengio, Y. (2012). Random search for hyper parameter optimization. Journal of machine learning research, 13(2012), 281-305. http://www.jmlr.org/papers/volume13/bergstra12a/bergstra12a.pdf
- [7] Snoek, J., Larochelle, H., & Adams, R. P. (2012). Practical bayesian optimization of machine learning algorithms. Advances in neural information processing systems, 25. https://papers.nips.cc/paper_files/paper/2012/hash/05311655a15b75fab86956663e1819cd-Abstract.html
- [8] Yang, X. S. (2010). Nature inspired metaheuristic algorithms. Luniver Press. https://dl.acm.org/doi/10.5555/1893084
- [9] Holland, J. H. (1975). Adaptation in natural and artificial systems: An introductory analysis with applications to biology, control, and artificial intelligence. University of Michigan Press. https://doi.org/10.7551/mitpress/1090.001.0001
- [10] Goldberg, D. E. (1989). Genetic algorithms in search, optimization, and machine learning. Addison-Wesley Publishing Company. https://www.amazon.com/Genetic-Algorithms-Optimization-Machine-Learning/dp/0201157675
- [11] Kennedy, J., & Eberhart, R. (1995). Particle swarm optimization. Proceedings of ICNN'95-international conference on neural networks (Vol. 4, pp. 1942-1948). IEEE. https://doi.org/10.1109/ICNN.1995.488968
- [12] Shi, Y., Eberhart, R. (1998). A modified particle swarm optimizer. Evolutionary computation proceedings (Vol. 890, pp. 69–73). IEEE. https://doi.org/10.1109/ICEC.1998.699146
- [13] Dorigo, M. (2007). Ant colony optimization. Scholarpedia, 2(3), 1461. http://www.scholarpedia.org/article/Ant_colony_optimization
- [14] Socha, K., & Dorigo, M. (2008). Ant colony optimization for continuous domains. European journal of operational research, 185(3), 1155–1173. https://doi.org/10.1016/j.ejor.2006.06.046
- [15] Storn, R., & Price, K. (1997). Differential evolution a simple and efficient heuristic for global optimization over continuous spaces. Journal of global optimization, 11(4), 341–359. https://doi.org/10.1023/A:1008202821328
- [16] Das, S., Suganthan, P. N. (2011). Differential evolution: A survey of the state-of-the-art. IEEE transactions on evolutionary computation, 15(1), 4-31. https://doi.org/10.1109/TEVC.2010.2059031
- [17] Wolpert, D. H., & Macready, W. G. (1997). No free lunch theorems for optimization. IEEE transactions on evolutionary computation, 1(1), 67–82. https://doi.org/10.1109/4235.585893
- [18] Young, S. R., Rose, D. C., Karnowski, T. P., Lim, S. H., & Patton, R. M. (2015). Optimizing deep learning hyper parameters through an evolutionary algorithm. Proceedings of the workshop on machine learning in hig performance computing environments (pp. 1–5). MLHPC. https://doi.org/10.1145/2834892.2834896
- [19] Lorenzo, P. R., Nalepa, J., Kawulok, M., Ramos, L. S., & Pastor, J. R. (2017). Particle swarm optimization for hyper parameter selection in deep neural networks. Proceedings of the genetic and evolutionary computation conference (pp. 481–488). GECCO. https://doi.org/10.1145/3071178.3071208