Hybrid Metaheuristic–Reinforcement Learning Framework for Personalized Treatment Policy Optimization

Authors

  • Reem Al-Qahtani * Department of Computer Science, College of Computer and Information Sciences, King Saud University, Riyadh, Saudi Arabia.
  • Arjun Nair Department of Computer Science and Engineering, Indian Institute of Technology Delhi, New Delhi, India.

https://doi.org/10.48313/maa.v1i4.108

Abstract

Personalized medicine increasingly demands dynamic, patient-specific treatment strategies that adapt to evolving clinical states, yet optimizing such strategies remains computationally intractable for conventional approaches. Dynamic Treatment Regimes (DTRs) formalize this challenge as sequential decision-making under uncertainty, but existing Reinforcement Learning (RL) methods for DTR optimization suffer from sample inefficiency, susceptibility to local optima, and safety concerns when trained on limited offline clinical data. This paper introduces Metaheuristic–Reinforcement learning (METRO-RL) optimizer, a novel two-stage hybrid framework that synergistically combines Differential Evolution (DE) for global policy search with a Deep Q-Network (DQN) variant for local policy refinement. In Stage 1, a modified DE algorithm with patient-state-aware mutation operators explores the policy parameter space globally, identifying promising policy regions while respecting clinical safety constraints. In Stage 2, a dueling DQN with Conservative Q-Learning (CQL) regularization fine-tunes the best candidate policies using prioritized experience replay from offline clinical datasets. A novel patient similarity kernel based on demographic features, clinical trajectory alignment via Dynamic Time Warping (DTW), and comorbidity profiles guides both stages, enabling effective knowledge transfer across patient subgroups. We evaluate METRO-RL on two clinically significant applications: Type 2 Diabetes (T2D) management using a simulated 10,000-patient cohort based on UVA/Padova simulator parameters and sepsis treatment in the Intensive Care Unit (ICU) using the MIMIC-IV database (approximately 18,500 sepsis episodes). On T2D management, METRO-RL achieves a composite glycemic control score of 0.782 ± 0.024, representing a 16.2% improvement over standard DQN and a 18.8% improvement over the clinician policy. For sepsis treatment, METRO-RL yields an estimated 90-day mortality of 18.7% ± 2.1%, an 11.4% relative reduction compared to the clinician policy. METRO-RL consistently outperforms standalone DQN, actor-critic, Batch-Constrained Q-Learning (BCQ), CQL, and evolutionary strategy baselines across all Off-Policy Evaluation (OPE) metrics while maintaining the lowest rate of clinically contraindicated actions (0.3%). 

Keywords:

Dynamic treatment regime, Personalized medicine, Metaheuristic optimization, Reinforcement learning, Clinical decision support, Type 2 diabetes, Sepsis management

References

  1. [1] Collins, F. S., & Varmus, H. (2015). A new initiative on precision medicine. The new england journal of medicine, 372(9), 793. https://doi.org/10.1056/NEJMp1500523

  2. [2] Rajkomar, A., Dean, J., & Kohane, I. (2019). Machine learning in medicine. New england journal of medicine, 380(14), 1347–1358. https://doi.org/10.1056/NEJMra1814259

  3. [3] Murphy, S. A. (2003). Optimal dynamic treatment regimes. Journal of the royal statistical society series b: Statistical methodology, 65(2), 331–355. https://doi.org/10.1111/1467-9868.00389

  4. [4] Robins, J. M. (2004). Optimal structural nested models for optimal sequential decisions. In Proceedings of the second seattle symposium in biostatistics: Analysis of correlated data (pp. 189-326). New York, NY: Springer New York. https://doi.org/10.1007/978-1-4419-9076-1_11

  5. [5] Chakraborty, B., & Moodie, E. E. (2013). Statistical methods for dynamic treatment regimes: Reinforcement learning, causal inference, and personalized medicine. New York: Springer. https://doi.org/10.1007/978-1-4614-7428-9

  6. [6] Murphy, S. A. (2005). A generalization error for Q-learning. https://jmlr.org/papers/volume6/murphy05a/murphy05a.pdf

  7. [7] Watkins, C. J. C. H., & Dayan, P. (1992). Q-learning. Machine learning, 8(3), 279–292. https://doi.org/10.1007/BF00992698

  8. [8] Robins, J. M., Hernan, M. A., & Brumback, B. (2000). Marginal structural models and causal inference in epidemiology. Epidemiology, 11(5), 550–560. https://doi.org/10.1097/00001648-200009000-00011

  9. [9] Zhao, Y., Zeng, D., Rush, A. J., & Kosorok, M. R. (2012). Estimating individualized treatment rules using outcome weighted learning. Journal of the american statistical association, 107(499), 1106–1118. https://doi.org/10.1080/01621459.2012.695674

  10. [10] Raghu, A., Komorowski, M., Celi, L. A., Szolovits, P., & Ghassemi, M. (2017). Continuous state-space models for optimal sepsis treatment: A deep reinforcement learning approach. Machine learning for healthcare conference (pp. 147-163). PMLR. https://doi.org/10.48550/arXiv.1705.08422

  11. [11] Padmanabhan, R., Meskin, N., & Haddad, W. M. (2017). Reinforcement learning-based control of drug dosing for cancer chemotherapy treatment. Mathematical biosciences, 293, 11–20. https://doi.org/10.1016/j.mbs.2017.08.004

  12. [12] Shortreed, S. M., Laber, E., Lizotte, D. J., Stroup, T. S., Pineau, J., & Murphy, S. A. (2011). Informing sequential clinical decision-making through reinforcement learning: An empirical study. Machine learning, 84(1), 109–136. https://doi.org/10.1007/s10994-010-5229-0

  13. [13] Daskalaki, E., Diem, P., & Mougiakakou, S. G. (2013). An actor-critic based controller for glucose regulation in type 1 diabetes. Computer methods and programs in biomedicine, 109(2), 116–125. https://doi.org/10.1016/j.cmpb.2012.03.002

  14. [14] Gottesman, O., Johansson, F., Komorowski, M., Faisal, A., Sontag, D., Doshi-Velez, F., & Celi, L. A. (2019). Guidelines for reinforcement learning in healthcare. Nature medicine, 25(1), 16–18. https://doi.org/10.1038/s41591-018-0310-5

  15. [15] Levine, S., Kumar, A., Tucker, G., & Fu, J. (2020). Offline reinforcement learning: Tutorial, review, and perspectives on open problems. https://doi.org/10.48550/arXiv.2005.01643

  16. [16] Fujimoto, S., Meger, D., & Precup, D. (2019). Off-policy deep reinforcement learning without exploration. International conference on machine learning (pp. 2052-2062). PMLR. https://proceedings.mlr.press/v97/fujimoto19a/fujimoto19a.pdf

  17. [17] Kumar, A., Zhou, A., Tucker, G., & Levine, S. (2020). Conservative q-learning for offline reinforcement learning. Advances in neural information processing systems, 33, 1179–1191. https://dl.acm.org/doi/abs/10.5555/3495724.3495824

  18. [18] Achiam, J., Held, D., Tamar, A., & Abbeel, P. (2017). Constrained policy optimization. International conference on machine learning (pp. 22-31). PMLR. https://proceedings.mlr.press/v70/achiam17a.html

  19. [19] Storn, R., & Price, K. (1997). Differential evolution-A simple and efficient heuristic for global optimization over continuous spaces. Journal of global optimization, 11(4), 341–359. https://doi.org/10.1023/A:1008202821328

  20. [20] Das, S., & Suganthan, P. N. (2010). Differential evolution: A survey of the state-of-the-art. IEEE transactions on evolutionary computation, 15(1), 4–31. https://doi.org/10.1109/TEVC.2010.2059031

  21. [21] Petrovski, A., & McCall, J. (2001). Multi-objective optimisation of cancer chemotherapy using evolutionary algorithms. International conference on evolutionary multi-criterion optimization (pp. 531-545). Berlin, Heidelberg: Springer Berlin Heidelberg. https://doi.org/10.1007/3-540-44719-9_37

  22. [22] Holdsworth, C., Kim, M., Liao, J., & Phillips, M. H. (2010). A hierarchical evolutionary algorithm for multiobjective optimization in IMRT. Medical physics, 37(9), 4986–4997. https://doi.org/10.1118/1.3478276

  23. [23] Abo-Hammour, Z. S., Samhouri, A. D., & Mubarak, Y. (2014). Continuous genetic algorithm as a novel solver for Stokes and nonlinear Navier Stokes problems. Mathematical problems in engineering, 2014(1), 649630. https://doi.org/10.1155/2014/649630

  24. [24] Stanley, K. O., Clune, J., Lehman, J., & Miikkulainen, R. (2019). Designing neural networks through neuroevolution. Nature machine intelligence, 1(1), 24–35. https://doi.org/10.1038/s42256-018-0006-z

  25. [25] Jaderberg, M., Dalibard, V., Osindero, S., Czarnecki, W. M., Donahue, J., Razavi, A., ... & Kavukcuoglu, K. (2017). Population based training of neural networks. https://doi.org/10.48550/arXiv.1711.09846

  26. [26] Salimans, T., Ho, J., Chen, X., Sidor, S., & Sutskever, I. (2017). Evolution strategies as a scalable alternative to reinforcement learning. https://doi.org/10.48550/arXiv.1703.03864

  27. [27] Khadka, S., & Tumer, K. (2018). Evolution-guided policy gradient in reinforcement learning. Advances in neural information processing systems, 31. https://dl.acm.org/doi/10.5555/3326943.3327053

  28. [28] Robins, J. (1986). A new approach to causal inference in mortality studies with a sustained exposure period—application to control of the healthy worker survivor effect. Mathematical modelling, 7(9–12), 1393–1512. https://doi.org/10.1016/0270-0255(86)90088-6

  29. [29] Zhao, Y., Zeng, D., Socinski, M. A., & Kosorok, M. R. (2011). Reinforcement learning strategies for clinical trials in nonsmall cell lung cancer. Biometrics, 67(4), 1422–1433. https://doi.org/10.1111/j.1541-0420.2011.01572.x

  30. [30] Moodie, E. E. M., Dean, N., & Sun, Y. R. (2014). Q-learning: Flexible learning about useful utilities. Statistics in biosciences, 6(2), 223–243. https://doi.org/10.1007/s12561-013-9103-z

  31. [31] Lei, H., Nahum-Shani, I., Lynch, K., Oslin, D., & Murphy, S. A. (2012). A" SMART" design for building individualized treatment sequences. Annual review of clinical psychology, 8(1), 21–48. https://doi.org/10.1146/annurev-clinpsy-032511-143152

  32. [32] Almirall, D., Nahum-Shani, I., Sherwood, N. E., & Murphy, S. A. (2014). Introduction to SMART designs for the development of adaptive interventions: With application to weight loss research. Translational behavioral medicine, 4(3), 260–274. https://doi.org/10.1007/s13142-014-0265-0

  33. [33] Zhang, B., Tsiatis, A. A., Laber, E. B., & Davidian, M. (2013). Robust estimation of optimal dynamic treatment regimes for sequential treatment decisions. Biometrika, 100(3), 681–694. https://doi.org/10.1093/biomet/ast014

  34. [34] Zhao, Y.-Q., Zeng, D., Laber, E. B., & Kosorok, M. R. (2015). New statistical learning methods for estimating optimal dynamic treatment regimes. Journal of the american statistical association, 110(510), 583–598. https://doi.org/10.1080/01621459.2014.937488

  35. [35] Laber, E. B., & Zhao, Y.Q. (2015). Tree-based methods for individualized treatment regimes. Biometrika, 102(3), 501–514. https://doi.org/10.1093/biomet/asv028

  36. [36] Zhang, Y., Laber, E. B., Davidian, M., & Tsiatis, A. A. (2018). Interpretable dynamic treatment regimes. Journal of the american statistical association, 113(524), 1541–1549. https://doi.org/10.1080/01621459.2017.1345743

  37. [37] Tsiatis, A. A., Davidian, M., Holloway, S. T., & Laber, E. B. (2019). Dynamic treatment regimes: Statistical methods for precision medicine. CRC Press. https://doi.org/10.1201/9780429192692

  38. [38] Komorowski, M., Celi, L. A., Badawi, O., Gordon, A. C., & Faisal, A. A. (2018). The artificial intelligence clinician learns optimal treatment strategies for sepsis in intensive care. Nature medicine, 24(11), 1716–1720. https://doi.org/10.1038/s41591-018-0213-5

  39. [39] Prasad, N., Cheng, L.F., Chivers, C., Draugelis, M., & Engelhardt, B. E. (2017). A reinforcement learning approach to weaning of mechanical ventilation in intensive care units. https://doi.org/10.48550/arXiv.1704.06300

  40. [40] Nemati, S., Ghassemi, M. M., & Clifford, G. D. (2016). Optimal medication dosing from suboptimal clinical examples: A deep reinforcement learning approach. 2016 38th annual international conference of the IEEE engineering in medicine and biology society (EMBC) (pp. 2978-2981). IEEE. https://doi.org/10.1109/EMBC.2016.7591355

  41. [41] Fox, I., Lee, J., Pop-Busui, R., & Wiens, J. (2020). Deep reinforcement learning for closed-loop blood glucose control. Machine learning for healthcare conference (pp. 508-536). PMLR. https://proceedings.mlr.press/v126/fox20a.html

  42. [42] Wu, Y., Tucker, G., & Nachum, O. (2019). Behavior regularized offline reinforcement learning. https://doi.org/10.48550/arXiv.1911.11361

  43. [43] Tessler, C., Mankowitz, D. J., & Mannor, S. (2018). Reward constrained policy optimization. https://doi.org/10.48550/arXiv.1805.11074

  44. [44] Huang, S., Papernot, N., Goodfellow, I., Duan, Y., & Abbeel, P. (2017). Adversarial attacks on neural network policies. https://doi.org/10.48550/arXiv.1702.02284

  45. [45] Hanset, A., Meskens, N., & Duvivier, D. (2010). Using constraint programming to schedule an operating theatre. 2010 IEEE workshop on health care management (WHCM) (pp. 1-6). IEEE. https://doi.org/10.1109/WHCM.2010.5441245

  46. [46] Anter, A. M., & Hassenian, A. E. (2019). CT liver tumor segmentation hybrid approach using neutrosophic sets, fast fuzzy c-means and adaptive watershed algorithm. Artificial intelligence in medicine, 97, 105–117. https://doi.org/10.1016/j.artmed.2018.11.007

  47. [47] Li, Y., Yao, J., & Yao, D. (2004). Automatic beam angle selection in IMRT planning using genetic algorithm. Physics in medicine & biology, 49(10), 1915–1932. https://doi.org/10.1088/0031-9155/49/10/007

  48. [48] Veng-Pedersen, P., Gobburu, J. V. S., Meyer, M. C., & Straughn, A. B. (2000). Carbamazepine level-A in vivo-in vitro correlation (IVIVC): A scaled convolution based predictive approach. Biopharmaceutics & drug disposition, 21(1), 1–6. https://doi.org/10.1002/1099-081X(200001)21:1%3C1::AID-BDD207%3E3.0.CO;2-D

  49. [49] Marchetti, G., Barolo, M., Jovanovic, L., Zisser, H., & Seborg, D. E. (2008). An improved PID switching control strategy for type 1 diabetes. IEEE transactions on biomedical engineering, 55(3), 857–865. https://doi.org/10.1109/TBME.2008.915665

  50. [50] Moles, C. G., Mendes, P., & Banga, J. R. (2003). Parameter estimation in biochemical pathways: A comparison of global optimization methods. Genome research, 13(11), 2467–2474. https://doi.org/10.1101/gr.1262503

  51. [51] Riquelme, N., Von Lücken, C., & Baran, B. (2015). Performance metrics in multi-objective optimization. 2015 Latin American computing conference (CLEI) (pp. 1-11). IEEE. https://doi.org/10.1109/CLEI.2015.7360024

  52. [52] Stanley, K. O., & Miikkulainen, R. (2002). Evolving neural networks through augmenting topologies. Evolutionary computation, 10(2), 99–127. https://doi.org/10.1162/106365602320169811

  53. [53] Stanley, K. O., D’Ambrosio, D. B., & Gauci, J. (2009). A hypercube-based encoding for evolving large-scale neural networks. Artificial life, 15(2), 185–212. https://doi.org/10.1162/artl.2009.15.2.15202

  54. [54] Pourchot, A., & Sigaud, O. (2018). CEM-RL: Combining evolutionary and gradient-based methods for policy search. https://doi.org/10.48550/arXiv.1810.01222

  55. [55] Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., & Meger, D. (2018). Deep reinforcement learning that matters. Proceedings of the AAAI conference on artificial intelligence (pp. 3207–3214). AAAI Press. https://doi.org/10.1609/aaai.v32i1.11694

  56. [56] Wang, Z., Schaul, T., Hessel, M., Hasselt, H., Lanctot, M., & Freitas, N. (2016). Dueling network architectures for deep reinforcement learning. International conference on machine learning (pp. 1995-2003). PMLR. https://proceedings.mlr.press/v48/wangf16.html

  57. [57] Schaul, T., Quan, J., Antonoglou, I., & Silver, D. (2015). Prioritized experience replay. https://doi.org/10.48550/arXiv.1511.05952

  58. [58] Precup, D., Sutton, R. S., & Singh, S. (2000). Eligibility traces for off-policy policy evaluation. Proceedings of the seventeenth international conference on machine learning (ICML) (pp. 759–766). Morgan Kaufmann. https://hdl.handle.net/20.500.14394/10401

  59. [59] Le, H., Voloshin, C., & Yue, Y. (2019). Batch policy learning under constraints. Proceedings of the 36th international conference on machine learning, 97, 3703–3712. https://proceedings.mlr.press/v97/le19a.html

  60. [60] Jiang, N., & Li, L. (2016). Doubly robust off-policy value evaluation for reinforcement learning. Proceedings of the 33rd international conference on machine learning (PP. 652–661 ). PMLR. https://proceedings.mlr.press/v48/jiang16.pdf

  61. [61] Kovatchev, B. P., Breton, M., Dalla Man, C., & Cobelli, C. (2009). In silico preclinical trials: A proof of concept in closed-loop control of type 1 diabetes. Journal of Diabetes science and technology, 3(1), 44–55. https://doi.org/10.1177/193229680900300106

  62. [62] Man, C. D., Micheletto, F., Lv, D., Breton, M., Kovatchev, B., & Cobelli, C. (2014). The UVA/PADOVA type 1 diabetes simulator: New features. Journal of Diabetes science and technology, 8(1), 26–34. https://doi.org/10.1177/1932296813514502

  63. [63] Holman, R. R., Paul, S. K., Bethel, M. A., Matthews, D. R., & Neil, H. A. W. (2008). 10-year follow-up of intensive glucose control in type 2 diabetes. New england journal of medicine, 359(15), 1577–1589. https://doi.org/10.1056/NEJMoa0806470

  64. [64] Johnson, A. E., Bulgarelli, L., Shen, L., Gayles, A., Shammout, A., Horng, S., ... & Mark, R. G. (2023). MIMIC-IV, a freely accessible electronic health record dataset. Scientific data, 10(1), 1. https://doi.org/10.1038/s41597-022-01899-x

  65. [65] Singer, M., Deutschman, C. S., Seymour, C. W., Shankar-Hari, M., Annane, D., Bauer, M., ... & Angus, D. C. (2016). The third international consensus definitions for sepsis and septic shock (Sepsis-3). JAMA, 315(8), 801-810. https://doi.org/10.1001/jama.2016.0287

  66. [66] Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., & Riedmiller, M. (2013). Playing Atari with deep reinforcement learning. https://doi.org/10.48550/arXiv.1312.5602

  67. [67] Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T. P., Harley, T., Silver, D., & Kavukcuoglu, K. (2016). Asynchronous methods for deep reinforcement learning. Proceedings of the 33rd international conference on machine learning (pp. 1928–1937). PMLR. https://proceedings.mlr.press/v48/mniha16.html

  68. [68] Hansen, N. (2006). The CMA evolution strategy: A comparing review. Towards a new evolutionary computation: Advances in the estimation of distribution algorithms, 75–102. https://doi.org/10.1007/3-540-32494-1_4

  69. [69] Thomas, P., & Brunskill, E. (2016). Data-efficient off-policy policy evaluation for reinforcement learning. International conference on machine learning (pp. 2139-2148). PMLR. https://proceedings.mlr.press/v48/thomasa16.html

  70. [70] Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464), 447–453. https://doi.org/10.1126/science.aax2342

  71. [71] Grote, T., & Berens, P. (2020). On the ethics of algorithmic decision-making in healthcare. Journal of medical ethics, 46(3), 205–211. https://doi.org/10.1136/medethics-2019-105586

  72. [72] Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., ... & Mordatch, I. (2021). Decision transformer: Reinforcement learning via sequence modeling. Advances in neural information processing systems, 34, 15084-15097. https://proceedings.neurips.cc/paper/2021/hash/7f489f642a0ddb10272b5c31057f0663-Abstract.html

  73. [73] Kostrikov, I., Nair, A., & Levine, S. (2021). Offline reinforcement learning with implicit q-learning. https://doi.org/10.48550/arXiv.2110.06169

  74. [74] Wang, Z., Hunt, J. J., & Zhou, M. (2022). Diffusion policies as an expressive policy class for offline reinforcement learning. https://doi.org/10.48550/arXiv.2208.06193

Published

2024-12-24

How to Cite

Al-Qahtani, R. ., & Nair, A. . (2024). Hybrid Metaheuristic–Reinforcement Learning Framework for Personalized Treatment Policy Optimization. Metaheuristic Algorithms With Applications, 1(4), 417-439. https://doi.org/10.48313/maa.v1i4.108

Similar Articles

41-50 of 68

You may also start an advanced similarity search for this article.