An Automated Framework for Self-Adaptive Quality Assurance in Software Systems Using Policy-Based Reinforcement Learning

Authors

  • Hayfaa Subhi Malallah Information Technology Department, Technical College of Duhok, Duhok Polytechnic University, Kurdistan Region, Iraq https://orcid.org/0009-0006-8999-3097
  • Ayad Abdulrahman Saleem Computer Information Systems Department, Technical College of Zakho, Duhok Polytechnic University, Kurdistan Region, Iraq

DOI:

https://doi.org/10.65542/djei.v2i3.57

Keywords:

Self-adaptation, reinforcement learning, software engineering, autonomic computing, quality assurance

Abstract

Self-adaptive information systems must keep their quality requirements while their environment changes at run time. Building the adaptation logic by hand is difficult, because design-time uncertainty makes it impossible to foresee every environmental change. Online Reinforcement Learning (RL) can build this logic automatically. However, the value-based RL methods used so far have two practical limits: the exploration rate must be tuned by hand, and continuous states must be discretised by hand. This paper presents a framework that removes both limits by using policy-based RL. The Analyze and Plan phases of the MAPE-K loop are redefined as a single policy-based decision step, and Proximal Policy Optimization (PPO) is applied for online adaptation in continuous and discrete action spaces. The framework is evaluated on two systems: a self-adaptive web application and a predictive process-monitoring system. Across four workload patterns and two concept drifts, the framework learns effective policies without exploration tuning or state discretisation. It improves on a value-based baseline (DQN) with statistical significance and performs on par with a maximum-entropy method (SAC) under the tested settings, while keeping good sample efficiency and stability.

References

1. Sivathapandi, P.; Sudharsanam, S.R.; Manivannan, P. Development of Adaptive Machine Learning-Based Testing Strategies for Dynamic Microservices Performance Optimization. Journal of Science & Technology 2023, 4, 102–137.

2. Sur David Reinforcement Learning for Adaptive Test Suite Management in Salesforce DevOps. 2025.

3. Klein, C.; Maggio, M.; Arzén, K.E.; Hernández-Rodriguez, F. Brownout: Building More Robust Cloud Applications. Proceedings of the 36th international conference on software engineering 2014, 700–711, doi:10.1145/2568225.2568227. DOI: https://doi.org/10.1145/2568225.2568227

4. Metzger, A.; Föcker, F. Predictive Business Process Monitoring Considering Reliability Estimates. International Conference on Advanced Information Systems Engineering 2017, 10253 LNCS, 445–460, doi:10.1007/978-3-319-59536-8_28. DOI: https://doi.org/10.1007/978-3-319-59536-8_28

5. Palm, A.; Metzger, A.; Pohl, K. Online Reinforcement Learning for Self-Adaptive Information Systems. International conference on advanced information systems engineering 2020, 12127 LNCS, 169–184, doi:10.1007/978-3-030-49435-3_11. DOI: https://doi.org/10.1007/978-3-030-49435-3_11

6. Sada, A. Ben; Khelloufi, A.; Naouri, A.; Ning, H.; N Aung Multi-Agent Deep Reinforcement Learning-Based Inference Task Scheduling and Offloading for Maximum Inference Accuracy under Time and Energy Constraints. Electronics (Basel). 2024.

7. Lora, C.; Pavan, P.; International, N.M.-2024 15th; 2024, undefined Scalable Multi-Agent Reinforcement Learning Architectures for Cloud-Based Data Centers. 2024 15th International Conference on Computing Communication and Networking Technologies (ICCCNT) 2024. DOI: https://doi.org/10.1109/ICCCNT61001.2024.10725653

8. Altin, N.; Eyimaya, S.; Energies, A.N.-; 2023, undefined Multi-Agent-Based Controller for Microgrids: An Overview and Case Study. Energies (Basel). 2023. DOI: https://doi.org/10.3390/en16052445

9. Li, Y.; Han, S.; Wang, S.; M Wang Collaborative Evolution of Intelligent Agents in Large-Scale Microservice Systems. 2025 4th International Conference on Electronic Information Technology (EIT) 2025. DOI: https://doi.org/10.1109/EIT67313.2025.11231817

10. Tong, G.; Meng, C.; Song, S.; … M.P.-2023 I. international; 2023, undefined Gma: Graph Multi-Agent Microservice Autoscaling Algorithm in Edge-Cloud Environment. 2023 IEEE international conference on web services (ICWS) 2023. DOI: https://doi.org/10.1109/ICWS60048.2023.00058

11. Zhang, W.; Guo, H.; Yang, J.; Tian, Z.; Zhang, Y.; Yan, C.; Li, Z.; Li, T.; Shi, X.; Zheng, L.; et al. MABC: Multi-Agent Blockchain-Inspired Collaboration for Root Cause Analysis in Micro-Services Architecture. Findings of the Association for Computational Linguistics: EMNLP 2024 2024, 4017–4033. DOI: https://doi.org/10.18653/v1/2024.findings-emnlp.232

12. Kurunthachalam, A. Enhancing AI-Driven Software Optimization with Attention-Based Memory Transformers and Graph Multi-Agent RL. International Journal of Scientific Engineering and Science 2025.

13. Shyam, G.; on, P.B.-2023 I.I.C.; 2023, undefined Multi-Agent Systems for Resource Allocation in Cloud Computing. 2023 IEEE International Conference on Contemporary Computing and Communications (InC4) 2023. DOI: https://doi.org/10.1109/InC457730.2023.10262945

14. Barua, B.; Kaiser, M.S. AI-Driven Resource Allocation Framework for Microservices in Hybrid Cloud Platforms. arXiv preprint arXiv:2412.02610 2024.

15. Wang, Y.; Xing, S. Reinforcement Learning for Dynamic and Predictive CPU Resource Management in Cloud Computing. Journal of Data Analysis and Information Processing 2025, 13, 255–268, doi:10.4236/JDAIP.2025.133015. DOI: https://doi.org/10.4236/jdaip.2025.133015

16. Rao, K.S.; Rao, A.A.; Raju, P.R. Design of an Iterative Adaptive Method for Volatility-Aware Test Case Prioritization in Rapidly Evolving Software Systems. MethodsX 2025, 15, doi:10.1016/J.MEX.2025.103582. DOI: https://doi.org/10.1016/j.mex.2025.103582

17. Rajasekar, V.R.; Santhi, G. Pervasive Auto-Scaling Method for Improving the Quality of Resource Allocation in Cloud Platforms. Big Data and Cognitive Computing 2025, 9, doi:10.3390/BDCC9110294. DOI: https://doi.org/10.3390/bdcc9110294

18. Kallel, A.; Rekik, M.; Khemakhem, M. A Deep Reinforcement Learning-Based Optimization Approach for Containerized Microservice Scheduling in Hybrid Fog/Cloud Environments. Eng. Appl. Artif. Intell. 2025, 141, doi:10.1016/J.ENGAPPAI.2024.109745. DOI: https://doi.org/10.1016/j.engappai.2024.109745

19. Zhang, R.; Hou, J.; Walter, F.; Gu, S.; Guan, J.; Röhrbein, F.; Du, Y.; Cai, P.; Chen, G.; Knoll, A. Multi-Agent Reinforcement Learning for Autonomous Driving: A Survey. arXiv preprint arXiv:2408.09675 2024.

20. Malki, A.; Malki, M.; Benslimane, S.-M. Data Microservice Composition Optimization Using Deep Reinforcement Learning. Future Generation Computer Systems 2026, 178, 108290, doi:10.1016/J.FUTURE.2025.108290. DOI: https://doi.org/10.1016/j.future.2025.108290

21. Amato, A.; Morelli, A.; M Fogli Multi-Agent Reinforcement Learning for Distributed Workflow Orchestration at the Tactical Edge. MILCOM 2024-2024 IEEE Military Communications Conference (MILCOM) 2024. DOI: https://doi.org/10.1109/MILCOM61039.2024.10773787

22. Hasselt, H. Van; Guez, A.; D Silver Deep Reinforcement Learning with Double Q-Learning. Proceedings of the AAAI conference on artificial intelligence 2016. DOI: https://doi.org/10.1609/aaai.v30i1.10295

23. Schulman, J.; Wolski, F.; Dhariwal, P. Proximal Policy Optimization Algorithms. arXiv preprint arXiv:1707.06347 2017.

24. Boiko Oleksandr Risk Management in High-Load Service Infrastructure Using AI-Based Predictive Models. Актуальні питання економічних наук 2025, doi:10.5281/zenodo.17141025.

25. Mehdi Syed, A.A.; Anazagasty, E. AI-Driven Infrastructure Automation: Leveraging AI and ML for Self-Healing and Auto-Scaling Cloud Environments. International Journal of Artificial Intelligence, Data Science, and Machine Learning 2024, 5, 32–43, doi:10.63282/3050-9262.IJAIDSML-V5I1P104. DOI: https://doi.org/10.63282/3050-9262.IJAIDSML-V5I1P104

26. Cui, L.; Shi, T.; Lu, R.; Zhang, T. Autoscaling in Mobile Edge Computing Based on Multi-Agent Reinforcement Learning. Proceedings of the 2023 9th International Conference on Communication and Information Processing 2023, 520–527, doi:10.1145/3638884.3638966. DOI: https://doi.org/10.1145/3638884.3638966

27. HUNKO, I. Adaptive Approaches to Software Testing with Embedded Artificial Intelligence in Dynamic Environments. International Journal of Current Science Research and Review 2025, 08, doi:10.47191/IJCSRR/V8-I5-10. DOI: https://doi.org/10.47191/ijcsrr/V8-i5-10

28. Kephart, J.O.; Chess, D.M. The Vision of Autonomic Computing. Computer (Long. Beach. Calif). 2003, 36, 41–50, doi:10.1109/MC.2003.1160055. DOI: https://doi.org/10.1109/MC.2003.1160055

29. Schulman, J.; Moritz, P.; Levine, S.; Jordan, M.I.; Abbeel, P. High-Dimensional Continuous Control Using Generalized Advantage Estimation. arXiv preprint arXiv:1506.02438 2016.

30. Mnih, V.; Puigdomènech Badia, A.; Mirza, M.; Graves, A.; Harley, T.; Lillicrap, T.P.; Silver, D.; Kavukcuoglu, K. Asynchronous Methods for Deep Reinforcement Learning. International conference on machine learning 2016.

31. R. Hogendoorn Workload Patterns for Self-Adaptive Systems, Delft University of Technology, 2012.

32. Urdaneta, G.; Pierre, G.; van Steen, M. Wikipedia Workload Analysis for Decentralized Hosting. Computer Networks 2009, 53, 1830–1845, doi:10.1016/J.COMNET.2009.02.019. DOI: https://doi.org/10.1016/j.comnet.2009.02.019

33. Dulac-Arnold, G.; Evans, R.; Van Hasselt, H.; Sunehag, P.; Lillicrap, T.; Hunt, J.; Mann, T.; Weber, T.; Degris, T.; Coppin, B.; et al. Deep Reinforcement Learning in Large Discrete Action Spaces. arXiv preprint arXiv:1512.07679 2016.

34. Jin, L.; Tang, M.; Pan, J.; Zhang, M. Asynchronous Fractional Multi-Agent Deep Reinforcement Learning for Age-Minimal Mobile Edge Computing. arXiv preprint arXiv:2409.16832 2024. DOI: https://doi.org/10.1609/aaai.v38i11.29192

Downloads

Published

2026-07-12

How to Cite

Subhi Malallah, H., & Abdulrahman Saleem, A. (2026). An Automated Framework for Self-Adaptive Quality Assurance in Software Systems Using Policy-Based Reinforcement Learning. Dasinya Journal for Engineering and Informatics, 2(3). https://doi.org/10.65542/djei.v2i3.57

Similar Articles

1 2 3 > >> 

You may also start an advanced similarity search for this article.