Risk management in high-load service infrastructure using AI-based predictive models
DOI:
https://doi.org/10.5281/zenodo.17141025Keywords:
failure prediction, system anomalies, machine learning, load balancing, cloud platforms, infrastructure resilience, resource management.Abstract
High-load information services require effective risk management to ensure business continuity and maintain stable user service. The increasing volume of data and the complexity of distributed computing systems pose potential threats to technical failures and performance degradation. This study aims to develop a methodology for risk management in high-load service infrastructure using artificial intelligence models to predict anomalies and potential failures. The methodology involves collecting and analyzing server metrics, application logs, and real-time user load. Machine learning algorithms, including time series analysis, clustering, and neural networks, are applied to identify patterns and forecast critical events. Mechanisms for automatic load balancing and scaling of computational resources in cloud environments based on predictive models have been developed. The results demonstrate that the application of AI-based predictive models significantly reduces response time to anomalies and failures, improves system component performance, and decreases the likelihood of technical outages during peak loads. The proposed algorithms provide more uniform resource allocation, enhance the fault tolerance of distributed databases and server clusters, and ensure continuous user access to services. Analysis shows that integrating predictive models into risk management reduces redundant resource costs and optimizes scaling in cloud infrastructures. The conclusions highlight the practical significance of intelligent approaches to risk management in high-load systems. The proposed methodology enhances service reliability and stability, enables a prompt response to potential failures, and optimizes the utilization of computational and network resources. The results serve as a foundation for developing risk management strategies in distributed systems, enhancing infrastructure resilience, and supporting business continuity in cloud platforms.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2025 Олександр Анатолійович Бойко

This work is licensed under a Creative Commons Attribution 4.0 International License.