Software System Performance and ReliabilitySafety Systems Engineering in AutonomyInfrastructure Resilience and Vulnerability Analysis

Niyazi Gökberk Gündüz, Florian Hofer, E. Guerra, Nabil El Ioini, Claus Pahl

2026.1.1IEEE SOFTWARE

DOI: 10.1109/ms.2026.3676644

Abstract

Enterprises that manage their cloud systems through <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Infrastructure-as-Code</i> (IaC) often push dozens of changes every day. Human operators create a bottleneck in managing incidents in these rapidly changing systems: incident post-mortems reveal detection latencies of tens of minutes and manual recoveries that stretch into hours. An autonomous controller is presented that provides reliable incident self-healing in a <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Development and Operations</i> (DevOps) environment. This self-healing controller architecture instantiates a common autonomic computing pattern. It provides end-to-end, rule-driven self-healing for an IaC pipeline without the need for specialized hardware or proprietary IT operation platforms. It contributes to handling infrastructure incidents in a reliable and automated fashion.

Citation format

GÜNDÜZ, Niyazi Gökberk, et al. Automated incident management using infrastructure-as-code. IEEE SOFTWARE, 2026: 1–7.