Abstract
Enterprises that manage their cloud systems through Infrastructure-as-Code (IaC) often push dozens of changes every day. Human operators create a bottleneck in managing incidents in these rapidly changing systems: incident post-mortems reveal detection latencies of tens of minutes and manual recoveries that stretch into hours. An autonomous controller is presented that provides reliable incident self-healing in a Development and Operations (DevOps) environment. This self-healing controller architecture instantiates a common autonomic computing pattern. It provides end-to-end, rule-driven self-healing for an IaC pipeline without the need for specialized hardware or proprietary IT operation platforms. It contributes to handling infrastructure incidents in a reliable and automated fashion.