Abstract
Infrastructure-As-Code (IaC) allows organizations to push dozens of software changes every day. However, human operators of that software remain a bottleneck when incidents must be addressed during software operation. Incident post-mortems show detection latencies of tens of minutes and manual recovery times of hours. The research question is therefore whether we can define an autonomous, rule-based self-healing system that addresses the auditability and safety guarantees in modern DevOps and incident management. We focus here on a rule catalog for self-healing that forms the core of the classical MAPE-K autonomic loop. The heart of the system is the reusable catalog of 20 event-condition-action (ECA) rules for incident management.