Abstract
Process mining aims to obtain insights from event logs through the automated analyses of recorded process data in information systems, with the ultimate aim to improve business processes running in organisations. However, real-life event logs are often incomplete, noisy, or ambiguous, such as missing timestamps or having ambiguous event labels, which traditional deterministic models cannot capture. Recent process mining developments have considered uncertainty in process mining artifacts more explicitly: in logs of recorded process behaviour, uncertainty may implicitly or explicitly influence process mining outcomes, while in process models, explicit uncertainty allows analysts to interpret and value outcomes. In this paper, we provide a conceptual foundation for uncertainty in process mining by introducing a four-level specification that separately addresses uncertainty in log attributes (e.g., activity labels of events, frequencies) and model elements (e.g., service times, read guards). For each type of uncertainty, we illustrate the levels with concrete examples to help understanding and application. We then provide a structured overview of the state of the art in stochastic process mining, classified using our specification, and present key open research challenges.