نوع مقاله : مقاله مروری
عنوان مقاله English
نویسندگان English
The scarcity of real-world data, driven by privacy constraints, data collection costs, and poor data quality, poses a major challenge to the development and evaluation of process mining algorithms. Synthetic event logs that preserve the statistical and structural characteristics of real processes while maintaining no direct linkage to real-world instances offer an effective means of addressing these limitations. This study presents a systematic review of machine learning–based methods for synthetic event log generation in process mining. A systematic search was conducted in ScienceDirect, PubMed, IEEE Xplore, and SpringerLink through August 2025. Of 1,990 initially identified records, 15 studies were ultimately included following duplicate removal and a three-stage screening process (title, title and abstract, and full text).
Machine learning has emerged as an advanced paradigm for synthetic event log generation, encompassing an evolutionary spectrum from traditional machine learning approaches, such as decision trees, to advanced deep learning and generative artificial intelligence methods, including language models and autoencoders, as well as hybrid approaches. The reviewed studies indicate a growing emphasis on deep learning- and generative AI–based methods. The primary objectives of synthetic event log generation were algorithm evaluation (5 studies, 33%), confidentiality preservation (5, 33%), negative event generation (2, 13%), multidimensional data generation (2, 13%), and scenario analysis (1, 7%). Fidelity, defined as the statistical similarity between synthetic and real data, was assessed in all studies. However, process mining utility, defined as the evaluation of process discovery or conformance checking using synthetic data, was reported in only eight studies (53%), while privacy was explicitly evaluated in only two (13%). The absence of an integrated evaluation framework and the limited assessment of privacy remain major challenges for the reliable comparison of existing approaches.
Future research should focus on integrating the generative capabilities of advanced models with structured process models, such as Petri nets, to ensure the structural validity of generated logs through the integration of data and domain knowledge. Comparative evaluation of different generation mechanisms is also needed to determine how the choice of generation mechanism affects the performance of downstream process mining algorithms. A fundamental question therefore remains unresolved: which generation mechanism(s) can produce synthetic event logs that are most effective for process mining?
کلیدواژهها English