Abstract
Modern cloud computing and distributed microservice architectures exhibit high degrees of heterogeneity and operational dynamism, making system failure prediction and reliability assurance notoriously challenging. Traditional reliability management frameworks largely rely on reactive mitigation strategies triggered after anomalies exceed critical thresholds, inevitably incurring system downtime and degraded Quality of Service (QoS). This paper presents a proactive fault prediction framework that couples deep sequential log anomaly detection with dynamic Bayesian network (DBN) modeling to preemptively identify, track, and mitigate failure propagation in distributed environments. Unstructured operational logs are first parsed into semantic event templates via an optimized clustering parser and mapped into vector representations using contextual log embeddings. A bidirectional sequential neural model subsequently detects subtle behavioral anomalies within rolling temporal windows. To resolve the causality of observed anomalies and capture cascading failure paths across interdependent services, detected anomaly vectors parameterize a dynamic Bayesian network calibrated against service call graphs. Evaluated on large-scale open distributed benchmarks (HDFS, OpenStack, and a containerized microservices testbed), our proposed approach achieves an average F1-score of 95.8% in anomaly identification and reliably forecasts impending system-level faults with an average lead time of 6.4 minutes prior to service failure. These empirical findings demonstrate that incorporating probabilistic causal reasoning into log-driven observability platforms substantially reduces Mean Time to Resolution (MTTR) and provides actionable diagnostic lead times for autonomous self-healing cloud systems.