Financial institutions are tapping AI to monitor systems for issues that could trigger outages.
“Outages in financial services are rarely caused by a single-point failure,” Andy Schmidt, vice president and global industry lead for banking at technology and professional services company CGI, told FinAi News.
Outages can result from:

- Issues across interconnected systems;
- Complex data flows;
- Configuration errors;
- Software changes;
- Issues at an outside vendor; or
- Capacity constraints.
CrowdStrike’s July 19, 2024, outage, for example, was caused by an error in a software update, or patch, according to previous reporting.
“With AI, financial institutions can get ahead of these issues by shifting from reactive to a more proactive, even predictive, approach,” Schmidt said.
In an interview with FinAi News, Schmidt discussed how AI can be used to prevent and predict outages at financial institutions. What follows is an edited version of that conversation:
FinAi News: What factors lead to outages at financial institutions?
Andy Schmidt: In some cases, an isolated incident like a bad software patch may act as the initial trigger, but it is the complexity of the environment that amplifies and propagates the disruption.
In our experience at CGI, institutions that are most resilient combine architectural discipline, strong change governance and advanced monitoring, including AI-driven capabilities, to detect and contain issues before they escalate.
FinAi News: What role does AI play in staying ahead of an outage?
AS: Techniques such as employing AI-powered predictive maintenance, for example, can analyze hardware and software performance over time and identify early signs of degradation. AI can continuously analyze system telemetry, identify anomalies and detect patterns that historically precede outages, allowing teams to intervene early before customers are impacted.
AI can also detect and catalog software versions so that you can identify affected servers proactively and triage accordingly. Combining AI-driven insights with human oversight is key to ensuring these signals are interpreted correctly and acted on in time.
FinAi News: What is AI monitoring for when predicting outages?
AS: AI monitors a wide range of signals to predict potential outages, focusing on deviations from normal system behavior. That includes performance metrics like latency, transaction throughput and error rates, as well as infrastructure-level signals such as CPU usage, memory pressure, network congestion and release version.
What makes AI particularly effective is its ability to correlate signals across multiple layers — including applications, infrastructure and third-party services — to identify gaps and patterns humans would struggle to detect manually. For example, a slight increase in response times combined with a configuration change and unusual traffic patterns might be a new software feature or it might signal an emerging issue.
Over time, AI establishes a baseline of what “normal” or “healthy” looks like for each specific environment. It can flag early warning indicators hours before a traditional monitoring threshold.
In our experience at CGI, this cross-layer visibility is critical to identifying warning signs early and preventing localized issues from escalating into broader disruptions.
FinAi News: How does AI change IT strategies when it comes to monitoring systems and overall resiliency plans?
AS: AI is reshaping IT strategies by shifting the focus from maintaining uptime to building true operational resilience and adaptability. It enables more dynamic and intelligent approaches to system monitoring, moving beyond static thresholds toward continuous, context-aware analysis.
From a resiliency standpoint, it also enables much more sophisticated scenario planning through stress testing and simulation. As a result, IT strategies become less about uptime alone and more about adaptability, allowing institutions to better anticipate and respond to disruptions across complex environments.
However, AI does not replace human decision-making. Its effectiveness depends on how well it is integrated into governance frameworks and operational processes. Organizations that successfully adopt AI embed it within broader resilience frameworks, aligning technology, governance and operations to respond effectively to disruption.
FinAi News: For FIs that are exploring AI for resilience, where should they start?
AS: It’s important that institutions focus on critical systems where downtime has the greatest impact on the business and customers. AI should be applied with clear intent, not as a standalone initiative, but as part of a broader resilience strategy with having a human-in-the-loop is non-negotiable.
From there, the priority should be building a strong data foundation. AI is only as good as the data it’s trained on, so integrating high-quality, real-time telemetry across systems is essential. Institutions can then expand that intelligence across the broader ecosystem, including third-party dependencies, including partners and marketplaces.
Register here for the FinAi Lending Summit, set for Oct. 7-8 in Las Vegas.






