Rapidly scaling technology can foster innovation at the bank level, but it can also create operational risks — from data breaches and environmental failures to other unexpected events.
Central to guarding against such threats is cultivating resiliency in bank infrastructure architecture and development. Following are four best practices for enabling operational technical resiliency:
Practice failure mode analysis

Critical to building resiliency at the bank is identifying points of disruption or failure. Failure mode analysis (FMA), a process that “raises questions” before it is too late, is a crucial philosophy when building architectures for resiliency, Neeleesh Prabhu, managing director of architecture and enterprise services in information technology at trade settlement giant Depository Trust and Clearing Corporation (DTCC), told Bank Automation News.
“What happens if the database your application relies on is no longer available? What happens if you are reliant on another application to provide you with certain data, and if that application is not available?” Prabhu said. “FMA is all about raising these questions, discussing what the response would be, and designing that into the application itself.”
Align architectural design with failure response
Banks should align the design with failure event responses when building the architectural pattern of the technology infrastructure. A high-quality pattern facilitates easy solutions for common failures while also enabling developers to custom-align responses to specific cases and contexts. Prabhu pointed to a data reconciliation pattern at DTCC as an example of successful pattern-response alignment.
“We leverage it [data reconciliation pattern] when an application has experienced a problem in one data center, and now needs to relocate to another data center,” he said. “The first thing you do is check the state of the data integrity within the application. And the data reconciliation pattern enables an easy solution to do that.”
Automate application testing
Automation is the “key component” of how DTCC tests the integrity of an application, and it can help build a resiliency scorecard that allows developers to analyze relevancy and effectiveness of failure responses, Prabhu said.
“To enable the scorecards, especially in today’s environment, where a lot of application development gets done via agile mechanisms, the automated nature of testing plays a very critical role there,” he said.
Another testing option to consider is chaos testing, a “highly automated” capability that cuts across different layers of a technology stack to simulate unpredictable scenarios, Prabhu noted.
Design responses for on-prem, cloud data centers
As hybrid cloud models become increasingly common across financial services, banks should focus on designing failure responses that work across on-premises and cloud data centers, Prabhu said. Operational rotation, where an application is rotated from its current data center to another data center, can help banks gain “resiliency power” against geographical and electrical issues, natural disasters and data breaches.
Observability is key when deploying operational rotation for resiliency in a hybrid cloud ecosystem, he told BAN.
“From an operating perspective, you need to be able to realize the pattern through alerting and monitoring capability. You could also call it observability,” Prabhu said. “One best practice is to leverage instrumentation and tooling that cuts across data centers and the hybrid cloud environment, so that when there is a failure occurring, you are notified quickly.”
When deploying operational rotation at the bank, automation is equally important. Banks should ensure that the runbook, or the system of operations and procedures that effectuates transfers of applications, is automated to avoid errors.
Bank Automation Summit Fall 2022, taking place Sept. 19-20 in Seattle, is a crucial event on automation and automation technology in banking. Learn more and register for Bank Automation Summit Fall 2022.






