Creating a central repository of data, or a “data lake,” can enhance bank systems’ interoperability and data utilization, but the practice can also backfire with the wrong approach and become a detriment.
That’s the message from Michael Hom, head of financial services solutions at database systems provider InterSystems, whose customers include $407.5 billion Bank of New York Mellon, $103.7 billion Credit Suisse and $245.4 billion HSBC. InterSystems’ solutions help financial services providers handle data “from a low-latency standpoint as well as a volume standpoint,” Hom told Bank Automation News.
At a high level, a data lake can help an organization improve access to its data, taking all data sources and combining them in one place. Banks and financial institutions (FIs) with this strategy “want to build, harness and harvest all that rich data that makes it valuable, and be able to make that accessible to the rest of their firm,” Hom said.
Other drivers for creating a data lake include meeting regulatory compliance and reporting needs or accomplishing data-driven projects, such as those involving automation and machine learning (ML).
InterSystems tries to ascertain whether clients have a holistic, inclusive view of data, Hom said. Even with a data lake, various departments of an organization can access, use and add to data independently and differently, which can create siloes, inefficiencies and other undesirable effects.
3 data lake elements to consider:
- The time element;
- What data is present; and
- Relationships between datasets.
Part of the challenge presented by a data lake is that data is dynamic and constantly evolving, Hom noted, and any static understanding of it will become inaccurate. “If you have a ‘snapshot’ view of the world where you just stored [the data] and that’s it, it becomes stale. It can also become outdated,” he said.
As a bank pools its data, not only is that data constantly changing, it also becomes increasingly diverse and complex, and the bank may not be connecting the dots between data points.
“The real root problem of data lakes is, OK, you have all your data in place — that’s great, but do you know what you have and how good it is?” said Hom. Not only must a bank or FI have that information, it also needs to understand the relationships between data elements.
For example, a bank may have a particular customer who has a loan and a credit card and owns government bonds. It can be a challenge to bring all that information together — which may be input in different ways and at different times — and to understand that it relates to the same customer.
“We sometimes say that you can have data, but then equally important is having all the metadata,” Hom said. “Metadata also includes relationships across your datasets. How do you get all that coming together?”
Keep it clean
Knowing how to efficiently access and manage a data lake is imperative, and must be done continuously. Data management is “not a one-time task,” and lack of control over data assets can turn a lake into a swamp.
A data lake, in its ideal form, is a unified, collective resource with standardized, well-managed, updated and accessible data.
On the other hand, a data swamp, while also a centralized location, is in many ways the opposite: a collection of unstructured, disparate, poorly managed and/or unusable data. A data swamp could be a data lake with data that lacks contextual information, is irrelevant, is ungoverned, and lacks automation and a data-cleaning strategy, according to Information Age.
“That’s why I talk about the cleanliness of your data,” Hom told BAN. “You could have a stale copy of data and someone made a modification to it,” he noted. This would muddy the data lake.
Proper data utilization and management should be viewed as critical, since data drives decisions and processing, and processes are increasingly becoming automated, he said. Banks and FIs should assume their competitors are combing through and using their data in a variety of ways to gain advantages.
“Everyone else is utilizing data trying to find an edge, and if you’re not incorporating the right data, or are using the wrong data — or you had it stored somewhere, copied the data and someone modified it by accident or on purpose — and you’re using that, you can generate wrong insights, or you potentially could report it to regulators incorrectly,” Hom told BAN. “So that can have larger ramifications.”






