Some vendors claim their products can automate data, eliminating the need for data scientists, but experts say that’s not entirely the case, even when using artificial intelligence.
Such claims sound “more like marketing terms than then actual real application,” RJ Sherman, vice president of innovation at $183.3 billion Citizens Bank, told Bank Automation News.

In banking, one key automation challenge is pulling together data from multiple applications to make the data easier to work with, Sherman said, adding that he has seen some applications do this well.
As with most forms of automation, the reality is more about automating aspects of the process, rather than providing an end-to-end automated solution. For example, it’s possible to automate quality checks, data profiling and the generation test data, said Karen Lopez, a data architect and an independent data analyst.
“All of those things can be automated,” Lopez said. “But the analysis of data is something that generally involves a human touch of some type.”
Yet, it remains the data scientist’s job to train data models to recognize anomalies, she added.
AI and data automation
When automating data, there is a place for AI, which can be used to “fill in the data gap,” Stuart Tarmy, global director of financial services industry solutions at Aerospike, told BAN. Aerospike is a NoSQL, or non-tabular, non-relational, database that powers real-time networks for PayPal, Visa and Charles Schwab. Aerospike also powers the recommendation engine for home decor company Wayfair, which combines transactional data with anonymized demographic data to offer recommendations to customers, Tarmy said. Often the data Aerospike leverages comes from third-party sources that have different levels of quality and gaps.
“Before you can really do the analytics, you have to bring [the data] in and then standardize it, cleanse it,” Tarmy said. “Oftentimes, there’s data gaps and you can use AI to fill in the data gap, so that is a very effective use for AI.”
Another potential use case for AI in data automation is reconciliation, Tarmy said, giving the example of an unspecified large multinational bank that uses more than 100 people to reconcile trades during the overnight hours.
“AI machine learning is a way that you can automate some of that reconciliation,” Tarmy said. “You’ll never automate all of it because there’s some things a person has to dig into … but you can automate probably a lot of that to do it faster and cleaner.”
Models that have been pre-trained on AI also can be used to address some easy problems, Lopez said. For instance, a pre-trained model could sort a batch of photos into subjects — e.g., this is a blueberry, this is a person. Lopez said a favorite demo exemplifying AI’s use is an algorithm that removes from a batch of photos any with a finger obscuring part of the picture.
Data for the non-data scientist?
Managing bank data can be a challenge because of legacy systems, mergers and acquisitions, and the mix of processes, said Christian Nentwich, CEO and founder of data engineering company Duco.
Duco, which uses machine learning to automate data quality and reconciliation for structured data, can take data in different formats from disparate systems and refine the data into a common format. This is typically a manual process, Nentwich said.
Approximately half of the company’s 100 clients are banks, which use the solution for data quality and data reconciliation. Capital market trading is a common use case, including trade data, position and reference data, prices and risk indicators. Retail banks use Docu primarily for their payments data, Nentwich said.
“We rely on machine learning and other modern techniques to make the data accessible to you and clean it up,” Nentwich said. “We just take the data as it is, and we rely on the machine learning and other algorithms to standardize the data for us, but without generating a project — just relying much more on machine learning and putting the control in the hands of the business.”
Lopez, the data architect, agreed that non-programmers can, to some extent, handle parts of the data pipelines, and some data transformations can be performed with low-code solutions.
Low-code tools and automation solutions are useful to data scientists, because better data prep and data sourcing free up time to do the real work of a data scientist, Lopez added. “Data scientists spend 80% of their time sourcing, cleansing and prepping data,” and the remaining 20% designing the model to analyze the data, Lopez said.
“Finding out which data you should be using, that’s the analysis part — that almost always requires a human to understand ‘What’s this data?’” Lopez said.






