As ROI questions continue to loom over AI in financial services, token optimization and strong change management are becoming fundamental.
Despite decreasing per-token costs, many financial institutions are not seeing a return on their AI investments as token inefficiencies and ineffective prompts turn many models into “slop generators,” Scott Simari, principal at Sendero Consulting, told FinAi News.
Deploying AI just to avoid falling behind competitors is one reason why FIs are struggling to get the most out of their token spend, said Simari, who works with banks, credit unions and private equity firms
“There’s a failure to really forecast that usage and that demand,” he said. “There’s still a level of change management and education that’s needed.”
Annie DeStefano, founder and chief executive of Ann DeStefano Advisory, which provides consulting to FIs and fintechs, told FinAi News that AI infrastructure spending and token usage are “not optimized” at many institutions, hindering ROI.
“In lieu of that being optimized and in lieu of the real revenue impacts being sorted out, how do you really justify that long-term business case?” she said. “I think that’s the piece that companies are grappling with, knowing they have to do something with AI.”
AI has yet to boost profitability for 66% of FIs surveyed by the University of Cambridge’s Judge Business School. The school surveyed 628 industry professionals for an April report.
Model routing
To ensure efficiency, FIs must develop an intelligence layer that evaluates prompts and routes them to the most appropriate LLM for the tasks, Senerdo’s Simari said.
Happen Bank is among the institutions that have taken this approach, CEO Scott Sanborn previously told FinAi News.
The bank is redesigning its infrastructure so employees can easily track token spend, leading to “more cognizant” model selection, he said.
“Do I really need Opus or Fable to do this silly research questionor could I go a few models back and spend one-fifth and get the same result?”
Simari said he urges banks to adopt flexible LLM strategies that use a mix of open-source and frontier models while being “agnostic of any model provider.”
Storing prompts in a separate repository also is important because “there’s not a migration project” when FIs implement new models or readjust their routing strategy, he said.
“You can literally just say, ‘Hey model, go look at the [repository]; this is where your instructions are.’”
The ‘big fat middle’
Concise prompts anchored in relevant context and data are crucial to token optimization, Simari said, adding that too much context can be a “double-edged sword.”
“If you’re sloppy in your context, you could be introducing slop into the model,” he said.
Regarding prompting, there are often large gaps between AI “power users” and skeptics at the same company, Simari said.
Thus, significant value lies in the middle of the spectrum because this typically represents most employees, he said.
“That big fat middle is where all the opportunity exists on lots of levels for the individual user to get a little more sophisticated on that AI theory of mind or what they could be doing with prompting and context,” he said.
Register here for the FinAi Lending Summit, set for Oct. 7-8 in Las Vegas.





